Paginate the nightly active-user export
Python · Python · intermediate · modification
Refactors the nightly export off the single fetch_active() query, which had started timing out as the active set grew, onto paged reads: count the active users, then pull them PAGE_SIZE at a time and concatenate the pages in order. Verified against the 1000-row staging snapshot, which exported cleanly.
collect_active_users feeds the nightly subscriber export; downstream billing reconciles against that file row for row, so a user missing from the export is a user who does not get invoiced.
Requirements
- collect_active_users must return every active user for the nightly export, in page order.
- The active set outgrew a single query, so page through the repository with repo.fetch_active_page(page, PAGE_SIZE), which returns the rows in the half-open window [page * PAGE_SIZE, page * PAGE_SIZE + PAGE_SIZE) for a zero-based page index, and an empty list past the end.
- repo.count_active() returns the total number of active users. That total is arbitrary — whatever the query matched — and is NOT guaranteed to be a multiple of PAGE_SIZE.
- The export runs against a read-replica snapshot, so the active count is stable for the duration of the run (count and fetches see the same rows).
Files touched
- app/jobs/export.py
--- app/jobs/export.py def collect_active_users(repo): - """Return every active user for the nightly export.""" - return repo.fetch_active() + """Return every active user for the nightly export. + The active set outgrew a single query, so we page through the + repository PAGE_SIZE rows at a time and concatenate the pages in order. + """ + total = repo.count_active() + page_count = total // PAGE_SIZE + users = [] + for page in range(page_count): + users.extend(repo.fetch_active_page(page, PAGE_SIZE)) + return users +