Filter already-ingested events in memory instead of per-row

Python · Python · intermediate · modification

Speeds up de-duplication. The old path called store.exists(id) once per incoming event — one DB round-trip each. This pulls the full set of stored ids up front with a single store.all_ids() call and filters the batch in memory, dropping tens of thousands of round-trips per ingest. Verified against a 500-event batch that de-duped correctly.

new_events runs synchronously inside the event-ingest request handler. Production scale: each ingest batch is ~50,000 incoming events, and store.all_ids() returns the ~1,000,000 ids already persisted, as a plain in-memory list. Order must be preserved because downstream processing is position-sensitive.

Requirements

Files touched

--- app/ingest/dedup.py
     """Return only the events whose id is not already in the store.
 
-    Calls store.exists(id) once per event — one DB round-trip each.
+    Pull the stored ids once, then filter in memory — no per-event round-trip.
     """
-    return [event for event in incoming if not store.exists(event["id"])]
+    existing_ids = store.all_ids()
+    return [event for event in incoming if event["id"] not in existing_ids]
 

Review this PR

Python practice