Filter already-ingested events in memory instead of per-row
Python · Python · intermediate · modification
Speeds up de-duplication. The old path called store.exists(id) once per incoming event — one DB round-trip each. This pulls the full set of stored ids up front with a single store.all_ids() call and filters the batch in memory, dropping tens of thousands of round-trips per ingest. Verified against a 500-event batch that de-duped correctly.
new_events runs synchronously inside the event-ingest request handler. Production scale: each ingest batch is ~50,000 incoming events, and store.all_ids() returns the ~1,000,000 ids already persisted, as a plain in-memory list. Order must be preserved because downstream processing is position-sensitive.
Requirements
- new_events(incoming, store) returns only the events from `incoming` whose id is not already stored, preserving their original order.
- The de-dup must not do a database round-trip per event — fetch what it needs once, then filter in memory.
- The in-memory filter must run in linear time in the two sizes (incoming batch and stored ids); this runs inside the ingest request path and cannot stall it.
Files touched
- app/ingest/dedup.py
--- app/ingest/dedup.py
"""Return only the events whose id is not already in the store.
- Calls store.exists(id) once per event — one DB round-trip each.
+ Pull the stored ids once, then filter in memory — no per-event round-trip.
"""
- return [event for event in incoming if not store.exists(event["id"])]
+ existing_ids = store.all_ids()
+ return [event for event in incoming if event["id"] not in existing_ids]