Add the batch post-record fetch handler
Python · Python · intermediate · greenfield
Adds the batch fetch handler: takes a list of post URIs, fetches each record from the store concurrently, and returns the ones that exist. Ran it against a handful of URIs and it returned the records quickly.
A high-throughput read path in a social API. `store.get_record` is a network round-trip against a backing store with a bounded connection pool. Batch sizes are driven by callers — a newly added internal service submits batches of 15,000-20,000 URIs at a time.
Requirements
- get_post_records(uris) fetches each URI's record concurrently via `await store.get_record(uri)` and returns {uri: record} for the URIs that have a record (skipping None).
- A batch may contain anywhere from 1 to 20,000 URIs — an internal service submits large batches. The store's connection pool is finite; the handler must not open more simultaneous round-trips than the pool can serve.
- The number of in-flight fetches must stay bounded regardless of batch size (a semaphore, a worker pool, or fixed-size chunks) so one large batch cannot exhaust connections.
Files touched
- app/records.py
--- app/records.py
+import asyncio
+from typing import Dict, List
+
+# `store` is the shared async backing-store client, injected at import time.
+from app.store import store
+
+
+async def get_post_records(uris: List[str]) -> Dict[str, Dict]:
+ """Fetch the record for each URI concurrently and return {uri: record}
+ for the URIs that have one (skipping the ones that return None)."""
+ if not uris:
+ return {}
+
+ # Create a list of coroutines, one for each URI fetch operation
+ fetch_tasks = [store.get_record(uri) for uri in uris]
+
+ # Run all fetch tasks concurrently
+ results = await asyncio.gather(*fetch_tasks, return_exceptions=False)
+
+ records = {}
+ for uri, record in zip(uris, results):
+ if record is not None:
+ records[uri] = record
+
+ return records
+