Add the overnight ticket summariser with provider-aware backoff
Python · LLM apps · beginner · greenfield
Adds the nightly summariser for the triage queue: one pinned-snapshot model call per ticket, a constant one-sentence system prompt, and app-owned backoff on 429s — the SDK's built-in retries are disabled so the loop honours the provider's own delay hint and fails loudly with a RuntimeError once the 4-attempt budget is exhausted. Verified the happy path against the staging key and the fallback backoff schedule (1 s / 2 s / 4 s) with mocked 429 responses. Tickets run sequentially to stay comfortably under the org rate limit.
Nightly batch job run by support ops: 50–500 tickets per run inside a 2-hour cron window, nobody waiting on it live. Several teams share the provider org, so 429s carrying a `retry-after-ms` hint are routine during batch hours; headerless 429s are rare.
Requirements
- The worker calls the provider through the `openai` SDK's async client with the pinned dated snapshot `gpt-4o-mini-2024-07-18`. The client is constructed with `max_retries=0`, so the retry loop in `summarize_ticket` is the only retry mechanism: at most `MAX_ATTEMPTS = 4` provider calls per ticket.
- On a `RateLimitError` (HTTP 429) the provider includes a `retry-after-ms` response header: a non-negative integer as a decimal string, in **milliseconds**, saying how long the caller should wait before retrying. The handler must wait exactly that many milliseconds — `retry-after-ms: 800` means a 0.8-second wait, not an 800-second one — except that hints above `MAX_HINT_MS = 60000` (one minute) are clamped to one minute. A negative value or one that is not a decimal integer counts as unparsable.
- When the header is absent or unparsable, the worker falls back to exponential backoff: `BASE_BACKOFF_SECONDS * 2 ** (attempt - 1)` seconds (1 s after the first failure, 2 s after the second, 4 s after the third).
- If the fourth attempt is also rate-limited, `summarize_ticket` raises a `RuntimeError` chained from the last `RateLimitError`. Only `RateLimitError` is retried; any other exception propagates to the caller immediately.
- A successful provider response carries exactly one choice, and its `content` is the summary (`""` when the content is `None`). `summarize_all` processes tickets sequentially in input order and returns a `ticket_id` → summary mapping; ticket ids are unique within a batch.
Files touched
- app/summariser.py
--- app/summariser.py +"""Overnight summarisation worker: support tickets -> one-line summaries. + +Backoff is owned by this module (the client is built with ``max_retries=0``), +so the retry loop below is the only retry mechanism in the app. +""" + +from __future__ import annotations + +import asyncio +import logging +from dataclasses import dataclass + +from openai import AsyncOpenAI, RateLimitError + +logger = logging.getLogger(__name__) + +MODEL = "gpt-4o-mini-2024-07-18" # pinned dated snapshot, per ADR-011 +MAX_ATTEMPTS = 4 # provider calls per ticket, hard cap +BASE_BACKOFF_SECONDS = 1.0 # fallback when the provider sends no delay hint +MAX_HINT_MS = 60_000 # never honour a delay hint longer than one minute + +SYSTEM_PROMPT = ( + "You summarise support tickets for the triage team. Reply with exactly " + "one sentence (at most 40 words) stating the customer's core problem." +) + +_client = AsyncOpenAI(max_retries=0) # reads OPENAI_API_KEY from the environment + +