Turn the invoice-note summariser into a batch endpoint with per-row outcomes
Python · LLM apps · beginner · greenfield
Turns the one-off invoice-note summariser into a batch endpoint: `POST /summaries/batch` walks the rows in order, streams each model call and accumulates the text deltas, and when the model asks for `normalize_amount` the streamed fragments are merged by index, the tool runs once per `tool_call_id` and the result goes back as a `role: "tool"` message before the follow-up call. Tool-call chunks are expected here — the notes are full of money amounts — and a chunk with no text contributes nothing to the summary. Off-schema tool arguments come back to the model as a tool error instead of raising, and `normalize_amount` itself never throws. Exercised against a seeded batch of 40 invoice notes; the amounts were normalized, the summaries came back one sentence each and `failed` was empty.
Back-office nightly job: an operator triggers it from the admin UI with batches of up to 50 invoice-note rows. Notes are free text and usually contain money amounts in mixed formats (`1 200,50 EUR`, `$1,200.50`, `EUR 1200`), so the model routinely calls `normalize_amount`. Batches run during shared-org peak hours, so a `RateLimitError` that survives the SDK's two retries, or a stream that times out mid-way, hits one row in most nightly runs. The caller reads `failed` to decide which rows to re-queue for the next run.
Requirements
- `POST /summaries/batch` accepts `{"rows": [Row, ...]}` with 1–50 rows; a `Row` is `{"row_id": int, "text": str}` where `row_id` is 1–10000, `text` is 1–4000 characters, and `row_id` values are unique within a request (all bounds and the uniqueness are enforced by the Pydantic request models, so the handler only ever sees valid rows). The endpoint answers with one JSON object `{"summaries": {<row_id as string>: <string>}, "failed": [<row_id>, ...]}`; nothing is streamed back to the client.
- Rows are processed sequentially in request order through the streaming API: `AsyncOpenAI` with the pinned dated snapshot `gpt-4o-mini-2024-07-18`, `stream=True`, `temperature=0` and `max_tokens=200`. Because the stream yields a chunk per delta, `chunk.choices[0].delta.content` can be `None` on chunks that carry no text — the chunks that carry `tool_calls` fragments and the final chunk that carries only `finish_reason` — and a row's summary is exactly the non-`None` text fragments of its calls concatenated in arrival order (the empty string when the model streamed no text at all).
- The batch is error-isolated per row: a row that fails for ANY reason — an exception while reading its stream, a provider error such as `RateLimitError` / `APITimeoutError` / `APIStatusError`, a malformed chunk — must have its `row_id` appended to `failed` and the loop must continue with the next row. The request must never fail because of a single row: no per-row error may propagate out of the handler, and the endpoint never answers 500 for one bad row. Every `row_id` of the request appears in exactly one of `summaries` and `failed`.
- The model has exactly one tool, `normalize_amount`, whose single required argument `raw` is declared as a string: it is a pure function with no side effects and no I/O, and it runs at most once per `tool_call_id`. Its JSON result goes back to the model as the content of the `role: "tool"` message with the matching `tool_call_id` and never becomes part of a summary. If the model made tool calls, exactly one follow-up streamed call — the same messages plus the assistant message carrying `tool_calls` and then the tool results — produces that row's summary text; if the model made no tool calls, the text of the first call is the summary.
- Tool-call fragments of a streamed response are merged by their `index` (0-based, assigned by the API): the `id`, `function.name` and `function.arguments` pieces of the same index are concatenated in arrival order into one call. Arguments are parsed with `json.loads`; malformed JSON, a non-object, a missing `raw`, a non-string `raw` or any unexpected property are not errors — they are returned to the model as the tool result `{"error": "…"}` so it can rephrase.
- The SDK's built-in retries (two, with backoff) are the only retry mechanism: the service has no retry loop of its own, and `asyncio.CancelledError` (a `BaseException`) must not be swallowed by a row's error handling.
Files touched
- app/batch_summarizer.py
--- app/batch_summarizer.py +"""Nightly batch summariser for invoice notes. + +One streamed model call per row, plus the `normalize_amount` tool the model +uses to turn a free-text money amount into integer cents. The batch endpoint +reports per-row outcomes: a row either lands in `summaries` or its id lands in +`failed`. +""" + +from __future__ import annotations + +import json +import re +from decimal import ROUND_HALF_UP, Decimal, InvalidOperation +from typing import Any + +from fastapi import APIRouter +from openai import AsyncOpenAI +from openai.types.chat import ChatCompletionMessageParam +from pydantic import BaseModel, Field, model_validator + +MODEL = "gpt-4o-mini-2024-07-18" # pinned dated snapshot, per ADR-011 +MAX_ROWS = 50 + +SYSTEM_PROMPT = ( + "You summarise invoice notes for the finance back office. Reply with one " + "sentence of plain text and nothing else. When a note mentions a money " + "amount, call normalize_amount to get the amount in cents before you " + "write the summary; never mention the tool or its raw result." +)