Add the meeting-notes summariser endpoint
Python · LLM apps · beginner · greenfield
Adds the meeting-notes summariser endpoint. `POST /summarize` counts the document's tokens with tiktoken, rejects anything over the 8,000-token input cap with a 422, and asks the pinned `gpt-4o-mini-2024-07-18` for decisions, action items and open questions. The reply budget is sized so the whole request stays inside the model's context window, and provider errors surface as a clean 502. Ran it against a few sample notes — summaries come back crisp.
An internal service; real transcripts range from a few hundred tokens up to the 8,000-token input cap, and clients render the returned summary verbatim.
Requirements
- `POST /summarize` accepts JSON `{"document": string}` and returns `{"summary": string}`. The document enters the model only as a `user` message; the system prompt is a constant string and no request content ever modifies it.
- The model call uses the pinned snapshot `gpt-4o-mini-2024-07-18` with `temperature=0` — never a floating alias. Its context window is 128,000 tokens, so any accepted document plus the reply budget always fits.
- Every summary gets the same fixed reply budget: the model call must set `max_tokens` to the constant `SUMMARY_MAX_TOKENS = 300` declared in `app/config.py`, whatever the document length. The reply budget must not be derived from the input size.
- Token counts are computed with `tiktoken`'s `cl100k_base` encoding (the service's required counting encoding, used for the input cap only), counting the document as ordinary text — special-token strings such as `<|endoftext|>` inside a document are plain characters, never an error. Documents longer than `MAX_INPUT_TOKENS = 8,000` tokens, and empty or whitespace-only documents, are rejected with HTTP 422 before any model call.
- If the provider call raises any `openai` error, the endpoint responds HTTP 502 with a generic detail message; stack traces and provider payloads never reach the client.
Files touched
- app/config.py
- app/summarizer.py
- app/main.py
--- app/config.py +"""Static settings for the meeting-notes summariser service.""" + +# ADR-011: always a dated snapshot, never a floating alias. The pinned +# model's context window is 128,000 tokens — far more than this service +# ever sends. +MODEL_NAME = "gpt-4o-mini-2024-07-18" + +# Every summary gets the same fixed reply budget: the model call sets +# max_tokens to this constant, whatever the document length. +SUMMARY_MAX_TOKENS = 300 + +# Documents longer than this (tiktoken cl100k_base) are rejected with 422 +# before any model call. +MAX_INPUT_TOKENS = 8_000 + +SYSTEM_PROMPT = ( + "You are a meeting-notes summariser. Summarise the user's document as " + "concise bullet points covering decisions, action items with owners, and " + "open questions. Base the summary only on the document text." +) + --- app/summarizer.py +"""Meeting-notes summariser: one LLM call per request, document in, summary out.""" + +import tiktoken +from openai import AsyncOpenAI + +from app.config import MAX_INPUT_TOKENS, MODEL_NAME, SYSTEM_PROMPT +