Stream the help-centre assistant's replies over SSE with a token-trimmed history
Python · LLM apps · advanced · greenfield
Adds the streaming answer path for the help-centre widget: `POST /chat/stream` relays model deltas as SSE frames and saves the completed answer to the per-user history store. The context window is capped by counting tokens with tiktoken and trimming the oldest messages first, so long-lived conversations can never push the call over the budget — and a message too big to ever fit is rejected with a 413 before a byte is streamed. Provider failures after the stream opened surface as one explicit `error` frame with the detail kept in the server log, and a failed or disconnected stream never stores its partial text as the answer. Exercised streamed replies, a killed provider connection mid-stream (error frame, history untouched), and a 500-turn conversation (request stayed under 7168 tokens with the oldest turns dropped).
Customer-facing help-centre widget. The endpoint is authenticated: `current_user_id` (in `app/auth.py`, outside this PR) is an existing FastAPI dependency that returns the signed-in user's id after JWT validation. `tiktoken` is already a pinned dependency of the service (the nightly summariser uses it). Deployment is a single uvicorn worker and conversation history is deliberately ephemeral — an in-process store is the documented product decision, not an oversight; a restart starting conversations fresh is acceptable. The widget sends one message at a time: it disables its input until the terminal frame arrives, so the endpoint never has to serialise concurrent requests from the same user. It renders `delta` frames as they arrive and shows a retry affordance on an `error` frame.
Requirements
- `POST /chat/stream` accepts `{"message": string}` (1–2000 characters, bounds enforced by FastAPI validation) behind the existing `current_user_id` auth dependency, and responds with Server-Sent Events (`media_type="text/event-stream"`). Every SSE frame is exactly one JSON object on a `data:` line followed by a blank line.
- The event vocabulary is closed: zero or more `{"type": "delta", "text": …}` frames carrying non-empty model text, then exactly one terminal frame — `{"type": "done"}` when the model's stream completed, or `{"type": "error", "message": …}` when anything failed after streaming opened. Chunks with an empty `choices` list, a `None` delta, or `None`/empty `delta.content` are skipped: they neither produce a delta frame nor fail the stream. `finish_reason` is not part of the contract; a reply truncated by the token limit counts as completed and is persisted as received.
- Per-user history lives in the in-process store (single uvicorn worker; history is ephemeral by a documented product decision — see context). The system prompt is a module constant and is never stored; user text only ever enters the model call as `user`/`assistant` messages. A new user message is persisted once the request is accepted, and it stays in history even if the stream later fails. The store keeps at most `MAX_STORED_MESSAGES = 200` messages per user, evicting the oldest first.
- The assistant reply is persisted exactly once, only after the model's stream completes without error, as the concatenation of all received delta text (an empty string when the model produced no content) — before the `done` frame is emitted. On any exception during the provider call or the stream, and on a client disconnect while the provider stream is still running, no assistant text is persisted: partial replies are never stored as completed answers. (A disconnect after the model's stream completed may leave the finished answer stored — it is a complete answer, not a partial one.)
- Any failure after the stream has opened yields exactly one `error` event with a generic message; the exception detail goes to the server log (`logger.exception`) and never on the wire. The SDK's two built-in retries are the only retry mechanism — the app itself never retries.
- Every model call must fit an app-set budget of `MAX_CONTEXT_TOKENS = 8192` tokens, counted with tiktoken's `o200k_base` encoding as `role` + `content` tokens per message (the budget sits far below the model's real window, so unmodelled per-message protocol overhead is absorbed safely). The reply budget `max_tokens = REPLY_TOKEN_BUDGET = 1024` is reserved out of it first, so the request-side budget is 7168 tokens.
- Before the call, history is trimmed oldest-non-system-message-first until the request fits the 7168-token budget; the system prompt and the newest message are never dropped. If even those two exceed the budget, the endpoint responds HTTP 413 before streaming starts and persists nothing — a rejected request leaves the stored history unchanged.
- The model is the pinned dated snapshot `gpt-4o-mini-2024-07-18`.
Files touched
- app/chat_history.py
- app/chat_stream.py
--- app/chat_history.py
+"""Per-user chat history and the tiktoken-counted context budget."""
+
+from __future__ import annotations
+
+import tiktoken
+
+# The encoding gpt-4o-mini uses. Pinned explicitly (rather than derived with
+# tiktoken.encoding_for_model) so the counts never shift underneath a running
+# service; MODEL in chat_stream.py and this constant are bumped together.
+_ENCODING = tiktoken.get_encoding("o200k_base")
+
+# App-set cost cap on a single request, deliberately far below the model's
+# real window so the per-message protocol overhead that tiktoken does not
+# model is absorbed safely. The reply budget is reserved out of it before
+# any history is trimmed.
+MAX_CONTEXT_TOKENS = 8192
+REPLY_TOKEN_BUDGET = 1024 # sent as max_tokens on every call
+REQUEST_TOKEN_BUDGET = MAX_CONTEXT_TOKENS - REPLY_TOKEN_BUDGET # 7168
+
+# Per-user cap on stored messages: far more than the request budget can
+# ever carry, so it only bounds memory over a long-lived conversation.
+MAX_STORED_MESSAGES = 200
+
+
+def count_tokens(text: str) -> int:
+ """Token count of one string under the pinned encoding.
+
+ `disallowed_special=()` counts the text as ordinary text: by default
+ tiktoken raises on strings that spell a special token such as