Stream the help-centre assistant's replies over SSE with a token-trimmed history

Python · LLM apps · advanced · greenfield

Adds the streaming answer path for the help-centre widget: `POST /chat/stream` relays model deltas as SSE frames and saves the completed answer to the per-user history store. The context window is capped by counting tokens with tiktoken and trimming the oldest messages first, so long-lived conversations can never push the call over the budget — and a message too big to ever fit is rejected with a 413 before a byte is streamed. Provider failures after the stream opened surface as one explicit `error` frame with the detail kept in the server log, and a failed or disconnected stream never stores its partial text as the answer. Exercised streamed replies, a killed provider connection mid-stream (error frame, history untouched), and a 500-turn conversation (request stayed under 7168 tokens with the oldest turns dropped).

Customer-facing help-centre widget. The endpoint is authenticated: `current_user_id` (in `app/auth.py`, outside this PR) is an existing FastAPI dependency that returns the signed-in user's id after JWT validation. `tiktoken` is already a pinned dependency of the service (the nightly summariser uses it). Deployment is a single uvicorn worker and conversation history is deliberately ephemeral — an in-process store is the documented product decision, not an oversight; a restart starting conversations fresh is acceptable. The widget sends one message at a time: it disables its input until the terminal frame arrives, so the endpoint never has to serialise concurrent requests from the same user. It renders `delta` frames as they arrive and shows a retry affordance on an `error` frame.

Requirements

Files touched

--- app/chat_history.py
+"""Per-user chat history and the tiktoken-counted context budget."""
+
+from __future__ import annotations
+
+import tiktoken
+
+# The encoding gpt-4o-mini uses. Pinned explicitly (rather than derived with
+# tiktoken.encoding_for_model) so the counts never shift underneath a running
+# service; MODEL in chat_stream.py and this constant are bumped together.
+_ENCODING = tiktoken.get_encoding("o200k_base")
+
+# App-set cost cap on a single request, deliberately far below the model's
+# real window so the per-message protocol overhead that tiktoken does not
+# model is absorbed safely. The reply budget is reserved out of it before
+# any history is trimmed.
+MAX_CONTEXT_TOKENS = 8192
+REPLY_TOKEN_BUDGET = 1024  # sent as max_tokens on every call
+REQUEST_TOKEN_BUDGET = MAX_CONTEXT_TOKENS - REPLY_TOKEN_BUDGET  # 7168
+
+# Per-user cap on stored messages: far more than the request budget can
+# ever carry, so it only bounds memory over a long-lived conversation.
+MAX_STORED_MESSAGES = 200
+
+
+def count_tokens(text: str) -> int:
+    """Token count of one string under the pinned encoding.
+
+    `disallowed_special=()` counts the text as ordinary text: by default
+    tiktoken raises on strings that spell a special token such as

Review this PR

Python practice