Stream the support-chat answer to the widget over SSE

Python · LLM apps · intermediate · greenfield

Streams the support-chat answer to the widget: `POST /chat/stream` relays model deltas as SSE frames and saves the completed answer to the per-user history store. Provider hiccups are logged and the stream is closed cleanly so the browser is never left hanging on a dead connection. Text-less chunks are skipped, and each turn sends the system prompt plus at most the last 20 stored messages. Exercised against a live key — frames arrive incrementally and the answer lands in history.

A customer-support chat widget. The browser renders each `delta` frame as the assistant's reply and treats the `done` frame as "this answer is complete"; the stored history is what the next turn sends to the model as context.

Requirements

Files touched

--- app/services/chat_stream.py
+"""Streaming chat answer: OpenAI deltas -> SSE frames -> history store."""
+
+import json
+import logging
+import os
+from collections.abc import AsyncIterator
+
+import httpx
+import openai
+from openai import AsyncOpenAI
+
+from app.services import history
+
+logger = logging.getLogger(__name__)
+
+MODEL = "gpt-4o-mini-2024-07-18"
+MAX_HISTORY_MESSAGES = 20
+
+client = AsyncOpenAI(api_key=os.environ["OPENAI_API_KEY"])
+
+SYSTEM_PROMPT = (
+    "You are the support assistant for Northwind Outfitters. "
+    "Answer the customer's question using the conversation so far. "
+    "Keep replies short and plain-text."
+)
+
+
+class StreamInterrupted(Exception):
+    """The stream ended without the model's normal `stop` completion chunk."""

Review this PR

Python practice