Add per-conversation memory to the team chat assistant
Python · LLM apps · intermediate · greenfield
Adds `/assistant/chat` with per-conversation memory so the assistant can follow a thread: every exchange is appended to the thread's history and the model sees the full conversation each call, replies reserved at 1024 tokens. Failed provider calls leave the stored history untouched so a retry doesn't double-record the turn. Tried multi-turn threads back to back and the assistant kept context across turns.
Threads are long-lived: an on-call channel pastes build logs and incident timelines into the same conversation for days, and a busy thread routinely runs into the hundreds of turns.
Requirements
- `POST /assistant/chat` accepts `{"conversation_id": string (1–64 chars), "message": string (1–4000 chars)}` (bounds enforced by FastAPI validation) and returns `{"answer": string}`. Each conversation has its own in-process history; after a successful reply both the user turn and the assistant turn are appended to it, and a failed request leaves the stored history unchanged. The chat widget sends at most one in-flight request per conversation, so turns of one conversation are processed sequentially and no per-conversation locking is required.
- The model is the pinned dated snapshot `gpt-4-0613`, whose context window is 8192 tokens. The reply reserve is `max_tokens=1024`, so the messages sent to the model must fit a message budget of 7168 tokens.
- The request must never exceed the context window. Fitness is measured conservatively with `tiktoken`'s `cl100k_base` encoding as: 3 (completion priming) + Σ over all messages of (4 per-message overhead + tokens of the message content), and that total must be ≤ 7168. Content is counted as ordinary text: special-token spellings such as `<|endoftext|>` inside a message are plain characters, never an error. (The 4 overhead tokens are the provider's 3 framing tokens per message plus 1 for the one-token role name of these role/content-only messages, so this measure never undercounts for the declared shape.)
- When the new exchange would push the request past the 7168-token budget, the oldest history entries are dropped until it fits. The system message and the new user message are never dropped. If the request still does not fit with the history fully dropped, the endpoint rejects it with HTTP 422 before calling the provider.
- Provider failures (`RateLimitError`, `APITimeoutError`, `APIStatusError`) surface as HTTP 502; there is no app-level retry loop on top of the SDK's two built-in retries. A `None` message content is returned as the empty string.
Files touched
- app/team_assistant.py
--- app/team_assistant.py
+"""Team chat assistant: one persistent conversation per channel thread."""
+
+from __future__ import annotations
+
+from fastapi import APIRouter, HTTPException
+from openai import APIStatusError, APITimeoutError, AsyncOpenAI, RateLimitError
+from pydantic import BaseModel, Field
+
+router = APIRouter()
+
+_client = AsyncOpenAI() # reads OPENAI_API_KEY from the environment
+
+MODEL = "gpt-4-0613" # pinned snapshot; 8192-token context window
+MAX_REPLY_TOKENS = 1024
+
+SYSTEM_PROMPT = (
+ "You are the assistant for an internal engineering team. Answer "
+ "concisely, using the conversation so far for context. When the "
+ "answer depends on a system you cannot see, say what you would need "
+ "to check instead of guessing."
+)
+
+# conversation_id -> turns, oldest first: {"role": "user"|"assistant",
+# "content": str}. In-process by design; a restart starts threads fresh.
+_HISTORY: dict[str, list[dict[str, str]]] = {}
+
+
+class ChatRequest(BaseModel):
+ conversation_id: str = Field(min_length=1, max_length=64)