Add per-conversation memory to the team chat assistant

Python · LLM apps · intermediate · greenfield

Adds `/assistant/chat` with per-conversation memory so the assistant can follow a thread: every exchange is appended to the thread's history and the model sees the full conversation each call, replies reserved at 1024 tokens. Failed provider calls leave the stored history untouched so a retry doesn't double-record the turn. Tried multi-turn threads back to back and the assistant kept context across turns.

Threads are long-lived: an on-call channel pastes build logs and incident timelines into the same conversation for days, and a busy thread routinely runs into the hundreds of turns.

Requirements

Files touched

--- app/team_assistant.py
+"""Team chat assistant: one persistent conversation per channel thread."""
+
+from __future__ import annotations
+
+from fastapi import APIRouter, HTTPException
+from openai import APIStatusError, APITimeoutError, AsyncOpenAI, RateLimitError
+from pydantic import BaseModel, Field
+
+router = APIRouter()
+
+_client = AsyncOpenAI()  # reads OPENAI_API_KEY from the environment
+
+MODEL = "gpt-4-0613"  # pinned snapshot; 8192-token context window
+MAX_REPLY_TOKENS = 1024
+
+SYSTEM_PROMPT = (
+    "You are the assistant for an internal engineering team. Answer "
+    "concisely, using the conversation so far for context. When the "
+    "answer depends on a system you cannot see, say what you would need "
+    "to check instead of guessing."
+)
+
+# conversation_id -> turns, oldest first: {"role": "user"|"assistant",
+# "content": str}. In-process by design; a restart starts threads fresh.
+_HISTORY: dict[str, list[dict[str, str]]] = {}
+
+
+class ChatRequest(BaseModel):
+    conversation_id: str = Field(min_length=1, max_length=64)

Review this PR

Python practice