Nimbus assistant: chat endpoint with weather and calendar tools
Python · LLM apps · advanced · greenfield
Adds the Nimbus personal-assistant chat endpoint with two internal tools. ConversationHistory (app/history.py) owns the transcript and documents the chat-completions message-order contract; the endpoint (app/agent.py) runs a bounded tool loop: model-produced arguments are validated before a tool runs, a lone tool call is executed inline, and a batch of tool calls in one turn runs concurrently with asyncio.gather. Invalid arguments, unknown tools and service outages come back to the model as tool errors; provider outages surface as a clean 502. Exercised single-lookup questions ("weather in Berlin?", "what's on my calendar tomorrow?") against the staging services and the assistant composed the answers from the returned data.
Nimbus is an internal personal assistant for employees. `weather-svc` and `calendar-svc` are trusted internal HTTP microservices; the model's tool arguments are NOT trusted and are validated before a tool runs. The model routinely batches independent lookups into one turn (parallel function calling) — e.g. comparing the weather in two cities arrives as two `get_weather` calls in a single assistant message — and the system prompt explicitly encourages that.
Requirements
- `POST /assistant/chat` accepts `{"message": string}` (1–2000 characters, bounds enforced by FastAPI validation) and returns `{"answer": string}`. The assistant has exactly two tools, declared in `TOOLS`: `get_weather` (required `city`, non-empty string) and `list_events` (required `date`, string in `YYYY-MM-DD` format). Both tools call trusted internal HTTP microservices (`weather-svc`, `calendar-svc`) with a 5-second timeout.
- Message order follows the chat-completions contract declared in `app/history.py`: the assistant message carrying `tool_calls` is appended BEFORE any of its `role: "tool"` results, and each result carries its own matching `tool_call_id`. A transcript in which a `role: "tool"` message does not follow the assistant message that requested it is rejected by the API with HTTP 400 on the next call.
- `tool_call.function.arguments` is an untrusted JSON **string**. Arguments are validated before the tool runs: `get_weather` requires a non-empty string `city`; `list_events` requires a `date` string matching `YYYY-MM-DD` that is a real calendar date. Malformed JSON, a non-object, missing or invalid values, unknown tool names, and tool-service failures (timeout, non-2xx, non-JSON body) all come back to the model as the content of the matching `role: "tool"` message in the form `{"error": "…"}`; none of them raises an exception or fails the HTTP request.
- The model may return several tool calls in one turn (parallel function calling) and the system prompt encourages it for independent lookups; all calls of a turn run concurrently and every result is appended to the history with its own `tool_call_id`. A model response without tool calls ends the loop and its content is returned (`""` when the content is `None`).
- At most `MAX_TOOL_ROUNDS = 4` model calls per request; if the model is still requesting tools after the fourth call the endpoint responds HTTP 502. Provider exceptions (`RateLimitError`, `APITimeoutError`, `APIStatusError`) also respond 502; the SDK's two built-in retries are the only retry mechanism.
- The model is the pinned dated snapshot `gpt-4o-mini-2024-07-18`.
Files touched
- app/history.py
- app/agent.py
--- app/history.py
+"""Conversation history for the assistant agent.
+
+Ordering contract (chat-completions API): within the transcript, every
+``role: "tool"`` message must be preceded by the assistant message that
+carries the matching ``tool_call_id`` in its ``tool_calls``. An assistant
+turn that requests tools is therefore recorded FIRST — via
+:meth:`ConversationHistory.add_assistant` — and only then its results, one
+``add_tool_result`` call per tool call. A transcript where a tool message
+comes before (or without) its assistant ``tool_calls`` message is rejected
+by the API with HTTP 400 on the next request.
+
+This class records messages in the caller's append order; it deliberately
+does not reorder or validate them, so the contract above is the caller's
+responsibility.
+"""
+
+from __future__ import annotations
+
+from openai.types.chat import ChatCompletionMessage
+
+
+class ConversationHistory:
+ """The message list sent with every model call, in append order."""
+
+ def __init__(self, system_prompt: str) -> None:
+ self._messages: list[dict] = [
+ {"role": "system", "content": system_prompt}
+ ]
+