Add parallel tool execution to the support assistant
Python · LLM apps · advanced · greenfield
Parallelises tool execution in the support assistant: when the model asks for several tools at once — say, a customer lookup and their orders — they run concurrently instead of one after the other, each on its own DB session. Model-produced arguments are validated against the declared schemas before a tool runs, failed tools return their error as a structured tool message so the model can adapt, and provider outages still surface as a clean 502. Exercised the multi-tool path against a seeded database, including a forced warehouse timeout, and the assistant wove the results into a single answer.
Internal support console. Agents look up customers, their orders, and stock levels; when a model response asks for two or three tools at once, parallel execution cuts the wall-clock time for the agent. The tools call separate internal services — a customer DB, an orders service, and a warehouse system — any of which can occasionally time out. `require_support_agent` (in `app/auth.py`, outside this PR) is the existing FastAPI dependency that validates the console's SSO session; `SessionLocal` (in `app/db.py`, outside this PR) is the existing SQLAlchemy `async_sessionmaker`.
Requirements
- `POST /support/chat` is mounted behind the console's existing `require_support_agent` dependency. It accepts `{"message": string (1–2000 chars)}` (bounds enforced by FastAPI) and returns `{"answer": string}`. The assistant has three tools: `lookup_customer` (by exact email), `lookup_orders` (by customer id, optionally filtered by status), and `check_inventory` (by sku). Their full JSON schemas are declared as module-level constants in `app/tools.py`.
- `tool_call.function.arguments` is an untrusted JSON string produced by the model. It is parsed and validated against the tool's declared schema (required keys, no extra keys, types, `minLength` / `minimum` / `enum`) before the tool runs; malformed JSON or off-schema arguments go back to the model as `{"error": "..."}` in that call's `role: "tool"` message and never reach the database.
- When the model returns multiple `tool_calls` in a single response, all of them are executed concurrently via `asyncio.gather(return_exceptions=True)`, each on its own database session. The results are then fed back to the model as individual `role: "tool"` messages — exactly one per tool call, each carrying that call's own provider-assigned `tool_call_id` (`tc.id`).
- A tool whose execution raises (network timeout, invalid upstream response) comes back to the model as a generic tool error — `{"error": "<tool> is unavailable"}` under its own `tool_call_id` — so the model can decide how to recover; the exception detail goes to the server log only. The other tools' results in the same batch are delivered normally. The exception never propagates to the HTTP response.
- Message order follows the chat-completions contract: the assistant message carrying `tool_calls` is appended before any of its `role: "tool"` results. A model response without tool calls ends the loop and its content is returned (the empty string when `content` is `None`).
- The model is the pinned dated snapshot `gpt-4o-mini-2024-07-18`. Provider exceptions (`RateLimitError`, `APITimeoutError`, `APIStatusError`) surface as HTTP 502 with no app-level retry beyond the SDK's two built-in retries.
- At most `MAX_TOOL_ROUNDS = 5` model calls per request; exceeding that responds HTTP 502.
Files touched
- app/tools.py
- app/agent.py
--- app/tools.py
+"""Support assistant tools: schemas and async implementations."""
+
+from __future__ import annotations
+
+from typing import Any
+
+from sqlalchemy import select
+from sqlalchemy.ext.asyncio import AsyncSession
+
+from app.models import Customer, InventoryItem, Order
+
+ORDER_STATUSES = ("placed", "shipped", "delivered", "cancelled")
+
+# ── Tool schemas ──────────────────────────────────────────────────────────
+
+LOOKUP_CUSTOMER_TOOL: dict[str, Any] = {
+ "type": "function",
+ "function": {
+ "name": "lookup_customer",
+ "description": "Look up a customer by exact email address.",
+ "parameters": {
+ "type": "object",
+ "properties": {
+ "email": {
+ "type": "string",
+ "minLength": 1,
+ "description": "The customer's exact email address.",
+ },
+ },