Add an AI refund agent to the support chat
Python · LLM apps · advanced · greenfield
Adds an AI refund agent for the support chat: one FastAPI endpoint runs a single turn against the pinned gpt-4o-mini snapshot with a single `issue_refund` tool backed by the payments service. Tool arguments are validated before execution, payments-side failures go back to the model as tool errors instead of 500s, and timed-out completion calls are retried with backoff so a transient provider blip does not fail the customer's request. Turns are capped at 5 model calls, and a refund can only ever be issued once per request. Exercised the happy path and the tool-error path by hand against a sandbox order.
A customer-support service. The payments team exposes refunds as an internal HTTP endpoint with no idempotency key — every accepted call moves money — and the OpenAI client runs with `max_retries=0`, so the agent loop in `refund_agent.py` is the only place a timed-out provider call is retried.
Requirements
- `POST /support/refund-chat` accepts JSON `{user_id, message}` (both non-empty strings; `message` at most 4000 characters) and returns `{reply}` — the assistant's final text; the empty string is the valid reply when the model produced no text on its final call.
- The agent uses the openai 1.x SDK against the pinned snapshot `gpt-4o-mini-2024-07-18` and offers exactly one tool, `issue_refund`, whose declared arguments are `order_id` (non-empty string) and `amount_cents` (positive integer). The client is constructed with `max_retries=0` and completion calls use a 20-second timeout; the agent's own retry loop is the service's only retry mechanism.
- A refund is executed at most once per request, whatever the model does: the agent refuses a second `issue_refund` execution in the same turn — including two tool calls in one assistant message, or a call in a later round — and returns `{"error": …}` as that call's tool result instead. An execution attempt counts whether it produced a receipt or a payments-side failure (`RefundError`): a transport error while receiving the receipt does not prove the money did not move, so the turn records the attempt (`refund_attempted`) and refuses further executions; only calls rejected by argument validation, which never reach the payments service, do not count. The payments service's refund endpoint has no idempotency key — every accepted call creates a new refund and moves money — so a retry after a completion call timed out mid-turn must never re-run `issue_refund` for a refund that was already attempted.
- On `openai.APITimeoutError` the timed-out completion call is retried once, provided completion budget remains (at most 2 execution attempts of the turn in total), with a backoff of `attempt number × 2 seconds`; a retry resumes the turn where it left off: all messages already exchanged in the failed attempt — including the assistant's tool calls and executed tool results — are preserved across attempts. If both attempts time out, the endpoint responds 504. A timeout on the fifth completion call is budget exhaustion, not a retry: no backoff, HTTP 502.
- Tool arguments are validated against the declared shapes before the tool runs (required `order_id` and `amount_cents`, no other properties); malformed JSON arguments, unknown tool names, failed validation, and payments-side failures (rejections and transport errors, which surface as `RefundError`) are returned to the model as the `role: "tool"` error content for that call — never raised to the caller.
- One request makes at most 5 completion calls in total, counted across retries and including calls that timed out; if the model is still requesting tools when the cap is reached, or the fifth call times out, the endpoint responds 502 rather than looping forever.
- The payments service always answers with JSON: a refund receipt (`refund_id`, `status`, `amount_cents`) on success and `{"detail": ...}` on errors; the `httpx` call to it uses a 10-second timeout.
Files touched
- app/services/refund_agent.py
- app/services/payments.py
- app/routers/support_chat.py
--- app/services/refund_agent.py +"""One-turn refund agent for the support chat. + +Runs a single agent turn against a pinned OpenAI snapshot: the model may call +the `issue_refund` tool, and the turn ends when the model replies with plain +text. A timed-out completion call is retried once (the SDK's own retries are +disabled) so a transient provider hiccup does not fail the request. +""" + +import json +import os +import time + +import openai +from openai import OpenAI + +from app.services import payments + +MODEL = "gpt-4o-mini-2024-07-18" +MAX_ATTEMPTS = 2 +RETRY_BACKOFF_SECONDS = 2.0 +COMPLETION_TIMEOUT_SECONDS = 20.0 +MAX_COMPLETION_CALLS = 5 # per request, timed-out calls included + +client = OpenAI(api_key=os.environ["OPENAI_API_KEY"], max_retries=0) + +SYSTEM_PROMPT = ( + "You are the support agent for Northwind Outfitters. " + "The customer's refund request arrives as the user message. " + "To move money, call issue_refund with the order id and the amount in "