Add the LLM ticket-triage classifier endpoint
Python · LLM apps · advanced · greenfield
Adds the triage classifier the mail-ingest worker calls for every incoming support email: one POST endpoint, one single-turn model call. The system prompt is a constant and the (untrusted) email travels only as a user message; the composed message is token-counted with tiktoken and truncated to the budget before the call; the model's JSON reply is validated against the declared category/priority/summary contract before it is returned, and anything off-contract or any provider outage comes back as a clean 502. Exercised against seeded mailboxes: billing/bug/feature emails classify into the right queues, oversized bodies get truncated instead of blowing the context window, and a stubbed off-contract reply (markdown-fenced JSON, out-of-enum priority) yields 502 rather than a queue assignment.
Internal helpdesk service: the endpoint is called only by the mail-ingest worker over the service mesh (mTLS, not publicly routed) and classifies ~5k emails/day into queue assignments. There is no per-end-user auth because there are no end users — the mesh identity of the ingest worker is the authorization. No state is kept between calls.
Requirements
- `POST /triage/classify` accepts `{"subject": string, "body": string}` and returns `{"category": string, "priority": string, "summary": string}`. Input bounds are enforced at the API layer by Pydantic validation: `subject` is 1–300 characters and `body` is 1–40,000 characters; anything out of bounds is a 422 from the framework and never reaches the model.
- The system prompt is a module-level constant. The customer's subject and body are untrusted text (an email can contain "ignore previous instructions" content) and only ever enter the model as a single `user` message — no part of the email is interpolated into the system prompt or any other instruction-shaped channel. Classification is single-turn: no conversation history is stored or resent.
- The model's reply must be a single JSON object — not wrapped in markdown fences — with `category` (one of `billing`, `bug`, `feature_request`, `account`, `other`), `priority` (one of `low`, `medium`, `high`, `urgent`) and `summary` (a string of 1–200 characters). The endpoint parses and validates the reply against exactly that contract before returning it; extra keys in the model's reply are ignored. Anything off-contract — `None` content, malformed JSON, a non-object, an out-of-enum category or priority, a missing/empty/over-long summary — responds HTTP 502 and is never guessed at or coerced.
- The request must never exceed 4,096 tokens total, counted with tiktoken's `o200k_base` encoding (the encoding the pinned model family uses): the reply budget is a fixed `max_tokens=200`, and the composed user message is truncated to the remaining input allowance — after subtracting the reply budget, the constant system prompt's own token count and a flat wire-format overhead reserve — before the call is made. Truncation must not raise on multi-byte content.
- Provider exceptions (`RateLimitError`, `APITimeoutError`, `APIStatusError`) respond HTTP 502. The openai SDK's two built-in retries with backoff are the only retry mechanism; the app adds no retry loop of its own (each classification is a billed call).
- The model is the pinned dated snapshot `gpt-4o-mini-2024-07-18` and the call uses `temperature=0`, so classifications are as stable as the provider allows.
- `max_tokens` is the fixed reply budget above — it is never derived from the input's token count.
Files touched
- app/ticket_triage.py
--- app/ticket_triage.py
+"""Ticket triage: classify incoming support emails with a pinned model.
+
+Single-turn classification — no conversation history, no tools. The customer
+email is untrusted text and travels to the model only inside a `user`
+message; the system prompt is a module constant.
+"""
+
+from __future__ import annotations
+
+import json
+
+import tiktoken
+from fastapi import APIRouter, HTTPException
+from openai import APIStatusError, APITimeoutError, AsyncOpenAI, RateLimitError
+from pydantic import BaseModel, Field
+
+router = APIRouter()
+
+_client = AsyncOpenAI() # reads OPENAI_API_KEY from the environment
+
+MODEL = "gpt-4o-mini-2024-07-18" # pinned dated snapshot, per ADR-011
+# o200k_base is the encoding the gpt-4o model family uses. get_encoding()
+# fetches the BPE file on first use; the deploy image pre-warms the tiktoken
+# cache so no request pays (or can fail on) that download.
+ENCODING = tiktoken.get_encoding("o200k_base")
+
+MAX_REQUEST_TOKENS = 4096 # hard cap: input messages + reply budget combined
+MAX_REPLY_TOKENS = 200 # the classification JSON is tiny; this is generous
+# Chat-completions wire-format overhead for 2 messages is ~9 tokens on this