Add the LLM ticket-triage classifier endpoint

Python · LLM apps · advanced · greenfield

Adds the triage classifier the mail-ingest worker calls for every incoming support email: one POST endpoint, one single-turn model call. The system prompt is a constant and the (untrusted) email travels only as a user message; the composed message is token-counted with tiktoken and truncated to the budget before the call; the model's JSON reply is validated against the declared category/priority/summary contract before it is returned, and anything off-contract or any provider outage comes back as a clean 502. Exercised against seeded mailboxes: billing/bug/feature emails classify into the right queues, oversized bodies get truncated instead of blowing the context window, and a stubbed off-contract reply (markdown-fenced JSON, out-of-enum priority) yields 502 rather than a queue assignment.

Internal helpdesk service: the endpoint is called only by the mail-ingest worker over the service mesh (mTLS, not publicly routed) and classifies ~5k emails/day into queue assignments. There is no per-end-user auth because there are no end users — the mesh identity of the ingest worker is the authorization. No state is kept between calls.

Requirements

Files touched

--- app/ticket_triage.py
+"""Ticket triage: classify incoming support emails with a pinned model.
+
+Single-turn classification — no conversation history, no tools. The customer
+email is untrusted text and travels to the model only inside a `user`
+message; the system prompt is a module constant.
+"""
+
+from __future__ import annotations
+
+import json
+
+import tiktoken
+from fastapi import APIRouter, HTTPException
+from openai import APIStatusError, APITimeoutError, AsyncOpenAI, RateLimitError
+from pydantic import BaseModel, Field
+
+router = APIRouter()
+
+_client = AsyncOpenAI()  # reads OPENAI_API_KEY from the environment
+
+MODEL = "gpt-4o-mini-2024-07-18"  # pinned dated snapshot, per ADR-011
+# o200k_base is the encoding the gpt-4o model family uses. get_encoding()
+# fetches the BPE file on first use; the deploy image pre-warms the tiktoken
+# cache so no request pays (or can fail on) that download.
+ENCODING = tiktoken.get_encoding("o200k_base")
+
+MAX_REQUEST_TOKENS = 4096  # hard cap: input messages + reply budget combined
+MAX_REPLY_TOKENS = 200     # the classification JSON is tiny; this is generous
+# Chat-completions wire-format overhead for 2 messages is ~9 tokens on this

Review this PR

Python practice