Add a prompt-suite eval harness for the support pipeline
Python · LLM apps · beginner · greenfield
Adds the prompt-suite eval harness for the support pipeline: runs each prompt once at temperature 0, scores by exact match and writes a timestamped run file that the dashboard diffs for regression tracking. Every result records the model the API reported, so runs carry their own provenance. Provider outages abort the run cleanly, and the run file is written atomically so a failed run never leaves a partial one.
The eval dashboard diffs successive run files in `runs/` after every prompt-suite change; a score movement is read as the effect of the change under review. The provider decides which snapshot serves a floating alias and can repoint it whenever a newer snapshot ships; it reports the actual serving model on every response.
Requirements
- `python -m app.eval_harness` runs the fixed suite `PROMPTS` (five prompts, a non-empty compile-time constant) sequentially - one chat completion per prompt - and on success writes exactly one run file `runs/<run_id>.json`, where `run_id` is a UTC timestamp with microseconds. A failed or interrupted run leaves no run file behind — not even a partial one.
- The model is the pinned dated snapshot `gpt-4o-2024-08-06` - never a floating alias such as `gpt-4o` (ADR-011). The dashboard compares scores across run files, so cross-run comparability requires that every completion of every run was requested from that same snapshot.
- Every result record carries `model_served`, the model the API reports on the response (`completion.model`), as provenance for the dashboard; the run record carries the requested model under `model`.
- A prompt scores 1 when the completion's text content, `strip()`-ed and `casefold()`-ed, equals the expected answer `strip()`-ed and `casefold()`-ed, and 0 otherwise. `None` content scores 0 and must never raise.
- Every completion is requested with `temperature=0` and `max_tokens=64`.
- Provider exceptions (`RateLimitError`, `APITimeoutError`, `APIStatusError`) abort the run with exit code 1 and a stderr message naming the failing prompt; the SDK's two built-in retries are the only retry mechanism.
Files touched
- app/eval_harness.py
--- app/eval_harness.py
+"""Eval harness: run the support-pipeline prompt suite and record scores.
+
+Each successful run writes runs/<run_id>.json; the eval dashboard diffs
+successive run files to attribute score movement to the prompt edit under
+review.
+"""
+
+from __future__ import annotations
+
+import asyncio
+import json
+import os
+import sys
+import time
+from datetime import datetime, timezone
+from pathlib import Path
+
+from openai import APIStatusError, APITimeoutError, AsyncOpenAI, RateLimitError
+
+client = AsyncOpenAI() # reads OPENAI_API_KEY from the environment
+
+MODEL = "gpt-4o"
+MAX_TOKENS = 64
+RUNS_DIR = Path("runs")
+
+SYSTEM_PROMPT = (
+ "You are the triage stage of the Northwind Supply support pipeline. "
+ "Reply with only the requested value - no sentences, no punctuation."
+)