Add a prompt-suite eval harness for the support pipeline

Python · LLM apps · beginner · greenfield

Adds the prompt-suite eval harness for the support pipeline: runs each prompt once at temperature 0, scores by exact match and writes a timestamped run file that the dashboard diffs for regression tracking. Every result records the model the API reported, so runs carry their own provenance. Provider outages abort the run cleanly, and the run file is written atomically so a failed run never leaves a partial one.

The eval dashboard diffs successive run files in `runs/` after every prompt-suite change; a score movement is read as the effect of the change under review. The provider decides which snapshot serves a floating alias and can repoint it whenever a newer snapshot ships; it reports the actual serving model on every response.

Requirements

Files touched

--- app/eval_harness.py
+"""Eval harness: run the support-pipeline prompt suite and record scores.
+
+Each successful run writes runs/<run_id>.json; the eval dashboard diffs
+successive run files to attribute score movement to the prompt edit under
+review.
+"""
+
+from __future__ import annotations
+
+import asyncio
+import json
+import os
+import sys
+import time
+from datetime import datetime, timezone
+from pathlib import Path
+
+from openai import APIStatusError, APITimeoutError, AsyncOpenAI, RateLimitError
+
+client = AsyncOpenAI()  # reads OPENAI_API_KEY from the environment
+
+MODEL = "gpt-4o"
+MAX_TOKENS = 64
+RUNS_DIR = Path("runs")
+
+SYSTEM_PROMPT = (
+    "You are the triage stage of the Northwind Supply support pipeline. "
+    "Reply with only the requested value - no sentences, no punctuation."
+)

Review this PR

Python practice