Building production AI features without an evaluation framework is like deploying backend code without CI/CD. A prompt tweak that improves customer summary formatting might inadvertently cause the model to hallucinate account numbers or miss JSON syntax requirements.
Because Large Language Models are non-deterministic, testing requires a hybrid approach: combining strict deterministic assertions with LLM-as-a-judge scoring functions.
Anatomy of a Hybrid Eval Harness
An effective LLM evaluation harness evaluates candidate model responses against a curated dataset of golden inputs and edge cases.
| Test Level | Method | Target Metric |
|---|---|---|
| Level 1: Structure | JSON Schema validation / Regex parsing | 100% syntactic compliance |
| Level 2: Groundedness | Fact verification against context payload | Zero hallucinated facts |
| Level 3: Quality | LLM Judge with strict rubric criteria | Semantic accuracy & tone adherence |
Building an Assertion Harness in TypeScript
Here is how we structure an automated evaluation test runner using Node/TypeScript:
src/evals/runner.tsimport { z } from 'zod'; const OutputSchema = z.object({ summary: z.string().min(10), sentiment: z.enum(['positive', 'neutral', 'negative']), actionItems: z.array(z.string()) }); export async function evaluateResponse(rawOutput: string) { // 1. Deterministic Schema Check const json = JSON.parse(rawOutput); const parsed = OutputSchema.safeParse(json); if (!parsed.success) { return { passed: false, reason: 'Schema violation' }; } // 2. Deterministic Boundary Check if (parsed.data.actionItems.length === 0) { return { passed: false, reason: 'Missing required action items' }; } return { passed: true }; }
Preventing Model Lock-In and Regressions
Running automated evals allows teams to safely swap model backends (e.g. switching between Claude 3.5 Sonnet, GPT-4o, and local open-weights models) while verifying that accuracy benchmarks stay consistent.
For more details on the hidden costs of AI code generation and verification, read our piece on the AI debugging tax.