Journal / AI

Evaluating LLMs in production: unit testing non-deterministic workflows.

In traditional software, modifying code either passes or fails your test suite. In AI applications, changing a system prompt or upgrading an LLM model version introduces silent regressions across edge cases.

RL
RBB LAB
Studio
Published 29 Sep 2026 8 min read
ai:eval RBB/LAB AI RBB LAB · JOURNAL 8 MIN READ

Building production AI features without an evaluation framework is like deploying backend code without CI/CD. A prompt tweak that improves customer summary formatting might inadvertently cause the model to hallucinate account numbers or miss JSON syntax requirements.

Because Large Language Models are non-deterministic, testing requires a hybrid approach: combining strict deterministic assertions with LLM-as-a-judge scoring functions.

100%
JSON Schema Enforcement
Pass/Fail
Deterministic Assertion Gate
>0.90
LLM Judge Similarity Score

Anatomy of a Hybrid Eval Harness

An effective LLM evaluation harness evaluates candidate model responses against a curated dataset of golden inputs and edge cases.

Test Level Method Target Metric
Level 1: Structure JSON Schema validation / Regex parsing 100% syntactic compliance
Level 2: Groundedness Fact verification against context payload Zero hallucinated facts
Level 3: Quality LLM Judge with strict rubric criteria Semantic accuracy & tone adherence
Never rely solely on subjective manual spot-checking. Continuous evaluation suites run in CI pipelines prior to merging prompt modifications.

Building an Assertion Harness in TypeScript

Here is how we structure an automated evaluation test runner using Node/TypeScript:

src/evals/runner.tsimport { z } from 'zod';

const OutputSchema = z.object({
  summary: z.string().min(10),
  sentiment: z.enum(['positive', 'neutral', 'negative']),
  actionItems: z.array(z.string())
});

export async function evaluateResponse(rawOutput: string) {
  // 1. Deterministic Schema Check
  const json = JSON.parse(rawOutput);
  const parsed = OutputSchema.safeParse(json);
  
  if (!parsed.success) {
    return { passed: false, reason: 'Schema violation' };
  }

  // 2. Deterministic Boundary Check
  if (parsed.data.actionItems.length === 0) {
    return { passed: false, reason: 'Missing required action items' };
  }

  return { passed: true };
}

Preventing Model Lock-In and Regressions

Running automated evals allows teams to safely swap model backends (e.g. switching between Claude 3.5 Sonnet, GPT-4o, and local open-weights models) while verifying that accuracy benchmarks stay consistent.

For more details on the hidden costs of AI code generation and verification, read our piece on the AI debugging tax.