Accuracy is useful when a task has one verifiable answer, but many language model features produce summaries, plans, explanations, or transformations with several acceptable outputs. Evaluating these systems requires a portfolio of evidence. Deterministic checks establish hard constraints, rubric graders assess bounded qualities, pairwise comparisons expose relative preferences, and human review resolves ambiguity. The goal is not to manufacture one universal score. It is to make release decisions repeatable, explainable, and sensitive to important failures.
Evaluation system and data flow
An evaluation pipeline should preserve the full chain from test case to release decision. A versioned dataset supplies inputs, expected facts, constraints, and metadata. A pinned application configuration produces candidate outputs. Several evaluators create structured judgments. An aggregator computes metrics by slice, compares them with a baseline, and applies regression gates. Samples then move to human review according to risk and uncertainty.
- Select a dataset version and freeze the prompt, model configuration, tools, retrieval snapshot, and decoding settings.
- Run candidates with isolated credentials and capture outputs, tool traces, latency, token counts, and errors.
- Apply deterministic validators before subjective grading.
- Run rubric based graders with explicit criteria and calibrated examples.
- Compare the candidate with a baseline using randomized pairwise presentation.
- Aggregate results by task, risk, language, and other meaningful slices.
- Apply release gates and route uncertain or consequential samples to reviewers.
Every result should reference immutable identifiers for the dataset, application build, prompt, evaluator, and model. Without this lineage, a score cannot be reproduced or compared safely. Store raw evaluator outputs as well as normalized scores so changes in parsing or aggregation can be audited.
Build datasets that represent decisions
A useful dataset contains more than a list of prompts. Each case should state why it exists, what behavior matters, what facts or constraints are known, and which slice labels apply. Include normal requests, boundary conditions, malformed inputs, adversarial instructions, refusals, long context, and cases where the correct behavior is to ask for clarification. Production derived examples can be valuable only after consent, redaction, deduplication, and retention controls are addressed.
Separate a stable regression set from an exploratory set. The stable set supports comparisons over time. The exploratory set can evolve as new failure modes appear. Avoid repeatedly tuning against a public or frequently inspected test set because it gradually becomes training data for the team. Maintain a private holdout for periodic checks, and document sampling so apparent progress is not caused by an easier mix of cases.
Coverage matters more than raw dataset size. A thousand similar summaries can be less informative than a smaller collection spanning factual conflict, missing context, formatting constraints, policy boundaries, and domain terminology. Weighting should reflect risk and product intent, but always publish unweighted slice results alongside any weighted aggregate.
Use deterministic checks first
Deterministic checks are fast, explainable, and stable. They should validate properties that do not require aesthetic judgment: schema conformance, required fields, allowed values, citation identifiers, exact calculations, tool argument types, prohibited secrets, output length, and whether requested records were preserved. If an output fails a hard contract, a model grader should not rescue it with a favorable subjective score.
Some semantic tasks also have deterministic or reference based checks. Code can be executed in a sandbox against tests. Structured transformations can be compared after normalization. Answers can be checked against a set of required claims without demanding exact wording. These checks need their own tests, because a faulty evaluator can create false confidence across every model run.
Rubric graders with bounded responsibilities
A rubric grader should assess one or a few clearly defined dimensions, such as factual support, completeness, instruction adherence, or clarity. Define score levels with observable criteria. Supply the input, relevant reference material, candidate output, and only the context needed for the judgment. Require a structured response containing the score, a short rationale, and evidence excerpts.
Model graders are variable measurement instruments. Pin their configuration where possible, use low randomness, repeat a subset of judgments, and compare agreement with trained human reviewers. Watch for position bias, verbosity bias, preference for the grader model's own style, and sensitivity to irrelevant formatting. Never place untrusted candidate text into a grader prompt as if it were an instruction. Delimit it as data and tell the grader to ignore embedded commands.
Pairwise comparisons
Pairwise evaluation asks whether a candidate, a baseline, or neither is better for a defined criterion. Reviewers often find this easier than assigning an absolute score. Randomize which output appears first, hide system identity, and include a tie option. Run both presentation orders for automated graders when position bias is material. Aggregate wins with uncertainty intervals rather than treating a narrow lead as certain improvement.
Pairwise results are relative. A candidate can beat a poor baseline while remaining unacceptable. Combine comparisons with absolute safety and quality thresholds. Also retain a fixed anchor output or periodically compare against older versions, because a chain of local wins can drift away from the original requirements.
Concrete evaluation runner
def evaluate_case(case, candidate, baseline, evaluators):
candidate_output = run_system(candidate, case.input)
baseline_output = run_system(baseline, case.input)
checks = run_deterministic_checks(
output=candidate_output,
schema=case.required_schema,
constraints=case.constraints
)
grades = []
if checks.hard_failures == 0:
for rubric in evaluators.rubrics:
grades.append(rubric.grade(
input=case.input,
reference=case.reference,
output=candidate_output
))
comparison = evaluators.pairwise.compare_randomized(
input=case.input,
reference=case.reference,
output_a=candidate_output,
output_b=baseline_output
)
return EvaluationRecord(
case_id=case.id,
slice_labels=case.slices,
checks=checks,
grades=grades,
comparison=comparison,
needs_review=is_uncertain(grades, comparison)
)
The runner should use bounded concurrency, per request timeouts, retry rules that distinguish transient transport errors from valid refusals, and cost limits. Preserve failed attempts rather than quietly retrying until a favorable output appears. For nondeterministic systems, evaluate repeated samples or report the configured sampling policy.
Regression gates and statistical care
A release gate should encode both nonnegotiable constraints and tolerated movement. Examples include zero critical policy violations, no schema regression, minimum pass rates for high risk slices, and a bounded decline in selected quality metrics. Compare paired outcomes on the same cases. Include confidence intervals or a suitable paired test, especially when differences are small. Averages alone can hide rare severe failures.
Define gates before examining the candidate result. Otherwise thresholds tend to shift to justify a desired release. Treat flaky infrastructure separately from model quality, but do not discard unexplained failures. Report missing results and evaluator errors as first class outcomes. If a model or dependency changes outside team control, rerun the baseline under the same conditions when feasible.
Human review
Human judgment remains important for nuanced correctness, tone, harmful ambiguity, and grader calibration. Give reviewers a concise rubric, examples near decision boundaries, and an escalation path. Blind system identity when possible. Measure agreement, discuss systematic disagreements, and revise unclear criteria rather than forcing consensus through averaging.
Review all critical safety failures and a stratified sample of ordinary cases. Also sample outputs where graders disagree, confidence is low, or a new slice lacks history. Protect reviewers from unnecessary sensitive data, apply access controls, and avoid storing personal annotations longer than needed.
Security, failure modes, and compromises
Evaluation infrastructure processes untrusted prompts and outputs. Sandbox executable code, restrict network and file access, use test credentials, redact secrets, and prevent tools from mutating real systems. Treat datasets as governed assets because they may contain personal, proprietary, or security sensitive material. Grader prompts and hidden references should not be exposed to the system under test.
Common mistakes include optimizing one aggregate score, using the same model as generator and sole judge, changing datasets between comparisons, ignoring evaluator failures, and accepting pairwise wins without an absolute quality floor. Another failure is adding every discovered example to the regression set, then repeatedly tuning until that fixed set no longer predicts unseen behavior.
Broader evaluation increases cost and release time. Automated graders improve scale but introduce bias and variance. Human review adds context but is slower and can be inconsistent. Deterministic gates are trustworthy for narrow contracts but cannot represent every quality dimension. A layered design uses each method only for judgments it can support.
Conclusion
Evaluation beyond accuracy is a disciplined decision system. Versioned datasets define coverage, deterministic checks enforce contracts, rubric graders measure bounded qualities, pairwise comparisons reveal relative change, and human review calibrates ambiguous judgments. Regression gates turn this evidence into release policy. With lineage, security controls, slice reporting, and explicit uncertainty, teams can improve language model systems without mistaking a convenient score for proof of quality.