Case Study 01 · AI Evaluation

LLM Evaluation & Error Analysis

A practice set focused on choosing the stronger model response, identifying the controlling error, and writing concise rationales that another evaluator can audit.

What this sample demonstrates

The task was not to generate a better-sounding answer. It was to compare two candidate responses and determine which one was more reliable. The evaluation process emphasized factual correctness, logical support, uncertainty calibration, practical usefulness, and whether the explanation actually followed from the evidence.

The practice also exposed a useful distinction: recognizing that an answer is wrong can be immediate; explaining why it is wrong in a concise, reviewer-friendly form is a separate skill. The exercises therefore focused on making implicit judgment explicit.

Evaluator standard

Name the decisive issue. Explain why it changes the ranking. Avoid padding the rationale with secondary observations unless they materially affect the judgment.

Evaluation method

1. IDENTIFY CLAIMS
2. CHECK SUPPORT
3. FIND DECISIVE ERROR
4. COMPARE RESPONSES
5. WRITE RATIONALE

Accuracy

Is the central claim true, or merely plausible-sounding?

Inference

Does the available evidence actually support the conclusion?

Calibration

Does the answer preserve uncertainty when multiple explanations remain possible?

Usefulness

Is the advice merely adequate, or does it account for relevant practical risk?

Relevance

Does the response address the controlling issue without unnecessary detours?

Auditability

Can another evaluator understand exactly why the ranking was made?

Representative evaluations

The two examples below preserve enough of the original practice context to present them without reconstructing missing prompts. Other exercises are summarized by error class rather than presented as invented A/B examples.

Example A · Unsupported diagnosis

Observation does not establish cause

Evaluation
“Response B is better because a fan making noise doesn't prove that it is broken. Response A jumps from an observation to a diagnosis without ruling out other plausible causes.”
Controlling error

Unsupported inference. Response A treats one possible explanation as established fact.

Why B wins

Response B preserves the distinction between symptom and diagnosis, which is the more defensible reasoning pattern.

logical validityuncertainty calibrationpremature conclusion
Example B · Risk and redundancy

Adequate in ordinary use is not the same as resilient

Evaluation
“For most people, keeping one external SSD may be enough and SSDs are reliable, but it is still a single point of failure. Best practice is to keep multiple backups.”
Controlling issue

Failure to distinguish typical adequacy from robust best practice. Reliability of one device does not eliminate the consequences of that device being the only copy.

Evaluation nuance

The stronger judgment does not require claiming the weaker response is absurd. It can be mostly reasonable while still omitting a material risk.

risk analysissingle point of failurepractical judgment

Error taxonomy used across the practice set

False factual claim

One response states something that is simply false while the competing response is factually correct. The ranking should center on correctness, not presentation quality.

False conclusion

A response may begin from a true observation but derive a conclusion that does not follow. The evaluator should distinguish premise accuracy from inference quality.

Premature certainty

The response selects one explanation before alternatives have been ruled out. Appropriate uncertainty is a quality feature, not evasiveness.

Incomplete practical reasoning

Advice can be technically defensible yet materially weaker because it ignores redundancy, failure modes, or real-world operating conditions.

Both answers need improvement

Pairwise evaluation does not require pretending the preferred answer is excellent. A reviewer can rank B over A while still noting that both could be clearer or more precise.

Over-explained rationale

A correct judgment can become less useful if the decisive reason is buried. The target is enough explanation to make the decision reproducible.

Reusable evaluator script

The practice converged on a compact structure that prevents having to reconstruct the entire reasoning chain every time:

[Preferred response] is better because [decisive reason]. [Other response] [states / assumes / concludes] [specific problem]. The evidence only supports [more accurate or appropriately limited conclusion].

This structure is intentionally short. It forces the evaluator to identify the controlling issue, distinguish evidence from conclusion, and produce a rationale that can be checked by another reviewer.

Human–AI division of labor

Human responsibility

  • choose the stronger response
  • detect the factual or logical defect
  • decide which issue is decisive
  • reject weak explanations
  • refine the rationale until it matches the actual judgment

AI contribution

  • generate practice comparisons
  • challenge or critique rationales
  • model clearer evaluator phrasing
  • provide repeated exercises
  • help convert implicit reasoning into explicit review language

Role relevance

This case study is directly relevant to work involving LLM response evaluation, AI QA, human rating, rubric-based review, model-output auditing, and reasoning-quality assessment. The central skill is not preference expression; it is making a defensible comparative judgment and stating the reason precisely enough that the decision can be reviewed.

LLM evaluationAI QAhuman ratingerror analysisrubric reasoningmodel quality