AI & MLDeep
Intermediate

Evaluating LLMs

11 min read

Learn
Deep Reading
Estimated 11 mins
Prereq
Intermediate
Basic ML concepts helpful
Interactive
Static Playbook
Static guide & reference tables

Why evaluation is the hard part

Training a capable LLM is only half the problem — knowing whether it's actually good, and better than the alternative, is the other half, and arguably harder. Unlike a classifier with a clear accuracy metric, LLM output is open-ended text: there's rarely one correct answer, quality is often subjective, and the same model can excel at one task and fail badly at a superficially similar one.

LLM evaluation has converged on a layered approach: automated benchmarks for broad capability signals, task-specific metrics for narrow correctness, LLM-as-judge for scalable quality assessment, and human evaluation as the final ground truth — each catching failure modes the others miss.

Standardized benchmarks: MMLU and friends

MMLU (Massive Multitask Language Understanding) is a 57-subject multiple-choice benchmark spanning topics from elementary math to professional law and medicine. It's become a standard headline number in model release announcements because it's broad, automatically gradeable, and correlates reasonably well with general capability — though models can be specifically optimized to perform well on it without proportionally improving real-world usefulness, a phenomenon known as benchmark contamination or overfitting.

HumanEval tests code generation specifically: the model is given a function signature and docstring, and its generated implementation is graded by actually running it against hidden unit tests — a rare case where LLM output can be verified objectively as correct or incorrect, rather than judged.

Other common benchmarks include HellaSwag (commonsense reasoning), GSM8K (grade-school math word problems), TruthfulQA (resistance to generating plausible-sounding falsehoods), and MT-Bench (multi-turn conversation quality, itself graded with LLM-as-judge).

Note

A model can top the MMLU leaderboard and still be a poor product experience — verbose, sycophantic, unreliable at following formatting instructions, or prone to hallucinating citations. Benchmarks are necessary but not sufficient. Treat a high benchmark score as evidence of general capability, not proof the model is right for your specific use case.

Perplexity: the classic language modeling metric

Perplexity measures how "surprised" a model is by a sequence of text — mathematically, it's the exponentiated average negative log-likelihood the model assigns to each token in a held-out test set: $\text{PPL} = \exp\left(-\frac{1}{N}\sum_{i=1}^{N} \log P(x_i \mid x_{<i})\right)$.

Lower perplexity means the model assigned higher probability to the actual next tokens — it's a good measure of raw language modeling fit during pre-training. Its major limitation: perplexity says nothing about whether the model's outputs are *useful*, *safe*, or *correct*. A model can achieve excellent perplexity while still hallucinating facts or failing to follow instructions, because perplexity only checks how well-calibrated the model's next-token distribution is against natural text — not whether the content is true.

LLM-as-judge: using a model to grade a model

As LLM outputs became too open-ended for exact-match scoring, LLM-as-judge emerged as the dominant scalable evaluation method: prompt a strong model (often GPT-4-class or Claude) with the original question, the response(s) to evaluate, and a rubric, and have it produce a score or a pairwise preference ("Response A is better than Response B because...").

This scales far better than human evaluation — you can grade thousands of outputs for the cost of API calls rather than paying annotators — and correlates surprisingly well with human judgment on many tasks. But it inherits the judge model's own biases: LLM judges are known to favor longer responses, favor responses stylistically similar to their own training distribution, and show position bias (preferring whichever answer appears first in the prompt) unless the evaluation is carefully randomized and controlled.

python

LLM evaluation methods at a glance

MethodWhat it measuresCostMain weakness
Benchmarks (MMLU, HumanEval)Broad capability on fixed tasksLow — automatedCan be gamed; poor proxy for real usage
PerplexityLanguage modeling fitLow — automatedIgnores correctness and usefulness
LLM-as-judgeResponse quality vs. a rubricMedium — API callsInherits judge's own biases
Human evaluationReal-world preference and correctnessHigh — human timeSlow, expensive, hard to scale

Human evaluation and hallucination metrics

Despite the rise of automated methods, human evaluation remains the gold standard, especially for subjective qualities like tone, helpfulness, and safety. Teams typically run structured human evaluations with clear rubrics and multiple annotators per sample to measure inter-annotator agreement — if humans disagree with each other a lot, the task itself may be ambiguous, not just the model's answer.

Hallucination measurement is its own subfield: common approaches include fact-checking generated claims against a trusted knowledge base or the retrieved source documents (particularly important in RAG systems), using a second LLM call to verify each factual claim is supported by the provided context (sometimes called faithfulness or groundedness scoring), and tracking rates of unsupported claims over a fixed evaluation set across model versions to catch regressions.

Note

Serious LLM teams don't evaluate once before shipping — they build an ongoing eval pipeline: a fixed regression test set that runs on every model or prompt change, live monitoring of user feedback signals (thumbs up/down, regeneration rate), and periodic spot-checks by humans. A model that scored well at launch can silently degrade in effective quality as usage patterns shift or an upstream provider updates the model without notice.

What's next

Evaluation matters most once you're building real systems — see Retrieval-Augmented Generation (RAG) for how groundedness and hallucination reduction work in practice, or Fine-Tuning LLMs to see how eval results feed back into training decisions.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.