Placeholder draft — replace with the real essay.
An eval suite is a claim about what matters, dressed up as a measurement. Most of the ones I’ve inherited were written in a hurry, against whatever examples were on hand that week, and then never revisited — which means the claim they’re making has quietly gone stale while everyone keeps trusting the number.
The cranky part
If nobody can tell you, off the top of their head, what a 3-point drop on your main eval actually means for a user, the eval isn’t measuring the thing you think it’s measuring. It’s measuring itself.
What I do instead now
Fewer aggregate scores, more named failure modes with real transcripts attached. A score is a summary; a transcript is evidence. I want people arguing about the transcript, not the summary.