PythonHub
Link
click to show
click to show
Bad evals, my own: five exercises from two LLM judges
The author uses five exercises from two real LLM judges to expose evaluation pitfalls, including inconsistent results, biased test sets, misleading metrics, and pass/fail thresholds that become unreliable as test suites grow. He shows why trustworthy evaluations require representative data, clearly defined metrics, repeated testing, and preserved run artifacts, revealing flaws in his own...
https://digline.dev/blog/bad-evals-my-own/
1 · 107 ·