Web appOpen in Telegram

PostBad evals, my own: five exercises from two LLM judges

30 September 2026
P
PythonHub
Link
click to show
Bad evals, my own: five exercises from two LLM judges The author uses five exercises from two real LLM judges to expose evaluation pitfalls, including inconsistent results, biased test sets, misleading metrics, and pass/fail thresholds that become unreliable as test suites grow. He shows why trustworthy evaluations require representative data, clearly defined metrics, repeated testing, and preserved run artifacts, revealing flaws in his own... https://digline.dev/blog/bad-evals-my-own/
1 · 107 ·

Nearby in the feed

PPythonHubHardware-Agnostic Models in vLLM The article explains how vLLM is introducing hardware-agnostic layers so it can keep supporting diverse models and acceleratorsPPythonHubPyPy v8.0.0 PyPy 8.0.0 introduces its first Python 3.12 interpreter as a beta, alongside Python 2.7 and 3.11 releases, and raises the minimum glibc requirement
this message
PPythonHubZeroModels ZeroModels is an open-source Keras 3 library of pretrained models spanning vision, language, speech, depth estimation, and multimodal tasks. https://PPythonHubKev tiny Jev-like family of decision models built on top of Qwen3.5 you can train and run on your own. https://github.com/jaredpalmer/kev
PPythonHubPythonHub@PythonHub · channel · Tech
2 649subscribers137average post reach
Venue feed Open in Telegram

An open public feed from the search index ChatCrawler — “Google for public Telegram”; refreshed as the venue is crawled. Times are UTC.

Public content only, official Telegram API. About · FAQ · What we do not do · Remove a page · Catalog · Search · How we count