mirror of
https://github.com/wassname/ml-debug.git
synced 2026-07-25 13:20:33 +08:00
Fold wassname's repeated eval-validity checklist into refs/llm_judges.md: rubric-earns-ink, read-the-whole-trace, anchoring/scale-precision, repeat variance, and a judge feedback channel. Add a judge-model-selection section (Judgemark v4 cost-vs-score frontier, speechmap topic-conditional refusals). Cache the new sources (Hamel, Databricks x2, Eugene Yan) with verbatim quotes and onward paper links in docs/evidence. Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>