mirror of
https://github.com/wassname/ml-debug.git
synced 2026-07-25 13:20:33 +08:00
Add judge-validity checklist and model-selection to llm_judges ref
Fold wassname's repeated eval-validity checklist into refs/llm_judges.md: rubric-earns-ink, read-the-whole-trace, anchoring/scale-precision, repeat variance, and a judge feedback channel. Add a judge-model-selection section (Judgemark v4 cost-vs-score frontier, speechmap topic-conditional refusals). Cache the new sources (Hamel, Databricks x2, Eugene Yan) with verbatim quotes and onward paper links in docs/evidence. Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>
This commit is contained in:
@@ -39,3 +39,51 @@ Self-preference is causally linked to self-recognition:
|
||||
> Constrained decoding methods for structured outputs (e.g., JSON-mode) impose an inductive bias on the model's output distribution.
|
||||
|
||||
Related: JudgeBench leaderboard for judge quality — https://huggingface.co/spaces/ScalerLab/JudgeBench
|
||||
|
||||
---
|
||||
|
||||
Source: Hamel Husain, Databricks (x2), Eugene Yan blog posts (via WebFetch)
|
||||
Title: practitioner rules on scale precision, rubric quality, and reading judge traces
|
||||
Fetched-via: WebFetch summarizer model, 2026-07-22; short quotes cross-checked by re-fetching
|
||||
Fetch-status: quotes below reproduced twice identically EXCEPT the Databricks scale range, where two fetches disagreed (0-3 or 0-4 vs 0-3 or 1-5) so it is paraphrased, not quoted, in the ref
|
||||
|
||||
## Hamel Husain, "Creating a LLM-as-a-Judge That Drives Business Results" — https://hamel.dev/blog/posts/llm-judge/
|
||||
|
||||
Critique-shadowing workflow; look at the data before writing the judge, prefer binary:
|
||||
|
||||
> You cannot write a good judge prompt until you've seen the data.
|
||||
|
||||
> If your evaluations consist of a bunch of metrics that LLMs score on a 1-5 scale (or any other scale), you're doing it wrong.
|
||||
|
||||
Onward links from this post: Shankar et al., "Who Validates the Validators?" (arXiv:2404.12272); Yan, "ALIGN Eval" (https://aligneval.com/); OpenAI Cookbook, "Custom LLM as a Judge to Detect Hallucinations with Braintrust" (https://cookbook.openai.com/examples/custom-llm-as-a-judge); "What We've Learned From A Year of Building with LLMs" (https://applied-llms.org/).
|
||||
|
||||
## Databricks, "Best Practices for LLM Evaluation of RAG Applications" — https://www.databricks.com/blog/LLM-auto-eval-best-practices-RAG
|
||||
|
||||
Use a low-precision integer scale:
|
||||
|
||||
> Scales like 0-10 are difficult to come up with distinguishing criteria between all scores.
|
||||
|
||||
The recommended range paraphrased (fetches disagreed on 0-4 vs 1-5): a coarse 0-3 or 1-5 Likert scale, with an example for each score in the range. Onward links: LMSYS MT-Bench (arXiv:2306.05685) and its FastChat judge prompts (https://github.com/lm-sys/FastChat/blob/main/fastchat/llm_judge/data/judge_prompts.jsonl).
|
||||
|
||||
## Databricks, "Enhancing LLM-as-a-Judge with Grading Notes" — https://www.databricks.com/blog/enhancing-llm-as-a-judge-with-grading-notes
|
||||
|
||||
Per-question domain rubrics fix the domain-knowledge gap:
|
||||
|
||||
> Without Grading Notes, both judge-LLMs overestimate the effectiveness by a significant margin, likely indicating the gap in domain knowledge to criticize.
|
||||
|
||||
> With Grading Notes introducing brief domain knowledge, both judge LLMs showed significant improvement in the alignment rate with humans, especially in the case of GPT-4: alignment rate increased to 96.3% for Llama3 and 93.1% for GPT-4o, which corresponds to 85% and 67.5% reduction in misalignment rate, respectively.
|
||||
|
||||
## Eugene Yan, "Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge)" — https://eugeneyan.com/writing/llm-evaluators/
|
||||
|
||||
A survey; the few-shot-instability line is attributed to Luo et al., "ChatGPT as a Factual Inconsistency Evaluator" (arXiv:2303.15621):
|
||||
|
||||
> performance unstable when changing the label, example order, and number of examples
|
||||
|
||||
Key papers the survey collects (worth reading past this file):
|
||||
- Zheng et al., MT-Bench — position/verbosity/self-enhancement bias — arXiv:2306.05685
|
||||
- Liu et al., G-Eval — CoT scoring with GPT-4 — arXiv:2303.16634
|
||||
- Doddapaneni et al., "Finding Blind Spots in Evaluator LLMs with Interpretable Checklists" — GPT-4 misses >50% of quality drops on several axes — arXiv:2406.13439
|
||||
- Huang et al., "On the Limitations of Fine-Tuned Judge Models" — finetuned judges fail out-of-domain — arXiv:2403.02839
|
||||
- Shankar et al., "Who Validates the Validators?" — aligning LLM eval with human preferences — arXiv:2404.12272
|
||||
|
||||
Judge leaderboards (live, not cached): Judgemark v4 (https://eqbench.com/judgemark-v4.html), SpeechMap refusal rates (https://speechmap.ai/).
|
||||
|
||||
Reference in New Issue
Block a user