mirror of
https://github.com/wassname/ml-debug.git
synced 2026-07-26 13:27:35 +08:00
Add a per-model, per-difficulty token-budget note to judge-selection: models vary ~6x in tokens/task (Artificial Analysis), reasoning models emit ~8x more and scale with difficulty (Epoch AI, DeepSeek-R1), and the depth-vs-breadth trade (cap reasoning low + N passes on easy tasks, which doubles as the repeat-variance check; don't truncate near model capability). Add Gu et al. "A Survey on LLM-as-a-Judge" as a second survey anchor. Cache new quotes. Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>