mirror of
https://github.com/wassname/ml_debug.git
synced 2026-08-25 11:21:20 +08:00
NoLiMa effective length was wrong in both files: most models fall below the 85% threshold at 1-4K tokens, not 8-16K (only GPT-4o 8K, GPT-4.1 16K). NoLiMa 32K count was 11 models, not 10. The Loo quote's first sentence was stitched from two places and is dropped. JudgeLRM's 8.14% is an Introduction number, not an abstract one, and the abstract quote had lost /14B and its trailing clause. Shi et al is 'The findings', not 'Our findings'. Body quotes now replace abstract quotes where the number matters; tags [ID] -> [FT].