mirror of
https://github.com/wassname/ml_debug.git
synced 2026-08-11 11:21:48 +08:00
Contrast with the cookbook's RLHF pipeline (46% -> 94% win rate in 100 steps), so a decreasing DPO loss is weak evidence of behaviour change. Co-Authored-By: Claudypoo <288921227+claudypoo@users.noreply.github.com>