mirror of
https://github.com/wassname/Open-Assistant.git
synced 2026-07-26 13:07:22 +08:00
76 B
76 B
Sections to train Reward Model (RM)
Currently we format