mirror of
https://github.com/wassname/Open-Assistant.git
synced 2026-08-11 11:13:12 +08:00
76 B
76 B
Sections to train Reward Model (RM)
Currently we format