Update README.md

This commit is contained in:
wassname
2023-04-22 11:21:51 +08:00
committed by GitHub
parent 24b79e585a
commit 2df58a68c4
+21 -4
View File
@@ -1,5 +1,5 @@
This is a list of resources for reinforcement learning from human feedback and other methods to instruct large language models.
This is a list of resources for reinforcement learning from human feedback (RLHF) and other methods to instruct large language models.
## Evaluation
@@ -9,9 +9,6 @@ There are multiple ways to formally evaluate LLM capabilities. Right now project
- [openai/evals: Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.](https://github.com/openai/evals)
- [stanford-crfm/helm: Holistic Evaluation of Language Models (HELM), a framework to increase the transparency of language models (https://arxiv.org/abs/2211.09110).](https://github.com/stanford-crfm/helm)
## Training
## Data
@@ -63,6 +60,26 @@ A great way to find new instruction datasets is to
- Look at compilations like - [OIG](https://laion.ai/blog/oig-dataset/)
- [github instruction-turning tag](https://github.com/topics/instruction-tuning)
## Training
### Libraries
- https://github.com/lucidrains/PaLM-rlhf-pytorch - Implementation of RLHF on top of the PaLM
- https://github.com/CarperAI/trlx - A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)
- https://huggingface.co/blog/stackllama - StackLLaMA: A hands-on guide to train LLaMA with RLHF
### Papers/Methods
- [RLHF](https://arxiv.org/pdf/2009.01325.pdf)
- Chain of Hindsight https://arxiv.org/abs/2302.02676 the model it trained to rank it's own output, so it's kind of like diffusion, letting the model operate iterativly.
- SFT - Supervised Fine Tuning this is normal fine tuning
- HIR: [Hindsight Instruction Relabeling](https://twitter.com/tianjun_zhang/status/1628180891368570881) 💩 offline RL reinvented with extra steps
FARL: SL
Algorithm Distillation: classical control problems. Offline RL
Similar lists
- https://github.com/yaodongC/awesome-instruction-dataset