mirror of
https://github.com/wassname/Open-Assistant.git
synced 2026-08-07 11:17:38 +08:00
@@ -0,0 +1,3 @@
|
||||
# Data
|
||||
|
||||
Resources related to data.
|
||||
@@ -0,0 +1,23 @@
|
||||
# Data Augmentation
|
||||
|
||||
(pull request welcome)
|
||||
|
||||
## What is data augmentation
|
||||
|
||||
Data augmentation is a technique we can use to get better data faster. Using
|
||||
machine learning models to analyze long data (like an essay) and compress it
|
||||
into instructions.
|
||||
|
||||
## How to contribute
|
||||
|
||||
To contribute to data augmentation you can write a short Python script that uses
|
||||
a model from HuggingFace to analyze the text.
|
||||
[Here](https://docs.google.com/document/d/13a188pPvqnlvuVa3e_suVz4YO5s-JWeiOOrpp0odImg/edit)
|
||||
are examples of what you can do.
|
||||
|
||||
And here are example implementations:
|
||||
[Idea 3](https://colab.research.google.com/drive/1GllCN5PgSYxBxINZsv3A2r0SpdznHlbT?usp=sharing),
|
||||
[Idea 4](https://colab.research.google.com/drive/1nZx5LRjO61fYprFyqtrwPDLOis6ctR4p#scrollTo=1EE8CriiaCXj)
|
||||
|
||||
To contribute simply choose one of many ideas from the document above and
|
||||
implement it.
|
||||
@@ -0,0 +1,426 @@
|
||||
# Datasets
|
||||
|
||||
The datasets for this project are currently hosted as loading scripts on the
|
||||
[Open-Assistant organization](https://huggingface.co/OpenAssistant) the Hugging
|
||||
Face Hub. Each of them can be loaded by first installing the 🤗 Datasets
|
||||
library:
|
||||
|
||||
```bash
|
||||
python -m pip install datasets
|
||||
```
|
||||
|
||||
and then running:
|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
|
||||
dataset = load_dataset("OpenAssistant/{dataset-name}")
|
||||
```
|
||||
|
||||
We use this GitHub repository to accept new submissions and standardize quality
|
||||
control. See the instructions below if you'd like to contribute a new dataset to
|
||||
the project.
|
||||
|
||||
## Adding a new dataset
|
||||
|
||||
### 0. Pre-Requisites
|
||||
|
||||
Install Git and create a GitHub account prior to implementing a dataset; you can
|
||||
follow instructions to install Git
|
||||
[here](https://git-scm.com/book/en/v2/Getting-Started-Installing-Git).
|
||||
|
||||
You will also need at least Python 3.8+. If you are installing Python, we
|
||||
recommend downloading
|
||||
[Anaconda](https://docs.anaconda.com/anaconda/install/index.html) to curate a
|
||||
python environment with necessary packages. **We strongly recommend Python 3.8+
|
||||
for stability**.
|
||||
|
||||
### 1. **Fork the OpenAssistant repository**
|
||||
|
||||
Fork the
|
||||
`OpenAssistant`[repository](https://github.com/LAION-AI/Open-Assistant). To do
|
||||
this, click the link to the repository and click "Fork" in the upper-right
|
||||
corner. You should get an option to fork to your account, provided you are
|
||||
signed into Github.
|
||||
|
||||
After you fork, clone the repository locally. You can do so as follows:
|
||||
|
||||
```bash
|
||||
git clone git@github.com:<your_github_username>/OpenAssistant.git
|
||||
cd OpenAssistant # enter the directory
|
||||
```
|
||||
|
||||
Next, you want to set your `upstream` location to enable you to push/pull (add
|
||||
or receive updates). You can do so as follows:
|
||||
|
||||
```bash
|
||||
git remote add upstream git@github.com:LAION-AI/Open-Assistant.git
|
||||
```
|
||||
|
||||
You can optionally check that this was set properly by running the following
|
||||
command:
|
||||
|
||||
```bash
|
||||
git remote -v
|
||||
```
|
||||
|
||||
The output of this command should look as follows:
|
||||
|
||||
```bash
|
||||
origin git@github.com:<your_github_username>/Open-Assistant.git (fetch)
|
||||
origin git@github.com:<your_github_username>/Open-Assistant.git (push)
|
||||
upstream git@github.com:LAION-AI/Open-Assistant.git (fetch)
|
||||
upstream git@github.com:LAION-AI/Open-Assistant.git (push)
|
||||
```
|
||||
|
||||
If you do NOT have an `origin` for whatever reason, then run:
|
||||
|
||||
```bash
|
||||
git remote add origin git@github.com:<your_github_username>/OpenAssistant.git
|
||||
```
|
||||
|
||||
The goal of `upstream` is to keep your repository up-to-date to any changes that
|
||||
are made officially to the OpenAssistant repo. You can do this as follows by
|
||||
running the following commands:
|
||||
|
||||
```
|
||||
git fetch upstream
|
||||
git pull
|
||||
```
|
||||
|
||||
Provided you have no _merge conflicts_, this will ensure the repo stays
|
||||
up-to-date as you make changes. However, before you make changes, you should
|
||||
make a custom branch to implement your changes.
|
||||
|
||||
You can make a new branch as such:
|
||||
|
||||
```
|
||||
git checkout -b <dataset_name>
|
||||
```
|
||||
|
||||
:::caution
|
||||
|
||||
Please do not make changes on the master branch!
|
||||
|
||||
:::
|
||||
|
||||
Always make sure you're on the right branch with the following command:
|
||||
|
||||
```
|
||||
git branch
|
||||
```
|
||||
|
||||
The correct branch will have a asterisk \* in front of it.
|
||||
|
||||
### 2. **Create a development environment**
|
||||
|
||||
You can make an environment in any way you choose to. We highlight two possible
|
||||
options:
|
||||
|
||||
#### 2a) Create a conda environment
|
||||
|
||||
The following instructions will create an Anaconda `openassistant` environment.
|
||||
|
||||
- Install [anaconda](https://docs.anaconda.com/anaconda/install/) for your
|
||||
appropriate operating system.
|
||||
- Run the following command while in the `biomedical` folder (you can pick your
|
||||
python version):
|
||||
|
||||
```bash
|
||||
conda create -n openassistant python=3.8 # Creates a conda env
|
||||
conda activate openassistant # Activate your conda environment
|
||||
cd openassistant
|
||||
pip install -r dev-requirements.txt # Install this while in the openassistant folder
|
||||
```
|
||||
|
||||
You can deactivate your environment at any time by either exiting your terminal
|
||||
or using `conda deactivate`.
|
||||
|
||||
#### 2b) Create a venv environment
|
||||
|
||||
Python 3.3+ has venv automatically installed; official information is found
|
||||
[here](https://packaging.python.org/en/latest/guides/installing-using-pip-and-virtual-environments/).
|
||||
|
||||
```
|
||||
python3 -m venv <your_env_name_here>
|
||||
source <your_env_name_here>/bin/activate # activate environment
|
||||
cd openassistant
|
||||
pip install -r dev-requirements.txt # Install this while in the openassistant folder
|
||||
```
|
||||
|
||||
Make sure your `pip` package points to your environment's source.
|
||||
|
||||
### 3. Prepare a folder in `datasets` for your dataloader
|
||||
|
||||
Make a new directory within the `openassistant/datasets` directory:
|
||||
|
||||
```bash
|
||||
mkdir openassistant/datasets/<dataset_name>
|
||||
```
|
||||
|
||||
**NOTE**: Please use snake_case, i.e. lowercase letters and underscores when
|
||||
choosing a `<dataset_name>`.
|
||||
|
||||
Add an `__init__.py` file to this directory:
|
||||
|
||||
```bash
|
||||
touch openassistant/datasets/<dataset_name>/__init__.py
|
||||
```
|
||||
|
||||
Next, copy the `template.py` script and `hub.py` module of `templates` into your
|
||||
dataset folder. The `template.py` script has "TODOs" to fill in for your
|
||||
dataloading script:
|
||||
|
||||
```bash
|
||||
cp templates/hub.py openassistant/datasets/<dataset_name>/
|
||||
cp templates/template.py openassistant/datasets/<dataset_name>/<dataset_name>.py
|
||||
```
|
||||
|
||||
#### (Optional) Prepare local dataset files
|
||||
|
||||
If your dataset files aren't publicly available via URLs (e.g. because you
|
||||
implemented a web scraper), you'll need to implement some extra logic to store
|
||||
and prepare the data locally prior to implementing a loading script in 🤗
|
||||
Datasets.
|
||||
|
||||
To do so, first copy the template script for dataset creation:
|
||||
|
||||
```bash
|
||||
cp templates/prepare.py openassistant/datasets/<dataset_name>/
|
||||
```
|
||||
|
||||
Next, implement any logic that is needed to prepare a local version of the
|
||||
dataset files (by convention we store them in `datasets/<dataset_name>/data/`).
|
||||
Add any extra dependencies to a `requirements.txt` file and provide instructions
|
||||
on how to prepare the dataset files in a README:
|
||||
|
||||
```bash
|
||||
touch openassistant/datasets/<dataset_name>/requirements.txt
|
||||
cp templates/README.py openassistant/datasets/<dataset_name>/
|
||||
```
|
||||
|
||||
**Note:** Do not commit any dataset files to the OpenAssistant repo - all data
|
||||
will be hosted on the Hugging Face Hub. This step is needed for the project's
|
||||
data admins to be able to replicate the dataset creation process before pushing
|
||||
to the Hub.
|
||||
|
||||
### 4. Implement your dataset
|
||||
|
||||
To implement your dataloader, you will need to follow `template.py` and fill in
|
||||
all necessary TODOs. There are three key methods that are important:
|
||||
|
||||
- `_info`: Specifies the schema of the expected dataloader
|
||||
- `_split_generators`: Downloads and extracts data for each split (e.g.
|
||||
train/val/test) or associate local data with each split.
|
||||
- `_generate_examples`: Create examples from data that conform to each schema
|
||||
defined in `_info`.
|
||||
|
||||
For the `_info_` function, you will need to define `features` for your
|
||||
`DatasetInfo` object. For each dataset config, choose the right schema from our
|
||||
list of examples. You can find the schemas in the
|
||||
[schemas directory](https://github.com/LAION-AI/Open-Assistant/tree/main/openassistant).
|
||||
|
||||
You will use this schema in the `_generate_examples` return value.
|
||||
|
||||
Populate the information in the dataset according to this schema; some fields
|
||||
may be empty.
|
||||
|
||||
#### Example scripts
|
||||
|
||||
TODO
|
||||
|
||||
#### Running & debugging
|
||||
|
||||
You can run your data loader script during development by appending the
|
||||
following statement to your code
|
||||
([templates/template.py](https://github.com/LAION-AI/Open-Assistant/blob/main/openassistant/templates/template.py)
|
||||
already includes this):
|
||||
|
||||
```python
|
||||
if __name__ == "__main__":
|
||||
datasets.load_dataset(__file__)
|
||||
```
|
||||
|
||||
If you want to use an interactive debugger during development, you will have to
|
||||
use `breakpoint()` instead of setting breakpoints directly in your IDE. Most
|
||||
IDEs will recognize the `breakpoint()` statement and pause there during
|
||||
debugging. If your preferred IDE doesn't support this, you can always run the
|
||||
script in your terminal and debug with `pdb`.
|
||||
|
||||
### 5. Check if your dataloader works
|
||||
|
||||
Make sure your dataset is implemented correctly by checking in python the
|
||||
following commands:
|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
|
||||
data = load_dataset("openassistant/datasets/<dataset_name>/<dataset_name>.py", name="<dataset_name>_<schema>")
|
||||
```
|
||||
|
||||
Run these commands from the top level of the `OpenAssistant` repo.
|
||||
|
||||
### 6. Create a dataset card
|
||||
|
||||
Copy and fill out the template dataset card:
|
||||
|
||||
```bash
|
||||
cp templates/dataset_card.md openassistant/datasets/<dataset_name>/README.md
|
||||
```
|
||||
|
||||
### 7. Format your code
|
||||
|
||||
From the main directory, run the code quality checks via the following command:
|
||||
|
||||
```
|
||||
pre-commit run --all-files
|
||||
```
|
||||
|
||||
This runs the black formatter, isort, and lints to ensure that the code is
|
||||
readable and looks nice. Flake8 linting errors may require manual changes.
|
||||
|
||||
### 8. Commit your changes
|
||||
|
||||
First, commit your changes to the branch to "add" the work:
|
||||
|
||||
```
|
||||
git add openassistant/datasets/<dataset_name>/*.py
|
||||
git commit -m "A message describing your commits"
|
||||
```
|
||||
|
||||
Then, run the following commands to incorporate any new changes in the master
|
||||
branch of datasets as follows:
|
||||
|
||||
```
|
||||
git fetch upstream
|
||||
git rebase upstream/main
|
||||
```
|
||||
|
||||
**Run these commands in your custom branch**.
|
||||
|
||||
Push these changes to **your fork** with the following command:
|
||||
|
||||
```
|
||||
git push -u origin <dataset_name>
|
||||
```
|
||||
|
||||
### 9. **Make a pull request**
|
||||
|
||||
Make a Pull Request to implement your changes on the main repository
|
||||
[here](https://github.com/LAION-AI/Open-Assistant/pulls). To do so, click "New
|
||||
Pull Request". Then, choose your branch from your fork to push into "base:main".
|
||||
|
||||
When opening a PR, please link the
|
||||
[issue](https://github.com/LAION-AI/Open-Assistant/issues) corresponding to your
|
||||
dataset using
|
||||
[closing keywords](https://docs.github.com/en/issues/tracking-your-work-with-issues/linking-a-pull-request-to-an-issue)
|
||||
in the PR's description, e.g. `resolves #17`.
|
||||
|
||||
## [Admins] Uploading a dataset to the Hugging Face Hub
|
||||
|
||||
Uploading a new dataset from `openassistant/datasets/<dataset_name>` to the
|
||||
Hugging Face Hub typically involves the following steps:
|
||||
|
||||
1. Setup
|
||||
2. Create a new dataset repository
|
||||
3. Copy a dataset loading script and dataset card
|
||||
4. Upload to the Hub
|
||||
|
||||
### 1. Setup
|
||||
|
||||
To upload a dataset to the OpenAssistant organization, you first need to:
|
||||
|
||||
- Create a [Hugging Face account](https://huggingface.co/join) (it's free)
|
||||
- Join the [OpenAssistant organization](https://huggingface.co/OpenAssistant) by
|
||||
clicking on the _Request to join this org_ button on the top right-hand side
|
||||
|
||||
Next, check that you're correctly logged in and that `git-lfs` is installed so
|
||||
that the dataset can be uploaded. To log in, create a **write access token**
|
||||
that can be found under your Hugging Face profile (icon in the top right corner
|
||||
on [hf.co](http://hf.co/), then Settings -> Access Tokens -> User Access Tokens
|
||||
-> New Token. Alternatively, you can go to
|
||||
[your token settings](https://huggingface.co/settings/tokens) directly.
|
||||
|
||||
Once you've created a token, run:
|
||||
|
||||
```bash
|
||||
huggingface-cli login
|
||||
```
|
||||
|
||||
in a terminal, or case you're working in a notebook
|
||||
|
||||
```python
|
||||
from huggingface_hub import notebook_login
|
||||
|
||||
notebook_login()
|
||||
```
|
||||
|
||||
You can then copy-paste your token to log in locally.
|
||||
|
||||
Next, let's make sure that `git-lfs` is correctly installed. To do so, simply
|
||||
run:
|
||||
|
||||
```bash
|
||||
git-lfs -v
|
||||
```
|
||||
|
||||
The output should show something like
|
||||
`git-lfs/2.13.2 (GitHub; linux amd64; go 1.15.4)`. If your console states that
|
||||
the `git-lfs` command was not found, please make sure to install it
|
||||
[here](https://git-lfs.github.com/) or simply via:
|
||||
|
||||
```bash
|
||||
sudo apt-get install git-lfs
|
||||
git config --global user.email "you@example.com"
|
||||
git config --global user.name "Your Name"
|
||||
```
|
||||
|
||||
The final step of the setup is to install the 🤗 Datasets library by running:
|
||||
|
||||
```bash
|
||||
python -m pip install datasets
|
||||
```
|
||||
|
||||
### 2. Create a new dataset repository
|
||||
|
||||
Follow [this guide](https://huggingface.co/docs/datasets/upload_dataset) for
|
||||
instructions on creating a new dataset repo on the Hub. Use the same snake_case
|
||||
name as the dataset in `openassistant/datasets/<dataset_name>`.
|
||||
|
||||
Once you've created the dataset repo, clone it by running:
|
||||
|
||||
```bash
|
||||
git clone https://huggingface.co/datasets/OpenAssistant/<dataset_name>
|
||||
cd <dataset_name>
|
||||
```
|
||||
|
||||
### 3. Copy a dataset loading script and dataset card
|
||||
|
||||
Next, copy the loading script and dataset card to your repo:
|
||||
|
||||
```bash
|
||||
cp openassistant/datasets/<dataset_name>/<dataset_name>.py .
|
||||
cp openassistant/datasets/<dataset_name>/README.md .
|
||||
```
|
||||
|
||||
#### (Optional) Prepare local dataset files
|
||||
|
||||
If the dataset files of `openassistant/datasets/<dataset_name>` aren't public,
|
||||
you'll need to run the `openassistant/datasets/<dataset_name>/prepare.py` script
|
||||
to create them. Store them in the same directory that is specified by the
|
||||
loading script (`data` by default).
|
||||
|
||||
### 4. Upload to the Hub
|
||||
|
||||
Once the dataset script and card are ready, use Git to push them to the Hub
|
||||
(along with any data files you may need).
|
||||
|
||||
At this point, you can load the dataset by running:
|
||||
|
||||
```python
|
||||
from datasets import load_dataset
|
||||
|
||||
load_dataset("OpenAssistant/{dataset_name}")
|
||||
```
|
||||
|
||||
Congratulations - you've now added a dataset to the OpenAssistant org!
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 201 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 61 KiB |
@@ -0,0 +1,238 @@
|
||||
import dbpng from "./img/db.png";
|
||||
import webdbpng from "./img/webdb.png";
|
||||
|
||||
# Data Schemas
|
||||
|
||||
## Introduction
|
||||
|
||||
This document describes the data schemas used by OpenAssistant. The schemas are
|
||||
defined as Python classes, but can be implemented in any format, be that Python,
|
||||
JSON, XML, SQL, Parquet files, etc.
|
||||
|
||||
Also, the schemas are leaning heavily on the
|
||||
[OpenAssistant Data Structures](https://docs.google.com/presentation/d/1iaX_nxasVWlvPiSNs0cllR9L_1neZq0RJxd6MFEalUY/edit?usp=sharing)
|
||||
presentation.
|
||||
|
||||
_Note on conformity: be pragmatic and decide what makes sense 🙂 , it's more
|
||||
important that we move forward than cramming everything into a uniform thing._
|
||||
|
||||
## Data Schemas
|
||||
|
||||
### Main structure: conversation trees
|
||||
|
||||
Conversation trees are the fundamental data structure. Many of the datasets we
|
||||
want to collect can be represented as conversation trees, such as QA datasets,
|
||||
chat logs, reddit dumps, etc. The main idea is that a conversation tree starts
|
||||
with a prompt and branches out from there. Every node can also have metadata,
|
||||
such as collected rankings, labels, or other information.
|
||||
|
||||
Datasets that just represent linear data, such as a list of questions and
|
||||
answers, can be represented as a conversation tree with just a single branch.
|
||||
|
||||
```python
|
||||
class ConversationTreeNode:
|
||||
text: str # The text of the node
|
||||
role: Literal['prompter', 'assistant'] # Whether the node is a user prompt/follow-up or an assistant response
|
||||
children: list[ConversationTreeNode] # The children of the node (if you have a linear conversation, this will be of length 0 or 1)
|
||||
metadata: dict[str, Any] # Node metadata (see below)
|
||||
|
||||
class ConversationTree:
|
||||
root: ConversationTreeNode # The node containing the initial prompt
|
||||
metadata: dict[str, Any] # Tree metadata, different from root node metadata.
|
||||
|
||||
```
|
||||
|
||||
### Metadata
|
||||
|
||||
Metadata encapsulates all the information that is not part of the conversation
|
||||
itself. This includes data about how the node was created (i.e. where it is
|
||||
from: crowd-sourced, templated, scraped, etc.), when it was created, its labels,
|
||||
tags, collected rankings, and other information.
|
||||
|
||||
## Example: Reddit AMA dataset
|
||||
|
||||
- Represent each question-follow-up set as a conversation tree.
|
||||
- Store things like usernames, timestamps, upvotes, etc. as metadata of the
|
||||
nodes.
|
||||
- Store things like the AMA title, the AMA author, the AMA subreddit, etc. as
|
||||
metadata of the tree.
|
||||
|
||||
## Example: QA dataset
|
||||
|
||||
- Represent each question-answer pair as a conversation tree.
|
||||
- The question is the prompt, the answer is the assistant response.
|
||||
- If the dataset contains multiple answers to each question, each answer can be
|
||||
a child of the question node.
|
||||
- If the dataset contains context text, it can be added as metadata to the
|
||||
question node.
|
||||
|
||||
## Example: Templated math problem dataset
|
||||
|
||||
- Represent each problem as a conversation tree with the problem text as the
|
||||
prompt and the solution as the assistant response.
|
||||
- Store the problem type (e.g. algebra, geometry, etc.) as metadata of the tree.
|
||||
- Store the template used also as metadata of the tree, as well as the source of
|
||||
the data used to fill the template.
|
||||
|
||||
## File Formats
|
||||
|
||||
The above data should be representable in most file formats, but some care has
|
||||
to be taken with respect to the recursive nature of the data.
|
||||
|
||||
Most row-major formats (JSON, Avro, Protobuf, etc.), as well as many databases,
|
||||
have no trouble with recursive (or arbitrary) schemas, but column-major formats,
|
||||
such as Parquet, do. For datasets with linear conversations, like many of the
|
||||
datasets we are collecting, this is not a problem. Instead of a tree of nodes,
|
||||
simply represent the conversation as a list of nodes. For true tree-like
|
||||
conversations, we should use a row-major format.
|
||||
|
||||
## Other considerations
|
||||
|
||||
- For text data of moderate size, it really doesn't matter much. It's more
|
||||
important to use consistent data structures and naming, than to worry about
|
||||
the exact file format.
|
||||
- For crowd-sourced data, we are collecting it into a SQL database already.
|
||||
- Parquet files are a good choice for large datasets, modulo the issues with
|
||||
recursive schemas.
|
||||
- If parquet can't be used, gzipped JSON-line files are a good choice. So are
|
||||
Avro files and protobufs. Keep in mind that column-major files are better for
|
||||
reading, filtering, and aggregating, but row-major files are better for
|
||||
writing.
|
||||
|
||||
# Task-Specific Data Schemas
|
||||
|
||||
The main tasks are a) generation of response text and b) ranking of responses.
|
||||
The following sections describe the data schemas for each of these tasks. Both
|
||||
should be implementable in parquet files.
|
||||
|
||||
Note: These files are meant to be consumed by ML algorithms and should ideally
|
||||
be produced from the above files.
|
||||
|
||||
## Common Data Structures
|
||||
|
||||
```python
|
||||
|
||||
class Message:
|
||||
text: str # The text of the message
|
||||
role: Literal['prompter', 'assistant'] # Whether the message is a user prompt/follow-up or an assistant response
|
||||
|
||||
class Thread:
|
||||
messages: list[Message] # The messages in the conversation
|
||||
|
||||
```
|
||||
|
||||
The corresponding parquet schemas are:
|
||||
|
||||
```parquet
|
||||
message Message {
|
||||
required binary text (UTF8);
|
||||
required binary role (UTF8);
|
||||
}
|
||||
|
||||
message Thread {
|
||||
required group messages (LIST) {
|
||||
repeated group list {
|
||||
required group element {
|
||||
required binary text (UTF8);
|
||||
required binary role (UTF8);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
## Generation
|
||||
|
||||
```python
|
||||
|
||||
class GenerationExample:
|
||||
thread: Thread # The conversation thread before the message to be generated
|
||||
message: Message # The message to be generated
|
||||
|
||||
```
|
||||
|
||||
The corresponding parquet schema is:
|
||||
|
||||
```parquet
|
||||
message GenerationExample {
|
||||
required group thread (LIST) {
|
||||
repeated group list {
|
||||
required group element {
|
||||
required binary text (UTF8);
|
||||
required binary role (UTF8);
|
||||
}
|
||||
}
|
||||
}
|
||||
required group message (LIST) {
|
||||
repeated group list {
|
||||
required group element {
|
||||
required binary text (UTF8);
|
||||
required binary role (UTF8);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
## Ranking
|
||||
|
||||
```python
|
||||
|
||||
class RankingExample:
|
||||
thread: Thread # The conversation thread before the message to be ranked
|
||||
messages: list[Message] # The messages to be ranked, in oder of decreasing preference
|
||||
|
||||
```
|
||||
|
||||
The corresponding parquet schema is:
|
||||
|
||||
```parquet
|
||||
message RankingExample {
|
||||
required group thread (LIST) {
|
||||
repeated group list {
|
||||
required group element {
|
||||
required binary text (UTF8);
|
||||
required binary role (UTF8);
|
||||
}
|
||||
}
|
||||
}
|
||||
required group messages (LIST) {
|
||||
repeated group list {
|
||||
required group element {
|
||||
required binary text (UTF8);
|
||||
required binary role (UTF8);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Databases
|
||||
|
||||
Open-Assistant uses two databases, one for the backend and one for the frontend.
|
||||
Both are [PostgreSQL](https://www.postgresql.org/) databases which run in docker
|
||||
containers.
|
||||
|
||||
### Backend ER-Diagram
|
||||
|
||||
ER-Diagram of backend Database
|
||||
|
||||
<img src={dbpng} />
|
||||
|
||||
**Notes**
|
||||
|
||||
- In order for the diagram to not be too messy, foreign key connection to
|
||||
`api_client` are not shown.
|
||||
- `frontend_message_id` references `id` of `taskInteraction` on the frontend
|
||||
|
||||
### Frontend ER-Diagram
|
||||
|
||||
ER-Diagram of frontend Database
|
||||
|
||||
<img src={webdbpng} />
|
||||
|
||||
**Notes**
|
||||
|
||||
- `id` of `registeredTask` references `id` of `message` on the backend
|
||||
@@ -0,0 +1,79 @@
|
||||
# Supervised Datasets
|
||||
|
||||
For discussion about usage of supervised data see issue
|
||||
<https://github.com/LAION-AI/Open-Assistant/issues/186>.
|
||||
|
||||
## Motivation
|
||||
|
||||
An important part of making the assistant useful is to teach it to understand
|
||||
and follow instructions, and to perform large set of tasks well.
|
||||
|
||||
While RLHF seems like the main ingredient, using existing supervised data might
|
||||
help.
|
||||
|
||||
There are two large-scale projects in the area of instruction-following /
|
||||
multitask learning: Promptsource and Natural Instructions - these projects
|
||||
crowdsourced templates and turned existing NLP datasets into
|
||||
instruction-following seq2seq form in natural langauge. They include both
|
||||
long-output training examples like generating a sentence that is a likely
|
||||
consequence of sentence in the prompt, and short-output, like rating prediction
|
||||
from review. (Pre-)training on such datasets should help model understand and
|
||||
follow instructions and teach it many abilities neccessary to perform a large
|
||||
set of tasks correctly. However, these data are not dialog-like - they do not
|
||||
look like a normal conversation.
|
||||
|
||||
There are also supervised dialog datasets such as Blended Skill Talk or SODA. In
|
||||
constrast to instruction-following datasets, dialog data is not as focused on
|
||||
"academic tasks" or correctness, but encourage the model to respond naturally
|
||||
like a person would.
|
||||
|
||||
### Promptsource
|
||||
|
||||
- GitHub: <https://github.com/bigscience-workshop/promptsource>
|
||||
- paper:
|
||||
[Multitask Prompted Training Enables Zero-Shot Task Generalization](https://arxiv.org/abs/2110.08207)
|
||||
- project for preparing templates and working with them
|
||||
- they generated a dataset using the templates:
|
||||
- <https://huggingface.co/datasets/bigscience/P3>
|
||||
- <https://huggingface.co/datasets/bigscience/xP3> (with multilingual data but
|
||||
English prompt)
|
||||
- <https://huggingface.co/datasets/bigscience/xP3mt> (with multilingual data
|
||||
and machine-translated prompt)
|
||||
- they trained zero-shot models (= models for following instructions in the
|
||||
input)
|
||||
- based on T5 architecture (encoder-decoder) called T0 family (and MT0 for
|
||||
multilingual)
|
||||
- and based on GPT architecture (decoder-only) called BloomZ family
|
||||
- Huggingface demo: [T0](https://huggingface.co/bigscience/T0pp),
|
||||
[MT0](https://huggingface.co/bigscience/mt0-large),
|
||||
[BloomZ](https://huggingface.co/bigscience/bloomz),
|
||||
- GitHub repo for T0: <https://github.com/bigscience-workshop/t-zero>
|
||||
- GitHub repo for BloomZ and MT0:
|
||||
<https://github.com/bigscience-workshop/xmtf>
|
||||
|
||||
### Natural instructions
|
||||
|
||||
- GitHub: <https://github.com/allenai/natural-instructions>
|
||||
- paper:
|
||||
[Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks](https://arxiv.org/abs/2204.07705)
|
||||
- they crowdsource directly the data prepared for instruction following (and
|
||||
learning from a few examples)
|
||||
- the GitHub repo = the dataset. It contains jsons
|
||||
- they trained zero-shot and in-context few-shot models (in multiple sizes):
|
||||
- mT5 architecture (encoder-decoder, multilingual pretraining)
|
||||
- Huggingface demo few-shot:
|
||||
<https://huggingface.co/allenai/tk-instruct-3b-def-pos>
|
||||
- Huggingface demo zero-shot:
|
||||
<https://huggingface.co/allenai/tk-instruct-3b-def>
|
||||
|
||||
### Blended Skill Talk
|
||||
|
||||
- used by Facebook in Blenderbot project
|
||||
- HuggingFace dataset: <https://huggingface.co/datasets/blended_skill_talk>
|
||||
- example model trained on it:
|
||||
<https://huggingface.co/facebook/blenderbot_small-90M>
|
||||
|
||||
### SODA
|
||||
|
||||
- GitHub: <https://github.com/skywalker023/sodaverse>
|
||||
- paper: <https://arxiv.org/abs/2212.10465>
|
||||
Reference in New Issue
Block a user