Michael Tu 5bf33bf8ea Update README to use Castor-models and Instructions for Internal Users (#113)
* Update instructions to use Castor-models
* Consolidate requirements.txt
* Refine README with convenience scripts
* Update internal instructions
* MP-CNN working dir minor edit
2018-05-25 18:14:10 -04:00
2018-05-25 11:52:25 -04:00
2017-10-29 15:55:36 -04:00
2018-05-25 11:52:25 -04:00
2017-11-04 18:22:40 -04:00
2017-04-18 12:32:43 -04:00
2017-11-25 14:39:39 -05:00

Castor

This is the common repo for PyTorch deep learning models by the Data Systems Group at the University of Waterloo.

Models

Predictions Over One Input Sequence

For sentiment analysis, topic classification, etc.

Predictions Over Two Input Sequences

For paraphrase detection, question answering, etc.

Each model directory has a README.md with further details.

Setting up PyTorch

If you are an internal Castor contributor and is planning to use the Data System Group's GPU machines in the lab, please follow the instructions here instead.

Copy and run the command at https://pytorch.org/ for your environment. PyTorch recommends the Anaconda environment, which we use in our lab. We are currently targeting PyTorch 0.4 for our codebase.

The typical installation command is

conda install pytorch torchvision -c pytorch

Other Python packages we use can be installed via pip:

pip install -r requirements.txt

Please also run the following inside the utils directory to build the trec_eval tool for evaluating certain datasets.

./get_trec_eval.sh

Data and Pre-Trained Models

If you are an internal Castor contributor and is planning to use the Data System Group's GPU machines in the lab, please follow the instructions here instead.

Data associated for use with this repository can be found at: https://git.uwaterloo.ca/jimmylin/Castor-data.git.

Pre-trained models can be found at: https://git.uwaterloo.ca/jimmylin/Castor-models.

Your directory structure should look like

.
├── Castor
├── Castor-data
└── Castor-models

For example (if you use HTTPS instead of SSH):

git clone https://github.com/castorini/Castor.git
git clone https://git.uwaterloo.ca/jimmylin/Castor-data.git
git clone https://git.uwaterloo.ca/jimmylin/Castor-models.git

After cloning the Castor-data repo, you need to unzip embeddings and run data pre-processing scripts. You can choose to follow instructions under each dataset / embedding directory separately, or just run the following script in Castor-data to do all of the steps for you:

./setup.sh
S
Description
PyTorch deep learning models for text processing
Readme Apache-2.0
1.2 MiB
0 Stars 1 Watchers 0 Forks
Languages
Python 93.7%
JavaScript 3.2%
Java 2.2%
HTML 0.5%
Shell 0.4%