Files
Castor/mp_cnn
Michael Tu 09b3a790a2 Use torchtext for MP-CNN (#76)
* Add SICK torchtext Dataset

* SICK dataset - torchtext postprocess into class probs

* Update model, driver, trainer, evaluator for SICK for torchtext

* MP-CNN: Fix bugs that prevent SICK from running on gpu 0

* MP-CNN: make SICK dataset w/ torchtext GPU-agnostic

* MP-CNN: support sparse features / idf overlap with torchtext

* Add MSRVID dataset with torchtext and update MP-CNN code to use it

* MP-CNN: Make torchtext deterministic by setting python random seed

* SICK and MSRVID datasets - add pair id for debug and build test vocab

* MP-CNN: Update readme to address potential module not found error

* MP-CNN: address review comments, can run on cpu
2017-11-01 12:13:30 -04:00
..
2017-11-01 12:13:30 -04:00
2017-11-01 12:13:30 -04:00
2017-11-01 12:13:30 -04:00
2017-11-01 12:13:30 -04:00
2017-11-01 12:13:30 -04:00
2017-11-01 12:13:30 -04:00

MP-CNN PyTorch Implementation

This is a PyTorch implementation of the following paper

The SICK and MSRVID datasets are available in https://github.com/castorini/data, as well as the GloVe word embeddings.

Directory layout should be like this:

├── Castor
│   ├── README.md
│   ├── ...
│   └── mp_cnn/
├── data
│   ├── README.md
│   ├── ...
│   ├── msrvid/
│   ├── sick/
│   └── GloVe/

To run MP-CNN on the SICK dataset, use the following command. --dropout 0 is for mimicking the original paper, although adding dropout can improve performance. If you have any problems running it check the Troubleshooting section below.

python main.py mpcnn.sick.model.castor --dataset sick --epochs 19 --epsilon 1e-7 --dropout 0
Implementation and config Pearson's r Spearman's p
Paper 0.8686 0.8047
PyTorch using above config 0.8763 0.8215

To run MP-CNN on the MSRVID dataset, use the following command:

python main.py mpcnn.msrvid.model.castor --dataset msrvid --batch-size 16 --epsilon 1e-7 --epochs 32 --dropout 0 --regularization 0.0025
Implementation and config Pearson's r
Paper 0.9090
PyTorch using above config 0.9050

These are not the optimal hyperparameters but they are decent. This README will be updated with more optimal hyperparameters and results in the future.

To see all options available, use

python main.py --help

Troubleshooting

ModuleNotFoundError: datasets

Traceback (most recent call last):
  File "main.py", line 9, in <module>
    from dataset import MPCNNDatasetFactory
  File "/u/z3tu/castorini/Castor/mp_cnn/dataset.py", line 12, in <module>
    from datasets.sick import SICK
ModuleNotFoundError: No module named 'datasets'

You need to make sure the repository root is in your PYTHONPATH environment variable. One way to do this is while you are in the repo root (Castor) as your current working directory, run export PYTHONPATH=$(pwd).

Optional Dependencies

To optionally visualize the learning curve during training, we make use of https://github.com/lanpa/tensorboard-pytorch to connect to TensorBoard. These projects require TensorFlow as a dependency, so you need to install TensorFlow before running the commands below. After these are installed, just add --tensorboard when running main.py and open TensorBoard in the browser.

pip install tensorboardX
pip install tensorflow-tensorboard