mirror of
https://github.com/wassname/Castor.git
synced 2026-10-08 11:38:45 +08:00
+ removed redundant loss regularization + added script to create torch word embedding file from word2vec model + updated README
SM model
References:
- Aliaksei _S_everyn and Alessandro _M_oschitti. 2015. Learning to Rank Short Text Pairs with Convolutional Deep Neural Networks. In Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '15). ACM, New York, NY, USA, 373-382. DOI: http://dx.doi.org/10.1145/2766462.2767738
Requirements
nltk==3.2.2
numpy==1.11.3
pytorch==0.1.12
gensim==1.0.1
The code uses torchtext for text processing. Set torchtext:
git clone https://github.com/pytorch/text.git
cd text
#use this commit number
git reset --hard 2980f1bc39ba6af332c5c2783da8bee109796d4c
python setup.py install
We use trec_eval for evaluation:
cd eval
tar -xvf trec_eval.9.0.tar.gz
cd trec_eval.9.0
make
cd ../..
Setup
Clone and create the dataset:
git clone https://github.com/castorini/data.git
git clone https://github.com/castorini/Castor.git
You should you see the following tree:
.
├── Castor
│ ├── README.md
│ ├── baseline_results.tsv
│ ├── idf_baseline
│ ├── kim_cnn
│ ├── mp_cnn
│ ├── setup.py
│ ├── sm_cnn
│ └── sm_modified_cnn
└── data
├── GloVe
├── ParagramEmbeddings
├── README.md
├── SimpleQuestions_v2
├── TrecQA
├── WikiQA
├── msrvid
├── requirements.txt
├── sick
├── twitterPPDB
├── utils
└── word2vec
To create the dataset:
cd Castor/sm_modified_cnn/
./create_dataset.sh
Training
Download the word2vec model from [here] (https://drive.google.com/file/d/0B2u_nClt6NbzUmhOZU55eEo4QWM/view?usp=sharing)
and copy it to the data/ folder.
You can train the SM model for the 4 following configurations:
- random - the word embedddings are initialized randomly and are tuned during training
- static - the word embeddings are static (Severyn and Moschitti, SIGIR'15)
- non-static - the word embeddings are tuned during training
- multichannel - contains static and non-static channels for question and answer conv layers
To train on GPU 0 with static configuration:
python train.py --mode static --gpu 0
NB: pass --no_cuda to use CPU
The trained model will be save to:
saves/static_best_model.pt
Testing the model
python main.py --trained_model saves/TREC/multichannel_best_model.pt
Evaluation
The performance on TrecQA dataset:
TrecQA:
Best dev
| Metric | rand | static | non-static | multichannel |
|---|---|---|---|---|
| MAP | 0.8096 | 0.8162 | 0.8387 | 0.8274 |
| MRR | 0.8560 | 0.8918 | 0.9058 | 0.8818 |
Test
| Metric | rand | static | non-static | multichannel |
|---|---|---|---|---|
| MAP | 0.7441 | 0.7524 | 0.7688 | 0.7641 |
| MRR | 0.8172 | 0.8012 | 0.8144 | 0.8174 |
WikiQA:
Best dev
| Metric | rand | static | non-static | multichannel |
|---|---|---|---|---|
| MAP | 0.7109 | 0.7204 | 0.7049 | 0.7245 |
| MRR | 0.7169 | 0.7234 | 0.7075 | 0.7259 |
Test
| Metric | rand | static | non-static | multichannel |
|---|---|---|---|---|
| MAP | 0.6313 | 0.6378 | 0.6455 | 0.6476 |
| MRR | 0.6522 | 0.6542 | 0.6689 | 0.6646 |
NB: The results on WikiQA are based on the SM model hyperparameters.
To create your own word2vec.pt file
- Download word2vec from here
to the
data/folder
python utils.py --input data/aquaint+wiki.txt.gz.ndim=50.bin