mirror of
https://github.com/wassname/Castor.git
synced 2026-09-25 13:10:11 +08:00
945b1fa6c018aabdf4c71c21907258b496ac93e9
With the line as-was the vocab cache was stored as b'the' rather than the, meaning that word2vec wasn't found for terms causing massive performance loss (AP 0.71 cf 0.77).
Castor
Pytorch deep learning models.
- SM model: Similarity between question and candidate answers.
Setting up Pytorch
You need Python 3.6 to use the models in this repository.
As per pytorch.org,
"Anaconda is our recommended package manager"
conda install pytorch torchvision -c soumith
Other pytorch installation modalities (e.g. via pip) can be seen at pytorch.org.
We also recommend gensim. We use some gensim modules to cache word embeddings.
conda install gensim
Pytorch has good support for GPU computations. CUDA installation guide for linux can be found here
NOTE: Install CUDA libraries before installing conda and pytorch.
data for models
Sourcing and pre-processing of input data for each model is described in respective model/README.md's
Baselines
- IDF Baseline: IDF overlap between question and candidate answers.
Languages
Python
93.7%
JavaScript
3.2%
Java
2.2%
HTML
0.5%
Shell
0.4%