Files
Castor/anserini_dependency/README.md
T
MeowFei 8fd1ddcb4c Added RetrieveSentences.py (#74)
* added RetrieveSentences.py

* removed index
2017-10-28 15:56:20 -04:00

1.5 KiB

Retrieve Sentences

1. Clone Anserini and Castor

git clone https://github.com/castorini/Anserini.git
git clone https://github.com/castorini/Castor.git

Your directory structure should look like

├── Anserini
└── Castor

2. Compile Anserini

cd Anserini
mvn package

This creates anserini-0.0.1-SNAPSHOT.jar at Anserini/target

3. Download Dependencies

  • Download the TrecQA lucene index
  • Download the Google word2vec file from here

4. Run the following command

python ./anserini_dependency/RetrieveSentences.py

Possible parameters are:

option input format default description
-index string N/A Path of the Lucene index
-embeddings string "" Path of the word2vec index
-topics string "" topics file
-query string "" a single query
-hits [1, inf) 100 max number of hits to return
-scorer string Idf passage scores (Idf or Wmd)
-k [1, inf) 1 top-k passages to be retrieved

Note: Either a query or a topic must be passed in as an argument; they can't be both empty.