Piotr Czapla
8a2fed41c5
Merge pull request #18 from n-waves/refactor
...
[WIP] Refactoring #17
2019-01-01 15:13:57 +01:00
Piotr Czapla
f784d7bcd2
Make the article detection code work with our wikitext
2019-01-01 15:13:14 +01:00
Piotr Czapla
0945c699c7
Fixes #25 by adding article title in markdown format to wiki text
2019-01-01 15:07:02 +01:00
Piotr Czapla
4a858f3572
Simplfy and unify input data parsing
2019-01-01 14:43:12 +01:00
Piotr Czapla
d269d53d7d
Use fastai tokens instead of <xxx>
...
f'xx{token_name}' are kept as one token by Moses tokenizer, which is sometimes required if you want to have moses in tokenizers pipeline, and we use that for imdb.
2019-01-01 14:40:59 +01:00
Piotr Czapla
514a9e6b86
Fix BiLM training after update to newest fastai
2018-12-31 12:13:53 +01:00
Piotr Czapla
f1b49a0d34
fix imports in learner (fastai adpatation)
2018-12-27 15:15:09 +01:00
Piotr Czapla
82d6a30b11
Fix SentencePiece implementation, to get 94.5% on imdb
...
- train on whole articles when tokenizer is sentencepice
- add moses tokenizer to get the same tokens on imdb as on wt103
- apply pre / post processing rules to moses tokenzier so that sentencepiece works on correctly preprocessed text (lowercased with html removed)
- add alpha implementation on xnli (no tests)
- remove vocab adaptation when swiching from wiki to imdb. As otherwise we get 50% of missing words and 93% accuracy on imdb
2018-12-27 15:14:41 +01:00
Piotr Czapla
ff30fb5642
Fastai upgrade
2018-12-27 14:57:15 +01:00
Piotr Czapla
dca502a398
Add docs how to train a classfier
2018-12-12 00:31:48 +01:00
Piotr Czapla
f9394b9af1
Add callbacks to save history and best weights remove bs & drop_mult
2018-12-12 00:31:16 +01:00
Piotr Czapla
fac3ce343b
Fix BiLM implementation after upgrade of fastai
2018-12-12 00:29:23 +01:00
Piotr Czapla
1b9d04d7ef
Expose use_test_for_validation as param to train
2018-12-09 22:45:21 +01:00
Piotr Czapla
34b1d18600
load wikipedia as articles for v & fv tok.
2018-12-09 22:42:17 +01:00
Piotr Czapla
2338618563
Make biclasifier head a hyperparameter.
2018-12-09 00:39:27 +01:00
Piotr Czapla
591393b96c
Change AvgPooling to BiPooling as a default
2018-12-08 23:48:02 +01:00
Piotr Czapla
978bbcc923
Fix issue with finetuning language model (it wasn't freezed)
2018-12-08 23:10:31 +01:00
Piotr Czapla
0b7dc7ec6b
Add two experiments showing how to train imdb classification to 94% accuracy
2018-12-07 16:10:32 +01:00
Piotr Czapla
50ab34eea9
Fix classification scripts and give a way to test accuracy on test set
2018-12-07 16:10:20 +01:00
Piotr Czapla
4ae2a158a2
Respect batch size in training lm model.
2018-12-06 23:29:43 +01:00
Piotr Czapla
6a899015ed
Generate imdb unsp.csv
2018-12-05 16:27:41 +01:00
Piotr Czapla
2acbf15555
Tweak hyper training params of cls (drop_mul, bs, true_wd=True)
...
I've set the same hyperparams as in lesson3
2018-12-05 16:27:14 +01:00
Piotr Czapla
4454af167f
Use more text during pretraining of imdb
2018-12-05 16:26:16 +01:00
Piotr Czapla
5524006d81
Respect max_vocab
2018-12-04 16:41:22 +01:00
Piotr Czapla
908c3d7e8a
bug fix
2018-12-04 16:41:07 +01:00
Piotr Czapla
f25deb3049
Use trn + tst for LM training
2018-12-04 01:43:01 +01:00
Piotr Czapla
9826f6881c
Fix the way trn & val set is created in imdb
2018-12-04 01:35:59 +01:00
Piotr Czapla
039624870a
Fix loading tokenized data set in train cls
2018-12-04 01:28:03 +01:00
Piotr Czapla
a499bf9a20
Make cls train work with relative paths
2018-12-04 01:24:44 +01:00
Piotr Czapla
35a5dcb75a
Fix issue when running cls training from command line
2018-12-04 00:59:36 +01:00
Piotr Czapla
82c955ce6a
Add different tokenization algorithms to train_clas
2018-12-04 00:51:21 +01:00
Piotr Czapla
83427aadc6
Add moses with fastai preprocessing
2018-12-02 21:43:57 +01:00
Piotr Czapla
b9eb7388f6
Fix bidir for fastai tokenizer
2018-12-01 23:58:14 +01:00
Piotr Czapla
0c4aed6d05
Really fix conversion from str to Tokenzier
2018-12-01 16:51:21 +01:00
Piotr Czapla
e14c967cc8
Fix string parsing in tokenzier
2018-12-01 16:46:25 +01:00
Piotr Czapla
887211137a
Add fastai tokenizer to pretrain_lm
2018-12-01 16:42:11 +01:00
Piotr Czapla
c17dcce75e
spelling
2018-12-01 15:24:30 +01:00
Piotr Czapla
ab9faa2ad9
Clean up Fire interface.
2018-12-01 13:42:10 +01:00
Piotr Czapla
4b29376b44
Rewrite classifier to use changed pretrain_lm
2018-12-01 10:58:46 +01:00
Piotr Czapla
b3f5ae1ad5
Clean the way we save models
2018-12-01 10:58:23 +01:00
Piotr Czapla
6d6ebef1ca
Update to newst fastai
2018-12-01 10:57:20 +01:00
Piotr Czapla
80d4d4da29
Extract params to an experiment data class
...
You can run this as follows:
`python -m ulmfit.pretrain_lm --dir-path 'data/wiki/wikitext-2' --qrnn=True train_lm --num_epochs=1`
2018-11-25 12:37:45 +01:00
Piotr Czapla
eea9be09db
Merge pull request #6 from n-waves/bilm
...
Bidirectional language model + fixes to the end-to-end tests
2018-11-24 23:58:00 +01:00
Piotr Czapla
be117abac4
Make the end to end test run correctly
2018-11-24 23:51:23 +01:00
Piotr Czapla
8da47324c2
Clean ups and fixes
2018-11-22 15:32:40 +01:00
Piotr Czapla
dbd4884228
Fixes after mergin with master and updateing to newset fastai
2018-11-22 01:14:38 +01:00
Piotr Czapla
934fc79384
Merge branch 'master' into bilm
2018-11-21 23:48:23 +01:00
Piotr Czapla
9aa877dcd0
Share trained LM between different classfiication runs
2018-11-21 18:51:42 +01:00
Piotr Czapla
979eb196d8
Change the classfication training learning rate to the one that was working te best in my exp. on bidirectional clasification
2018-11-21 18:46:05 +01:00
Piotr Czapla
6e3ef21b1f
Add Avg BiClassifier
2018-11-21 18:44:49 +01:00