Commit Graph
125 Commits
Author SHA1 Message Date
Piotr Czapla f1b49a0d34 fix imports in learner (fastai adpatation) 2018-12-27 15:15:09 +01:00
Piotr Czapla 82d6a30b11 Fix SentencePiece implementation, to get 94.5% on imdb
- train on whole articles when tokenizer is sentencepice
- add moses tokenizer to get the same tokens on imdb as on wt103
- apply pre / post processing rules to moses tokenzier so that sentencepiece works on correctly preprocessed text (lowercased with html removed)
- add alpha implementation on xnli (no tests)
- remove vocab adaptation when swiching from wiki to imdb. As otherwise we get 50% of missing words and 93% accuracy on imdb
2018-12-27 15:14:41 +01:00
Piotr Czapla ff30fb5642 Fastai upgrade 2018-12-27 14:57:15 +01:00
Piotr Czapla dca502a398 Add docs how to train a classfier 2018-12-12 00:31:48 +01:00
Piotr Czapla f9394b9af1 Add callbacks to save history and best weights remove bs & drop_mult 2018-12-12 00:31:16 +01:00
Piotr Czapla fac3ce343b Fix BiLM implementation after upgrade of fastai 2018-12-12 00:29:23 +01:00
Piotr Czapla 1b9d04d7ef Expose use_test_for_validation as param to train 2018-12-09 22:45:21 +01:00
Piotr Czapla 34b1d18600 load wikipedia as articles for v & fv tok. 2018-12-09 22:42:17 +01:00
Piotr Czapla 2338618563 Make biclasifier head a hyperparameter. 2018-12-09 00:39:27 +01:00
Piotr Czapla 591393b96c Change AvgPooling to BiPooling as a default 2018-12-08 23:48:02 +01:00
Piotr Czapla 978bbcc923 Fix issue with finetuning language model (it wasn't freezed) 2018-12-08 23:10:31 +01:00
Piotr Czapla 0b7dc7ec6b Add two experiments showing how to train imdb classification to 94% accuracy 2018-12-07 16:10:32 +01:00
Piotr Czapla 50ab34eea9 Fix classification scripts and give a way to test accuracy on test set 2018-12-07 16:10:20 +01:00
Piotr Czapla 4ae2a158a2 Respect batch size in training lm model. 2018-12-06 23:29:43 +01:00
Piotr Czapla 6a899015ed Generate imdb unsp.csv 2018-12-05 16:27:41 +01:00
Piotr Czapla 2acbf15555 Tweak hyper training params of cls (drop_mul, bs, true_wd=True)
I've set the same hyperparams as in lesson3
2018-12-05 16:27:14 +01:00
Piotr Czapla 4454af167f Use more text during pretraining of imdb 2018-12-05 16:26:16 +01:00
Piotr Czapla 5524006d81 Respect max_vocab 2018-12-04 16:41:22 +01:00
Piotr Czapla 908c3d7e8a bug fix 2018-12-04 16:41:07 +01:00
Piotr Czapla f25deb3049 Use trn + tst for LM training 2018-12-04 01:43:01 +01:00
Piotr Czapla 9826f6881c Fix the way trn & val set is created in imdb 2018-12-04 01:35:59 +01:00
Piotr Czapla 039624870a Fix loading tokenized data set in train cls 2018-12-04 01:28:03 +01:00
Piotr Czapla a499bf9a20 Make cls train work with relative paths 2018-12-04 01:24:44 +01:00
Piotr Czapla 35a5dcb75a Fix issue when running cls training from command line 2018-12-04 00:59:36 +01:00
Piotr Czapla 82c955ce6a Add different tokenization algorithms to train_clas 2018-12-04 00:51:21 +01:00
Piotr Czapla 83427aadc6 Add moses with fastai preprocessing 2018-12-02 21:43:57 +01:00
Piotr Czapla b9eb7388f6 Fix bidir for fastai tokenizer 2018-12-01 23:58:14 +01:00
Piotr Czapla 0c4aed6d05 Really fix conversion from str to Tokenzier 2018-12-01 16:51:21 +01:00
Piotr Czapla e14c967cc8 Fix string parsing in tokenzier 2018-12-01 16:46:25 +01:00
Piotr Czapla 887211137a Add fastai tokenizer to pretrain_lm 2018-12-01 16:42:11 +01:00
Piotr Czapla c17dcce75e spelling 2018-12-01 15:24:30 +01:00
Piotr Czapla ab9faa2ad9 Clean up Fire interface. 2018-12-01 13:42:10 +01:00
Piotr Czapla 4b29376b44 Rewrite classifier to use changed pretrain_lm 2018-12-01 10:58:46 +01:00
Piotr Czapla b3f5ae1ad5 Clean the way we save models 2018-12-01 10:58:23 +01:00
Piotr Czapla 6d6ebef1ca Update to newst fastai 2018-12-01 10:57:20 +01:00
Piotr Czapla 80d4d4da29 Extract params to an experiment data class
You can run this as follows:

`python -m ulmfit.pretrain_lm --dir-path 'data/wiki/wikitext-2'  --qrnn=True train_lm --num_epochs=1`
2018-11-25 12:37:45 +01:00
Piotr Czapla eea9be09db Merge pull request #6 from n-waves/bilm
Bidirectional language model + fixes to the end-to-end tests
2018-11-24 23:58:00 +01:00
Piotr Czapla be117abac4 Make the end to end test run correctly 2018-11-24 23:51:23 +01:00
Piotr Czapla 8da47324c2 Clean ups and fixes 2018-11-22 15:32:40 +01:00
Piotr Czapla dbd4884228 Fixes after mergin with master and updateing to newset fastai 2018-11-22 01:14:38 +01:00
Piotr Czapla 934fc79384 Merge branch 'master' into bilm 2018-11-21 23:48:23 +01:00
Piotr Czapla 9aa877dcd0 Share trained LM between different classfiication runs 2018-11-21 18:51:42 +01:00
Piotr Czapla 979eb196d8 Change the classfication training learning rate to the one that was working te best in my exp. on bidirectional clasification 2018-11-21 18:46:05 +01:00
Piotr Czapla 6e3ef21b1f Add Avg BiClassifier 2018-11-21 18:44:49 +01:00
Piotr Czapla cbed02d5e0 Merge branch 'master' into bilm 2018-11-21 18:23:54 +01:00
NAUSICAA\Julian 8cb867b066 Compatibility with new fastai version 2018-11-20 19:21:02 -03:00
NAUSICAA\Julian 79691791b3 Typo fix 2018-11-20 18:53:23 -03:00
Julian Eisenschlos 88d94f38ba Merge pull request #16 from n-waves/sentencepiece_fixes
Sentencepiece Fixes
2018-11-20 17:14:41 -03:00
NAUSICAA\Julian 40a2990322 Add suggested changes to rules system 2018-11-20 09:06:03 -03:00
NAUSICAA\Julian a36c0518c6 Merge branch 'master' of https://github.com/n-waves/ulmfit-multilingual into sentencepiece_fixes 2018-11-20 08:55:37 -03:00