Commit Graph
85 Commits
Author SHA1 Message Date
Piotr Czapla 0085c18ae0 Fix splitting by article for local languages and make it more memory efficient
The issue was that the code assumed that empty lines have space followed by a new  line, which isn't the case for datasets generated by our scripts.
2019-01-03 12:16:41 +01:00
Piotr Czapla f784d7bcd2 Make the article detection code work with our wikitext 2019-01-01 15:13:14 +01:00
Piotr Czapla 0945c699c7 Fixes #25 by adding article title in markdown format to wiki text 2019-01-01 15:07:02 +01:00
Piotr Czapla 4a858f3572 Simplfy and unify input data parsing 2019-01-01 14:43:12 +01:00
Piotr Czapla 82d6a30b11 Fix SentencePiece implementation, to get 94.5% on imdb
- train on whole articles when tokenizer is sentencepice
- add moses tokenizer to get the same tokens on imdb as on wt103
- apply pre / post processing rules to moses tokenzier so that sentencepiece works on correctly preprocessed text (lowercased with html removed)
- add alpha implementation on xnli (no tests)
- remove vocab adaptation when swiching from wiki to imdb. As otherwise we get 50% of missing words and 93% accuracy on imdb
2018-12-27 15:14:41 +01:00
Piotr Czapla ff30fb5642 Fastai upgrade 2018-12-27 14:57:15 +01:00
Piotr Czapla f9394b9af1 Add callbacks to save history and best weights remove bs & drop_mult 2018-12-12 00:31:16 +01:00
Piotr Czapla 1b9d04d7ef Expose use_test_for_validation as param to train 2018-12-09 22:45:21 +01:00
Piotr Czapla 34b1d18600 load wikipedia as articles for v & fv tok. 2018-12-09 22:42:17 +01:00
Piotr Czapla 2338618563 Make biclasifier head a hyperparameter. 2018-12-09 00:39:27 +01:00
Piotr Czapla 978bbcc923 Fix issue with finetuning language model (it wasn't freezed) 2018-12-08 23:10:31 +01:00
Piotr Czapla 50ab34eea9 Fix classification scripts and give a way to test accuracy on test set 2018-12-07 16:10:20 +01:00
Piotr Czapla 4ae2a158a2 Respect batch size in training lm model. 2018-12-06 23:29:43 +01:00
Piotr Czapla 2acbf15555 Tweak hyper training params of cls (drop_mul, bs, true_wd=True)
I've set the same hyperparams as in lesson3
2018-12-05 16:27:14 +01:00
Piotr Czapla 4454af167f Use more text during pretraining of imdb 2018-12-05 16:26:16 +01:00
Piotr Czapla 5524006d81 Respect max_vocab 2018-12-04 16:41:22 +01:00
Piotr Czapla 908c3d7e8a bug fix 2018-12-04 16:41:07 +01:00
Piotr Czapla f25deb3049 Use trn + tst for LM training 2018-12-04 01:43:01 +01:00
Piotr Czapla 9826f6881c Fix the way trn & val set is created in imdb 2018-12-04 01:35:59 +01:00
Piotr Czapla 039624870a Fix loading tokenized data set in train cls 2018-12-04 01:28:03 +01:00
Piotr Czapla a499bf9a20 Make cls train work with relative paths 2018-12-04 01:24:44 +01:00
Piotr Czapla 35a5dcb75a Fix issue when running cls training from command line 2018-12-04 00:59:36 +01:00
Piotr Czapla 82c955ce6a Add different tokenization algorithms to train_clas 2018-12-04 00:51:21 +01:00
Piotr Czapla 83427aadc6 Add moses with fastai preprocessing 2018-12-02 21:43:57 +01:00
Piotr Czapla b9eb7388f6 Fix bidir for fastai tokenizer 2018-12-01 23:58:14 +01:00
Piotr Czapla 0c4aed6d05 Really fix conversion from str to Tokenzier 2018-12-01 16:51:21 +01:00
Piotr Czapla e14c967cc8 Fix string parsing in tokenzier 2018-12-01 16:46:25 +01:00
Piotr Czapla 887211137a Add fastai tokenizer to pretrain_lm 2018-12-01 16:42:11 +01:00
Piotr Czapla c17dcce75e spelling 2018-12-01 15:24:30 +01:00
Piotr Czapla ab9faa2ad9 Clean up Fire interface. 2018-12-01 13:42:10 +01:00
Piotr Czapla 4b29376b44 Rewrite classifier to use changed pretrain_lm 2018-12-01 10:58:46 +01:00
Piotr Czapla b3f5ae1ad5 Clean the way we save models 2018-12-01 10:58:23 +01:00
Piotr Czapla 80d4d4da29 Extract params to an experiment data class
You can run this as follows:

`python -m ulmfit.pretrain_lm --dir-path 'data/wiki/wikitext-2'  --qrnn=True train_lm --num_epochs=1`
2018-11-25 12:37:45 +01:00
Piotr Czapla be117abac4 Make the end to end test run correctly 2018-11-24 23:51:23 +01:00
Piotr Czapla 8da47324c2 Clean ups and fixes 2018-11-22 15:32:40 +01:00
Piotr Czapla dbd4884228 Fixes after mergin with master and updateing to newset fastai 2018-11-22 01:14:38 +01:00
Piotr Czapla 934fc79384 Merge branch 'master' into bilm 2018-11-21 23:48:23 +01:00
Piotr Czapla 9aa877dcd0 Share trained LM between different classfiication runs 2018-11-21 18:51:42 +01:00
Piotr Czapla 979eb196d8 Change the classfication training learning rate to the one that was working te best in my exp. on bidirectional clasification 2018-11-21 18:46:05 +01:00
Piotr Czapla cbed02d5e0 Merge branch 'master' into bilm 2018-11-21 18:23:54 +01:00
NAUSICAA\Julian 8cb867b066 Compatibility with new fastai version 2018-11-20 19:21:02 -03:00
Aayush 429a3fe4c2 Merge pull request #13 from n-waves/models_path_fix
Models path fix, default_rules and sentencepiece support for train_clas
2018-11-20 12:25:25 +05:30
Aayush 5ffbe8ba5c add vocab_size to sentencepiece 2018-11-19 19:13:55 +05:30
Piotr Czapla 7f1f8efcc3 Working version of biclassfier 2018-11-19 12:58:58 +01:00
Piotr Czapla c821d2e783 first version of bi classifier 2018-11-19 09:59:08 +01:00
Nirant 56a9feec7d Removed dependency note, use requirements.txt 2018-11-19 12:22:41 +05:30
Piotr Czapla 37b73e262f Fix bug where de-all was de-100 2018-11-17 17:14:13 +01:00
Piotr Czapla 2674a713fc fix resuming training of classifier 2018-11-17 17:13:15 +01:00
Piotr Czapla e97085337e Fix dropout and classification accuracy. 0.91 on imdb 2018-11-17 16:47:17 +01:00
Sebastian 7fe059c9a2 Fixed models path for vocabulary 2018-11-17 14:47:27 +00:00