Piotr Czapla
0085c18ae0
Fix splitting by article for local languages and make it more memory efficient
...
The issue was that the code assumed that empty lines have space followed by a new line, which isn't the case for datasets generated by our scripts.
2019-01-03 12:16:41 +01:00
Piotr Czapla
f784d7bcd2
Make the article detection code work with our wikitext
2019-01-01 15:13:14 +01:00
Piotr Czapla
0945c699c7
Fixes #25 by adding article title in markdown format to wiki text
2019-01-01 15:07:02 +01:00
Piotr Czapla
4a858f3572
Simplfy and unify input data parsing
2019-01-01 14:43:12 +01:00
Piotr Czapla
82d6a30b11
Fix SentencePiece implementation, to get 94.5% on imdb
...
- train on whole articles when tokenizer is sentencepice
- add moses tokenizer to get the same tokens on imdb as on wt103
- apply pre / post processing rules to moses tokenzier so that sentencepiece works on correctly preprocessed text (lowercased with html removed)
- add alpha implementation on xnli (no tests)
- remove vocab adaptation when swiching from wiki to imdb. As otherwise we get 50% of missing words and 93% accuracy on imdb
2018-12-27 15:14:41 +01:00
Piotr Czapla
ff30fb5642
Fastai upgrade
2018-12-27 14:57:15 +01:00
Piotr Czapla
f9394b9af1
Add callbacks to save history and best weights remove bs & drop_mult
2018-12-12 00:31:16 +01:00
Piotr Czapla
1b9d04d7ef
Expose use_test_for_validation as param to train
2018-12-09 22:45:21 +01:00
Piotr Czapla
34b1d18600
load wikipedia as articles for v & fv tok.
2018-12-09 22:42:17 +01:00
Piotr Czapla
2338618563
Make biclasifier head a hyperparameter.
2018-12-09 00:39:27 +01:00
Piotr Czapla
978bbcc923
Fix issue with finetuning language model (it wasn't freezed)
2018-12-08 23:10:31 +01:00
Piotr Czapla
50ab34eea9
Fix classification scripts and give a way to test accuracy on test set
2018-12-07 16:10:20 +01:00
Piotr Czapla
4ae2a158a2
Respect batch size in training lm model.
2018-12-06 23:29:43 +01:00
Piotr Czapla
2acbf15555
Tweak hyper training params of cls (drop_mul, bs, true_wd=True)
...
I've set the same hyperparams as in lesson3
2018-12-05 16:27:14 +01:00
Piotr Czapla
4454af167f
Use more text during pretraining of imdb
2018-12-05 16:26:16 +01:00
Piotr Czapla
5524006d81
Respect max_vocab
2018-12-04 16:41:22 +01:00
Piotr Czapla
908c3d7e8a
bug fix
2018-12-04 16:41:07 +01:00
Piotr Czapla
f25deb3049
Use trn + tst for LM training
2018-12-04 01:43:01 +01:00
Piotr Czapla
9826f6881c
Fix the way trn & val set is created in imdb
2018-12-04 01:35:59 +01:00
Piotr Czapla
039624870a
Fix loading tokenized data set in train cls
2018-12-04 01:28:03 +01:00
Piotr Czapla
a499bf9a20
Make cls train work with relative paths
2018-12-04 01:24:44 +01:00
Piotr Czapla
35a5dcb75a
Fix issue when running cls training from command line
2018-12-04 00:59:36 +01:00
Piotr Czapla
82c955ce6a
Add different tokenization algorithms to train_clas
2018-12-04 00:51:21 +01:00
Piotr Czapla
83427aadc6
Add moses with fastai preprocessing
2018-12-02 21:43:57 +01:00
Piotr Czapla
b9eb7388f6
Fix bidir for fastai tokenizer
2018-12-01 23:58:14 +01:00
Piotr Czapla
0c4aed6d05
Really fix conversion from str to Tokenzier
2018-12-01 16:51:21 +01:00
Piotr Czapla
e14c967cc8
Fix string parsing in tokenzier
2018-12-01 16:46:25 +01:00
Piotr Czapla
887211137a
Add fastai tokenizer to pretrain_lm
2018-12-01 16:42:11 +01:00
Piotr Czapla
c17dcce75e
spelling
2018-12-01 15:24:30 +01:00
Piotr Czapla
ab9faa2ad9
Clean up Fire interface.
2018-12-01 13:42:10 +01:00
Piotr Czapla
4b29376b44
Rewrite classifier to use changed pretrain_lm
2018-12-01 10:58:46 +01:00
Piotr Czapla
b3f5ae1ad5
Clean the way we save models
2018-12-01 10:58:23 +01:00
Piotr Czapla
80d4d4da29
Extract params to an experiment data class
...
You can run this as follows:
`python -m ulmfit.pretrain_lm --dir-path 'data/wiki/wikitext-2' --qrnn=True train_lm --num_epochs=1`
2018-11-25 12:37:45 +01:00
Piotr Czapla
be117abac4
Make the end to end test run correctly
2018-11-24 23:51:23 +01:00
Piotr Czapla
8da47324c2
Clean ups and fixes
2018-11-22 15:32:40 +01:00
Piotr Czapla
dbd4884228
Fixes after mergin with master and updateing to newset fastai
2018-11-22 01:14:38 +01:00
Piotr Czapla
934fc79384
Merge branch 'master' into bilm
2018-11-21 23:48:23 +01:00
Piotr Czapla
9aa877dcd0
Share trained LM between different classfiication runs
2018-11-21 18:51:42 +01:00
Piotr Czapla
979eb196d8
Change the classfication training learning rate to the one that was working te best in my exp. on bidirectional clasification
2018-11-21 18:46:05 +01:00
Piotr Czapla
cbed02d5e0
Merge branch 'master' into bilm
2018-11-21 18:23:54 +01:00
NAUSICAA\Julian
8cb867b066
Compatibility with new fastai version
2018-11-20 19:21:02 -03:00
Aayush
429a3fe4c2
Merge pull request #13 from n-waves/models_path_fix
...
Models path fix, default_rules and sentencepiece support for train_clas
2018-11-20 12:25:25 +05:30
Aayush
5ffbe8ba5c
add vocab_size to sentencepiece
2018-11-19 19:13:55 +05:30
Piotr Czapla
7f1f8efcc3
Working version of biclassfier
2018-11-19 12:58:58 +01:00
Piotr Czapla
c821d2e783
first version of bi classifier
2018-11-19 09:59:08 +01:00
Nirant
56a9feec7d
Removed dependency note, use requirements.txt
2018-11-19 12:22:41 +05:30
Piotr Czapla
37b73e262f
Fix bug where de-all was de-100
2018-11-17 17:14:13 +01:00
Piotr Czapla
2674a713fc
fix resuming training of classifier
2018-11-17 17:13:15 +01:00
Piotr Czapla
e97085337e
Fix dropout and classification accuracy. 0.91 on imdb
2018-11-17 16:47:17 +01:00
Sebastian
7fe059c9a2
Fixed models path for vocabulary
2018-11-17 14:47:27 +00:00