200 Commits
Author SHA1 Message Date
Piotr Czapla a630242f97 Expose validate_cls in ulmfit module 2019-02-11 11:01:08 +01:00
Piotr Czapla c63fe258d2 Fix validataion and add option to add noise to training labels 2019-02-11 10:53:25 +01:00
Piotr Czapla 22e40d81f3 Additional DE results and first results on FR and JA 2019-02-11 10:51:42 +01:00
Piotr Czapla e7ec9f0233 fix markdown formatting 2019-02-10 10:52:03 +01:00
Piotr Czapla 07ab9ccaf8 Add first MLDoc results 2019-02-10 09:53:42 +01:00
Piotr Czapla 672ef2d59f Remove old result logs 2019-02-10 09:53:26 +01:00
Piotr Czapla 26736d95de Add code to test & train mldoc classifier 2019-02-10 09:52:58 +01:00
Tomasz Pietruszka 950daac903 Merge branch 'backwards-lm' 2019-01-23 00:08:22 +01:00
Tomasz Pietruszka c6cf44a7b9 Fixed test_bilm_classifier_loads_encoder - inconsistent settings, problematic input (too short) 2019-01-23 00:05:45 +01:00
Tomasz Pietruszka 386fc49431 Adapted test_text_data to batch dimension being the first, and only xxbos token at the start 2019-01-22 20:31:30 +01:00
Tomasz Pietruszka 0e864fb291 test_end_to_end now working, added instructions to README, sentencepiece dependency 2019-01-22 20:10:38 +01:00
Tomasz Pietruszka 6fa30829c4 Merge branch 'master' into backwards-lm 2019-01-14 17:20:48 +01:00
Tomasz Pietruszka dd296f3088 prepare_wiki.sh made executable 2019-01-13 17:30:35 +01:00
Tomasz Pietruszka 5ba8ae4c59 non-ascii char removed from code. Caused display bugs 2019-01-13 17:29:55 +01:00
Tomasz Pietruszka b00410e0cb LM save with_opt fix 2019-01-13 17:28:39 +01:00
Tomasz Pietruszka a1e66c39d4 tokenzier->tokenizer typo 2019-01-13 17:27:29 +01:00
Tomasz Pietruszka 593c5661bf Added alpha and beta params for RNNTrainer 2019-01-13 17:26:16 +01:00
Tomasz Pietruszka e6407fe0c9 Added the param and model type for BwdLM 2019-01-10 00:47:22 +01:00
Piotr Czapla 0085c18ae0 Fix splitting by article for local languages and make it more memory efficient
The issue was that the code assumed that empty lines have space followed by a new  line, which isn't the case for datasets generated by our scripts.
2019-01-03 12:16:41 +01:00
Piotr Czapla 8a2fed41c5 Merge pull request #18 from n-waves/refactor
[WIP] Refactoring #17
2019-01-01 15:13:57 +01:00
Piotr Czapla f784d7bcd2 Make the article detection code work with our wikitext 2019-01-01 15:13:14 +01:00
Piotr Czapla 0945c699c7 Fixes #25 by adding article title in markdown format to wiki text 2019-01-01 15:07:02 +01:00
Piotr Czapla 4a858f3572 Simplfy and unify input data parsing 2019-01-01 14:43:12 +01:00
Piotr Czapla d269d53d7d Use fastai tokens instead of <xxx>
f'xx{token_name}' are kept as one token by Moses tokenizer, which is sometimes required if you want to have moses in tokenizers pipeline, and we use that for imdb.
2019-01-01 14:40:59 +01:00
Piotr Czapla 514a9e6b86 Fix BiLM training after update to newest fastai 2018-12-31 12:13:53 +01:00
Piotr Czapla f1b49a0d34 fix imports in learner (fastai adpatation) 2018-12-27 15:15:09 +01:00
Piotr Czapla 82d6a30b11 Fix SentencePiece implementation, to get 94.5% on imdb
- train on whole articles when tokenizer is sentencepice
- add moses tokenizer to get the same tokens on imdb as on wt103
- apply pre / post processing rules to moses tokenzier so that sentencepiece works on correctly preprocessed text (lowercased with html removed)
- add alpha implementation on xnli (no tests)
- remove vocab adaptation when swiching from wiki to imdb. As otherwise we get 50% of missing words and 93% accuracy on imdb
2018-12-27 15:14:41 +01:00
Piotr Czapla ff30fb5642 Fastai upgrade 2018-12-27 14:57:15 +01:00
Piotr Czapla dca502a398 Add docs how to train a classfier 2018-12-12 00:31:48 +01:00
Piotr Czapla f9394b9af1 Add callbacks to save history and best weights remove bs & drop_mult 2018-12-12 00:31:16 +01:00
Piotr Czapla fac3ce343b Fix BiLM implementation after upgrade of fastai 2018-12-12 00:29:23 +01:00
Piotr Czapla 1b9d04d7ef Expose use_test_for_validation as param to train 2018-12-09 22:45:21 +01:00
Piotr Czapla 34b1d18600 load wikipedia as articles for v & fv tok. 2018-12-09 22:42:17 +01:00
Piotr Czapla 2338618563 Make biclasifier head a hyperparameter. 2018-12-09 00:39:27 +01:00
Piotr Czapla 591393b96c Change AvgPooling to BiPooling as a default 2018-12-08 23:48:02 +01:00
Piotr Czapla 978bbcc923 Fix issue with finetuning language model (it wasn't freezed) 2018-12-08 23:10:31 +01:00
Piotr Czapla 0b7dc7ec6b Add two experiments showing how to train imdb classification to 94% accuracy 2018-12-07 16:10:32 +01:00
Piotr Czapla 50ab34eea9 Fix classification scripts and give a way to test accuracy on test set 2018-12-07 16:10:20 +01:00
Piotr Czapla 4ae2a158a2 Respect batch size in training lm model. 2018-12-06 23:29:43 +01:00
Piotr Czapla 6a899015ed Generate imdb unsp.csv 2018-12-05 16:27:41 +01:00
Piotr Czapla 2acbf15555 Tweak hyper training params of cls (drop_mul, bs, true_wd=True)
I've set the same hyperparams as in lesson3
2018-12-05 16:27:14 +01:00
Piotr Czapla 4454af167f Use more text during pretraining of imdb 2018-12-05 16:26:16 +01:00
Piotr Czapla 5524006d81 Respect max_vocab 2018-12-04 16:41:22 +01:00
Piotr Czapla 908c3d7e8a bug fix 2018-12-04 16:41:07 +01:00
Piotr Czapla f25deb3049 Use trn + tst for LM training 2018-12-04 01:43:01 +01:00
Piotr Czapla 9826f6881c Fix the way trn & val set is created in imdb 2018-12-04 01:35:59 +01:00
Piotr Czapla 039624870a Fix loading tokenized data set in train cls 2018-12-04 01:28:03 +01:00
Piotr Czapla a499bf9a20 Make cls train work with relative paths 2018-12-04 01:24:44 +01:00
Piotr Czapla 35a5dcb75a Fix issue when running cls training from command line 2018-12-04 00:59:36 +01:00
Piotr Czapla 82c955ce6a Add different tokenization algorithms to train_clas 2018-12-04 00:51:21 +01:00