Commit Graph
137 Commits
Author SHA1 Message Date
Ubuntu 38e422ff0e fix dropout 2019-05-11 22:00:34 +00:00
Ubuntu 19507f1d38 seed + new tokens for poleval2018 2019-05-11 19:51:27 +00:00
Cahya Wirawan 616bec18dc Added an option to set the minimal limit of tokens per article 2019-04-30 11:48:14 +02:00
Piotr Czapla 9f494b3ec6 Remove old bilm code that wasn't working 2019-04-19 10:32:35 +02:00
Piotr Czapla 8facc21cfa Add F1 metrics to evaluate 2019-04-14 19:13:43 +02:00
Piotr Czapla a948d7df91 Merge branch 'master' of https://github.com/n-waves/ulmfit-multilingual 2019-03-26 21:29:33 +01:00
Piotr Czapla 0a823be17b Maki ti possible to not load unusp 2019-03-26 21:29:23 +01:00
Piotr Czapla 24f1d741f0 Add generating of pseudo labels 2019-03-26 21:27:56 +01:00
NAUSICAA\Julian 227f3b5dff Merge branch 'master' of https://github.com/n-waves/ulmfit-multilingual into text_cols 2019-03-07 13:18:52 -03:00
Piotr Czapla 837925ff53 Imporved validate_cls & eval to pick the best model based on val accuracy 2019-03-03 13:29:49 +01:00
Piotr Czapla 7b2ac9e94b Add ability to use random-init=True 2019-03-01 17:41:07 +01:00
Piotr Czapla 11b2b14523 Fix -m ulmfit tar method 2019-03-01 17:39:58 +01:00
NAUSICAA\Julian fd022c6826 Merge branch 'master' of https://github.com/n-waves/ulmfit-multilingual into text_cols 2019-02-27 23:30:59 -03:00
Marcin f750586114 Remove redundant bptt param 2019-02-26 18:39:38 +01:00
Piotr Czapla c6e0373170 Make models use bptt parameter 2019-02-26 18:06:23 +01:00
NAUSICAA\Julian dd74a6d282 Remove previous multi field patch 2019-02-24 23:15:20 -03:00
NAUSICAA\Julian 06cd4d4d0b Merge branch 'master' of https://github.com/n-waves/ulmfit-multilingual into text_cols 2019-02-24 22:58:46 -03:00
Piotr Czapla 260faa703c Correct the label smoothing implementation 2019-02-22 16:55:35 +01:00
Piotr Czapla 015f04ec08 Training with noise & label smoothing 2019-02-22 12:02:18 +01:00
Piotr Czapla 99d6b22447 Correct noise generation training + convenience functions 2019-02-20 10:24:30 +01:00
NAUSICAA\Julian 13ae29a95d Set mark_fields in True 2019-02-18 22:20:24 -03:00
Piotr Czapla 7dc7aac327 Merge all columns in classification task into first column
This should fix CLS issues.
2019-02-18 21:50:04 +01:00
Piotr Czapla 9fbcf56df3 Add different learning schedules, with default to the old schedule
use --lr-sched=1cycle for better results
2019-02-18 21:49:16 +01:00
NAUSICAA\Julian 458c06f779 Fix df name 2019-02-18 17:05:11 -03:00
NAUSICAA\Julian dbb929e3ca Adding more text cols to use all CLS data 2019-02-18 16:54:00 -03:00
Piotr Czapla ce6cc607ae Add saving itos.pkl so that the LM can be used to finetuning 2019-02-17 23:18:42 +01:00
Piotr Czapla c3276da062 Fixing Imdb loading 2019-02-17 23:18:12 +01:00
Piotr Czapla 119417fb6e Disable early stopping as it was causing OOMs 2019-02-17 23:04:25 +01:00
Piotr Czapla 490c792278 Upgrade to the recent the todays version of Fastai 2019-02-17 23:03:54 +01:00
Piotr Czapla 8733487d55 Make the validate vs train decision based on the existance of cls_last.pth istead of a model directory 2019-02-15 01:17:53 +01:00
Piotr Czapla 5dced1e488 Remove bidir 2019-02-15 01:16:37 +01:00
Piotr Czapla 0e6534ad7b Expose num_lm_epochs in ulmfit eval 2019-02-15 01:11:39 +01:00
Piotr Czapla 0f084168c1 Merge branch 'master' of https://github.com/n-waves/ulmfit-multilingual 2019-02-14 22:35:29 +01:00
Piotr Czapla cd47b3b5dc Fix use_moses=True for mldoc so that it is identical to wiki with uses_moses=False
The issue was that Moses was executed after pre_rules when use_moses = True, But when data set was pre tokenized with Moses (use_moses=False) the pre_rules were executed  after.
So our wikipedia had the following processing:
- raw text
- Moses
- pre_rules
- split(' ') # fastai BaseTokenizer
- post_rules
- sentence piece

While mldoc had the following tokenziation
- raw text
- pre_rules
- Moses
- post_rules
- sentence piece

After fix I've retrained the classfiers (without finetuning) and I haven't notice huge changes in the performance. 4 languages received slight improvment 4 got a slight decrease in performance.
2019-02-14 22:35:20 +01:00
Piotr Czapla c28c0fde16 Make ulmfit eval more secure and give more flexibility in dataset_template
The dataset_template can use lang as additional token to construct globs patterns.
2019-02-14 22:28:25 +01:00
Marcin fdac9f7ccd Save only the best LM model 2019-02-14 14:01:33 +01:00
Marcin b609951561 Fix path of pretrained model 2019-02-14 00:00:56 +01:00
Piotr Czapla 1340f4235c Improve ulmfit eval to allow for zeroshot laser evaluation 2019-02-13 15:29:55 +01:00
Marcin 0e2211dfaf Merge branch 'master' into use-configs 2019-02-13 14:10:59 +01:00
Marcin 2eee051c67 Expose max length parameter 2019-02-13 13:40:05 +01:00
Piotr Czapla 3586a7dcf9 Merge branch 'pr/29' 2019-02-12 20:43:44 +01:00
Piotr Czapla e7271f2a29 Add ability to evalulate multiple models at once 2019-02-12 15:01:01 +01:00
Marcin 8497cc1e5c Move LM and classifier parameters to configs 2019-02-12 03:09:51 +01:00
Piotr Czapla a630242f97 Expose validate_cls in ulmfit module 2019-02-11 11:01:08 +01:00
Piotr Czapla c63fe258d2 Fix validataion and add option to add noise to training labels 2019-02-11 10:53:25 +01:00
Piotr Czapla 26736d95de Add code to test & train mldoc classifier 2019-02-10 09:52:58 +01:00
Tomasz Pietruszka 6fa30829c4 Merge branch 'master' into backwards-lm 2019-01-14 17:20:48 +01:00
Tomasz Pietruszka 5ba8ae4c59 non-ascii char removed from code. Caused display bugs 2019-01-13 17:29:55 +01:00
Tomasz Pietruszka b00410e0cb LM save with_opt fix 2019-01-13 17:28:39 +01:00
Tomasz Pietruszka a1e66c39d4 tokenzier->tokenizer typo 2019-01-13 17:27:29 +01:00