Ubuntu
38e422ff0e
fix dropout
2019-05-11 22:00:34 +00:00
Ubuntu
19507f1d38
seed + new tokens for poleval2018
2019-05-11 19:51:27 +00:00
Cahya Wirawan
616bec18dc
Added an option to set the minimal limit of tokens per article
2019-04-30 11:48:14 +02:00
Piotr Czapla
9f494b3ec6
Remove old bilm code that wasn't working
2019-04-19 10:32:35 +02:00
Piotr Czapla
8facc21cfa
Add F1 metrics to evaluate
2019-04-14 19:13:43 +02:00
Piotr Czapla
a948d7df91
Merge branch 'master' of https://github.com/n-waves/ulmfit-multilingual
2019-03-26 21:29:33 +01:00
Piotr Czapla
0a823be17b
Maki ti possible to not load unusp
2019-03-26 21:29:23 +01:00
Piotr Czapla
24f1d741f0
Add generating of pseudo labels
2019-03-26 21:27:56 +01:00
NAUSICAA\Julian
227f3b5dff
Merge branch 'master' of https://github.com/n-waves/ulmfit-multilingual into text_cols
2019-03-07 13:18:52 -03:00
Piotr Czapla
837925ff53
Imporved validate_cls & eval to pick the best model based on val accuracy
2019-03-03 13:29:49 +01:00
Piotr Czapla
7b2ac9e94b
Add ability to use random-init=True
2019-03-01 17:41:07 +01:00
Piotr Czapla
11b2b14523
Fix -m ulmfit tar method
2019-03-01 17:39:58 +01:00
NAUSICAA\Julian
fd022c6826
Merge branch 'master' of https://github.com/n-waves/ulmfit-multilingual into text_cols
2019-02-27 23:30:59 -03:00
Marcin
f750586114
Remove redundant bptt param
2019-02-26 18:39:38 +01:00
Piotr Czapla
c6e0373170
Make models use bptt parameter
2019-02-26 18:06:23 +01:00
NAUSICAA\Julian
dd74a6d282
Remove previous multi field patch
2019-02-24 23:15:20 -03:00
NAUSICAA\Julian
06cd4d4d0b
Merge branch 'master' of https://github.com/n-waves/ulmfit-multilingual into text_cols
2019-02-24 22:58:46 -03:00
Piotr Czapla
260faa703c
Correct the label smoothing implementation
2019-02-22 16:55:35 +01:00
Piotr Czapla
015f04ec08
Training with noise & label smoothing
2019-02-22 12:02:18 +01:00
Piotr Czapla
99d6b22447
Correct noise generation training + convenience functions
2019-02-20 10:24:30 +01:00
NAUSICAA\Julian
13ae29a95d
Set mark_fields in True
2019-02-18 22:20:24 -03:00
Piotr Czapla
7dc7aac327
Merge all columns in classification task into first column
...
This should fix CLS issues.
2019-02-18 21:50:04 +01:00
Piotr Czapla
9fbcf56df3
Add different learning schedules, with default to the old schedule
...
use --lr-sched=1cycle for better results
2019-02-18 21:49:16 +01:00
NAUSICAA\Julian
458c06f779
Fix df name
2019-02-18 17:05:11 -03:00
NAUSICAA\Julian
dbb929e3ca
Adding more text cols to use all CLS data
2019-02-18 16:54:00 -03:00
Piotr Czapla
ce6cc607ae
Add saving itos.pkl so that the LM can be used to finetuning
2019-02-17 23:18:42 +01:00
Piotr Czapla
c3276da062
Fixing Imdb loading
2019-02-17 23:18:12 +01:00
Piotr Czapla
119417fb6e
Disable early stopping as it was causing OOMs
2019-02-17 23:04:25 +01:00
Piotr Czapla
490c792278
Upgrade to the recent the todays version of Fastai
2019-02-17 23:03:54 +01:00
Piotr Czapla
8733487d55
Make the validate vs train decision based on the existance of cls_last.pth istead of a model directory
2019-02-15 01:17:53 +01:00
Piotr Czapla
5dced1e488
Remove bidir
2019-02-15 01:16:37 +01:00
Piotr Czapla
0e6534ad7b
Expose num_lm_epochs in ulmfit eval
2019-02-15 01:11:39 +01:00
Piotr Czapla
0f084168c1
Merge branch 'master' of https://github.com/n-waves/ulmfit-multilingual
2019-02-14 22:35:29 +01:00
Piotr Czapla
cd47b3b5dc
Fix use_moses=True for mldoc so that it is identical to wiki with uses_moses=False
...
The issue was that Moses was executed after pre_rules when use_moses = True, But when data set was pre tokenized with Moses (use_moses=False) the pre_rules were executed after.
So our wikipedia had the following processing:
- raw text
- Moses
- pre_rules
- split(' ') # fastai BaseTokenizer
- post_rules
- sentence piece
While mldoc had the following tokenziation
- raw text
- pre_rules
- Moses
- post_rules
- sentence piece
After fix I've retrained the classfiers (without finetuning) and I haven't notice huge changes in the performance. 4 languages received slight improvment 4 got a slight decrease in performance.
2019-02-14 22:35:20 +01:00
Piotr Czapla
c28c0fde16
Make ulmfit eval more secure and give more flexibility in dataset_template
...
The dataset_template can use lang as additional token to construct globs patterns.
2019-02-14 22:28:25 +01:00
Marcin
fdac9f7ccd
Save only the best LM model
2019-02-14 14:01:33 +01:00
Marcin
b609951561
Fix path of pretrained model
2019-02-14 00:00:56 +01:00
Piotr Czapla
1340f4235c
Improve ulmfit eval to allow for zeroshot laser evaluation
2019-02-13 15:29:55 +01:00
Marcin
0e2211dfaf
Merge branch 'master' into use-configs
2019-02-13 14:10:59 +01:00
Marcin
2eee051c67
Expose max length parameter
2019-02-13 13:40:05 +01:00
Piotr Czapla
3586a7dcf9
Merge branch 'pr/29'
2019-02-12 20:43:44 +01:00
Piotr Czapla
e7271f2a29
Add ability to evalulate multiple models at once
2019-02-12 15:01:01 +01:00
Marcin
8497cc1e5c
Move LM and classifier parameters to configs
2019-02-12 03:09:51 +01:00
Piotr Czapla
a630242f97
Expose validate_cls in ulmfit module
2019-02-11 11:01:08 +01:00
Piotr Czapla
c63fe258d2
Fix validataion and add option to add noise to training labels
2019-02-11 10:53:25 +01:00
Piotr Czapla
26736d95de
Add code to test & train mldoc classifier
2019-02-10 09:52:58 +01:00
Tomasz Pietruszka
6fa30829c4
Merge branch 'master' into backwards-lm
2019-01-14 17:20:48 +01:00
Tomasz Pietruszka
5ba8ae4c59
non-ascii char removed from code. Caused display bugs
2019-01-13 17:29:55 +01:00
Tomasz Pietruszka
b00410e0cb
LM save with_opt fix
2019-01-13 17:28:39 +01:00
Tomasz Pietruszka
a1e66c39d4
tokenzier->tokenizer typo
2019-01-13 17:27:29 +01:00