Piotr Czapla
|
70c74a1cc5
|
Refactor tokenization
So that we can instantiate tokenization before we know what dataset we want to use it on. Previously it was tidly copuled.
|
2019-10-15 04:17:54 +02:00 |
|
Piotr Czapla
|
77a2780a6b
|
Clean up datasets
|
2019-10-15 00:57:40 +02:00 |
|
Piotr Czapla
|
ad50c73377
|
Make ULMFiT name consistant.
|
2019-10-15 00:12:31 +02:00 |
|
Piotr Czapla
|
43f5c1282d
|
Add sotabench scripts
|
2019-10-15 00:09:18 +02:00 |
|
Piotr Czapla
|
f846bf8a4b
|
Add OpenFIle preproc to SentencePiece preproc
|
2019-10-14 17:20:26 +02:00 |
|
Piotr Czapla
|
b362f2a7cb
|
from_pretrained
|
2019-10-14 17:18:28 +02:00 |
|
Piotr Czapla
|
5a3fbcecb5
|
Ability to use fastai databunch directly
|
2019-10-14 17:18:13 +02:00 |
|
Piotr Czapla
|
6336f5ebb8
|
Make it possible to create databunch out of dataframes
|
2019-10-14 13:14:07 +02:00 |
|
Piotr Czapla
|
91dbd9bc84
|
Clean up configurations
So we use the name of function automatically
|
2019-10-14 13:13:09 +02:00 |
|
Piotr Czapla
|
7847331751
|
Remove cupy from requirements.txt
qrnn isn't using it anymore
|
2019-10-14 13:11:52 +02:00 |
|
Piotr Czapla
|
1fe2bd46a6
|
finall clean ups
|
2019-10-07 14:50:03 +02:00 |
|
Piotr Czapla
|
3fe6c19af2
|
fix spelling eeror in finetune_lm
|
2019-09-08 08:43:15 +02:00 |
|
Piotr Czapla
|
cac4df27bd
|
Add paper version configuration and orignal ulmfit
|
2019-09-07 22:37:26 +02:00 |
|
Piotr Czapla
|
c390b4ed35
|
Add code to JA-multifit
|
2019-09-07 21:47:43 +02:00 |
|
Piotr Czapla
|
dd473aa98d
|
Add repoduction of Multifit result for JA using newest hyper params
|
2019-09-07 21:18:07 +02:00 |
|
Piotr Czapla
|
1ff9e01766
|
Remove unused code
|
2019-09-07 21:17:32 +02:00 |
|
Piotr Czapla
|
ee7c23a3be
|
Remove old multifit logs
|
2019-09-07 19:12:44 +02:00 |
|
Piotr Czapla
|
02ee52d0ef
|
Update README.md
|
2019-09-07 19:11:52 +02:00 |
|
Piotr Czapla
|
26e54a9c7d
|
Use the pervious datset_path from finetuning for classsificator training
|
2019-09-07 19:11:29 +02:00 |
|
Piotr Czapla
|
2cdf380adf
|
New bolerplate code
|
2019-09-07 16:19:09 +02:00 |
|
Piotr Czapla
|
8162cc5648
|
Clean up the old multfit training code
|
2019-09-07 15:22:01 +02:00 |
|
Piotr Czapla
|
2fe5c8a588
|
Update to fastai v1.0.57 - use new sentence piece implementaiton & sizes of hidden layers
|
2019-08-29 15:51:17 +02:00 |
|
Piotr Czapla
|
68d6b1c829
|
fix bug in the hack for poleval reddit
|
2019-06-10 20:31:04 +02:00 |
|
Piotr Czapla
|
a7ac4f5170
|
missing file
|
2019-06-10 20:14:57 +02:00 |
|
Piotr Czapla
|
7d59c4ecf2
|
Sentence piece has new option -fix that changes coverage to 99.95%
|
2019-06-10 20:14:11 +02:00 |
|
Piotr Czapla
|
ce0a29f385
|
Make it possible to set ftseed in poleval19_init
|
2019-06-06 17:45:13 +02:00 |
|
Piotr Czapla
|
854391a130
|
Merge branch 'reproduce-poleval' of https://github.com/n-waves/ulmfit-multilingual into reproduce-poleval
|
2019-05-19 00:17:27 +02:00 |
|
Piotr Czapla
|
38e131f802
|
Make the ulmfit compatible with old fastai (ulmfit_multilingual)
|
2019-05-19 00:17:03 +02:00 |
|
Marcin
|
7326df2c98
|
Train sentencepiece on unsup
|
2019-05-19 00:12:34 +02:00 |
|
Piotr Czapla
|
6f4db9de1a
|
Fix ensemble output (by removing index column)
|
2019-05-18 22:50:55 +02:00 |
|
Piotr Czapla
|
06c32ae2dc
|
Add ability to create folders when saving ensemble output
|
2019-05-18 21:59:12 +02:00 |
|
Piotr Czapla
|
f38f8c0670
|
fix issue in the databunch name generation
|
2019-05-18 21:48:32 +02:00 |
|
Piotr Czapla
|
8927e45fb2
|
Add ensemble commnad
|
2019-05-18 21:48:08 +02:00 |
|
Marcin
|
15c2d0663a
|
Save predictions for a proper dataset
|
2019-05-18 14:42:35 +02:00 |
|
Piotr Czapla
|
7a33ea5d1f
|
Add no-test + ability to start training from the classifcation data set without wiki
|
2019-05-18 13:10:29 +02:00 |
|
Piotr Czapla
|
75b934bb83
|
Merge branch 'reproduce-poleval' of https://github.com/n-waves/ulmfit-multilingual into reproduce-poleval
|
2019-05-15 18:00:45 +02:00 |
|
Piotr Czapla
|
7c9129670e
|
Poleval 19 ablation experiments
|
2019-05-15 18:00:40 +02:00 |
|
Marcin
|
30f5a7901c
|
Save predictions
|
2019-05-14 15:34:40 +02:00 |
|
Piotr Czapla
|
7955db37ad
|
Merge branch 'reproduce-poleval' of https://github.com/n-waves/ulmfit-multilingual into reproduce-poleval
|
2019-05-14 13:27:03 +02:00 |
|
Piotr Czapla
|
ad3c11edb8
|
Add ability to turnoff weighted CrossEntropy
|
2019-05-14 13:26:57 +02:00 |
|
Piotr Czapla
|
6dc4a5f102
|
Add Kappa and Mathew score calcualtion + ls command to main
|
2019-05-14 11:53:16 +02:00 |
|
Marcin
|
746c6ca185
|
Don't reset lmseed
|
2019-05-14 06:57:02 +02:00 |
|
Marcin
|
fe31778d70
|
Add option to remove duplicates
|
2019-05-14 02:30:25 +02:00 |
|
Marcin
|
84405b6e6c
|
Allow bwd in model name
|
2019-05-13 20:11:11 +02:00 |
|
Piotr Czapla
|
b065e36f1b
|
Fix double name exception
|
2019-05-13 20:02:48 +02:00 |
|
Piotr Czapla
|
9101a0313c
|
Let poleval19_full accept name
|
2019-05-13 19:53:45 +02:00 |
|
Piotr Czapla
|
29b600b8b5
|
Add ability to check different model save
|
2019-05-13 17:47:57 +02:00 |
|
Piotr Czapla
|
b55715e4f6
|
Fix ftseed used in pretrain_lm so ulmfit lm works again
|
2019-05-13 17:47:28 +02:00 |
|
Piotr Czapla
|
6ef75e4dd7
|
Fix issue in poleval_eval + add valid metrics
|
2019-05-13 09:17:59 +02:00 |
|
Piotr Czapla
|
1c3044f5df
|
fix issue with cls_best add lmseed to poleval
|
2019-05-13 03:44:35 +02:00 |
|