mirror of
https://github.com/wassname/multifit.git
synced 2026-08-21 11:18:10 +08:00
70c74a1cc5f9338fbd05232d6ffec00253df9240
So that we can instantiate tokenization before we know what dataset we want to use it on. Previously it was tidly copuled.
Retraining a multifit model from wikipedia
The whole training process from wikipedia to mldoc can be run as follows:
python -m ulmfit new multifit_fp16 \
pretrain-lm train- data/wiki/de-100 - \
finetune-lm train- data/mldoc/de-1 - \
classifier train- data/mldoc/de-1
You can evaulate any model with the following command:
python -m ulmfit load data/mldoc/de-1/models/fsp15k/multfit_fp16 classifier validate data/mldoc/de-1
python -m ulmfit new multifit_fp16_nl3 pretrain-lm train- data/wiki/wikitext-103
Description
The code to reproduce results from paper "MultiFiT: Efficient Multi-lingual Language Model Fine-tuning" https://arxiv.org/abs/1909.04761
1.5 MiB
Languages
Jupyter Notebook
67.3%
Python
31.1%
Shell
1.6%