mirror of
https://github.com/wassname/multifit.git
synced 2026-08-26 11:22:17 +08:00
So that we can instantiate tokenization before we know what dataset we want to use it on. Previously it was tidly copuled.
561 B
561 B
Retraining a multifit model from wikipedia
The whole training process from wikipedia to mldoc can be run as follows:
python -m ulmfit new multifit_fp16 \
pretrain-lm train- data/wiki/de-100 - \
finetune-lm train- data/mldoc/de-1 - \
classifier train- data/mldoc/de-1
You can evaulate any model with the following command:
python -m ulmfit load data/mldoc/de-1/models/fsp15k/multfit_fp16 classifier validate data/mldoc/de-1
python -m ulmfit new multifit_fp16_nl3 pretrain-lm train- data/wiki/wikitext-103