Aayush 3e8f9f8b5f [WIP] Sub-word tokenization with sentencepiece
For resolving issue #3.

@eisenjulian let's use this branch. I've written some skeleton code (untested currently) to be used around your tokenizer. You can commit that into fastai_contrib for our purpose.
2018-11-15 23:01:49 +05:30
2018-11-08 20:33:10 +01:00
2018-11-14 21:07:15 +05:30

ulmfit-multilingual

Temporary repository used for collaboration on application of for multiple languages.

how to contribute

We have a fork of fastai to propose changes to fastai.text, with a branch for this project: https://github.com/n-waves/fastai/tree/ulmfit_multilingual

Let us know that you want to start collaboration on fastai forum thread: Multilingual ULMFIT and you will get access to both repositories.

Here is what I did:

$ cd fastai
$ git remote add n-waves https://github.com/n-waves/fastai.git
$ git remote -v 
n-waves	https://github.com/n-waves/fastai.git (fetch)
n-waves	https://github.com/n-waves/fastai.git (push)
origin	https://github.com/fastai/fastai.git (fetch)
origin	https://github.com/fastai/fastai.git (push)

$ git fetch n-waves
$ git checkout ulmfit_multilingual
Branch 'ulmfit_multilingual' set up to track remote branch 'ulmfit_multilingual' from 'n-waves'.
Switched to a new branch 'ulmfit_multilingual'

$ git push --set-upstream n-waves ulmfit_multilingual  # to automatically push ulmfit_multilingual branch to the n-waves repo

Repo structure

  • fastai_contrib -- anything that can be ported to fastai once we finish the project like: NLI models, Sentence Piece tok.,
  • ulmfit
    • data -- scripts to fetch and prepare data: wikipedia, xnli, classification data sets
    • lm -- scripts to train language models
    • bilm -- scripts to train biLM ELMo style, Bert style
    • class -- scripts to test classifiers on multiple languages
    • xnli -- scripts to test nli
S
Description
The code to reproduce results from paper "MultiFiT: Efficient Multi-lingual Language Model Fine-tuning" https://arxiv.org/abs/1909.04761
Readme MIT
1.5 MiB
0 Stars 1 Watchers 0 Forks
Languages
Jupyter Notebook 67.3%
Python 31.1%
Shell 1.6%