From ed31417b265aa96d1b089d960cee7dd2f03bef1f Mon Sep 17 00:00:00 2001 From: William Falcon Date: Thu, 27 Jun 2019 11:59:27 -0400 Subject: [PATCH] renamed options --- docs/Trainer/Training Loop.md | 51 ++++++++++++++++++++++++++++++++++- docs/Trainer/index.md | 26 +++++++++--------- 2 files changed, 62 insertions(+), 15 deletions(-) diff --git a/docs/Trainer/Training Loop.md b/docs/Trainer/Training Loop.md index 471e6055..946b8ac9 100644 --- a/docs/Trainer/Training Loop.md +++ b/docs/Trainer/Training Loop.md @@ -5,10 +5,21 @@ The asdf Accumulated gradients runs K small batches of size N before doing a backwards pass. The effect is a large effective batch size of size KxN. ``` {.python} -# default 1 (ie: no accumulated grads) +# DEFAULT (ie: no accumulated grads) trainer = Trainer(accumulate_grad_batches=1) ``` +--- +#### Anneal Learning rate +Cut the learning rate by 10 at every epoch listed in this list. +``` {.python} +# DEFAULT (don't anneal) +trainer = Trainer(lr_scheduler_milestones=None) + +# cut LR by 10 at 100, 200, and 300 epochs +trainer = Trainer(lr_scheduler_milestones=[100, 200, 300]) +``` + --- #### Check GPU usage Lightning automatically logs gpu usage to the test tube logs. It'll only do it at the metric logging interval, so it doesn't slow down training. @@ -17,6 +28,7 @@ Lightning automatically logs gpu usage to the test tube logs. It'll only do it a #### Check which gradients are nan This option prints a list of tensors with nan gradients. ``` {.python} +# DEFAULT trainer = Trainer(print_nan_grads=False) ``` @@ -24,12 +36,14 @@ trainer = Trainer(print_nan_grads=False) #### Check validation every n epochs If you have a small dataset you might want to check validation every n epochs ``` {.python} +# DEFAULT trainer = Trainer(check_val_every_n_epoch=1) ``` --- #### Display metrics in progress bar ``` {.python} +# DEFAULT trainer = Trainer(progress_bar=True) ``` @@ -41,5 +55,40 @@ By default lightning prints a list of parameters *and submodules* when it starts #### Force training for min or max epochs It can be useful to force training for a minimum number of epochs or limit to a max number ``` {.python} +# DEFAULT trainer = Trainer(min_nb_epochs=1, max_nb_epochs=1000) ``` + +--- +#### Inspect gradient norms +Looking at grad norms can help you figure out where training might be going wrong. +``` {.python} +# DEFAULT (-1 doesn't track norms) +trainer = Trainer(track_grad_norm=-1) + +# track the LP norm (P=2 here) +trainer = Trainer(track_grad_norm=2) +``` + + +--- +#### Make model overfit on subset of data +A useful debugging trick is to make your model overfit a tiny fraction of the data. +``` {.python} +# DEFAULT don't overfit (ie: normal training) +trainer = Trainer(overfit_pct=0.0) + +# overfit on 1% of data +trainer = Trainer(overfit_pct=0.01) +``` + +--- +#### Set how much of the training set to check +If you don't want to check 100% of the validation set (for debugging or if it's huge), set this flag +``` {.python} +# DEFAULT +trainer = Trainer(train_percent_check=1.0) + +# check 10% only +trainer = Trainer(train_percent_check=0.1) +``` diff --git a/docs/Trainer/index.md b/docs/Trainer/index.md index ac04f17c..a8c054bc 100644 --- a/docs/Trainer/index.md +++ b/docs/Trainer/index.md @@ -18,20 +18,18 @@ But of course the fun is in all the advanced things it can do: **Training loop** -- Accumulate gradients -- Check GPU usage -- Check which gradients are nan -- Check validation every n epochs -- Display metrics in progress bar -- Display the parameter count by layer -- Force training for min or max epochs -- Inspect gradient norms -- Learning rate annealing -- Make model overfit on subset of data -- Multiple optimizers (like GANs) -- Set how much of the training set to check (1-100%) -- Show progress bar -- training_step function +- [Accumulate gradients](Training%20Loop/#accumulated-gradients) +- [Anneal Learning rate](Training%20Loop/#anneal-learning-rate) +- [Check GPU usage](Training%20Loop/#Check-gpu-usage) +- [Check which gradients are nan](Training%20Loop/#check-which-gradients-are-nan) +- [Check validation every n epochs](Training%20Loop/#check-validation-every-n-epochs) +- [Display metrics in progress bar](Training%20Loop/#display-metrics-in-progress-bar) +- [Display the parameter count by layer](Training%20Loop/#display-the-parameter-count-by-layer) +- [Force training for min or max epochs](Training%20Loop/#force-training-for-min-or-max-epochs) +- [Inspect gradient norms](Training%20Loop/#inspect-gradient-norms) +- [Make model overfit on subset of data](Training%20Loop/#make-model-overfit-on-subset-of-data) +- [Use multiple optimizers (like GANs)](../Pytorch-lightning/LightningModule/#configure_optimizers) +- [Set how much of the training set to check (1-100%)](Training%20Loop/#set-how-much-of-the-training-set-to-check) **Validation loop**