mirror of
https://github.com/wassname/pytorch-lightning.git
synced 2026-08-21 11:20:03 +08:00
* add doctest to circleci * Revert "add doctest to circleci" This reverts commit c45b34ea911a81f87989f6c3a832b1e8d8c471c6. * Revert "Revert "add doctest to circleci"" This reverts commit 41fca97fdcfe1cf4f6bdb3bbba75d25fa3b11f70. * doctest docs rst files * Revert "doctest docs rst files" This reverts commit b4a2e83e3da5ed1909de500ec14b6b614527c07f. * doctest only rst * doctest debugging.rst * doctest apex * doctest callbacks * doctest early stopping * doctest for child modules * doctest experiment reporting * indentation * doctest fast training * doctest for hyperparams * doctests for lr_finder * doctests multi-gpu * more doctest * make doctest drone * fix label build error * update fast training * update invalid imports * fix problem with int device count * rebase stuff * wip * wip * wip * intro guide * add missing code block * circleci * logger import for doctest * test if doctest runs on drone * fix mnist download * also run install deps for building docs * install cmake * try sudo * hide output * try pip stuff * try to mock horovod * Tranfer -> Transfer * add torchvision to extras * revert pip stuff * mlflow file location * do not mock torch * torchvision * drone extra req. * try higher sphinx version * Revert "try higher sphinx version" This reverts commit 490ac28e46d6fd52352640dfdf0d765befa56988. * try coverage command * try coverage command * try undoc flag * newline * undo drone * report coverage * review Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com> * remove torchvision from extras * skip tests only if torchvision not available * fix testoutput torchvision Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com>
82 lines
2.5 KiB
ReStructuredText
82 lines
2.5 KiB
ReStructuredText
.. testsetup:: *
|
|
|
|
from pytorch_lightning.trainer.trainer import Trainer
|
|
|
|
Debugging
|
|
=========
|
|
The following are flags that make debugging much easier.
|
|
|
|
Fast dev run
|
|
------------
|
|
This flag runs a "unit test" by running 1 training batch and 1 validation batch.
|
|
The point is to detect any bugs in the training/validation loop without having to wait for
|
|
a full epoch to crash.
|
|
|
|
(See: :paramref:`~pytorch_lightning.trainer.trainer.Trainer.fast_dev_run`
|
|
argument of :class:`~pytorch_lightning.trainer.trainer.Trainer`)
|
|
|
|
.. testcode::
|
|
|
|
trainer = Trainer(fast_dev_run=True)
|
|
|
|
Inspect gradient norms
|
|
----------------------
|
|
Logs (to a logger), the norm of each weight matrix.
|
|
|
|
(See: :paramref:`~pytorch_lightning.trainer.trainer.Trainer.track_grad_norm`
|
|
argument of :class:`~pytorch_lightning.trainer.trainer.Trainer`)
|
|
|
|
.. testcode::
|
|
|
|
# the 2-norm
|
|
trainer = Trainer(track_grad_norm=2)
|
|
|
|
Log GPU usage
|
|
-------------
|
|
Logs (to a logger) the GPU usage for each GPU on the master machine.
|
|
|
|
(See: :paramref:`~pytorch_lightning.trainer.trainer.Trainer.log_gpu_memory`
|
|
argument of :class:`~pytorch_lightning.trainer.trainer.Trainer`)
|
|
|
|
.. testcode::
|
|
|
|
trainer = Trainer(log_gpu_memory=True)
|
|
|
|
Make model overfit on subset of data
|
|
------------------------------------
|
|
|
|
A good debugging technique is to take a tiny portion of your data (say 2 samples per class),
|
|
and try to get your model to overfit. If it can't, it's a sign it won't work with large datasets.
|
|
|
|
(See: :paramref:`~pytorch_lightning.trainer.trainer.Trainer.overfit_pct`
|
|
argument of :class:`~pytorch_lightning.trainer.trainer.Trainer`)
|
|
|
|
.. testcode::
|
|
|
|
trainer = Trainer(overfit_pct=0.01)
|
|
|
|
Print the parameter count by layer
|
|
----------------------------------
|
|
Whenever the .fit() function gets called, the Trainer will print the weights summary for the lightningModule.
|
|
To disable this behavior, turn off this flag:
|
|
|
|
(See: :paramref:`~pytorch_lightning.trainer.trainer.Trainer.weights_summary`
|
|
argument of :class:`~pytorch_lightning.trainer.trainer.Trainer`)
|
|
|
|
.. testcode::
|
|
|
|
trainer = Trainer(weights_summary=None)
|
|
|
|
|
|
Set the number of validation sanity steps
|
|
-----------------------------------------
|
|
Lightning runs a few steps of validation in the beginning of training.
|
|
This avoids crashing in the validation loop sometime deep into a lengthy training loop.
|
|
|
|
(See: :paramref:`~pytorch_lightning.trainer.trainer.Trainer.num_sanity_val_steps`
|
|
argument of :class:`~pytorch_lightning.trainer.trainer.Trainer`)
|
|
|
|
.. testcode::
|
|
|
|
# DEFAULT
|
|
trainer = Trainer(num_sanity_val_steps=5) |