Ashwin Bharambe
0e71705a0a
[checkpoint logic] Fix bug which doesn't account for NoneType for model.hparams ( #1817 )
...
The intention of the code is to output a warning message when `hparams`
is null or not set. Instead the code now fatals when
`model.hparams = None`. Prevent that.
2020-05-13 17:14:11 -04:00
William Falcon
12138ced7c
Update __init__.py
2020-05-13 14:42:50 -04:00
663b90035c
Bugfix: accumulation and suggestion for learning rate finder ( #1801 )
...
* fix suggestion being too naive
* fix accumulation error and added new tests
* fix styling
* update CHANGELOG.md
* update based on review
* fix tests
* Apply suggestions from code review
* Apply suggestions from code review
* Apply suggestions from code review
* Apply suggestions from code review
Co-authored-by: Nicki Skafte <nugginea@gmail.com >
Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com >
2020-05-13 14:40:44 -04:00
Ashwin Bharambe
aefc5314bc
[ddp] Support multi-node distributed execution under torchelastic ( #1811 )
...
The changes are quite local and limited in nature -- viz., checking for
some indicator environment variables. We check for (SLURM_LOCALID,
NODE_RANK, GROUP_RANK) in order. If multiple are found set, a warning is
logged.
This patch also fixes a minor bug with comparing the `WORLD_SIZE`
environment variable. This can be a string type.
2020-05-13 14:06:59 -04:00
b1d9656470
Update README.md ( #1798 )
...
* Update README.md
* Update README.md
committed suggestion
Co-authored-by: William Falcon <waf2107@columbia.edu >
* Update README.md
Co-authored-by: Justus Schock <12886177+justusschock@users.noreply.github.com >
* Update README.md
Co-authored-by: Justus Schock <12886177+justusschock@users.noreply.github.com >
Co-authored-by: William Falcon <waf2107@columbia.edu >
Co-authored-by: Justus Schock <12886177+justusschock@users.noreply.github.com >
2020-05-13 12:22:12 -04:00
22d7d03118
Replace meta_tags.csv with hparams.yaml ( #1271 )
...
* Add support for hierarchical dict
* Support nested Namespace
* Add docstring
* Migrate hparam flattening to each logger
* Modify URLs in CHANGELOG
* typo
* Simplify the conditional branch about Namespace
Co-Authored-By: Jirka Borovec <Borda@users.noreply.github.com >
* Update CHANGELOG.md
Co-Authored-By: Jirka Borovec <Borda@users.noreply.github.com >
* added examples section to docstring
* renamed _dict -> input_dict
* mata_tags.csv -> hparams.yaml
* code style fixes
* add pyyaml
* remove unused import
* create the member NAME_HPARAMS_FILE
* improve tests
* Update tensorboard.py
* pass the local test w/o relavents of Horovod
* formatting
* update dependencies
* fix dependencies
* Apply suggestions from code review
* add savings
* warn
* docstrings
* tests
* Apply suggestions from code review
* saving
* Apply suggestions from code review
* use default
* remove logging
* typo fixes
* update docs
* update CHANGELOG
* clean imports
* add blank lines
* Update pytorch_lightning/core/lightning.py
Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com >
* Update pytorch_lightning/core/lightning.py
Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com >
* back to namespace
* add docs
* test fix
* update dependencies
* add space
Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com >
Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com >
2020-05-13 15:05:15 +02:00
35fe2efe27
added override for hparams in load_from_ckpt ( #1797 )
...
* added override for hparams in load_from_ckpt
* override hparams
* override hparams
* Apply suggestions from code review
Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com >
* update doctest
* typo
* chlog
Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com >
Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com >
Co-authored-by: Jirka <jirka.borovec@seznam.cz >
2020-05-13 10:27:22 +02:00
Jirka Borovec and Justus Schock
10ce1c0256
device property ( #1791 )
...
* device property
* add/copy properties
* inherit
* rename
* Apply suggestions from code review
Co-authored-by: Justus Schock <12886177+justusschock@users.noreply.github.com >
* dtype
* prop
* pt api
Co-authored-by: Justus Schock <12886177+justusschock@users.noreply.github.com >
2020-05-12 23:18:39 -04:00
Adrian Wälchli
8978794730
add missing flag ( #1805 )
2020-05-12 17:06:38 -04:00
William Falcon
5be1bc48e9
Update README.md
2020-05-12 11:43:53 -04:00
William Falcon
d70d86985e
Update README.md
2020-05-12 11:43:01 -04:00
William Falcon
98cb7c2ce2
Update README.md
2020-05-12 08:59:23 -04:00
William Falcon
087bb34c68
Update README.md
2020-05-12 08:56:32 -04:00
Oliver Neumann
9059d21042
Missing profiler attribute in add_argparse_args() ArgumentParser ( #1794 )
...
* Fixed typing annotation by adding boolean type. After that Profiler flag will be added to argparse.
* Updated CHANGELOG.md
* Updated git_init_arguments_and_types() to pass doctests.
* Added doctest example to add_argparse_parser()
2020-05-12 08:53:26 -04:00
William Falcon
c52382f547
Update README.md
2020-05-12 08:52:43 -04:00
William Falcon
8584df54e9
Update README.md
2020-05-12 08:52:11 -04:00
William Falcon
a5c19ea784
Update README.md
2020-05-12 08:49:29 -04:00
William Falcon
423b82ea6c
Update README.md
2020-05-12 08:46:55 -04:00
William Falcon
39584d08ad
Update README.md
2020-05-12 08:46:22 -04:00
619f984c36
Option to provide seed to random generators to ensure reproducibility ( #1572 )
...
* Option to provide seed to random generators to ensure reproducibility
I added small function in utilities which imports torch, numpy, python
random and sets seed for all of the libraries to ensure reproducibility
of results.
* Apply recommendations from core contributors on seeding
1. Moved the seeding code to another file
2. Make deterministic as a parameter for trainer class
3. Add assertions for seeding numpy
4. Added warnings
5. torch.manual_seed should be enough for seeding torch
* Revert "Apply recommendations from core contributors on seeding"
This reverts commit a213c8e6882eec8a9e7408b9418926d2db7c5461.
* Revert "Revert "Apply recommendations from core contributors on seeding""
This reverts commit 59b2da53c62878de7aab0aa3feb3115e105eea06.
* Change in test, for correct seeding
* Allow seed equal to 0
* Allow seed to be uint32.max
* Added deterministic to benchmarks
* Cuda manual seed as in benchmark seeding
* Seeding should be done before model initialization
* cuda manual_seed is not necessary
* Fixing seed test_cpu_lbfgs
On some seeds seems like lbfgs doesn't converge.
So I fixed the seed during testing.
* rebasing issue with old reproducibility.py
* Improved documentation and ability to seed before initializing Train
class
* Change in docs
* Removed seed from trainer, update for documentation
* Typo in the docs
* Added seed_everything to _all_
* Fixing old changes
* Model initialization should be earlier then Trainer
* Update pytorch_lightning/trainer/__init__.py
From Example to testcode
Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com >
* Fixing according to the contributors suggestions
* Moving horovod deterministic to Trainer class
* deterministic flag affects horovod docs update
* Improved static typing
* Added deterministic to test runners of horovod
It is failing on some versions, not very predictable
* static seeds for horovod tests
* Change for reset_seed function in tests
* Seeding horovod using reset_seed from tutils
* Update pytorch_lightning/trainer/__init__.py
* chlog
* Update trainer.py
* change "testcode" to "Example" in trainer init documentation
* Update pytorch_lightning/trainer/seed.py, first line in comment
Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com >
Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com >
Co-authored-by: Jirka <jirka.borovec@seznam.cz >
Co-authored-by: William Falcon <waf2107@columbia.edu >
2020-05-12 07:53:20 -04:00
William Falcon
7af4505519
Update README.md
2020-05-12 07:50:23 -04:00
William Falcon
a4fc4ffa6e
Update README.md
2020-05-12 07:49:17 -04:00
William Falcon
6216501455
Update README.md
2020-05-12 07:47:57 -04:00
William Falcon
6517d1cf5c
Add files via upload
2020-05-12 07:46:55 -04:00
Justus Schock
5f292390fd
Bug fix hparam logging with metrics ( #1647 )
...
* add metric logging
* Use pytorch built-in method
* Update tensorboard.py
* Update tensorboard.py
2020-05-12 07:25:12 -04:00
Jirka Borovec
35ac30e688
Fix build Docker releases ( #1783 )
...
* gh act - if
* gh act - if
* gh act - steps
* gh act - steps
* gh act - steps
* name
* name
* reorder
* docker
* timeout
* repo
* show
* show
* ver
* rc
* tag
2020-05-12 06:54:59 -04:00
William Falcon and Jirka
10b16dbfab
made ddp the default if no backend specified with multiple GPUs ( #1789 )
...
* made ddp the default if no backend specified with multiple GPUs
* fix
* spawn
Co-authored-by: Jirka <jirka.borovec@seznam.cz >
2020-05-12 06:54:23 -04:00
Travis Addair and Jirka
acab068c74
Join Horovod workers at the end of trainer.fit() to prevent race conditions following training ( #1786 )
...
* Join Horovod workers at the end of trainer.fit() to prevent race conditions following training
* flake8
* flake8
Co-authored-by: Jirka <jirka.borovec@seznam.cz >
2020-05-12 09:15:25 +00:00
William Falcon
7b60d49432
fixed native amp + ddp ( #1788 )
...
* fixed native amp + ddp
* fixed native amp + ddp
2020-05-12 00:25:06 -04:00
Jeremy Jordan
1df0d2dc97
set logger level for package ( #1718 )
...
* move logging config to trainer class init
* alternate logging config
2020-05-12 00:14:35 -04:00
William Falcon
4b30ef6480
Device ( #1790 )
...
* added self.device
* added docs
2020-05-12 00:09:48 -04:00
Kevin Chen
de1fdd8d3b
Removed test_dataloader call in check_testing_model_configuration ( #1670 )
...
* Removed test_dataloader call
* Check if test_dataloader is actually overriden
* Fixed method spelling
* Replaced lambdas
* Replaced None with super method
* Fixed testpass
2020-05-12 00:08:07 -04:00
William Falcon
5bb6b41b78
dataloaders with fast_dev_run ( #1787 )
...
* dataloaders with fast_dev_run
* dataloaders with fast_dev_run
* dataloaders with fast_dev_run
* fix
* pep 8
2020-05-11 23:32:44 -04:00
Jirka Borovec and Adrian Wälchli
9d2df24d6b
RC & Docs/changelog ( #1776 )
...
* missing
* RC
* tol
* Apply suggestions from code review
Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com >
* test
Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com >
2020-05-11 21:57:53 -04:00
Fabio Natanael Kepler
d120f97896
Fix saving native AMP scaler state ( #1777 )
...
Saving was introduced in #1561 .
2020-05-11 21:38:37 -04:00
William Falcon
eeb411144f
enable fast_dev_run without a validation loop ( #1779 )
...
* fix val dataloader
* Update evaluation_loop.py
2020-05-11 11:30:22 -04:00
William Falcon
88c086bbd2
Update progress.py
2020-05-11 09:48:15 -04:00
William Falcon
15c11fc848
fixes no val loader
2020-05-11 09:47:33 -04:00
William Falcon
d9bc8a978a
Update README.md
2020-05-10 17:09:09 -04:00
Rohit Gupta
d962ab5d89
Fix lr key name in case of param groups ( #1719 )
...
* Fix lr key name in case of param groups
* Add tests
* Update test and added configure_optimizers__param_groups
* Update CHANGELOG
2020-05-10 17:05:34 -04:00
Justus Schock and Jirka Borovec
7f64ad7a33
Fix Docker Pipeline ( #1765 )
...
* Update and rename docker_builds.yml to docker_nightly_builds.yml
* Update and rename docker_nightly_builds.yml to docker_builds.yml
* Update docker_builds.yml
* Update .github/workflows/docker_builds.yml
Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com >
2020-05-10 17:04:51 -04:00
Piotr Łusakowski
0cb6767465
Fix NeptuneLogger to work in ddp mode ( #1753 )
2020-05-10 13:19:18 -04:00
Alexander Kreuzer and Alexander Kreuzer
ee17c7c9c8
Fixed error message and test docstring ( #1698 )
...
training_dataloader -> train_dataloader
Co-authored-by: Alexander Kreuzer <alexander.kreuzer@sap.com >
2020-05-10 13:16:16 -04:00
Anthony Bisulco
76af84718a
Group argument wandb ( #1760 )
...
* group argument wandb
* formatting fix
2020-05-10 13:15:51 -04:00
Jirka Borovec
134eb61e1a
Tests: refactor cleanup ( #1744 )
...
* wip
* cleaning
* optim imports
* -
* default hparams
* fix restore
* fix imports
2020-05-10 13:15:28 -04:00
4970927ec8
Feature: auto scale batch size ( #1638 )
...
* auto batch finder
* fix styling
* add description
* add different modes
* fix copy paste error
* better organised code
* fix styling
* add tests
* fix
* fix
* add some documentation
* added CHANGELOG.md
* some documentation
* update based on review
* Update trainer.py
* Update docs/source/training_tricks.rst
Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com >
* Update tests/trainer/test_trainer_tricks.py
Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com >
* Update tests/trainer/test_trainer_tricks.py
Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com >
* Apply suggestions from code review
Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com >
* use EvalModelTemplate
* param tests
* rename
* wrap params
* rename function
* rename
* rename param
* fix
* abs
* rename
* refactor code
* add docs
* try
* arg
* loop
* exept
* loop
* drop bool
* docs
* docs
* added check and test for passing dataloader to fit
* styling fix
* update based on review
Co-authored-by: Nicki Skafte <nugginea@gmail.com >
Co-authored-by: William Falcon <waf2107@columbia.edu >
Co-authored-by: Adrian Wälchli <aedu.waelchli@gmail.com >
Co-authored-by: Jirka Borovec <Borda@users.noreply.github.com >
Co-authored-by: Jirka <jirka.borovec@seznam.cz >
2020-05-09 08:28:36 -04:00
Adrian Wälchli and Jirka
25bbd059df
Also update progress_bar in training_epoch_end ( #1724 )
...
* update prog. bar metrics on train epoch end
* changelog
* wip test
* more thorough testing
* comments
* update docs
* move test
Co-authored-by: Jirka <jirka.borovec@seznam.cz >
2020-05-08 23:31:56 -04:00
Yuri Brovman and ybrovman
3a642601e8
added warning for None dataloader ( #1745 )
...
* added warning for None dataloader
* fixed variable style
* updated warning message
* remove unused import
Co-authored-by: ybrovman <ybrovman@ebay.com >
2020-05-07 09:26:41 -04:00
Shunta Komatsu
f656882942
Fix typo ( #1750 )
2020-05-07 09:25:54 -04:00
Pavel Grunt
b9364f96b1
lr_finder: Fix typo in docstring ( #1746 )
2020-05-06 12:39:22 -04:00