[tune] Cluster Fault Tolerance (#3309)

This PR introduces cluster-level fault tolerance for Tune by checkpointing global state. This occurs with relatively high frequency and allows users to easily resume experiments when the cluster crashes.

Note that this PR may affect automated workflows due to auto-prompting, but this is resolvable.
This commit is contained in:
Richard Liaw authored and GitHub committed 2018-12-29 11:42:25 +08:00
1 parent 382b138fc7
commit aad3c50e2d
16 files changed
+806 -128

No files matched your search

+1
View File
@@ -394,6 +394,7 @@ def train_func(config, reporter): # add a reporter arg
time.sleep(0.1)
reporter(timesteps_total=i, mean_accuracy=i+97) # report metrics
os.environ["TUNE_RESUME_PROMPT_OFF"] = "True"
ray.init(redis_address="{}")
ray.tune.register_trainable("train_func", train_func)