mirror of
https://github.com/wassname/ray.git
synced 2026-08-04 13:14:14 +08:00
[tune] Node Fault Tolerance (#3238)
This PR introduces single-node fault tolerance for Tune. ## Previous behavior: - Actors will be restarted without checking if resources are available. This can lead to problems if we lose resources. ## New behavior: - RUNNING trials will be resumed on another node on a best effort basis (meaning they will run if resources available). - If the cluster is saturated, RUNNING trials on that failed node will become PENDING and queued. - During recovery, TrialSchedulers and SearchAlgorithms should receive notification of this (via `trial_runner.stop_trial`) so that they don’t wait/block for a trial that isn’t running. Remaining questions: - Should `last_result` be consistent during restore? Yes; but not for earlier trials (trials that are yet to be checkpointed). - Waiting for some PRs to merge first (#3239) Closes #2851.
This commit is contained in:
@@ -278,6 +278,7 @@ class RayTrialExecutor(TrialExecutor):
|
||||
def save(self, trial, storage=Checkpoint.DISK):
|
||||
"""Saves the trial's state to a checkpoint."""
|
||||
trial._checkpoint.storage = storage
|
||||
trial._checkpoint.last_result = trial.last_result
|
||||
if storage == Checkpoint.MEMORY:
|
||||
trial._checkpoint.value = trial.runner.save_to_object.remote()
|
||||
else:
|
||||
@@ -301,6 +302,8 @@ class RayTrialExecutor(TrialExecutor):
|
||||
ray.get(trial.runner.restore_from_object.remote(value))
|
||||
else:
|
||||
ray.get(trial.runner.restore.remote(value))
|
||||
trial.last_result = checkpoint.last_result
|
||||
|
||||
return True
|
||||
except Exception:
|
||||
logger.exception("Error restoring runner.")
|
||||
|
||||
Reference in New Issue
Block a user