mirror of
https://github.com/wassname/ray.git
synced 2026-08-17 11:25:34 +08:00
Option to retry failed actor tasks (#8330)
* Python * Consolidate state in the direct actor transport, set the caller starts at * todo * Remove unused * Update and unit tests * Doc * Remove unused * doc * Remove debug * Update src/ray/core_worker/transport/direct_actor_transport.h Co-authored-by: Eric Liang <ekhliang@gmail.com> * Update src/ray/core_worker/transport/direct_actor_transport.cc Co-authored-by: Eric Liang <ekhliang@gmail.com> * lint and fix build * Update * Fix build * Fix tests * Unit test for max_task_retries=0 * Fix java? * Fix bad test * Cross language fix * fix java Co-authored-by: Eric Liang <ekhliang@gmail.com>
This commit is contained in:
co-authored by
Eric Liang
parent
41d8c2bd0a
commit
bd169749e0
@@ -78,12 +78,86 @@ You can experiment with this behavior by running the following code.
|
||||
actor = Actor.remote()
|
||||
|
||||
# The actor will be restarted up to 5 times. After that, methods will
|
||||
# raise exceptions. The actor is restarted by rerunning its
|
||||
# constructor. Methods that were executing when the actor died will also
|
||||
# raise exceptions.
|
||||
# always raise a `RayActorError` exception. The actor is restarted by
|
||||
# rerunning its constructor. Methods that were sent or executing when the
|
||||
# actor died will also raise a `RayActorError` exception.
|
||||
for _ in range(100):
|
||||
try:
|
||||
counter = ray.get(actor.increment_and_possibly_fail.remote())
|
||||
print(counter)
|
||||
except ray.exceptions.RayActorError:
|
||||
print('FAILURE')
|
||||
|
||||
By default, actor tasks execute with at-most-once semantics
|
||||
(``max_task_retries=0`` in the ``@ray.remote`` decorator). This means that if an
|
||||
actor task is submitted to an actor that is unreachable, Ray will report the
|
||||
error with ``RayActorError``, a Python-level exception that is thrown when
|
||||
``ray.get`` is called on the future returned by the task. Note that this
|
||||
exception may be thrown even though the task did indeed execute successfully.
|
||||
For example, this can happen if the actor dies immediately after executing the
|
||||
task.
|
||||
|
||||
Ray also offers at-least-once execution semantics for actor tasks
|
||||
(``max_task_retries=-1`` or ``max_task_retries > 0``). This means that if an
|
||||
actor task is submitted to an actor that is unreachable, the system will
|
||||
automatically retry the task until it receives a reply from the actor. With
|
||||
this option, the system will only throw a ``RayActorError`` to the application
|
||||
if one of the following occurs: (1) the actor’s ``max_restarts`` limit has been
|
||||
exceeded and the actor cannot be restarted anymore, or (2) the
|
||||
``max_task_retries`` limit has been exceeded for this particular task. The
|
||||
limit can be set to infinity with ``max_task_retries = -1``.
|
||||
|
||||
You can experiment with this behavior by running the following code.
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
import os
|
||||
import ray
|
||||
|
||||
ray.init(ignore_reinit_error=True)
|
||||
|
||||
@ray.remote(max_restarts=5, max_task_retries=-1)
|
||||
class Actor:
|
||||
def __init__(self):
|
||||
self.counter = 0
|
||||
|
||||
def increment_and_possibly_fail(self):
|
||||
# Exit after every 10 tasks.
|
||||
if self.counter == 10:
|
||||
os._exit(0)
|
||||
self.counter += 1
|
||||
return self.counter
|
||||
|
||||
actor = Actor.remote()
|
||||
|
||||
# The actor will be reconstructed up to 5 times. The actor is
|
||||
# reconstructed by rerunning its constructor. Methods that were
|
||||
# executing when the actor died will be retried and will not
|
||||
# raise a `RayActorError`. Retried methods may execute twice, once
|
||||
# on the failed actor and a second time on the restarted actor.
|
||||
for _ in range(50):
|
||||
counter = ray.get(actor.increment_and_possibly_fail.remote())
|
||||
print(counter) # Prints the sequence 1-10 5 times.
|
||||
|
||||
# After the actor has been restarted 5 times, all subsequent methods will
|
||||
# raise a `RayActorError`.
|
||||
for _ in range(10):
|
||||
try:
|
||||
counter = ray.get(actor.increment_and_possibly_fail.remote())
|
||||
print(counter) # Unreachable.
|
||||
except ray.exceptions.RayActorError:
|
||||
print('FAILURE') # Prints 10 times.
|
||||
|
||||
For at-least-once actors, the system will still guarantee execution ordering
|
||||
according to the initial submission order. For example, any tasks submitted
|
||||
after a failed actor task will not execute on the actor until the failed actor
|
||||
task has been successfully retried. The system also will not attempt to
|
||||
re-execute any tasks that executed successfully before the failure.
|
||||
|
||||
At-least-once execution is best suited for read-only actors or actors with
|
||||
ephemeral state that does not need to be rebuilt after a failure. For actors
|
||||
that have critical state, it is best to take periodic checkpoints and either
|
||||
manually restart the actor or automatically restart the actor with at-most-once
|
||||
semantics. If the actor’s exact state at the time of failure is needed, the
|
||||
application is responsible for resubmitting all tasks since the last
|
||||
checkpoint.
|
||||
|
||||
Reference in New Issue
Block a user