[rllib] Cleanup RNN support and make it work with multi-GPU optimizer (#2394)

Cleanup: TFPolicyGraph now automatically adds loss input entries for state_in_*, so that graph sub-classes don't need to worry about it.

Multi-GPU support:

Allow setting up model tower replicas with existing state input tensors

Truncate the per-device minibatch slices so that they are always a multiple of max_seq_len.
This commit is contained in:
Eric Liang
2018-07-17 06:55:46 +02:00
committed by GitHub
parent 1b645fcc8b
commit 0cecf6b79c
14 changed files with 163 additions and 138 deletions
+8 -1
View File
@@ -30,7 +30,14 @@ docker run --rm --shm-size=10G --memory=10G $DOCKER_SHA \
--env CartPole-v1 \
--run PPO \
--stop '{"training_iteration": 2}' \
--config '{"simple_optimizer": true, "model": {"use_lstm": true}}'
--config '{"simple_optimizer": false, "num_sgd_iter": 2, "model": {"use_lstm": true}}'
docker run --rm --shm-size=10G --memory=10G $DOCKER_SHA \
python /ray/python/ray/rllib/train.py \
--env CartPole-v1 \
--run PPO \
--stop '{"training_iteration": 2}' \
--config '{"simple_optimizer": true, "num_sgd_iter": 2, "model": {"use_lstm": true}}'
docker run --rm --shm-size=10G --memory=10G $DOCKER_SHA \
python /ray/python/ray/rllib/train.py \