Update README

This commit is contained in:
Shangtong Zhang
2018-04-24 23:33:05 -06:00
parent be4f2720fe
commit d0b4f496f3
+2 -2
View File
@@ -11,7 +11,7 @@ Implemented algorithms:
* (Continuous/Discrete) Synchronous Proximal Policy Optimization (PPO)
* Action Conditional Video Prediction
Asynchronous algorithms below are removed in this repo but can be found in [v0.1](https://github.com/ShangtongZhang/DeepRL/releases/tag/v0.1)
Asynchronous algorithms below are removed in current version but can be found in [v0.1](https://github.com/ShangtongZhang/DeepRL/releases/tag/v0.1).
* Async Advantage Actor Critic (A3C)
* Async One-Step Q-Learning
* Async One-Step Sarsa
@@ -20,7 +20,7 @@ Asynchronous algorithms below are removed in this repo but can be found in [v0.1
* Distributed Deep Deterministic Policy Gradient (Distributed DDPG, aka D3PG)
* Parallelized Proximal Policy Optimization (P3O, similar to DPPO)
Support for Pytorch v0.3.x can be found in [v0.2](https://github.com/ShangtongZhang/DeepRL/releases/tag/v0.2). Note all the figures are generated via this version. After the upgrade to PyTorch v0.4.0, I have only tested the classical control tasks.
Support for PyTorch v0.3.x can be found in [v0.2](https://github.com/ShangtongZhang/DeepRL/releases/tag/v0.2). Note all the figures are generated via this version. After the upgrade to PyTorch v0.4.0, I have only tested the classical control tasks.
# Curves
> Curves for CartPole are trivial so I didn't place it here. And there isn't any fixed random seed. The curves are generated in the same manner as OpenAI baselines (one run and smoothed by recent 100 episodes)