mirror of
https://github.com/wassname/DeepRL.git
synced 2026-09-09 11:13:47 +08:00
Update README
This commit is contained in:
@@ -94,6 +94,10 @@ Prediction is sampled after 110K iterations and I only implemented one-step trai
|
||||

|
||||
A deterministic test episode is triggered every 10 episodes. 2.5M steps and 14 hours in total.
|
||||
|
||||
## A2C & N-Step DQN
|
||||

|
||||
Online training progression of a single run. Entropy regularization is used for A2C, resulting in the variance in the curve.
|
||||
|
||||
# Dependency
|
||||
> Tested in macOS 10.12 and CentO/S 6.8
|
||||
* Open AI gym
|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 223 KiB |
Reference in New Issue
Block a user