2018-09-12 08:34:28 +05:30
2018-08-31 17:23:01 +05:30
2018-09-12 08:34:28 +05:30
2018-09-04 00:55:19 +05:30
2018-08-31 17:25:08 +05:30
2018-08-31 17:25:08 +05:30
2018-09-05 16:02:23 +05:30
2018-08-31 17:25:08 +05:30
2018-09-04 16:36:40 +05:30
2018-08-31 17:25:08 +05:30

Description


Reimplementation of Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Contributions are welcome. If you find any mistake (very likely) or know how to make it more stable, don't hesitate to send a pull request.

Requirements


Run


Use the default hyperparameters.

For SAC (Gaussian Policy):

python main.py --algo SAC --env-name HalfCheetah-v2

For SAC (Gaussian Mixture Policy):

python main.py --algo SAC(GMM) --env-name HalfCheetah-v2 --k 4

TODO


  • Gaussian Policy
  • Reparameterization
  • Gaussian Mixture Model
  • Use 2 Q-functions
  • Evaluate the trained Policy
  • Deterministic Policy (hard target update)
S
Description
PyTorch implementation of soft actor critic
Readme MIT
1.4 MiB
Languages
Python 96%
Makefile 4%