mirror of
https://github.com/wassname/pytorch-soft-actor-critic.git
synced 2026-07-27 11:26:22 +08:00
63a2be2b8edb2591c32cba6c6e2f62f839537b26
Description
Reimplementation of Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.
Contributions are welcome. If you find any mistake (very likely) or know how to make it more stable, don't hesitate to send a pull request.
Requirements
Run
Use the default hyperparameters.
For SAC (Gaussian Policy):
python main.py --algo SAC --env-name HalfCheetah-v2
For SAC (Gaussian Mixture Policy):
python main.py --algo SAC(GMM) --env-name HalfCheetah-v2 --k 4
TODO
- Gaussian Policy
- Reparameterization
- Gaussian Mixture Model
- Use 2 Q-functions
- Evaluate the trained Policy
- Deterministic Policy (hard target update)
Languages
Python
96%
Makefile
4%