{"id":20936763,"url":"https://github.com/hoverslam/rl-proximal-policy-optimization","last_synced_at":"2026-04-27T13:32:32.163Z","repository":{"id":249614406,"uuid":"827860814","full_name":"hoverslam/rl-proximal-policy-optimization","owner":"hoverslam","description":"Proximal Policy Optimization (PPO) implemented in PyTorch","archived":false,"fork":false,"pushed_at":"2024-07-24T00:46:49.000Z","size":64903,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-01-19T20:17:11.625Z","etag":null,"topics":["ppo","procgen","proximal-policy-optimization","pytorch"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/hoverslam.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-07-12T14:34:30.000Z","updated_at":"2024-07-24T00:46:52.000Z","dependencies_parsed_at":"2024-07-22T08:40:35.164Z","dependency_job_id":null,"html_url":"https://github.com/hoverslam/rl-proximal-policy-optimization","commit_stats":null,"previous_names":["hoverslam/rl-proximal-policy-optimization"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hoverslam%2Frl-proximal-policy-optimization","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hoverslam%2Frl-proximal-policy-optimization/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hoverslam%2Frl-proximal-policy-optimization/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hoverslam%2Frl-proximal-policy-optimization/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/hoverslam","download_url":"https://codeload.github.com/hoverslam/rl-proximal-policy-optimization/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":243330325,"owners_count":20274039,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ppo","procgen","proximal-policy-optimization","pytorch"],"created_at":"2024-11-18T22:25:34.297Z","updated_at":"2025-12-28T13:24:58.950Z","avatar_url":"https://github.com/hoverslam.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Proximal Policy Optimization\n\nThis repository contains a from-scratch implementation of Proximal Policy Optimization (PPO) using PyTorch. The main objective is to provide a clear and concise implementation of PPO and validate its performance by replicating results from the paper:\n\n\u003e [Leveraging Procedural Generation to Benchmark Reinforcement Learning](https://arxiv.org/abs/1912.01588) by Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman (2020).\n\nAs in the paper the agents are trained and benchmarked against the [Procgen Benchmark](https://github.com/openai/procgen) which provides 16 procedurally-generated gym environments.\n\n\n## What is PPO?\nProximal Policy Optimization is a popular reinforcement learning algorithm. It falls under the category of policy gradient methods, which optimize the policy directly by computing gradients of expected rewards with respect to policy parameters. The algorithm is designed to address some of the limitations of standard policy gradient methods:\n\n- **Stability and Reliability**: Traditional policy gradient methods can be unstable due to large policy updates. PPO improves stability by using a clipped objective function, ensuring that policy updates do not deviate too much from the previous policy.\n\n- **Sample Efficiency**: PPO uses multiple epochs of minibatch updates, making better use of collected data and improving sample efficiency.\n\n- **Simplicity**: Unlike more complex algorithms like Trust Region Policy Optimization (TRPO), PPO is simpler to implement while retaining many of the same benefits. It achieves a good balance between performance and computational complexity.\n\nThe algorithm introduces a clipping mechanism to the policy objective, which prevents large policy updates and improves training stability. This is in contrast to traditional policy gradient methods that may take large and potentially harmful updates. The PPO objective function penalizes changes that move the new policy too far away from the old policy, ensuring smoother and more reliable updates. \n\n\n## Results\n\n\u003cdiv align=\"center\"\u003e\n\u003ctable\u003e\n\u003ctbody\u003e \n    \u003ctr\u003e\n        \u003ctd\u003e\u003cimg src=\"img/coinrun.png\" style=\"height: 300px; width:380px;\"/\u003e \u003c/td\u003e\n        \u003ctd\u003e\u003cimg src=\"img/coinrun.gif\" style=\"height: 300px; width:300px;\"/\u003e\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n        \u003ctd\u003e\u003cimg src=\"img/starpilot.png\" style=\"height: 300px; width:380px;\"/\u003e\u003c/td\u003e\n        \u003ctd\u003e\u003cimg src=\"img/starpilot.gif\" style=\"height: 300px; width:300px;\"/\u003e\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n        \u003ctd\u003e\u003cimg src=\"img/leaper.png\" style=\"height: 300px; width:380px;\"/\u003e\u003c/td\u003e\n        \u003ctd\u003e\u003cimg src=\"img/leaper.gif\" style=\"height: 300px; width:300px\"/\u003e\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n        \u003ctd\u003e\u003cimg src=\"img/bigfish.png\" style=\"height: 300px; width:380px;\"/\u003e\u003c/td\u003e\n        \u003ctd\u003e\u003cimg src=\"img/bigfish.gif\" style=\"height: 300px; width:300px\"/\u003e\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n        \u003ctd\u003e\u003cimg src=\"img/ninja.png\" style=\"height: 300px; width:380px;\"/\u003e\u003c/td\u003e\n        \u003ctd\u003e\u003cimg src=\"img/ninja.gif\" style=\"height: 300px; width:300px\"/\u003e\u003c/td\u003e\n    \u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\n---\n\n\u003cdiv align=\"center\"\u003e\n    \u003cimg src=\"./img/results_procgen_paper.png\" alt=\"Results from the Procgen paper\"/\u003e\n    \u003cp\u003e\u003cem\u003eResults from Leveraging Procedural Generation to Benchmark Reinforcement Learning by Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman (2020).\u003c/em\u003e\u003c/p\u003e\n\u003c/div\u003e\n\n\n## Installation\n\n1. Clone the repository:\n\n```\ngit clone https://github.com/hoverslam/rl-proximal-policy-optimization\n```\n\n2. Navigate to the directory:\n\n```\ncd rl-proximal-policy-optimization\n```\n\n3. Set up a virtual environment:\n\n```bash\n# Create a virtual environment\npython -3.10 -m venv .venv\n\n# Activate the virtual environment\n.venv\\Scripts\\activate\n```\n\n4. (Optional) Install PyTorch with CUDA support:\n\n```\npip install torch==2.3.1 --index-url https://download.pytorch.org/whl/cu118\n```\n\n5. Install the dependencies:\n\n```\npip install -r requirements.txt\n```\n\n\n## Usage\n\n### Training\n\n```powershell\npython -m train --env_name=\"starpilot\"\n```\n\nFor possible options use ```--help``` :\n```powershell\npython -m train --help\n```\n\n\u003e [!NOTE]  \n\u003e Results were obtained with the default settings and `num_evals=60`.\n\nTraining, including evaluation, consumed approximately 3 GPU hours on the following system:\n\n* Processor: 13th Gen Intel Core i5-13600KF, 3.50 GHz\n* GPU: GeForce RTX 3060\n* Memory: 32 GB\n* Operating System: Windows 10 Pro, 64-bit\n\n### Play\n\n```powershell\npython -m play --env_name=\"starpilot\"\n```\n\u003e [!WARNING]\n\u003e Closing the window will not terminate the script. To stop the script, manually press Ctrl+C.\n\nFor possible options use ```--help``` :\n```powershell\npython -m play --help\n```\n\n\n## File structure\n\n* :file_folder: `img/`: Contains images and GIFs.\n\n* :file_folder: `ppo/`\n\n    * :page_facing_up: `agent.py`: Contains the `PPOAgent` and `ImpalaCNN` classes.\n\n    * :page_facing_up: `trainer.py`: Contains the `PPOTrainer` class.\n\n    * :page_facing_up: `utils.py`: Contains some helper functions and classes.\n\n* :file_folder: `pretrained/`: Directory to store trained models.\n\n* :file_folder: `results/`: Directory to store evaluation results.\n\n* :page_facing_up: `train.py`: Script to train the agent.\n\n* :page_facing_up: `play.py`: Script to let a pretrained agent play a game.\n\n* :orange_book: `results.ipynb`: Interactive notebook for visualizing training results.\n\n\n## License\n\nThe code in this project is licensed under the [MIT License](LICENSE.txt).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhoverslam%2Frl-proximal-policy-optimization","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fhoverslam%2Frl-proximal-policy-optimization","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhoverslam%2Frl-proximal-policy-optimization/lists"}