{"id":13435294,"url":"https://github.com/oxwhirl/pymarl","last_synced_at":"2025-05-15T15:02:13.345Z","repository":{"id":37736119,"uuid":"154679202","full_name":"oxwhirl/pymarl","owner":"oxwhirl","description":"Python Multi-Agent Reinforcement Learning framework","archived":false,"fork":false,"pushed_at":"2022-12-08T02:58:39.000Z","size":283,"stargazers_count":2013,"open_issues_count":62,"forks_count":400,"subscribers_count":29,"default_branch":"master","last_synced_at":"2025-05-15T15:01:59.005Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/oxwhirl.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2018-10-25T13:48:43.000Z","updated_at":"2025-05-14T18:17:32.000Z","dependencies_parsed_at":"2023-01-24T07:15:47.355Z","dependency_job_id":null,"html_url":"https://github.com/oxwhirl/pymarl","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/oxwhirl%2Fpymarl","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/oxwhirl%2Fpymarl/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/oxwhirl%2Fpymarl/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/oxwhirl%2Fpymarl/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/oxwhirl","download_url":"https://codeload.github.com/oxwhirl/pymarl/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254364267,"owners_count":22058877,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-07-31T03:00:34.628Z","updated_at":"2025-05-15T15:02:13.285Z","avatar_url":"https://github.com/oxwhirl.png","language":"Python","funding_links":[],"categories":["Other reading material (blogs, websites, videos)","Reinforcement Learning (RL) and Deep Reinforcement Learning (DRL)","Python","Multi Agent in Real World"],"sub_categories":["Opponent Modelling","RL/DRL Algorithm Implementations and Software Frameworks"],"readme":"```diff\n- Please pay attention to the version of SC2 you are using for your experiments. \n- Performance is *not* always comparable between versions. \n- The results in SMAC (https://arxiv.org/abs/1902.04043) use SC2.4.6.2.69232 not SC2.4.10.\n```\n\n# Python MARL framework\n\nPyMARL is [WhiRL](http://whirl.cs.ox.ac.uk)'s framework for deep multi-agent reinforcement learning and includes implementations of the following algorithms:\n- [**QMIX**: QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning](https://arxiv.org/abs/1803.11485)\n- [**COMA**: Counterfactual Multi-Agent Policy Gradients](https://arxiv.org/abs/1705.08926)\n- [**VDN**: Value-Decomposition Networks For Cooperative Multi-Agent Learning](https://arxiv.org/abs/1706.05296) \n- [**IQL**: Independent Q-Learning](https://arxiv.org/abs/1511.08779)\n- [**QTRAN**: QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning](https://arxiv.org/abs/1905.05408)\n\nPyMARL is written in PyTorch and uses [SMAC](https://github.com/oxwhirl/smac) as its environment.\n\n## Installation instructions\n\nBuild the Dockerfile using \n```shell\ncd docker\nbash build.sh\n```\n\nSet up StarCraft II and SMAC:\n```shell\nbash install_sc2.sh\n```\n\nThis will download SC2 into the 3rdparty folder and copy the maps necessary to run over.\n\nThe requirements.txt file can be used to install the necessary packages into a virtual environment (not recomended).\n\n## Run an experiment \n\n```shell\npython3 src/main.py --config=qmix --env-config=sc2 with env_args.map_name=2s3z\n```\n\nThe config files act as defaults for an algorithm or environment. \n\nThey are all located in `src/config`.\n`--config` refers to the config files in `src/config/algs`\n`--env-config` refers to the config files in `src/config/envs`\n\nTo run experiments using the Docker container:\n```shell\nbash run.sh $GPU python3 src/main.py --config=qmix --env-config=sc2 with env_args.map_name=2s3z\n```\n\nAll results will be stored in the `Results` folder.\n\nThe previous config files used for the SMAC Beta have the suffix `_beta`.\n\n## Saving and loading learnt models\n\n### Saving models\n\nYou can save the learnt models to disk by setting `save_model = True`, which is set to `False` by default. The frequency of saving models can be adjusted using `save_model_interval` configuration. Models will be saved in the result directory, under the folder called *models*. The directory corresponding each run will contain models saved throughout the experiment, each within a folder corresponding to the number of timesteps passed since starting the learning process.\n\n### Loading models\n\nLearnt models can be loaded using the `checkpoint_path` parameter, after which the learning will proceed from the corresponding timestep. \n\n## Watching StarCraft II replays\n\n`save_replay` option allows saving replays of models which are loaded using `checkpoint_path`. Once the model is successfully loaded, `test_nepisode` number of episodes are run on the test mode and a .SC2Replay file is saved in the Replay directory of StarCraft II. Please make sure to use the episode runner if you wish to save a replay, i.e., `runner=episode`. The name of the saved replay file starts with the given `env_args.save_replay_prefix` (map_name if empty), followed by the current timestamp. \n\nThe saved replays can be watched by double-clicking on them or using the following command:\n\n```shell\npython -m pysc2.bin.play --norender --rgb_minimap_size 0 --replay NAME.SC2Replay\n```\n\n**Note:** Replays cannot be watched using the Linux version of StarCraft II. Please use either the Mac or Windows version of the StarCraft II client.\n\n## Documentation/Support\n\nDocumentation is a little sparse at the moment (but will improve!). Please raise an issue in this repo, or email [Tabish](mailto:tabish.rashid@cs.ox.ac.uk)\n\n## Citing PyMARL \n\nIf you use PyMARL in your research, please cite the [SMAC paper](https://arxiv.org/abs/1902.04043).\n\n*M. Samvelyan, T. Rashid, C. Schroeder de Witt, G. Farquhar, N. Nardelli, T.G.J. Rudner, C.-M. Hung, P.H.S. Torr, J. Foerster, S. Whiteson. The StarCraft Multi-Agent Challenge, CoRR abs/1902.04043, 2019.*\n\nIn BibTeX format:\n\n```tex\n@article{samvelyan19smac,\n  title = {{The} {StarCraft} {Multi}-{Agent} {Challenge}},\n  author = {Mikayel Samvelyan and Tabish Rashid and Christian Schroeder de Witt and Gregory Farquhar and Nantas Nardelli and Tim G. J. Rudner and Chia-Man Hung and Philiph H. S. Torr and Jakob Foerster and Shimon Whiteson},\n  journal = {CoRR},\n  volume = {abs/1902.04043},\n  year = {2019},\n}\n```\n\n## License\n\nCode licensed under the Apache License v2.0\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Foxwhirl%2Fpymarl","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Foxwhirl%2Fpymarl","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Foxwhirl%2Fpymarl/lists"}