{"id":2642,"url":"https://github.com/cor3bit/awesome-marl-engineering","name":"awesome-marl-engineering","description":"A (very subjective) curated list of Safe Multi-agent Reinforcement Learning resources","projects_count":36,"last_synced_at":"2026-08-31T04:00:42.687Z","repository":{"id":111744737,"uuid":"557348118","full_name":"cor3bit/awesome-marl-engineering","owner":"cor3bit","description":"A (very subjective) curated list of Safe Multi-agent Reinforcement Learning resources","archived":false,"fork":false,"pushed_at":"2023-01-11T20:35:41.000Z","size":5,"stargazers_count":3,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2026-08-11T14:08:53.363Z","etag":null,"topics":["awesome","distributed-control","multi-agent-systems","reinforcement-learning"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/cor3bit.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null}},"created_at":"2022-10-25T14:23:34.000Z","updated_at":"2024-12-20T16:06:08.000Z","dependencies_parsed_at":"2024-01-04T20:14:49.986Z","dependency_job_id":"df58e23e-5a92-4af6-9773-f04ef6f1c62b","html_url":"https://github.com/cor3bit/awesome-marl-engineering","commit_stats":null,"previous_names":["cor3bit/awesome-marl-engineering","cor3bit/awesome-safe-marl"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/cor3bit/awesome-marl-engineering","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cor3bit%2Fawesome-marl-engineering","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cor3bit%2Fawesome-marl-engineering/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cor3bit%2Fawesome-marl-engineering/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cor3bit%2Fawesome-marl-engineering/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/cor3bit","download_url":"https://codeload.github.com/cor3bit/awesome-marl-engineering/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cor3bit%2Fawesome-marl-engineering/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":36991523,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-08-31T02:00:07.497Z","response_time":119,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2024-01-04T20:14:47.980Z","updated_at":"2026-08-31T04:00:42.688Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["MARL","Preliminaries","Engineering","Related awesome lists"],"sub_categories":["Papers (research)","Safe RL","Books","Papers (survey)","Building Multi-agent Environments"],"readme":"# Awesome Safe MARL [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)\r\n\r\nA (very subjective) curated list of Safe Multi-agent\r\nReinforcement Learning resources.\r\n\r\n\r\n## Preliminaries\r\n\r\n### Constrained Optimization\r\n\r\nTODO\r\n\r\n\r\n### Control Theory\r\n\r\nTODO\r\n\r\n\r\n### MBRL\r\n\r\nTODO\r\n\r\n\r\n### Safe RL\r\n\r\n- [Constrained Policy Optimization](https://arxiv.org/abs/1705.10528), Achiam et al, 2017. Algorithm: CPO.\r\n\r\n- [Saute RL: Almost Surely Safe Reinforcement Learning Using State Augmentation](https://arxiv.org/abs/2202.06558), Sootla et al, 2022. Algorithm: Saute MDP augmentation.\r\n\r\n\r\n\r\n### Graph Theory\r\n\r\nTODO\r\n\r\n### Federated Learning\r\n\r\nTODO\r\n\r\n\u003c!-- ## Building Blocks\r\n\r\nTODO\r\n\r\n## Safe MARL --\u003e\r\n\r\n## MARL\r\n\r\n### Books\r\n\r\n- Part V of [Algorithms for Decision Making](https://algorithmsbook.com/) by \r\nMykel J. Kochenderfer, Tim A. Wheeler and Kyle H. Wray\r\n- [Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations](http://www.masfoundations.org/index.html) by Yoav Shoham and Kevin Leyton-Brown\r\n- [Multi-Agent Coordination: A Reinforcement Learning Approach](https://www.wiley.com/en-us/Multi+Agent+Coordination:+A+Reinforcement+Learning+Approach-p-9781119699033) by Arup \r\nKumar Sadhu and Amit Konar :dollar:\r\n- [Distributed Optimization-Based Control of Multi-Agent Networks in Complex Environments](https://link.springer.com/book/10.1007/978-3-319-19072-3) by Minghui Zhu and Sonia Martinez :dollar:\r\n\r\n\r\n### Papers (survey)\r\n\r\n- [Multi-agent Reinforcement Learning: An Overview](https://citeseerx.ist.psu.edu/document?repid=rep1\u0026type=pdf\u0026doi=8851953ef486615fce803bda2e40aec97cbb5547), Busoniu et al, 2010.\r\n- [Multi-agent deep reinforcement learning: a survey](https://link.springer.com/content/pdf/10.1007/s10462-021-09996-w.pdf), Sven Gronauer and Klaus Diepold, 2021.\r\n- [A Survey of Multi-Agent Reinforcement Learning with Communication](https://arxiv.org/abs/2203.08975), Zhu et al, 2022.\r\n\r\n\r\n\r\n### Papers (research)\r\n\r\n#### POMDP for Multiple Agents\r\n\r\nTODO\r\n\r\n\r\n#### Value Factorization\r\n\r\n- [Multiagent Cooperation and Competition with Deep Reinforcement Learning](https://arxiv.org/abs/1511.08779), Tampuu et al, 2015.\r\n- [Value-Decomposition Networks For Cooperative Multi-Agent Learning](https://arxiv.org/abs/1706.05296), Sunehag et al, 2017.\r\n- [QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning](https://arxiv.org/abs/1803.11485), Rashid et al, 2018. Algorithm: QMIX.\r\n- [QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning](https://proceedings.mlr.press/v97/son19a.html), Son et al, 2019.\r\n- [QPLEX: Duplex Dueling Multi-Agent Q-Learning](https://arxiv.org/abs/2008.01062), Wang et al, 2020.\r\n\r\n\r\n#### Policy Gradient\r\n\r\n- [Counterfactual Multi-Agent Policy Gradients](https://ojs.aaai.org/index.php/AAAI/article/view/11794), Foerster et al, 2018.\r\n- [Distributed Multi-Agent Reinforcement Learning by Actor-Critic Method](https://www.sciencedirect.com/science/article/pii/S240589631932035X), P. Heredia, S. Mou, 2019.\r\n\r\n\r\n\r\n\r\n#### Communication\r\n\r\n- [Learning to Communicate with Deep Multi-Agent Reinforcement Learning](https://proceedings.neurips.cc/paper/2016/hash/c7635bfd99248a2cdef8249ef7bfbef4-Abstract.html), Foerster et al, 2016.\r\n- [Learning Multiagent Communication with Backpropagation](https://proceedings.neurips.cc/paper/2016/hash/55b1927fdafef39c48e5b73b5d61ea60-Abstract.html), Sukhbaatar et al, 2016.\r\n- [TarMAC: Targeted Multi-Agent Communication](https://proceedings.mlr.press/v97/das19a.html), Das et al, 2019.\r\n- [Learning to Ground Multi-Agent Communication with Autoencoders](https://proceedings.neurips.cc/paper/2021/hash/80fee67c8a4c4989bf8a580b4bbb0cd2-Abstract.html), Lin et al, 2021.\r\n- [Differentially Private and Communication Efficient Collaborative Learning](https://ojs.aaai.org/index.php/AAAI/article/view/16887), Ding et al, 2021.\r\n\r\n\r\n\r\n\r\n#### DLQR\r\n\r\n- [Distributed Q-Learning for Dynamically Decoupled Systems](https://arxiv.org/abs/1809.08745), Siavash Alemzadeh, Mehran Mesbahi, 2018.\r\n- [D3PI: Data-Driven Distributed Policy Iteration for Homogeneous Interconnected Systems](https://arxiv.org/abs/2103.11572), Alemzadeh et al, 2021.\r\n- [Data-Driven Distributed Optimal Consensus Control for Unknown Multiagent Systems With Input-Delay](https://ieeexplore.ieee.org/document/8334323), Zhang et al, 2018.\r\n- [Distributed Q-Learning with State Tracking for Multi-agent Networked Control](https://arxiv.org/abs/2012.12383), Wang et al, 2020.\r\n- [Distributed Adaptive Linear Quadratic Control using Distributed Reinforcement Learning](https://www.sciencedirect.com/science/article/pii/S2405896319307748), Daniel Goerges, 2019.\r\n- [Distributed Linear-Quadratic Control with Graph Neural Networks](https://arxiv.org/abs/2103.08417), Fernando Gama, Somayeh Sojoudi, 2021.\r\n- [Distributed Online Linear Quadratic Control for Linear Time-invariant Systems](https://arxiv.org/abs/2009.13749), Ting-Jui Chang, Shahin Shahrampour, 2020.\r\n- [Learning Distributed Stabilizing Controllers for Multi-Agent Systems](https://arxiv.org/abs/2009.13749), Jing et al, 2021. \r\n\r\n\r\n#### Vehicle Platoon\r\n\r\n- [Distributed Nonlinear Model Predictive Control for Connected Autonomous Electric Vehicles Platoon with Distance-Dependent Air Drag Formulation](https://www.mdpi.com/1996-1073/14/16/5122), Caiazzo et al, 2021.\r\n\r\n\r\n\r\n## Engineering\r\n\r\n### Building Multi-agent Environments\r\n\r\n- [Multi-Agent Learning Environments](https://agents.inf.ed.ac.uk/blog/multiagent-learning-environments/), blog post by Lukas Schäfer\r\n- [PettingZoo Neurips'21 paper](https://proceedings.neurips.cc/paper/2021/hash/7ed2d3454c5eea71148b11d0c25104ff-Abstract.html)\r\n- [PettingZoo Documentation](https://pettingzoo.farama.org/)\r\n- [ma-gym](https://github.com/koulanurag/ma-gym) with [minimal-marl](https://github.com/koulanurag/minimal-marl), set of baselines for vanilla MARL problems by Anurag Koul\r\n- [Ray Multi-Agent and Hierarchical Environments](https://docs.ray.io/en/latest/rllib/rllib-env.html#multi-agent-and-hierarchical)\r\n\r\n\r\n## Related awesome lists\r\n- [Awesome MARL](https://github.com/instadeepai/awesome-marl)\r\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/cor3bit%2Fawesome-marl-engineering/projects"}