{"id":19401062,"url":"https://github.com/google-research/rl-reliability-metrics","last_synced_at":"2025-04-24T07:30:29.163Z","repository":{"id":54634420,"uuid":"226411041","full_name":"google-research/rl-reliability-metrics","owner":"google-research","description":"The RL Reliability Metrics library provides a set of metrics for measuring the reliability of reinforcement learning (RL) algorithms, as well as statistical tools for comparing algorithms and for computing confidence intervals on these metrics.","archived":true,"fork":false,"pushed_at":"2023-07-31T15:06:56.000Z","size":171,"stargazers_count":166,"open_issues_count":0,"forks_count":22,"subscribers_count":9,"default_branch":"master","last_synced_at":"2025-03-13T13:44:07.287Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/google-research.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null}},"created_at":"2019-12-06T21:07:45.000Z","updated_at":"2025-03-08T02:18:01.000Z","dependencies_parsed_at":"2023-01-30T16:15:47.552Z","dependency_job_id":"0946b4b6-3613-4eb7-8aa2-6a1ddf1761f0","html_url":"https://github.com/google-research/rl-reliability-metrics","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/google-research%2Frl-reliability-metrics","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/google-research%2Frl-reliability-metrics/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/google-research%2Frl-reliability-metrics/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/google-research%2Frl-reliability-metrics/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/google-research","download_url":"https://codeload.github.com/google-research/rl-reliability-metrics/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":250582773,"owners_count":21453911,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-11-10T11:16:59.398Z","updated_at":"2025-04-24T07:30:28.878Z","avatar_url":"https://github.com/google-research.png","language":"Python","funding_links":[],"categories":["Reinforcement Learning (RL) and Deep Reinforcement Learning (DRL)"],"sub_categories":["RL/DRL Benchmarking"],"readme":"# RL Reliability Metrics\n\nThe RL Reliability Metrics library provides a set of metrics for measuring the\nreliability of reinforcement learning (RL) algorithms. The library also provides\nstatistical tools for computing confidence intervals and for comparing\nalgorithms on these metrics.\n\nAs input, this library accepts a set of RL training curves, or a set of rollouts\nof an already trained RL algorithm. The library computes reliability metrics\nacross different dimensions (additionally, it can also analyze non-reliability\nmetrics like median performance), and outputs plots presenting the reliability\nmetrics for each algorithm, aggregated across tasks or on a per-task basis. The\nlibrary also provides statistical tests for comparing algorithms based on these\nmetrics, and provides bootstrapped confidence intervals of the metric values.\n\n## Table of contents\n\n\u003ca href='#Paper'\u003ePaper\u003c/a\u003e\u003cbr\u003e\n\u003ca href='#Installation'\u003eInstallation\u003c/a\u003e\u003cbr\u003e\n\u003ca href='#Examples'\u003eExamples\u003c/a\u003e\u003cbr\u003e\n\u003ca href='#Datasets'\u003eDatasets\u003c/a\u003e\u003cbr\u003e\n\u003ca href='#Contributing'\u003eContributing\u003c/a\u003e\u003cbr\u003e\n\u003ca href='#Principles'\u003ePrinciples\u003c/a\u003e\u003cbr\u003e\n\u003ca href='#Disclaimer'\u003eDisclaimer\u003c/a\u003e\u003cbr\u003e\n\n## Paper\n\nPlease see the paper for a detailed description of the metrics and statistical\ntools implemented by the RL Reliability Metrics library, and for examples of\napplying the methods to common tasks and algorithms:\n[Measuring the Reliability of Reinforcement Learning Algorithms.](https://arxiv.org/abs/1912.05663)\n\nIf you use this code or reference the paper, please cite it as:\n\n```\n@conference{rl_reliability_metrics,\n  title = {Measuring the Reliability of Reinforcement Learning Algorithms},\n  author = {Stephanie CY Chan, Sam Fishman, John Canny, Anoop Korattikara, and Sergio Guadarrama},\n  booktitle = {International Conference on Learning Representations, Addis Ababa, Ethiopia},\n  year = 2020,\n}\n```\n\n## Installation\n\n```bash\ngit clone https://github.com/google-research/rl-reliability-metrics\ncd rl-reliability-metrics\npip3 install -r requirements.txt\n```\n\nNote: Only Python 3.x is supported.\n\n## Examples\n\nSee\n[`rl_reliability_metrics/examples/tf_agents_mujoco_subset`](rl_reliability_metrics/examples/tf_agents_mujoco_subset)\nfor an example of applying the full pipeline to a small example dataset.\n\n## Datasets\n\nFor the continuous control dataset that was analyzed in the\n[Measuring the Reliability of Reinforcement Learning Algorithms](https://arxiv.org/abs/1912.05663)\npaper (TF-Agents algorithm implementations evaluated on OpenAI MuJoCo\nbaselines), please download using\n[this URL](https://storage.googleapis.com/rl-reliability-metrics/data/tf_agents_full_dataset.tgz).\n\n## Contributing\n\nSee [CONTRIBUTING](CONTRIBUTING.md) for a guide on how to contribute.\n\n## Principles\n\nThis project adheres to [Google's AI principles](PRINCIPLES.md). By\nparticipating, using or contributing to this project you are expected to adhere\nto these principles.\n\n## Acknowledgements\n\nMany thanks to Toby Boyd for his assistance in the open-sourcing process, Oscar\nRamirez for code reviews, and Pablo Castro for his help with running experiments\nusing the Dopamine baselines data. Thanks also to the following people for\nhelpful discussions during the formulation of these metrics and the writing of\nthe paper: Mohammad Ghavamzadeh, Yinlam Chow, Danijar Hafner, Rohan Anil, Archit\nSharma, Vikas Sindhwani, Krzysztof Choromanski, Joelle Pineau, Hal Varian,\nShyue-Ming Loh, and Tim Hesterberg.\n\n## Disclaimer\n\nThis is not an official Google product.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fgoogle-research%2Frl-reliability-metrics","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fgoogle-research%2Frl-reliability-metrics","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fgoogle-research%2Frl-reliability-metrics/lists"}