{"id":34065973,"url":"https://github.com/amenra/guardbench","last_synced_at":"2026-03-12T05:30:56.197Z","repository":{"id":260092225,"uuid":"837144095","full_name":"AmenRa/GuardBench","owner":"AmenRa","description":"A Python library for guardrail models evaluation.","archived":false,"fork":false,"pushed_at":"2025-04-01T10:36:52.000Z","size":112,"stargazers_count":22,"open_issues_count":1,"forks_count":1,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-09-09T05:18:37.837Z","etag":null,"topics":["ai-safety","benchmark","evaluation","guardrail-models","guardrails","llm"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"eupl-1.2","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/AmenRa.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":"NOTICE.txt","maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2024-08-02T09:59:12.000Z","updated_at":"2025-09-02T09:17:20.000Z","dependencies_parsed_at":null,"dependency_job_id":"f000ee16-b590-4321-8739-18d759030b7b","html_url":"https://github.com/AmenRa/GuardBench","commit_stats":null,"previous_names":["amenra/guardbench"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/AmenRa/GuardBench","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AmenRa%2FGuardBench","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AmenRa%2FGuardBench/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AmenRa%2FGuardBench/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AmenRa%2FGuardBench/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/AmenRa","download_url":"https://codeload.github.com/AmenRa/GuardBench/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AmenRa%2FGuardBench/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":27719114,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-12-14T02:00:11.348Z","response_time":56,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai-safety","benchmark","evaluation","guardrail-models","guardrails","llm"],"created_at":"2025-12-14T06:05:30.853Z","updated_at":"2025-12-14T06:05:31.988Z","avatar_url":"https://github.com/AmenRa.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"https://repository-images.githubusercontent.com/837144095/8190ad0e-e9ff-4dda-9116-644d62d6b886\"\u003e\n\u003c/div\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003c!-- Python --\u003e\n  \u003ca href=\"https://www.python.org\" alt=\"Python\"\u003e\u003cimg src=\"https://badges.aleen42.com/src/python.svg\"\u003e\u003c/a\u003e\n  \u003c!-- Version --\u003e\n  \u003ca href=\"https://pypi.org/project/guardbench/\"\u003e\u003cimg src=\"https://img.shields.io/pypi/v/guardbench?color=light-green\" alt=\"PyPI version\"\u003e\u003c/a\u003e\n  \u003c!-- Docs --\u003e\n  \u003ca href=\"https://github.com/AmenRa/guardbench/tree/main/docs\"\u003e\u003cimg src=\"https://img.shields.io/badge/docs-passing-\u003cCOLOR\u003e.svg\" alt=\"Documentation Status\"\u003e\u003c/a\u003e\n  \u003c!-- Black --\u003e\n  \u003ca href=\"https://github.com/psf/black\" alt=\"Code style: black\"\u003e\u003cimg src=\"https://img.shields.io/badge/code%20style-black-000000.svg\"\u003e\u003c/a\u003e\n  \u003c!-- License --\u003e\n  \u003ca href=\"https://interoperable-europe.ec.europa.eu/sites/default/files/custom-page/attachment/2020-03/EUPL-1.2%20EN.txt\"\u003e\u003cimg src=\"https://img.shields.io/badge/license-EUPL-blue.svg\" alt=\"License: EUPL-1.2\"\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n# GuardBench\n\n## 🔥 News\n\n- [October 9, 2025] GuardBench now supports four additional datasets: [JBB Behaviors](https://huggingface.co/datasets/JailbreakBench/JBB-Behaviors), [NicheHazardQA](https://huggingface.co/datasets/SoftMINER-Group/NicheHazardQA), [HarmEval](https://huggingface.co/datasets/SoftMINER-Group/HarmEval), and [TechHazardQA](https://huggingface.co/datasets/SoftMINER-Group/TechHazardQA). Also, it now allows for choosing the metrics to show at the end of the evaluation. Supported metrics are: `precision` (Precision), `recall` (Recall), `f1` (F1), `mcc` (Matthews Correlation Coefficient), `auprc` (AUPRC), `sensitivity` (Sensitivity), `specificity` (Specificity), `g_mean` (G-Mean), `fpr` (False Positive Rate), `fnr` (False Negative Rate).\n\n## ⚡️ Introduction\n[`GuardBench`](https://github.com/AmenRa/guardbench) is a Python library for the evaluation of guardrail models, i.e., LLMs fine-tuned to detect unsafe content in human-AI interactions.\n[`GuardBench`](https://github.com/AmenRa/guardbench) provides a common interface to 40 evaluation datasets, which are downloaded and converted into a [standardized format](docs/data_format.md) for improved usability.\nIt also allows to quickly [compare results and export](docs/report.md) `LaTeX` tables for scientific publications.\n[`GuardBench`](https://github.com/AmenRa/guardbench)'s benchmarking pipeline can also be leveraged on [custom datasets](docs/custom_dataset.md).\n\n[`GuardBench`](https://github.com/AmenRa/guardbench) was featured in [EMNLP 2024](https://2024.emnlp.org).\nThe related paper is available [here](https://aclanthology.org/2024.emnlp-main.1022.pdf).\n\n[`GuardBench`](https://github.com/AmenRa/guardbench) has a public [leaderboard](https://huggingface.co/spaces/AmenRa/guardbench-leaderboard) available on HuggingFace.\n\nYou can find the list of supported datasets [here](docs/datasets.md).\nA few of them requires authorization. Please, read [this](docs/get_datasets.md).\n\nIf you use [`GuardBench`](https://github.com/AmenRa/guardbench) to evaluate guardrail models for your scientific publications, please consider [citing our work](#-citation).\n\n## ✨ Features\n- [40 datasets](docs/datasets.md) for guardrail models evaluation.\n- Automated evaluation pipeline.\n- User-friendly.\n- [Extendable](docs/custom_dataset.md).\n- Reproducible and sharable evaluation.\n- Exportable [evaluation reports](docs/report.md).\n\n## 🔌 Requirements\n```bash\npython\u003e=3.10\n```\n\n## 💾 Installation \n```bash\npip install guardbench\n```\n\n## 💡 Usage\n```python\nfrom guardbench import benchmark\n\ndef moderate(\n    conversations: list[list[dict[str, str]]],  # MANDATORY!\n    # additional `kwargs` as needed\n) -\u003e list[float]:\n    # do moderation\n    # return list of floats (unsafe probabilities)\n\nbenchmark(\n    moderate=moderate,  # User-defined moderation function\n    model_name=\"My Guardrail Model\",\n    batch_size=1,              # Default value\n    datasets=\"all\",            # Default value\n    metrics=[\"f1\", \"recall\"],  # Default value\n    # Note: you can pass additional `kwargs` for `moderate`\n)\n```\n\n### 📖 Examples\n- Follow our [tutorial](docs/llama_guard.md) on benchmarking [`Llama Guard`](https://arxiv.org/pdf/2312.06674) with [`GuardBench`](https://github.com/AmenRa/guardbench).\n- More examples are available in the [`scripts`](scripts/effectiveness) folder.\n\n## 📚 Documentation\nBrowse the documentation for more details about:\n- The [datasets](docs/datasets.md) and how to [obtain them](docs/get_datasets.md).\n- The [data format](data_format.md) used by [`GuardBench`](https://github.com/AmenRa/guardbench).\n- How to use the [`Report`](docs/report.md) class to compare models and export results as `LaTeX` tables.\n- How to leverage [`GuardBench`](https://github.com/AmenRa/guardbench)'s benchmarking pipeline on [custom datasets](docs/custom_dataset.md).\n\n## 🏆 Leaderboard\nYou can find [`GuardBench`](https://github.com/AmenRa/guardbench)'s leaderboard [here](https://huggingface.co/spaces/AmenRa/guardbench-leaderboard). If you want to submit your results, please contact us.\n\u003c!-- All results can be reproduced using the provided [`scripts`](scripts/effectiveness).   --\u003e\n\n## 👨‍💻 Authors\n- Elias Bassani (European Commission - Joint Research Centre)\n\n## 🎓 Citation\n```bibtex\n@inproceedings{guardbench,\n    title = \"{G}uard{B}ench: A Large-Scale Benchmark for Guardrail Models\",\n    author = \"Bassani, Elias  and\n      Sanchez, Ignacio\",\n    editor = \"Al-Onaizan, Yaser  and\n      Bansal, Mohit  and\n      Chen, Yun-Nung\",\n    booktitle = \"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing\",\n    month = nov,\n    year = \"2024\",\n    address = \"Miami, Florida, USA\",\n    publisher = \"Association for Computational Linguistics\",\n    url = \"https://aclanthology.org/2024.emnlp-main.1022\",\n    doi = \"10.18653/v1/2024.emnlp-main.1022\",\n    pages = \"18393--18409\",\n}\n```\n\n## 🎁 Feature Requests\nWould you like to see other features implemented? Please, open a [feature request](https://github.com/AmenRa/guardbench/issues/new?assignees=\u0026labels=enhancement\u0026template=feature_request.md\u0026title=%5BFeature+Request%5D+title).\n\n## 📄 License\n[GuardBench](https://github.com/AmenRa/guardbench) is provided as open-source software licensed under [EUPL v1.2](https://github.com/AmenRa/guardbench/blob/master/LICENSE).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Famenra%2Fguardbench","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Famenra%2Fguardbench","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Famenra%2Fguardbench/lists"}