{"id":13678558,"url":"https://github.com/st-tech/zr-obp","last_synced_at":"2026-01-17T01:13:30.776Z","repository":{"id":37878166,"uuid":"272587927","full_name":"st-tech/zr-obp","owner":"st-tech","description":"Open Bandit Pipeline: a python library for bandit algorithms and off-policy evaluation","archived":false,"fork":false,"pushed_at":"2024-06-03T11:04:44.000Z","size":30041,"stargazers_count":677,"open_issues_count":32,"forks_count":96,"subscribers_count":87,"default_branch":"master","last_synced_at":"2025-09-08T14:51:24.634Z","etag":null,"topics":["contextual-bandits","datasets","multi-armed-bandits","off-policy-evaluation","research"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/st-tech.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2020-06-16T02:10:15.000Z","updated_at":"2025-09-07T17:05:51.000Z","dependencies_parsed_at":"2023-01-25T12:01:00.444Z","dependency_job_id":"5e828a61-9389-4b87-9a7c-ab260948684e","html_url":"https://github.com/st-tech/zr-obp","commit_stats":{"total_commits":851,"total_committers":15,"mean_commits":"56.733333333333334","dds":"0.40423031727379555","last_synced_commit":"8cbd5fa4558b7ad2ba4781546d6604e4cc3e07c4"},"previous_names":[],"tags_count":14,"template":false,"template_full_name":null,"purl":"pkg:github/st-tech/zr-obp","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/st-tech%2Fzr-obp","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/st-tech%2Fzr-obp/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/st-tech%2Fzr-obp/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/st-tech%2Fzr-obp/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/st-tech","download_url":"https://codeload.github.com/st-tech/zr-obp/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/st-tech%2Fzr-obp/sbom","scorecard":{"id":845188,"data":{"date":"2025-08-11","repo":{"name":"github.com/st-tech/zr-obp","commit":"8cbd5fa4558b7ad2ba4781546d6604e4cc3e07c4"},"scorecard":{"version":"v5.2.1-40-gf6ed084d","commit":"f6ed084d17c9236477efd66e5b258b9d4cc7b389"},"score":2.9,"checks":[{"name":"Packaging","score":-1,"reason":"packaging workflow not detected","details":["Warn: no GitHub/GitLab publishing workflow detected."],"documentation":{"short":"Determines if the project is published as a package that others can easily download, install, easily update, and uninstall.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#packaging"}},{"name":"Maintained","score":0,"reason":"0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project is \"actively maintained\".","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#maintained"}},{"name":"Code-Review","score":4,"reason":"Found 2/5 approved changesets -- score normalized to 4","details":null,"documentation":{"short":"Determines if the project requires human code review before pull requests (aka merge requests) are merged.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#code-review"}},{"name":"Token-Permissions","score":0,"reason":"detected GitHub workflow tokens with excessive permissions","details":["Warn: no topLevel permission defined: .github/workflows/lints.yml:1","Warn: no topLevel permission defined: .github/workflows/tests.yml:1","Info: no jobLevel write permissions found"],"documentation":{"short":"Determines if the project's workflows follow the principle of least privilege.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#token-permissions"}},{"name":"Dangerous-Workflow","score":10,"reason":"no dangerous workflow patterns detected","details":null,"documentation":{"short":"Determines if the project's GitHub Action workflows avoid dangerous patterns.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#dangerous-workflow"}},{"name":"CII-Best-Practices","score":0,"reason":"no effort to earn an OpenSSF best practices badge detected","details":null,"documentation":{"short":"Determines if the project has an OpenSSF (formerly CII) Best Practices Badge.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#cii-best-practices"}},{"name":"Binary-Artifacts","score":10,"reason":"no binaries found in the repo","details":null,"documentation":{"short":"Determines if the project has generated executable (binary) artifacts in the source repository.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#binary-artifacts"}},{"name":"Security-Policy","score":0,"reason":"security policy file not detected","details":["Warn: no security policy file detected","Warn: no security file to analyze","Warn: no security file to analyze","Warn: no security file to analyze"],"documentation":{"short":"Determines if the project has published a security policy.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#security-policy"}},{"name":"License","score":10,"reason":"license file detected","details":["Info: project has a license file: LICENSE:0","Info: FSF or OSI recognized license: Apache License 2.0: LICENSE:0"],"documentation":{"short":"Determines if the project has defined a license.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#license"}},{"name":"Signed-Releases","score":-1,"reason":"no releases found","details":null,"documentation":{"short":"Determines if the project cryptographically signs release artifacts.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#signed-releases"}},{"name":"Fuzzing","score":0,"reason":"project is not fuzzed","details":["Warn: no fuzzer integrations found"],"documentation":{"short":"Determines if the project uses fuzzing.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#fuzzing"}},{"name":"Branch-Protection","score":0,"reason":"branch protection not enabled on development/release branches","details":["Warn: branch protection not enabled for branch 'master'"],"documentation":{"short":"Determines if the default and release branches are protected with GitHub's branch protection settings.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#branch-protection"}},{"name":"SAST","score":0,"reason":"SAST tool is not run on all commits -- score normalized to 0","details":["Warn: 0 commits out of 30 are checked with a SAST tool"],"documentation":{"short":"Determines if the project uses static code analysis.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#sast"}},{"name":"Pinned-Dependencies","score":0,"reason":"dependency not pinned by hash detected -- score normalized to 0","details":["Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/lints.yml:15: update your workflow using https://app.stepsecurity.io/secureworkflow/st-tech/zr-obp/lints.yml/master?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/lints.yml:18: update your workflow using https://app.stepsecurity.io/secureworkflow/st-tech/zr-obp/lints.yml/master?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/lints.yml:23: update your workflow using https://app.stepsecurity.io/secureworkflow/st-tech/zr-obp/lints.yml/master?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/tests.yml:22: update your workflow using https://app.stepsecurity.io/secureworkflow/st-tech/zr-obp/tests.yml/master?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/tests.yml:25: update your workflow using https://app.stepsecurity.io/secureworkflow/st-tech/zr-obp/tests.yml/master?enable=pin","Warn: pipCommand not pinned by hash: .github/workflows/lints.yml:29","Warn: pipCommand not pinned by hash: .github/workflows/lints.yml:30","Warn: pipCommand not pinned by hash: .github/workflows/tests.yml:31","Warn: pipCommand not pinned by hash: .github/workflows/tests.yml:32","Warn: pipCommand not pinned by hash: .github/workflows/tests.yml:35","Warn: pipCommand not pinned by hash: .github/workflows/tests.yml:37","Info:   0 out of   4 GitHub-owned GitHubAction dependencies pinned","Info:   0 out of   1 third-party GitHubAction dependencies pinned","Info:   0 out of   6 pipCommand dependencies pinned"],"documentation":{"short":"Determines if the project has declared and pinned the dependencies of its build process.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#pinned-dependencies"}},{"name":"Vulnerabilities","score":0,"reason":"42 existing vulnerabilities detected","details":["Warn: Project is vulnerable to: PYSEC-2024-48 / GHSA-fj7x-q9j7-g6q6","Warn: Project is vulnerable to: PYSEC-2024-230 / GHSA-248v-346w-9cwc","Warn: Project is vulnerable to: PYSEC-2022-42986 / GHSA-43fp-rhv2-5gv8","Warn: Project is vulnerable to: PYSEC-2023-135 / GHSA-xqr8-7jwr-rhp7","Warn: Project is vulnerable to: PYSEC-2024-60 / GHSA-jjg7-2v4v-x38h","Warn: Project is vulnerable to: PYSEC-2022-288 / GHSA-6hrg-qmvc-2xh8","Warn: Project is vulnerable to: GHSA-fpfv-jqm9-f5jm","Warn: Project is vulnerable to: GHSA-3f63-hfp8-52jq","Warn: Project is vulnerable to: GHSA-44wm-f244-xhp3","Warn: Project is vulnerable to: GHSA-4fx9-vc88-q2xc","Warn: Project is vulnerable to: PYSEC-2023-227 / GHSA-8ghj-p4vj-mr35","Warn: Project is vulnerable to: PYSEC-2022-10 / GHSA-8vj2-vxx3-667w","Warn: Project is vulnerable to: PYSEC-2022-168 / GHSA-9j59-75qj-795w","Warn: Project is vulnerable to: GHSA-j7hp-h8jx-5ppr","Warn: Project is vulnerable to: PYSEC-2022-42979 / GHSA-m2vv-5vj5-2hm7","Warn: Project is vulnerable to: PYSEC-2022-8 / GHSA-pw3c-h7wp-cvhx","Warn: Project is vulnerable to: PYSEC-2022-9 / GHSA-xrcv-f9gm-v42c","Warn: Project is vulnerable to: PYSEC-2023-175","Warn: Project is vulnerable to: GHSA-9hjg-9r4m-mvj7","Warn: Project is vulnerable to: GHSA-9wx4-h78v-vm56","Warn: Project is vulnerable to: PYSEC-2023-74 / GHSA-j8r2-6x86-q33q","Warn: Project is vulnerable to: PYSEC-2024-110 / GHSA-jw8x-6495-233v","Warn: Project is vulnerable to: GHSA-jxfp-4rvq-9h9m","Warn: Project is vulnerable to: PYSEC-2023-102","Warn: Project is vulnerable to: PYSEC-2023-114","Warn: Project is vulnerable to: GHSA-3749-ghw9-m3mg","Warn: Project is vulnerable to: PYSEC-2022-43015 / GHSA-47fc-vmwq-366v","Warn: Project is vulnerable to: PYSEC-2025-41 / GHSA-53q9-r3pm-6pq6","Warn: Project is vulnerable to: PYSEC-2024-252 / GHSA-5pcm-hx3q-hm94","Warn: Project is vulnerable to: GHSA-887c-mr87-cxwp","Warn: Project is vulnerable to: PYSEC-2024-251 / GHSA-pg7h-5qx3-wjr3","Warn: Project is vulnerable to: PYSEC-2024-250","Warn: Project is vulnerable to: PYSEC-2024-259","Warn: Project is vulnerable to: GHSA-g7vv-2v7x-gj9p","Warn: Project is vulnerable to: GHSA-34jh-p97f-mpxf","Warn: Project is vulnerable to: PYSEC-2023-212 / GHSA-g4mx-q9vg-27p4","Warn: Project is vulnerable to: GHSA-pq67-6m6q-mj2v","Warn: Project is vulnerable to: PYSEC-2023-192 / GHSA-v845-jxx5-vc9f","Warn: Project is vulnerable to: OSV-2022-1074","Warn: Project is vulnerable to: OSV-2022-715","Warn: Project is vulnerable to: PYSEC-2022-42969","Warn: Project is vulnerable to: GHSA-jfmj-5v4g-7637"],"documentation":{"short":"Determines if the project has open, known unfixed vulnerabilities.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#vulnerabilities"}}]},"last_synced_at":"2025-08-23T21:15:06.901Z","repository_id":37878166,"created_at":"2025-08-23T21:15:06.901Z","updated_at":"2025-08-23T21:15:06.901Z"},"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":28491127,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-01-17T00:50:05.742Z","status":"ssl_error","status_checked_at":"2026-01-17T00:43:11.982Z","response_time":107,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.5:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["contextual-bandits","datasets","multi-armed-bandits","off-policy-evaluation","research"],"created_at":"2024-08-02T13:00:55.105Z","updated_at":"2026-01-17T01:13:30.767Z","avatar_url":"https://github.com/st-tech.png","language":"Python","funding_links":[],"categories":["Python","Open Source Software/Implementations"],"sub_categories":["Off-Policy Evaluation and Learning: Applications"],"readme":"\u003cdiv align=\"center\"\u003e\u003cimg src=\"https://raw.githubusercontent.com/st-tech/zr-obp/master/images/logo.png\" width=\"60%\"/\u003e\u003c/div\u003e\n\n[![pypi](https://img.shields.io/pypi/v/obp.svg)](https://pypi.python.org/pypi/obp)\n[![Python](https://img.shields.io/badge/python-3.7%20%7C%203.8%20%7C%203.9-blue)](https://www.python.org)\n[![Downloads](https://pepy.tech/badge/obp)](https://pepy.tech/project/obp)\n![GitHub commit activity](https://img.shields.io/github/commit-activity/m/st-tech/zr-obp)\n![GitHub last commit](https://img.shields.io/github/last-commit/st-tech/zr-obp)\n[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)\n[![arXiv](https://img.shields.io/badge/arXiv-2008.07146-b31b1b.svg)](https://arxiv.org/abs/2008.07146)\n\n[[arXiv]](https://arxiv.org/abs/2008.07146)\n[[NeurIPS2021 Proceedings]](https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/33e75ff09dd601bbe69f351039152189-Abstract-round2.html)\n# Open Bandit Pipeline: a research framework for off-policy evaluation and learning\n\n**[Docs](https://zr-obp.readthedocs.io/en/latest/)** | **[Google Group](https://groups.google.com/g/open-bandit-project)** | **[Tutorial](https://sites.google.com/cornell.edu/recsys2021tutorial)** | **[Installation](#installation)** | **[Usage](#usage)** | **[Slides](./slides/slides_EN.pdf)** | **[Quickstart](./examples/quickstart)** | **[Open Bandit Dataset](./obd)** | **[日本語](./README_JN.md)**\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003eTable of Contents\u003c/strong\u003e\u003c/summary\u003e\n\n- [Open Bandit Pipeline: a research framework for off-policy evaluation and learning](#open-bandit-pipeline-a-research-framework-for-bandit-algorithms-and-off-policy-evaluation)\n- [Overview](#overview)\n  - [Open Bandit Dataset (OBD)](#open-bandit-dataset-obd)\n  - [Open Bandit Pipeline (OBP)](#open-bandit-pipeline-obp)\n    - [Algorithms and OPE Estimators Supported](#algorithms-and-ope-estimators-supported)\n- [Installation](#installation)\n- [Usage](#usage)\n  - [(1) Data loading and preprocessing](#1-data-loading-and-preprocessing)\n  - [(2) Off-Policy Learning](#2-off-policy-learning)\n  - [(3) Off-Policy Evaluation](#3-off-policy-evaluation)\n- [Citation](#citation)\n- [Google Group](#google-group)\n- [Contribution](#contribution)\n- [License](#license)\n- [Project Team](#project-team)\n- [Contact](#contact)\n- [References](#references)\n\n\u003c/details\u003e\n\n# Overview\n\n## Open Bandit Dataset (OBD)\n\n*Open Bandit Dataset* is a public real-world logged bandit dataset.\nThis dataset is provided by [ZOZO, Inc.](https://corp.zozo.com/en/about/profile/), the largest fashion e-commerce company in Japan.\nThe company uses some multi-armed bandit algorithms to recommend fashion items to users in a large-scale fashion e-commerce platform called [ZOZOTOWN](https://zozo.jp/).\nThe following figure presents the displayed fashion items as actions where there are three *positions* in the recommendation interface.\n\n\u003cdiv align=\"center\"\u003e\u003cimg src=\"https://raw.githubusercontent.com/st-tech/zr-obp/master/images/recommended_fashion_items.png\" width=\"45%\"/\u003e\u003c/div\u003e\n\u003cfigcaption\u003e\n\u003cp align=\"center\"\u003e\n  Recommended fashion items as actions in the ZOZOTOWN recommendation interface\n\u003c/p\u003e\n\u003c/figcaption\u003e\n\nThe dataset was collected during a 7-day experiment on three “campaigns,” corresponding to all, men's, and women's items, respectively.\nEach campaign randomly used either the Uniform Random policy or the Bernoulli Thompson Sampling (Bernoulli TS) policy for the data collection.\nOpen Bandit Dataset is unique in that it contains a set of *multiple* logged bandit datasets collected by running different policies on the same platform. This enables realistic and reproducible experimental comparisons of different OPE estimators for the first time (see Section 5 of the reference [paper](https://arxiv.org/abs/2008.07146) for the details of the evaluation of OPE protocol using Open Bandit Dataset).\n\n\u003cdiv align=\"center\"\u003e\u003cimg src=\"https://raw.githubusercontent.com/st-tech/zr-obp/master/images/obd_stats.png\" width=\"90%\"/\u003e\u003c/div\u003e\n\nThe small size version of our data is available at [obd](./obd).\nWe release the full size version of our data at [https://research.zozo.com/data.html](https://research.zozo.com/data.html).\nPlease download the full size version for research uses.\nPlease also see [obd/README.md](./obd/README.md) for the detailed dataset description.\n\n## Open Bandit Pipeline (OBP)\n\n*Open Bandit Pipeline* is an open-source Python software including a series of modules for implementing dataset preprocessing, policy learning methods, and OPE estimators. Our software provides a complete, standardized experimental procedure for OPE research, ensuring that performance comparisons are fair and reproducible. It also enables fast and accurate OPE implementation through a single unified interface, simplifying the practical use of OPE.\n\n\u003cdiv align=\"center\"\u003e\u003cimg src=\"https://raw.githubusercontent.com/st-tech/zr-obp/master/images/overview.png\" width=\"80%\"/\u003e\u003c/div\u003e\n\u003cfigcaption\u003e\n\u003cp align=\"center\"\u003e\n  Overview of the Open Bandit Pipeline\n\u003c/p\u003e\n\u003c/figcaption\u003e\n\nOpen Bandit Pipeline consists of the following main modules.\n\n- [**dataset module**](./obp/dataset/): This module provides a data loader for Open Bandit Dataset and a flexible interface for handling logged bandit data. It also provides tools to generate synthetic bandit data and transform multi-class classification data to bandit data.\n- [**policy module**](./obp/policy/): This module provides interfaces for implementing new online and offline bandit policies. It also implements several standard policy learning methods.\n- [**simulator module**](./obp/simulator/): This module provides functions for conducting offline bandit simulation. This module is necessary only when you use the ReplayMethod to evaluate online bandit policies. Please refer to [examples/quickstart/online.ipynb](./examples/quickstart/replay.ipynb) for a quickstart guide of implementing OPE of online bandit algorithms.\n- [**ope module**](./obp/ope/): This module provides generic abstract interfaces to support custom implementations so that researchers can evaluate their own estimators easily. It also implements several basic and advanced OPE estimators.\n\n### Supported Bandit Algorithms and OPE Estimators\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003eBandit Algorithms \u003c/strong\u003e(click to expand)\u003c/summary\u003e\n\u003cbr\u003e\n\n- Online\n  - Non-Contextual (Context-free)\n    - Random\n    - Epsilon Greedy\n    - Bernoulli Thompson Sampling\n  - Contextual (Linear)\n    - Linear Epsilon Greedy\n    - [Linear Thompson Sampling](http://proceedings.mlr.press/v28/agrawal13)\n    - [Linear Upper Confidence Bound](https://dl.acm.org/doi/pdf/10.1145/1772690.1772758)\n  - Contextual (Logistic)\n    - Logistic Epsilon Greedy\n    - [Logistic Thompson Sampling](https://papers.nips.cc/paper/4321-an-empirical-evaluation-of-thompson-sampling)\n    - [Logistic Upper Confidence Bound](https://dl.acm.org/doi/10.1145/2396761.2396767)\n- Offline (Off-Policy Learning)\n  - [Inverse Probability Weighting (IPW) Learner](https://arxiv.org/abs/1503.02834)\n  - Neural Network-based Policy Learner\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003eOPE Estimators \u003c/strong\u003e(click to expand)\u003c/summary\u003e\n\u003cbr\u003e\n\n- OPE of Online Bandit Algorithms\n  - [Replay Method (RM)](https://arxiv.org/abs/1003.5956)\n- OPE of Offline Bandit Algorithms\n  - [Direct Method (DM)](https://arxiv.org/abs/0812.4044)\n  - [Inverse Probability Weighting (IPW)](https://scholarworks.umass.edu/cgi/viewcontent.cgi?article=1079\u0026context=cs_faculty_pubs)\n  - [Self-Normalized Inverse Probability Weighting (SNIPW)](https://papers.nips.cc/paper/5748-the-self-normalized-estimator-for-counterfactual-learning)\n  - [Doubly Robust (DR)](https://arxiv.org/abs/1503.02834)\n  - [Switch Estimators](https://arxiv.org/abs/1612.01205)\n  - [More Robust Doubly Robust (MRDR)](https://arxiv.org/abs/1802.03493)\n  - [Doubly Robust with Optimistic Shrinkage (DRos)](https://arxiv.org/abs/1907.09623)\n  - [Sub-Gaussian Inverse Probability Weighting (SGIPW)](https://proceedings.neurips.cc/paper/2021/hash/4476b929e30dd0c4e8bdbcc82c6ba23a-Abstract.html)\n  - [Sub-Gaussian Doubly Robust (SGDR)](https://proceedings.neurips.cc/paper/2021/hash/4476b929e30dd0c4e8bdbcc82c6ba23a-Abstract.html)\n  - [Double Machine Learning (DML)](https://arxiv.org/abs/2002.08536)\n- OPE of Offline Slate Bandit Algorithms\n  - [Independent Inverse Propensity Scoring (IIPS)](https://arxiv.org/abs/1804.10488)\n  - [Reward Interaction Inverse Propensity Scoring (RIPS)](https://arxiv.org/abs/2007)\n  - Cascade Doubly Robust (Cascade-DR)\n- OPE of Offline Bandit Algorithms with Continuous Actions\n  - [Kernelized Inverse Probability Weighting](https://arxiv.org/abs/1802.06037)\n  - [Kernelized Self-Normalized Inverse Probability Weighting](https://arxiv.org/abs/1802.06037)\n  - [Kernelized Doubly Robust](https://arxiv.org/abs/1802.06037)\n\n\u003c/details\u003e\n\nPlease refer to Section 2 and the Appendix of the reference [paper](https://arxiv.org/abs/2008.07146) for the standard formulation of OPE and the definitions of a range of OPE estimators.\nNote that, in addition to the above algorithms and estimators, Open Bandit Pipeline provides flexible interfaces.\nTherefore, researchers can easily implement their own algorithms or estimators and evaluate them with our data and pipeline.\nMoreover, Open Bandit Pipeline provides an interface for handling real-world logged bandit data.\nThus, practitioners can combine their own real-world data with Open Bandit Pipeline and easily evaluate bandit algorithms' performance in their settings with OPE.\n\n\n# Installation\n\nYou can install OBP using Python's package manager `pip`.\n\n```\npip install obp\n```\n\nYou can also install OBP from source.\n```bash\ngit clone https://github.com/st-tech/zr-obp\ncd zr-obp\npython setup.py install\n```\n\nOpen Bandit Pipeline supports Python 3.7 or newer. See [pyproject.toml](./pyproject.toml) for other requirements.\n\n# Usage\n\n## Example with Synthetic Bandit Data\n\nHere is an example of conducting OPE of the performance of IPWLearner as an evaluation policy using Direct Method (DM), Inverse Probability Weighting (IPW), Doubly Robust (DR) as OPE estimators.\n\n```python\n# implementing OPE of the IPWLearner using synthetic bandit data\nfrom sklearn.linear_model import LogisticRegression\n# import open bandit pipeline (obp)\nfrom obp.dataset import SyntheticBanditDataset\nfrom obp.policy import IPWLearner\nfrom obp.ope import (\n    OffPolicyEvaluation,\n    RegressionModel,\n    InverseProbabilityWeighting as IPW,\n    DirectMethod as DM,\n    DoublyRobust as DR,\n)\n\n# (1) Generate Synthetic Bandit Data\ndataset = SyntheticBanditDataset(n_actions=10, reward_type=\"binary\")\nbandit_feedback_train = dataset.obtain_batch_bandit_feedback(n_rounds=1000)\nbandit_feedback_test = dataset.obtain_batch_bandit_feedback(n_rounds=1000)\n\n# (2) Off-Policy Learning\neval_policy = IPWLearner(n_actions=dataset.n_actions, base_classifier=LogisticRegression())\neval_policy.fit(\n    context=bandit_feedback_train[\"context\"],\n    action=bandit_feedback_train[\"action\"],\n    reward=bandit_feedback_train[\"reward\"],\n    pscore=bandit_feedback_train[\"pscore\"]\n)\naction_dist = eval_policy.predict(context=bandit_feedback_test[\"context\"])\n\n# (3) Off-Policy Evaluation\nregression_model = RegressionModel(\n    n_actions=dataset.n_actions,\n    base_model=LogisticRegression(),\n)\nestimated_rewards_by_reg_model = regression_model.fit_predict(\n    context=bandit_feedback_test[\"context\"],\n    action=bandit_feedback_test[\"action\"],\n    reward=bandit_feedback_test[\"reward\"],\n)\nope = OffPolicyEvaluation(\n    bandit_feedback=bandit_feedback_test,\n    ope_estimators=[IPW(), DM(), DR()]\n)\nope.visualize_off_policy_estimates(\n    action_dist=action_dist,\n    estimated_rewards_by_reg_model=estimated_rewards_by_reg_model,\n)\n```\n\n\u003cdiv align=\"center\"\u003e\u003cimg src=\"https://raw.githubusercontent.com/st-tech/zr-obp/master/images/ope_results_example.png\" width=\"60%\"/\u003e\u003c/div\u003e\n\u003cfigcaption\u003e\n\u003cp align=\"center\"\u003e\n  Performance of IPWLearner estimated by OPE\n\u003c/p\u003e\n\u003c/figcaption\u003e\n\n\nA formal quickstart example with synthetic bandit data is available at [examples/quickstart/synthetic.ipynb](./examples/quickstart/synthetic.ipynb). We also prepare a script to conduct the evaluation of OPE experiment with synthetic bandit data in [examples/synthetic](./examples/synthetic/).\n\n## Example with Multi-Class Classification Data\n\nResearchers often use multi-class classification data to evaluate the estimation accuracy of OPE estimators.\nOpen Bandit Pipeline facilitates this kind of OPE experiments with multi-class classification data as follows.\n\n```python\n# implementing an experiment to evaluate the accuracy of OPE using classification data\nfrom sklearn.datasets import load_digits\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.linear_model import LogisticRegression\n# import open bandit pipeline (obp)\nfrom obp.dataset import MultiClassToBanditReduction\nfrom obp.ope import OffPolicyEvaluation, InverseProbabilityWeighting as IPW\n\n# (1) Data Loading and Bandit Reduction\nX, y = load_digits(return_X_y=True)\ndataset = MultiClassToBanditReduction(X=X, y=y, base_classifier_b=LogisticRegression(random_state=12345))\ndataset.split_train_eval(eval_size=0.7, random_state=12345)\nbandit_feedback = dataset.obtain_batch_bandit_feedback(random_state=12345)\n\n# (2) Evaluation Policy Derivation\n# obtain action choice probabilities of an evaluation policy\naction_dist = dataset.obtain_action_dist_by_eval_policy(base_classifier_e=RandomForestClassifier(random_state=12345))\n# calculate the ground-truth performance of the evaluation policy\nground_truth = dataset.calc_ground_truth_policy_value(action_dist=action_dist)\nprint(ground_truth)\n0.9634340222575517\n\n# (3) Off-Policy Evaluation and Evaluation of OPE\nope = OffPolicyEvaluation(bandit_feedback=bandit_feedback, ope_estimators=[IPW()])\n# evaluate the estimation performance (accuracy) of IPW by the relative estimation error (relative-ee)\nrelative_estimation_errors = ope.evaluate_performance_of_estimators(\n        ground_truth_policy_value=ground_truth,\n        action_dist=action_dist,\n        metric=\"relative-ee\",\n)\nprint(relative_estimation_errors)\n{'ipw': 0.01827255896321327} # the accuracy of IPW in OPE\n```\n\nA formal quickstart example with multi-class classification data is available at [examples/quickstart/multiclass.ipynb](./examples/quickstart/multiclass.ipynb).\nWe also prepare a script to conduct the evaluation of OPE experiment with multi-class classification data in [examples/multiclass](./examples/multiclass/).\n\n## Example with Open Bandit Dataset\n\nHere is an example of conducting OPE of the performance of BernoulliTS as an evaluation policy using Inverse Probability Weighting (IPW) and logged bandit data generated by the Random policy (behavior policy) on the ZOZOTOWN platform.\n\n```python\n# implementing OPE of the BernoulliTS policy using log data generated by the Random policy\nfrom obp.dataset import OpenBanditDataset\nfrom obp.policy import BernoulliTS\nfrom obp.ope import OffPolicyEvaluation, InverseProbabilityWeighting as IPW\n\n# (1) Data Loading and Preprocessing\ndataset = OpenBanditDataset(behavior_policy='random', campaign='all')\nbandit_feedback = dataset.obtain_batch_bandit_feedback()\n\n# (2) Production Policy Replication\nevaluation_policy = BernoulliTS(\n    n_actions=dataset.n_actions,\n    len_list=dataset.len_list,\n    is_zozotown_prior=True, # replicate the policy in the ZOZOTOWN production\n    campaign=\"all\",\n    random_state=12345\n)\naction_dist = evaluation_policy.compute_batch_action_dist(\n    n_sim=100000, n_rounds=bandit_feedback[\"n_rounds\"]\n)\n\n# (3) Off-Policy Evaluation\nope = OffPolicyEvaluation(bandit_feedback=bandit_feedback, ope_estimators=[IPW()])\nestimated_policy_value = ope.estimate_policy_values(action_dist=action_dist)\n\n# estimated performance of BernoulliTS relative to the ground-truth performance of Random\nrelative_policy_value_of_bernoulli_ts = estimated_policy_value['ipw'] / bandit_feedback['reward'].mean()\nprint(relative_policy_value_of_bernoulli_ts)\n1.198126...\n```\n\nA formal quickstart example with Open Bandit Dataset is available at [examples/quickstart/obd.ipynb](./examples/quickstart/obd.ipynb). We also prepare a script to conduct the evaluation of OPE using Open Bandit Dataset in [examples/obd](./examples/obd). Please see [our documentation](https://zr-obp.readthedocs.io/en/latest/evaluation_ope.html) for the details of the evaluation of OPE protocol based on Open Bandit Dataset.\n\n\n# Citation\nIf you use our dataset and pipeline in your work, please cite our paper:\n\nYuta Saito, Shunsuke Aihara, Megumi Matsutani, Yusuke Narita.\u003cbr\u003e\n**Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation**\u003cbr\u003e\n[https://arxiv.org/abs/2008.07146](https://arxiv.org/abs/2008.07146)\n\nBibtex:\n```\n@article{saito2020open,\n  title={Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation},\n  author={Saito, Yuta and Shunsuke, Aihara and Megumi, Matsutani and Yusuke, Narita},\n  journal={arXiv preprint arXiv:2008.07146},\n  year={2020}\n}\n```\n\nThe paper has been accepted at *NeurIPS2021 Datasets and Benchmarks Track*. The camera-ready version of the paper is available [here](https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/33e75ff09dd601bbe69f351039152189-Abstract-round2.html).\n\n# Sister Package: pyIEOE\n\nIn addition to OBP, we develop a Python package called [**pyIEOE**](https://github.com/sony/pyIEOE), which allows practitioners to easily evaluate and compare the robustness of OPE estimators.\n\nPlease also see the following reference paper about IEOE (accepted at RecSys'21).\n\nYuta Saito, Takuma Udagawa, Haruka Kiyohara, Kazuki Mogi, Yusuke Narita, Kei Tateno.\u003cbr\u003e\n**Evaluating the Robustness of Off-Policy Evaluation**\u003cbr\u003e\n[https://arxiv.org/abs/2108.13703](https://arxiv.org/abs/2108.13703)\n\n# Google Group\nIf you are interested in the Open Bandit Project, please follow its updates via the google group: https://groups.google.com/g/open-bandit-project\n\n# Contribution\nAny contributions to Open Bandit Pipeline are more than welcome!\nPlease refer to [CONTRIBUTING.md](./CONTRIBUTING.md) for general guidelines how to contribute to the project.\n\n# License\nThis project is licensed under the Apache 2.0 License - see the [LICENSE](LICENSE) file for details.\n\n# Project Team\n\n- [Yuta Saito](https://usait0.com/en/) (**Main Contributor**; Cornell University)\n- [Shunsuke Aihara](https://www.linkedin.com/in/shunsukeaihara/) (ZOZO Research)\n- Megumi Matsutani (ZOZO Research)\n- [Yusuke Narita](https://www.yusuke-narita.com/) (Hanjuku-kaso Co., Ltd. / Yale University)\n\n## Developers\n- [Masahiro Nomura](https://twitter.com/nomuramasahir0) (CyberAgent, Inc. / Hanjuku-kaso Co., Ltd.)\n- [Koichi Takayama](https://fullflu.hatenablog.com/) (Hanjuku-kaso Co., Ltd.)\n- [Ryo Kuroiwa](https://kurorororo.github.io) (University of Toronto / Hanjuku-kaso Co., Ltd.)\n- [Haruka Kiyohara](https://sites.google.com/view/harukakiyohara) (Tokyo Institute of Technology / Hanjuku-kaso Co., Ltd.)\n\n# Contact\nFor any question about the paper, data, and pipeline, feel free to contact: ys552@cornell.edu\n\n# References\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003ePapers \u003c/strong\u003e(click to expand)\u003c/summary\u003e\n\n1. Alina Beygelzimer and John Langford. [The offset tree for learning with partial labels](https://arxiv.org/abs/0812.4044). In *Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery\u0026Data Mining*, 129–138, 2009.\n\n2. Olivier Chapelle and Lihong Li. [An empirical evaluation of thompson sampling](https://papers.nips.cc/paper/4321-an-empirical-evaluation-of-thompson-sampling). In *Advances in Neural Information Processing Systems*, 2249–2257, 2011.\n\n3. Lihong Li, Wei Chu, John Langford, and Xuanhui Wang. [Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms](https://arxiv.org/abs/1003.5956). In *Proceedings of the Fourth ACM International Conference on Web Search and Data Mining*, 297–306, 2011.\n\n4. Alex Strehl, John Langford, Lihong Li, and Sham M Kakade. [Learning from Logged Implicit Exploration Data](https://arxiv.org/abs/1003.0120). In *Advances in Neural Information Processing Systems*, 2217–2225, 2010.\n\n5.  Doina Precup, Richard S. Sutton, and Satinder Singh. [Eligibility Traces for Off-Policy Policy Evaluation](https://scholarworks.umass.edu/cgi/viewcontent.cgi?article=1079\u0026context=cs_faculty_pubs). In *Proceedings of the 17th International Conference on Machine Learning*, 759–766. 2000.\n\n6.  Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li. [Doubly Robust Policy Evaluation and Optimization](https://arxiv.org/abs/1503.02834). *Statistical Science*, 29:485–511, 2014.\n\n7. Adith Swaminathan and Thorsten Joachims. [The Self-normalized Estimator for Counterfactual Learning](https://papers.nips.cc/paper/5748-the-self-normalized-estimator-for-counterfactual-learning). In *Advances in Neural Information Processing Systems*, 3231–3239, 2015.\n\n8. Dhruv Kumar Mahajan, Rajeev Rastogi, Charu Tiwari, and Adway Mitra. [LogUCB: An Explore-Exploit Algorithm for Comments Recommendation](https://dl.acm.org/doi/10.1145/2396761.2396767). In *Proceedings of the 21st ACM international conference on Information and knowledge management*, 6–15. 2012.\n\n9.  Lihong Li, Wei Chu, John Langford, Taesup Moon, and Xuanhui Wang. [An Unbiased Offline Evaluation of Contextual Bandit Algorithms with Generalized Linear Models](http://proceedings.mlr.press/v26/li12a.html). In *Journal of Machine Learning Research: Workshop and Conference Proceedings*, volume 26, 19–36. 2012.\n\n10. Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudik. [Optimal and Adaptive Off-policy Evaluation in Contextual Bandits](https://arxiv.org/abs/1612.01205). In *Proceedings of the 34th International Conference on Machine Learning*, 3589–3597. 2017.\n\n11. Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh. [More Robust Doubly Robust Off-policy Evaluation](https://arxiv.org/abs/1802.03493). In *Proceedings of the 35th International Conference on Machine Learning*, 1447–1456. 2018.\n\n12. Nathan Kallus and Masatoshi Uehara. [Intrinsically Efficient, Stable, and Bounded Off-Policy Evaluation for Reinforcement Learning](https://arxiv.org/abs/1906.03735). In *Advances in Neural Information Processing Systems*. 2019.\n\n13. Yi Su, Lequn Wang, Michele Santacatterina, and Thorsten Joachims. [CAB: Continuous Adaptive Blending Estimator for Policy Evaluation and Learning](https://proceedings.mlr.press/v97/su19a). In *Proceedings of the 36th International Conference on Machine Learning*, 6005-6014, 2019.\n\n14. Yi Su, Maria Dimakopoulou, Akshay Krishnamurthy, and Miroslav Dudík. [Doubly Robust Off-policy Evaluation with Shrinkage](https://proceedings.mlr.press/v119/su20a.html). In *Proceedings of the 37th International Conference on Machine Learning*, 9167-9176, 2020.\n\n15. Nathan Kallus and Angela Zhou. [Policy Evaluation and Optimization with Continuous Treatments](https://arxiv.org/abs/1802.06037). In *International Conference on Artificial Intelligence and Statistics*, 1243–1251. PMLR, 2018.\n\n16. Aman Agarwal, Soumya Basu, Tobias Schnabel, and Thorsten Joachims. [Effective Evaluation using Logged Bandit Feedback from Multiple Loggers](https://arxiv.org/abs/1703.06180). In *Proceedings of the 23rd ACM SIGKDD international conference on Knowledge discovery and data mining*, 687–696, 2017.\n\n17. Nathan Kallus, Yuta Saito, and Masatoshi Uehara. [Optimal Off-Policy Evaluation from Multiple Logging Policies](http://proceedings.mlr.press/v139/kallus21a.html). In *Proceedings of the 38th International Conference on Machine Learning*, 5247-5256, 2021.\n\n18. Shuai Li, Yasin Abbasi-Yadkori, Branislav Kveton, S Muthukrishnan, Vishwa Vinay, and Zheng Wen. [Offline Evaluation of Ranking Policies with Click Models](https://arxiv.org/pdf/1804.10488). In *Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery\u0026Data Mining*, 1685–1694, 2018.\n\n19. James McInerney, Brian Brost, Praveen Chandar, Rishabh Mehrotra, and Benjamin Carterette. [Counterfactual Evaluation of Slate Recommendations with Sequential Reward Interactions](https://arxiv.org/abs/2007.12986). In *Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery\u0026Data Mining*, 1779–1788, 2020.\n\n20. Yusuke Narita, Shota Yasui, and Kohei Yata. [Debiased Off-Policy Evaluation for Recommendation Systems](https://dl.acm.org/doi/10.1145/3460231.3474231). In *Proceedings of the Fifteenth ACM Conference on Recommender Systems*, 372-379, 2021.\n\n21. Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. [Open Graph Benchmark: Datasets for Machine Learning on Graphs](https://arxiv.org/abs/2005.00687). In *Advances in Neural Information Processing Systems*. 2020.\n\n22. Noveen Sachdeva, Yi Su, and Thorsten Joachims. [Off-policy Bandits with Deficient Support](https://dl.acm.org/doi/10.1145/3394486.3403139). In *Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery \u0026 Data Mining*, 965-975, 2021.\n\n23. Yi Su, Pavithra Srinath, and Akshay Krishnamurthy. [Adaptive Estimator Selection for Off-Policy Evaluation](https://proceedings.mlr.press/v119/su20d.html). In *Proceedings of the 38th International Conference on Machine Learning*, 9196-9205, 2021.\n\n24. Haruka Kiyohara, Yuta Saito, Tatsuya Matsuhiro, Yusuke Narita, Nobuyuki Shimizu, Yasuo Yamamoto. [Doubly Robust Off-Policy Evaluation for Ranking Policies under the Cascade Behavior Model](https://dl.acm.org/doi/10.1145/3488560.3498380). In *Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining*, 487-497, 2022.\n\n25. Yuta Saito and Thorsten Joachims. [Off-Policy Evaluation for Large Action Spaces via Embeddings](https://arxiv.org/abs/2202.06317). In *Proceedings of the 39th International Conference on Machine Learning*, 2022.\n\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003eProjects \u003c/strong\u003e(click to expand)\u003c/summary\u003e\n\n\u003cbr\u003e\n\nThe Open Bandit Project is strongly inspired by **Open Graph Benchmark** --a collection of benchmark datasets, data loaders, and evaluators for graph machine learning:\n[[github](https://github.com/snap-stanford/ogb)] [[project page](https://ogb.stanford.edu)] [[paper](https://arxiv.org/abs/2005.00687)].\n\n\u003c/details\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fst-tech%2Fzr-obp","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fst-tech%2Fzr-obp","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fst-tech%2Fzr-obp/lists"}