{"id":24513642,"url":"https://github.com/andrewtavis/causeinfer","last_synced_at":"2025-09-11T13:40:40.935Z","repository":{"id":52381461,"uuid":"222944517","full_name":"andrewtavis/causeinfer","owner":"andrewtavis","description":"Machine learning based causal inference/uplift in Python","archived":false,"fork":false,"pushed_at":"2023-11-24T20:01:00.000Z","size":15037,"stargazers_count":61,"open_issues_count":5,"forks_count":11,"subscribers_count":6,"default_branch":"main","last_synced_at":"2025-09-08T20:56:29.112Z","etag":null,"topics":["ab-testing","causal-inference","causality","data-analysis","data-science","data-visualization","dataset","econometrics","machine-learning","open-source","python","python3","statistics","treatment-effects","uplift","uplift-modeling"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"bsd-3-clause","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/andrewtavis.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE.txt","code_of_conduct":".github/CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2019-11-20T13:31:54.000Z","updated_at":"2025-06-26T22:08:51.000Z","dependencies_parsed_at":"2025-01-22T00:55:50.764Z","dependency_job_id":"f07ad1d2-a27d-40ac-95ad-8e638b0042d4","html_url":"https://github.com/andrewtavis/causeinfer","commit_stats":null,"previous_names":[],"tags_count":4,"template":false,"template_full_name":null,"purl":"pkg:github/andrewtavis/causeinfer","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/andrewtavis%2Fcauseinfer","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/andrewtavis%2Fcauseinfer/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/andrewtavis%2Fcauseinfer/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/andrewtavis%2Fcauseinfer/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/andrewtavis","download_url":"https://codeload.github.com/andrewtavis/causeinfer/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/andrewtavis%2Fcauseinfer/sbom","scorecard":{"id":194179,"data":{"date":"2025-08-11","repo":{"name":"github.com/andrewtavis/causeinfer","commit":"ecbe82cc6d62881d0b648b506cb408f300444a4a"},"scorecard":{"version":"v5.2.1-40-gf6ed084d","commit":"f6ed084d17c9236477efd66e5b258b9d4cc7b389"},"score":2.6,"checks":[{"name":"Code-Review","score":0,"reason":"Found 0/29 approved changesets -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project requires human code review before pull requests (aka merge requests) are merged.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#code-review"}},{"name":"Dangerous-Workflow","score":10,"reason":"no dangerous workflow patterns detected","details":null,"documentation":{"short":"Determines if the project's GitHub Action workflows avoid dangerous patterns.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#dangerous-workflow"}},{"name":"Packaging","score":-1,"reason":"packaging workflow not detected","details":["Warn: no GitHub/GitLab publishing workflow detected."],"documentation":{"short":"Determines if the project is published as a package that others can easily download, install, easily update, and uninstall.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#packaging"}},{"name":"Token-Permissions","score":0,"reason":"detected GitHub workflow tokens with excessive permissions","details":["Warn: no topLevel permission defined: .github/workflows/ci.yml:1","Info: no jobLevel write permissions found"],"documentation":{"short":"Determines if the project's workflows follow the principle of least privilege.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#token-permissions"}},{"name":"Maintained","score":0,"reason":"0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project is \"actively maintained\".","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#maintained"}},{"name":"Binary-Artifacts","score":10,"reason":"no binaries found in the repo","details":null,"documentation":{"short":"Determines if the project has generated executable (binary) artifacts in the source repository.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#binary-artifacts"}},{"name":"CII-Best-Practices","score":0,"reason":"no effort to earn an OpenSSF best practices badge detected","details":null,"documentation":{"short":"Determines if the project has an OpenSSF (formerly CII) Best Practices Badge.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#cii-best-practices"}},{"name":"Pinned-Dependencies","score":0,"reason":"dependency not pinned by hash detected -- score normalized to 0","details":["Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/ci.yml:22: update your workflow using https://app.stepsecurity.io/secureworkflow/andrewtavis/causeinfer/ci.yml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/ci.yml:24: update your workflow using https://app.stepsecurity.io/secureworkflow/andrewtavis/causeinfer/ci.yml/main?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/ci.yml:28: update your workflow using https://app.stepsecurity.io/secureworkflow/andrewtavis/causeinfer/ci.yml/main?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/ci.yml:46: update your workflow using https://app.stepsecurity.io/secureworkflow/andrewtavis/causeinfer/ci.yml/main?enable=pin","Info:   0 out of   2 GitHub-owned GitHubAction dependencies pinned","Info:   0 out of   2 third-party GitHubAction dependencies pinned"],"documentation":{"short":"Determines if the project has declared and pinned the dependencies of its build process.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#pinned-dependencies"}},{"name":"Security-Policy","score":0,"reason":"security policy file not detected","details":["Warn: no security policy file detected","Warn: no security file to analyze","Warn: no security file to analyze","Warn: no security file to analyze"],"documentation":{"short":"Determines if the project has published a security policy.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#security-policy"}},{"name":"Fuzzing","score":0,"reason":"project is not fuzzed","details":["Warn: no fuzzer integrations found"],"documentation":{"short":"Determines if the project uses fuzzing.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#fuzzing"}},{"name":"License","score":10,"reason":"license file detected","details":["Info: project has a license file: LICENSE.txt:0","Info: FSF or OSI recognized license: BSD 3-Clause \"New\" or \"Revised\" License: LICENSE.txt:0"],"documentation":{"short":"Determines if the project has defined a license.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#license"}},{"name":"Branch-Protection","score":0,"reason":"branch protection not enabled on development/release branches","details":["Warn: branch protection not enabled for branch 'main'"],"documentation":{"short":"Determines if the default and release branches are protected with GitHub's branch protection settings.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#branch-protection"}},{"name":"Signed-Releases","score":-1,"reason":"no releases found","details":null,"documentation":{"short":"Determines if the project cryptographically signs release artifacts.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#signed-releases"}},{"name":"Vulnerabilities","score":1,"reason":"9 existing vulnerabilities detected","details":["Warn: Project is vulnerable to: PYSEC-2024-48 / GHSA-fj7x-q9j7-g6q6","Warn: Project is vulnerable to: GHSA-6p56-wp2h-9hxr","Warn: Project is vulnerable to: GHSA-fpfv-jqm9-f5jm","Warn: Project is vulnerable to: GHSA-9hjg-9r4m-mvj7","Warn: Project is vulnerable to: GHSA-9wx4-h78v-vm56","Warn: Project is vulnerable to: PYSEC-2023-74 / GHSA-j8r2-6x86-q33q","Warn: Project is vulnerable to: PYSEC-2024-110 / GHSA-jw8x-6495-233v","Warn: Project is vulnerable to: GHSA-jxfp-4rvq-9h9m","Warn: Project is vulnerable to: GHSA-g7vv-2v7x-gj9p"],"documentation":{"short":"Determines if the project has open, known unfixed vulnerabilities.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#vulnerabilities"}},{"name":"SAST","score":0,"reason":"SAST tool is not run on all commits -- score normalized to 0","details":["Warn: 0 commits out of 2 are checked with a SAST tool"],"documentation":{"short":"Determines if the project uses static code analysis.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#sast"}}]},"last_synced_at":"2025-08-16T21:29:32.714Z","repository_id":52381461,"created_at":"2025-08-16T21:29:32.714Z","updated_at":"2025-08-16T21:29:32.714Z"},"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":274450820,"owners_count":25287532,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-09-10T02:00:12.551Z","response_time":83,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ab-testing","causal-inference","causality","data-analysis","data-science","data-visualization","dataset","econometrics","machine-learning","open-source","python","python3","statistics","treatment-effects","uplift","uplift-modeling"],"created_at":"2025-01-22T00:55:48.057Z","updated_at":"2025-09-11T13:40:40.914Z","avatar_url":"https://github.com/andrewtavis.png","language":"Python","funding_links":[],"categories":["🚀 GitHub Repositories"],"sub_categories":["🌟 **Real-World Magic**"],"readme":"\u003cdiv align=\"center\"\u003e\n  \u003ca href=\"https://github.com/andrewtavis/causeinfer\"\u003e\u003cimg src=\"https://raw.githubusercontent.com/andrewtavis/causeinfer/main/.github/resources/logo/causeinfer_logo_transparent.png\" width=612 height=164\u003e\u003c/a\u003e\n\u003c/div\u003e\n\n\u003col\u003e\u003c/ol\u003e\n\n[![rtd](https://img.shields.io/readthedocs/causeinfer.svg?logo=read-the-docs)](http://causeinfer.readthedocs.io/en/latest/)\n[![ci](https://img.shields.io/github/actions/workflow/status/andrewtavis/causeinfer/.github/workflows/ci.yml?branch=main?logo=github)](https://github.com/andrewtavis/causeinfer/actions?query=workflow%3ACI)\n[![codecov](https://codecov.io/gh/andrewtavis/causeinfer/branch/main/graphs/badge.svg)](https://codecov.io/gh/andrewtavis/causeinfer)\n[![pyversions](https://img.shields.io/pypi/pyversions/causeinfer.svg?logo=python\u0026logoColor=FFD43B\u0026color=306998)](https://pypi.org/project/causeinfer/)\n[![pypi](https://img.shields.io/pypi/v/causeinfer.svg?color=4B8BBE)](https://pypi.org/project/causeinfer/)\n[![pypistatus](https://img.shields.io/pypi/status/causeinfer.svg)](https://pypi.org/project/causeinfer/)\n[![license](https://img.shields.io/github/license/andrewtavis/causeinfer.svg)](https://github.com/andrewtavis/causeinfer/blob/main/LICENSE.txt)\n[![coc](https://img.shields.io/badge/coc-Contributor%20Covenant-ff69b4.svg)](https://github.com/andrewtavis/causeinfer/blob/main/.github/CODE_OF_CONDUCT.md)\n[![codestyle](https://img.shields.io/badge/code%20style-black-000000.svg)](https://github.com/psf/black)\n[![colab](https://img.shields.io/badge/%20-Open%20in%20Colab-097ABB.svg?logo=google-colab\u0026color=097ABB\u0026labelColor=525252)](https://colab.research.google.com/github/andrewtavis/causeinfer)\n\n## Machine learning based causal inference/uplift in Python\n\n**causeinfer** is a Python package for estimating average and conditional average treatment effects using machine learning. The goal is to compile causal inference models both standard and advanced, as well as demonstrate their usage and efficacy - all this with the overarching ambition to help people learn causal inference techniques across business, medical, and socioeconomic fields. See the [documentation](https://causeinfer.readthedocs.io/en/latest/index.html) for a full outline of the package including the available models and datasets.\n\n\u003ca id=\"contents\"\u003e\u003c/a\u003e\n\n# **Contents**\n\n- [Installation](#installation)\n- [Application](#application)\n  - [Two Model Approach](#two-model-approach)\n  - [Interaction Term Approach](#interaction-term-approach)\n  - [Class Transformation Approaches](#class-transformation-approaches)\n  - [Reflective and Pessimistic Uplift](#reflective-and-pessimistic-uplift)\n- [Evaluation Methods](#evaluation-methods)\n  - [Visualization](#visualization)\n  - [Model Iteration](#model-iteration)\n- [Data and Examples](#data-and-examples)\n  - [Business Analytics](#business-analytics)\n  - [Medical Trials](#medical-trials)\n  - [Socioeconomic Analysis](#socioeconomic-analysis)\n- [To-Do](#to-do)\n- [References](#references)\n\n\u003ca id=\"installation\"\u003e\u003c/a\u003e\n\n# Installation [`⇧`](#contents)\n\ncauseinfer can be downloaded from PyPI via pip or sourced directly from this repository:\n\n```bash\npip install causeinfer\n```\n\n```bash\ngit clone https://github.com/andrewtavis/causeinfer.git\ncd causeinfer\npython setup.py install\n```\n\n```python\nimport causeinfer\n```\n\n\u003ca id=\"application\"\u003e\u003c/a\u003e\n\n# Application [`⇧`](#contents)\n\n## Standard Algorithms\n\n\u003ca id=\"two-model-approach\"\u003e\u003c/a\u003e\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eTwo Model Approach\u003c/strong\u003e\u003c/summary\u003e\n\u003c/p\u003e\n\nSeparate models for treatment and control groups are trained and combined to derive average treatment effects (Hansotia, 2002).\n\n```python\nfrom causeinfer.standard_algorithms.two_model import TwoModel\nfrom sklearn.ensemble import RandomForestClassifier, RandomForestRegressor\n\ntm_pred = TwoModel(\n    treatment_model=RandomForestRegressor(**kwargs),\n    control_model=RandomForestRegressor(**kwargs),\n)\ntm_pred.fit(X=X_train, y=y_train, w=w_train)\n\n# An array of predictions given a treatment and control model\ntm_preds = tm_pred.predict(X=X_test)\n\ntm_proba = TwoModel(\n    treatment_model=RandomForestClassifier(**kwargs),\n    control_model=RandomForestClassifier(**kwargs),\n)\ntm_proba.fit(X=X_train, y=y_train, w=w_train)\n\n# An array of predicted treatment class probabilities given models\ntm_probas = tm.predict_proba(X=X_test)\n```\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003ca id=\"interaction-term-approach\"\u003e\u003c/a\u003e\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eInteraction Term Approach\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\nAn interaction term between treatment and covariates is added to the data to allow for a basic single model application (Lo, 2002).\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"https://raw.githubusercontent.com/andrewtavis/causeinfer/main/.github/resources/images/interaction_term_data.png\" width=\"720\" height=\"282\"\u003e\n\u003c/div\u003e\n\n```python\nfrom causeinfer.standard_algorithms.interaction_term import InteractionTerm\nfrom sklearn.ensemble import RandomForestClassifier, RandomForestRegressor\n\nit_pred = InteractionTerm(model=RandomForestRegressor(**kwargs))\nit_pred.fit(X=X_train, y=y_train, w=w_train)\n\n# An array of predictions given a treatment and control interaction term\nit_preds = it_pred.predict(X=X_test)\n\nit_proba = InteractionTerm(model=RandomForestClassifier(**kwargs))\nit_proba.fit(X=X_train, y=y_train, w=w_train)\n\n# An array of predicted treatment class probabilities given interaction terms\nit_probas = it_proba.predict_proba(X=X_test)\n```\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003ca id=\"class-transformation-approaches\"\u003e\u003c/a\u003e\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eClass Transformation Approaches\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\nUnits are categorized into two or four classes to derive treatment effects from favorable class attributes (Lai, 2006; Kane, et al, 2014; Shaar, et al, 2016).\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"https://raw.githubusercontent.com/andrewtavis/causeinfer/main/.github/resources/images/new_known_unknown_classes.png\" width=\"720\" height=\"405\"\u003e\n\u003c/div\u003e\n\n```python\n# Binary Class Transformation\nfrom causeinfer.standard_algorithms.binary_transformation import BinaryTransformation\nfrom sklearn.ensemble import RandomForestClassifier\n\nbt = BinaryTransformation(model=RandomForestClassifier(**kwargs), regularize=True)\nbt.fit(X=X_train, y=y_train, w=w_train)\n\n# An array of predicted probabilities (P(Favorable Class), P(Unfavorable Class))\nbt_probas = bt.predict_proba(X=X_test)\n```\n\n```python\n# Quaternary Class Transformation\nfrom causeinfer.standard_algorithms.quaternary_transformation import (\n    QuaternaryTransformation,\n)\nfrom sklearn.ensemble import RandomForestClassifier\n\nqt = QuaternaryTransformation(model=RandomForestClassifier(**kwargs), regularize=True)\nqt.fit(X=X_train, y=y_train, w=w_train)\n\n# An array of predicted probabilities (P(Favorable Class), P(Unfavorable Class))\nqt_probas = qt.predict_proba(X=X_test)\n```\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003ca id=\"reflective-and-pessimistic-uplift\"\u003e\u003c/a\u003e\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eReflective and Pessimistic Uplift\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\nWeighted versions of the binary class transformation approach that are meant to dampen the original model's inherently noisy results (Shaar, et al, 2016).\n\n```python\n# Reflective Uplift Transformation\nfrom causeinfer.standard_algorithms.reflective import ReflectiveUplift\nfrom sklearn.ensemble import RandomForestClassifier\n\nru = ReflectiveUplift(model=RandomForestClassifier(**kwargs))\nru.fit(X=X_train, y=y_train, w=w_train)\n\n# An array of predicted probabilities (P(Favorable Class), P(Unfavorable Class))\nru_probas = ru.predict_proba(X=X_test)\n```\n\n```python\n# Pessimistic Uplift Transformation\nfrom causeinfer.standard_algorithms.pessimistic import PessimisticUplift\nfrom sklearn.ensemble import RandomForestClassifier\n\npu = PessimisticUplift(model=RandomForestClassifier(**kwargs))\npu.fit(X=X_train, y=y_train, w=w_train)\n\n# An array of predicted probabilities (P(Favorable Class), P(Unfavorable Class))\npu_probas = pu.predict_proba(X=X_test)\n```\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n## Advanced Algorithms\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eModels to Consider\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\n- Under consideration for inclusion in causeinfer:\n  - Generalized Random Forest via the R/C++ [grf](https://github.com/grf-labs/grf) - Athey, Tibshirani, and Wager (2019)\n  - The X-Learner - Kunzel, et al (2019)\n  - The R-Learner - Nie and Wager (2017)\n  - Double Machine Learning - Chernozhukov, et al (2018)\n  - Information Theory Trees/Forests - Soltys, et al (2015)\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003ca id=\"evaluation-methods\"\u003e\u003c/a\u003e\n\n# Evaluation Methods [`⇧`](#contents)\n\n\u003ca id=\"visualization\"\u003e\u003c/a\u003e\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eVisualization Metrics and Coefficients\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\nComparisons across stratified, ordered treatment response groups are used to derive model efficiency.\n\n```python\nfrom causeinfer.evaluation import plot_cum_gain, plot_qini\n\nvisual_eval_dict = {\n    \"y_test\": y_test,\n    \"w_test\": w_test,\n    \"two_model\": tm_effects,\n    \"interaction_term\": it_effects,\n    \"binary_trans\": bt_effects,\n    \"quaternary_trans\": qt_effects,\n}\n\ndf_visual_eval = pd.DataFrame(visual_eval_dict, columns=visual_eval_dict.keys())\nmodel_pred_cols = [\n    col for col in visual_eval_dict.keys() if col not in [\"y_test\", \"w_test\"]\n]\n```\n\n```python\nfig, (ax1, ax2) = plt.subplots(ncols=2, sharey=False, figsize=(20, 5))\n\nplot_cum_effect(\n    df=df_visual_eval,\n    n=100,\n    models=models,\n    percent_of_pop=True,\n    outcome_col=\"y_test\",\n    treatment_col=\"w_test\",\n    normalize=True,\n    random_seed=42,\n    axis=ax1,\n    legend_metrics=True,\n)\n\nplot_qini(  # or plot_cum_gain\n    df=df_visual_eval,\n    n=100,\n    models=models,\n    percent_of_pop=True,\n    outcome_col=\"y_test\",\n    treatment_col=\"w_test\",\n    normalize=True,\n    random_seed=42,\n    axis=ax2,\n    legend_metrics=True,\n)\n```\n\nHillstrom Metrics\n\n\u003cp align=\"middle\"\u003e\n  \u003cimg src=\"https://raw.githubusercontent.com/andrewtavis/causeinfer/main/.github/resources/images/hillstrom_cum_effect.png\" width=\"400\" /\u003e\n  \u003cimg src=\"https://raw.githubusercontent.com/andrewtavis/causeinfer/main/.github/resources/images/hillstrom_qini.png\" width=\"400\" /\u003e\n\u003c/p\u003e\n\nMayo PBC Metrics\n\n\u003cp align=\"middle\"\u003e\n  \u003cimg src=\"https://raw.githubusercontent.com/andrewtavis/causeinfer/main/.github/resources/images/mayo_cum_effect.png\" width=\"400\" /\u003e\n  \u003cimg src=\"https://raw.githubusercontent.com/andrewtavis/causeinfer/main/.github/resources/images/mayo_auuc.png\" width=\"400\" /\u003e\n\u003c/p\u003e\n\nCMF Microfinance Metrics\n\n\u003cp align=\"middle\"\u003e\n  \u003cimg src=\"https://raw.githubusercontent.com/andrewtavis/causeinfer/main/.github/resources/images/cmf_cum_effect.png\" width=\"400\" /\u003e\n  \u003cimg src=\"https://raw.githubusercontent.com/andrewtavis/causeinfer/main/.github/resources/images/cmf_qini.png\" width=\"400\" /\u003e\n\u003c/p\u003e\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003ca id=\"model-iteration\"\u003e\u003c/a\u003e\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eIterated Model Variance Analysis\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\nEasily iterate models to derive their average effects and prediction variances. See a full example across all datasets and models in [examples/model_iteration](https://github.com/andrewtavis/causeinfer/blob/main/examples/model_iteration.ipynb), with the results being shown below:\n\n|                  | TwoModel               | InteractionTerm       | BinaryTransformation   | QuaternaryTransformation | ReflectiveUplift         | PessimisticUplift        |\n| :--------------- | :--------------------- | :-------------------- | :--------------------- | :----------------------- | :----------------------- | :----------------------- |\n| Hillstrom        | -5.4762 ± 13.589\\*\\*\\* | -5.047 ± 15.417\\*\\*\\* | 0.5178 ± 15.7252\\*\\*\\* | 0.7397 ± 14.7509\\*\\*\\*   | 4.4872 ± 18.5918\\*\\*\\*\\* | -6.0052 ± 17.936\\*\\*\\*\\* |\n| Mayo PBC         | -0.145 ± 0.29          | -0.1335 ± 0.4471      | 0.5542 ± 0.4268        | 0.5315 ± 0.4424          | -0.8774 ± 0.233          | 0.1392 ± 0.3587          |\n| CMF Microfinance | 18.7289 ± 5.9138\\*\\*   | 17.0616 ± 6.6993\\*\\*  | nan                    | nan                      | nan                      | nan                      |\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003ca id=\"data-and-examples\"\u003e\u003c/a\u003e\n\n# Data and Examples [`⇧`](#contents)\n\n\u003ca id=\"business-analytics\"\u003e\u003c/a\u003e\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eBusiness Analytics\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\n- [Hillstrom Email Marketing](https://blog.minethatdata.com/2008/03/minethatdata-e-mail-analytics-and-data.html)\n  - Is directly downloaded and formatted with causeinfer (see [causeinfer.data.hillstrom](https://github.com/andrewtavis/causeinfer/blob/main/src/causeinfer/data/hillstrom.py))\n  - How to use this dataset is shown in [examples/business_hillstrom](https://github.com/andrewtavis/causeinfer/blob/main/examples/business_hillstrom.ipynb) and below\n\n```python\nfrom causeinfer.data import hillstrom\n\nhillstrom.download_hillstrom()\ndata_hillstrom = hillstrom.load_hillstrom(\n    user_file_path=\"datasets/hillstrom.csv\", format_covariates=True, normalize=True\n)\n\ndf = pd.DataFrame(\n    data_hillstrom[\"dataset_full\"], columns=data_hillstrom[\"dataset_full_names\"]\n)\n```\n\n- [Criteo Uplift](https://ailab.criteo.com/criteo-uplift-prediction-dataset/)\n  - Needed [(see issue)](https://github.com/andrewtavis/causeinfer/issues/18):\n    - Download and formatting script\n    - Example notebook\n    - Tests\n    - Documentation\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003ca id=\"medical-trials\"\u003e\u003c/a\u003e\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eMedical Trials\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\n- [Mayo Clinic PBC](https://www.mayo.edu/research/documents/pbchtml/DOC-10027635)\n  - Is directly downloaded and formatted with causeinfer (see [causeinfer.data.mayo_pbc](https://github.com/andrewtavis/causeinfer/blob/main/src/causeinfer/data/mayo_pbc.py))\n  - Also included in the [datasets directory](https://github.com/andrewtavis/causeinfer/blob/main/src/causeinfer/data/datasets) for direct download\n  - How to use this dataset is shown in [examples/medical_mayo_pbc](https://github.com/andrewtavis/causeinfer/blob/main/examples/medical_mayo_pbc.ipynb) and below\n\n```python\nfrom causeinfer.data import mayo_pbc\n\nmayo_pbc.download_mayo_pbc()\ndata_mayo_pbc = mayo_pbc.load_mayo_pbc(\n    user_file_path=\"datasets/mayo_pbc.text\", format_covariates=True, normalize=True\n)\n\ndf = pd.DataFrame(\n    data_mayo_pbc[\"dataset_full\"], columns=data_mayo_pbc[\"dataset_full_names\"]\n)\n```\n\n- [Pintilie Tamoxifen](https://onlinelibrary.wiley.com/doi/book/10.1002/9780470870709)\n  - Accompanied the linked text, but is now unavailable, so it is included in the [datasets directory](https://github.com/andrewtavis/causeinfer/blob/main/src/causeinfer/data/datasets) for direct download\n  - Needed [(see issue)](https://github.com/andrewtavis/causeinfer/issues/19):\n    - Formatting script\n    - Example notebook\n    - Tests\n    - Documentation\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003ca id=\"socioeconomic-analysis\"\u003e\u003c/a\u003e\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eSocioeconomic Analysis\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\n- [CMF Microfinance](https://www.aeaweb.org/articles?id=10.1257/app.20130533)\n  - Accompanied the linked text, but is now unavailable. It is included in the [datasets directory](https://github.com/andrewtavis/causeinfer/blob/main/src/causeinfer/data/datasets) for direct download\n  - Is formatted with causeinfer (see [causeinfer.data.cmf_micro](https://github.com/andrewtavis/causeinfer/blob/main/src/causeinfer/data/cmf_micro.py))\n  - How to use this dataset is shown in [examples/socioeconomic_cmf_micro](https://github.com/andrewtavis/causeinfer/blob/main/examples/socioeconomic_cmf_micro.ipynb) and below\n\n```python\nfrom causeinfer.data import cmf_micro\n\ndata_cmf_micro = cmf_micro.load_cmf_micro(\n    user_file_path=\"datasets/cmf_micro\", format_covariates=True, normalize=True\n)\n\ndf = pd.DataFrame(\n    data_cmf_micro[\"dataset_full\"], columns=data_cmf_micro[\"dataset_full_names\"]\n)\n```\n\n- [Lalonde Job Training](https://users.nber.org/~rdehejia/data/.nswdata2.html)\n  - Needed [(see issue)](https://github.com/andrewtavis/causeinfer/issues/20):\n    - Download and formatting script\n    - Example notebook\n    - Tests\n    - Documentation\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003ca id=\"to-do\"\u003e\u003c/a\u003e\n\n# To-Do [`⇧`](#contents)\n\nPlease see the [contribution guidelines](https://github.com/andrewtavis/causeinfer/blob/main/.github/CONTRIBUTING.md) if you are interested in contributing to this project. Work that is in progress or could be implemented includes:\n\n- Adding more baseline models and datasets [(see issues)](https://github.com/andrewtavis/causeinfer/issues)\n\n- Converting GRF files to Python and connecting them to the C++ boiler plate\n\n- Adding a data simulator [(see issue)](https://github.com/andrewtavis/causeinfer/issues/23)\n\n- Finding more causal inference datasets to be added [(see issue)](https://github.com/andrewtavis/causeinfer/issues/17)\n\n- Adding a `predict` method to [binary_transformation](https://github.com/andrewtavis/causeinfer/blob/main/src/causeinfer/standard_algorithms/binary_transformation.py) and [quaternary_transformation](https://github.com/andrewtavis/causeinfer/blob/main/src/causeinfer/standard_algorithms/quaternary_transformation.py)\n\n- Updating and refining the [documentation](https://causeinfer.readthedocs.io/en/latest/)\n\n- Improving [tests](https://github.com/andrewtavis/causeinfer/blob/main/tests) for greater [code coverage](https://codecov.io/gh/andrewtavis/causeinfer)\n\n- Improving [code quality](https://img.shields.io/codacy/grade/4ad05b30365d4097927d6f87ea273cf9?logo=codacy) by refactoring large functions and checking conventions\n\n# Similar Projects\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eSimilar packages and modules to causeinfer\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\n\u003cb\u003ePython\u003c/b\u003e\n\n- https://github.com/uber/causalml\n- https://github.com/Minyus/causallift\n- https://github.com/maks-sh/scikit-uplift\n- https://github.com/duketemon/pyuplift\n- https://github.com/microsoft/EconML\n- https://github.com/Microsoft/dowhy\n- https://github.com/wayfair/pylift/\n- https://github.com/jszymon/uplift_sklearn\n\n\u003cb\u003eOther Languages\u003c/b\u003e\n\n- https://github.com/grf-labs/grf (R/C++)\n- [https://github.com/soerenkuenzel/causalToolbox/X-Learner](https://github.com/soerenkuenzel/causalToolbox/blob/a06d81d74f4d575a8b34dc6b718db2778cfa0be9/R/XRF.R) (R/C++)\n- https://github.com/xnie/rlearner (R)\n\n\u003cb\u003eData and Misc\u003c/b\u003e\n\n- https://github.com/rguo12/awesome-causality-data\n- https://github.com/rguo12/awesome-causality-algorithms\n- https://github.com/zhaoxiliang/causalinference\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003ca id=\"references\"\u003e\u003c/a\u003e\n\n# References [`⇧`](#contents)\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eList of referenced codes\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\n- [pyuplift](https://github.com/duketemon/pyuplift) by [duketemon](https://github.com/duketemon) ([License](https://github.com/duketemon/pyuplift/blob/master/LICENSE))\n- [Causal ML](https://github.com/uber/causalml) by [Uber](https://github.com/uber) ([License](https://github.com/uber/causalml/blob/master/LICENSE))\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eList of theoretical references\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\n\u003cstrong\u003eBig Data and Machine Learning\u003c/strong\u003e\n\n- Athey, S. (2017). Beyond prediction: Using big data for policy problems. Science, Vol. 355, No. 6324, February 3, 2017, pp. 483-485.\n- Athey, S. \u0026 Imbens, G. (2015). Machine Learning Methods for Estimating Heterogeneous Causal Effects. Draft version submitted April 5th, 2015, arXiv:1504.01132v1, pp. 1-25.\n- Athey, S. \u0026 Imbens, G. (2019). Machine Learning Methods That Economists Should Know About. Annual Review of Economics, Vol. 11, August 2019, pp. 685-725.\n- Chernozhukov, V. et al. (2018). Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal, Vol. 21, No. 1, February 1, 2018, pp. C1–C68.\n- Mullainathan, S. \u0026 Spiess, J. (2017). Machine Learning: An Applied Econometric Approach. Journal of Economic Perspectives, Vol. 31, No. 2, Spring 2017, pp. 87-106.\n\n\u003cstrong\u003eCausal Inference\u003c/strong\u003e\n\n- Athey, S. \u0026 Imbens, G. (2017). The State of Applied Econometrics: Causality and Policy Evaluation. Journal of Economic Perspectives, Vol. 31, No. 2, Spring 2017, pp. 3-32.\n- Athey, S., Tibshirani, J. \u0026 Wager, S. (2019) Generalized random forests. The Annals of Statistics, Vol. 47, No. 2 (2019), pp. 1148-1178.\n- Athey, S. \u0026 Wager, S. (2019). Efficient Policy Learning. Draft version submitted on 9 Feb 2017, last revised 16 Sep 2019, arXiv:1702.02896v5, pp. 1-10.\n- Banerjee, A, et al. (2015) The Miracle of Microfinance? Evidence from a Randomized Evaluation. American Economic Journal: Applied Economics, Vol. 7, No. 1, January 1, 2015, pp. 22-53.\n- Ding, P. \u0026 Li, F. (2018). Causal Inference: A Missing Data Perspective. Statistical Science, Vol. 33, No. 2, 2018, pp. 214-237.\n- Farrell, M., Liang, T. \u0026 Misra S. (2018). Deep Neural Networks for Estimation and Inference: Application to Causal Effects and Other Semiparametric Estimands. Draft version submitted December 2018, arXiv:1809.09953, pp. 1-54.\n- Gutierrez, P. \u0026 Gérardy, JY. (2016). Causal Inference and Uplift Modeling: A review of the literature. JMLR: Workshop and Conference Proceedings 67, 2016, pp. 1–14.\n- Hitsch, G J. \u0026 Misra, S. (2018). Heterogeneous Treatment Effects and Optimal Targeting Policy Evaluation. January 28, 2018, Available at SSRN: ssrn.com/abstract=3111957 or dx.doi.org/10.2139/ssrn.3111957, pp. 1-64.\n- Powers, S. et al. (2018). Some methods for heterogeneous treatment effect estimation in high dimensions. Statistics in Medicine, Vol. 37, No. 11, May 20, 2018, pp. 1767-1787.\n- Rosenbaum, P. \u0026 Rubin, D. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika, Vol. 70, pp. 41-55.\n- Sekhon, J. (2007). The Neyman-Rubin Model of Causal Inference and Estimation via Matching Methods. The Oxford Handbook of Political Methodology, Winter 2017, pp. 1-46.\n- Wager, S. \u0026 Athey, S. (2018). Estimation and Inference of Heterogeneous Treatment Effects using Random Forests. Journal of the American Statistical Association, Vol. 113, 2018 - Issue 523, pp. 1228-1242.\n\n\u003cstrong\u003eUplift\u003c/strong\u003e\n\n- Devriendt, F. et al. (2018). A Literature Survey and Experimental Evaluation of the State-of-the-Art in Uplift Modeling: A Stepping Stone Toward the Development of Prescriptive Analytics. Big Data, Vol. 6, No. 1, March 1, 2018, pp. 1-29. Codes found at: data-lab.be/downloads.php.\n- Hansotia, B. \u0026 Rukstales, B. (2002). Incremental value modeling. Journal of Interactive Marketing, Vol. 16, No. 3, Summer 2002, pp. 35-46.\n- Haupt, J., Jacob, D., Gubela, R. \u0026 Lessmann, S. (2019). Affordable Uplift: Supervised Randomization in Controlled Experiments. Draft version submitted on October 1, 2019, arXiv:1910.00393v1, pp. 1-15.\n- Jaroszewicz, S. \u0026 Rzepakowski, P. (2014). Uplift modeling with survival data. Workshop on Health Informatics (HI-KDD) New York City, August 2014, pp. 1-8.\n- Jaśkowski, M. \u0026 Jaroszewicz, S. (2012). Uplift modeling for clinical trial data. In: ICML, 2012, Workshop on machine learning for clinical data analysis. Edinburgh, Scotland, June 2012, 1-8.\n- Kane, K., Lo, VSY. \u0026 Zheng, J. (2014). Mining for the truly responsive customers and prospects using true-lift modeling: Comparison of new and existing methods. Journal of Marketing Analytics, Vol. 2, No. 4, December 2014, pp 218–238.\n- Lai, L.Y.-T. (2006). Influential marketing: A new direct marketing strategy addressing the existence of voluntary buyers. Master of Science thesis, Simon Fraser University School of Computing Science, Burnaby, BC, Canada, pp. 1-68.\n- Lo, VSY. (2002). The true lift model: a novel data mining approach to response modeling in database marketing. SIGKDD Explor 4(2), pp. 78–86.\n- Lo, VSY. \u0026 Pachamanova, D. (2016). From predictive uplift modeling to prescriptive uplift analytics: A practical approach to treatment optimization while accounting for estimation risk. Journal of Marketing Analytics Vol. 3, No. 2, pp. 79–95.\n- Radcliffe N.J. \u0026 Surry, P.D. (1999). Differential response analysis: Modeling true response by isolating the effect of a single action. In Proceedings of Credit Scoring and Credit Control VI. Credit Research Centre, University of Edinburgh Management School.\n- Radcliffe N.J. \u0026 Surry, P.D. (2011). Real-World Uplift Modelling with Significance-Based Uplift Trees. Technical Report TR-2011-1, Stochastic Solutions, 2011, pp. 1-33.\n- Rzepakowski, P. \u0026 Jaroszewicz, S. (2012). Decision trees for uplift modeling with single and multiple treatments. Knowledge and Information Systems, Vol. 32, pp. 303–327.\n- Rzepakowski, P. \u0026 Jaroszewicz, S. (2012). Uplift modeling in direct marketing. Journal of Telecommunications and Information Technology, Vol. 2, 2012, pp. 43–50.\n- Rudaś, K. \u0026 Jaroszewicz, S. (2018). Linear regression for uplift modeling. Data Mining and Knowledge Discovery, Vol. 32, No. 5, September 2018, pp. 1275–1305.\n- Shaar, A., Abdessalem, T. and Segard, O (2016). “Pessimistic Uplift Modeling”. ACM SIGKDD, August 2016, San Francisco, California, USA.\n- Sołtys, M., Jaroszewicz, S. \u0026 Rzepakowski, P. (2015). Ensemble methods for uplift modeling. Data Mining and Knowledge Discovery, Vol. 29, No. 6, November 2015, pp. 1531–1559.\n\n\u003c/p\u003e\n\u003c/details\u003e\n\n\u003cdetails\u003e\u003csummary\u003e\u003cstrong\u003eList of data references\u003c/strong\u003e\u003c/summary\u003e\n\u003cp\u003e\n\n- Banerjee, A., Duflo, E., Glennerster, R., and Kinnan, C (2015). \"The Miracle of Microfinance? Evidence from a Randomized Evaluation.\" American Economic Journal: Applied Economics, 7 (1), pp. 22-53. URL: https://www.aeaweb.org/articles?id=10.1257/app.20130533.\n- K. Hillstrom. “The MineThatData E-Mail Analytics And Data Mining Challenge”. 2008. URL: https://blog.minethatdata.com/2008/03/minethatdata-e-mail-analytics-and-data.html.\n- Mayo Clinic. “Primary Biliary Cirrhosis”. 1991. URL: https://www.mayo.edu/research/documents/pbchtml/DOC-10027635.\n\n\u003c/p\u003e\n\u003c/details\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fandrewtavis%2Fcauseinfer","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fandrewtavis%2Fcauseinfer","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fandrewtavis%2Fcauseinfer/lists"}