{"id":31570441,"url":"https://github.com/singhsidhukuldeep/contextual-bandits","last_synced_at":"2026-03-27T03:10:04.587Z","repository":{"id":257308243,"uuid":"857876950","full_name":"singhsidhukuldeep/contextual-bandits","owner":"singhsidhukuldeep","description":"A comprehensive Python library implementing a variety of contextual and non-contextual multi-armed bandit algorithms—including LinUCB, Epsilon-Greedy, Upper Confidence Bound (UCB), Thompson Sampling, KernelUCB, NeuralLinearBandit, and DecisionTreeBandit—designed for reinforcement learning applications","archived":false,"fork":false,"pushed_at":"2024-12-31T17:41:56.000Z","size":91,"stargazers_count":13,"open_issues_count":0,"forks_count":1,"subscribers_count":1,"default_branch":"main","last_synced_at":"2026-01-06T16:34:05.329Z","etag":null,"topics":["algorithms","bandit-algorithms","contextual-bandits","epsilon-greedy","linucb","machine-learning","multi-armed-bandit","python","reinforcement-learning"],"latest_commit_sha":null,"homepage":"https://pypi.org/project/contextual-bandits-algos/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"gpl-3.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/singhsidhukuldeep.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-09-15T20:43:04.000Z","updated_at":"2025-12-08T06:41:12.000Z","dependencies_parsed_at":null,"dependency_job_id":"9b751600-184b-4397-847e-88a3adf663ab","html_url":"https://github.com/singhsidhukuldeep/contextual-bandits","commit_stats":null,"previous_names":["singhsidhukuldeep/contextual-bandits"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/singhsidhukuldeep/contextual-bandits","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/singhsidhukuldeep%2Fcontextual-bandits","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/singhsidhukuldeep%2Fcontextual-bandits/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/singhsidhukuldeep%2Fcontextual-bandits/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/singhsidhukuldeep%2Fcontextual-bandits/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/singhsidhukuldeep","download_url":"https://codeload.github.com/singhsidhukuldeep/contextual-bandits/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/singhsidhukuldeep%2Fcontextual-bandits/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":31013961,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-03-27T02:58:54.984Z","status":"ssl_error","status_checked_at":"2026-03-27T02:58:46.993Z","response_time":164,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["algorithms","bandit-algorithms","contextual-bandits","epsilon-greedy","linucb","machine-learning","multi-armed-bandit","python","reinforcement-learning"],"created_at":"2025-10-05T12:23:10.922Z","updated_at":"2026-03-27T03:10:04.582Z","avatar_url":"https://github.com/singhsidhukuldeep.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\n\u003ch1 align=\"center\"\u003eContextual Multi-Armed Bandits Library\u003c/h1\u003e\n\n\u003cp align=\"center\"\u003e\n\u003ca href=\"https://github.com/singhsidhukuldeep/contextual-bandits\"\u003e\u003cimg src=\"./Contextual Bandit Algorithms.png\" alt=\"contextual-bandits\" width =\"75%\" /\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\nA comprehensive Python library implementing a variety of contextual and non-contextual multi-armed bandit algorithms\u003cbr\u003e\n\u003ca href=\"https://pypi.org/project/contextual-bandits-algos/\"\u003e\u003cimg src=\"https://img.shields.io/pypi/pyversions/contextual-bandits-algos\" alt=\"Go to https://pypi.org/project/contextual-bandits-algos/\"/\u003e\u003c/a\u003e\n\u003ca href=\"https://pypi.org/project/contextual-bandits-algos/\"\u003e\u003cimg src=\"https://img.shields.io/pypi/v/contextual-bandits-algos\" alt=\"Go to https://pypi.org/project/contextual-bandits-algos/\"/\u003e\u003c/a\u003e\n\u003ca href=\"https://pypi.org/project/contextual-bandits-algos/\"\u003e\u003cimg src=\"https://img.shields.io/pypi/status/contextual-bandits-algos\" alt=\"Go to https://pypi.org/project/contextual-bandits-algos/\"/\u003e\u003c/a\u003e\n\u003c!-- \u003ca href=\"https://pypi.org/project/contextual-bandits-algos/\"\u003e\u003cimg src=\"https://img.shields.io/pypi/format/contextual-bandits-algos\" alt=\"Go to https://pypi.org/project/contextual-bandits-algos/\"/\u003e\u003c/a\u003e --\u003e\n\u003c!-- \u003ca href=\"https://lgtm.com/projects/g/singhsidhukuldeep/contextual-bandits-algos/context:python\"\u003e\u003cimg alt=\"Language grade: Python\" src=\"https://img.shields.io/lgtm/grade/python/g/singhsidhukuldeep/contextual-bandits-algos.svg?logo=lgtm\u0026logoWidth=18\"/\u003e\u003c/a\u003e --\u003e\n\u003ca href=\"https://pypistats.org/packages/contextual-bandits-algos\"\u003e\u003cimg src=\"https://img.shields.io/pypi/dm/contextual-bandits-algos\"/\u003e\u003c/a\u003e\n\u003c!-- \u003cimg src=\"https://visitor-badge.glitch.me/badge?page_id=request_boost\" alt=\"Go to https://pypi.org/project/contextual-bandits-algos/\"/\u003e --\u003e\n\u003cimg src=\"https://static.pepy.tech/personalized-badge/contextual-bandits-algos?period=total\u0026units=none\u0026left_color=black\u0026right_color=brightgreen\u0026left_text=Total%20Downloads\" alt=\"Go to https://pypi.org/project/contextual-bandits-algos/\"/\u003e\n\u003c/p\u003e\n\n## Overview\n\nA comprehensive Python library implementing a variety of contextual and non-contextual multi-armed bandit algorithms—including LinUCB, Epsilon-Greedy, Upper Confidence Bound (UCB), Thompson Sampling, KernelUCB, NeuralLinearBandit, and DecisionTreeBandit—designed for reinforcement learning applications that require decision-making under uncertainty with or without contextual information.\n\n\n\n## Features\n\n-   **Contextual Algorithms**:\n    -   **LinUCB**: Balances exploration and exploitation using linear regression with upper confidence bounds.\n    -   **Epsilon-Greedy**: Explores randomly with probability epsilon and exploits the best-known option otherwise.\n    -   **KernelUCB**: Uses kernel methods to capture non-linear relationships in the context space.\n    -   **NeuralLinearBandit**: Combines neural networks for feature extraction with linear models for prediction.\n    -   **DecisionTreeBandit**: Employs decision trees to model complex relationships between context and rewards.\n\n-   **Non-Contextual Algorithms**:\n    -   **Upper Confidence Bound (UCB)**: Selects arms based on upper confidence bounds of estimated rewards.\n    -   **Thompson Sampling**: Uses Bayesian methods to balance exploration and exploitation.\n\n\n## Installation\n\n```bash\npip install contextual-bandits-algos\n```\n\n### Instructions to Use the Updated Library\n\n\n1.  **Clone the Repository**\n\n    ```bash\n    git clone https://github.com/singhsidhukuldeep/contextual-bandits.git\n    cd contextual-bandits\n    ```\n\n2.  **Install the Requirements**\n\n    ```bash\n    pip install -r requirements.txt\n    ```\n\n3.  **Install the Package**\n\n    ```bash\n    pip install .\n    ```\n\n4.  **Run the Examples**\n\n    ```bash\n    python examples/example_linucb.py\n    python examples/example_epsilon_greedy.py\n    python examples/example_ucb.py\n    python examples/example_thompson_sampling.py\n    python examples/example_kernelucb.py\n    python examples/example_neurallinear.py\n    python examples/example_tree_bandit.py\n    ```\n\n    To run both algorithms in a single script:\n\n    ```bash\n    python examples/example_usage.py\n    ```\n\n    Run all examples *.py files\n    \n    ```bash\n    cd examples\n    for f in *.py; do echo \"$f\" \u0026 python \"$f\"; done\n    ```\n\n## Algorithms and Examples\n\n\nBelow are detailed descriptions of each algorithm, indicating whether they are contextual or non-contextual, how they work, and examples demonstrating their usage.\n\n### 1. LinUCB\n\n-   **Type**: Contextual\n-   **Description**: The LinUCB algorithm is a contextual bandit algorithm that uses linear regression to predict the expected reward for each arm given the current context. It balances exploration and exploitation by adding an upper confidence bound to the estimated rewards.\n\n#### How It Works\n\n-   **Model**: Assumes that the reward is a linear function of the context features.\n-   **Exploration**: Incorporates uncertainty in the estimation by adding a confidence interval (scaled by `alpha`).\n-   **Exploitation**: Chooses the arm with the highest upper confidence bound.\n\n\n\n### 2. Epsilon-Greedy\n\n-   **Type**: Contextual\n-   **Description**: The Epsilon-Greedy algorithm selects a random arm with probability `epsilon` (exploration) and the best-known arm with probability `1 - epsilon` (exploitation). It uses the context to predict rewards for each arm.\n\n#### How It Works\n\n-   **Model**: Uses linear models or other estimators to predict rewards based on context.\n-   **Exploration**: With probability `epsilon`, selects a random arm.\n-   **Exploitation**: With probability `1 - epsilon`, selects the arm with the highest predicted reward.\n\n\n\n### 3. Upper Confidence Bound (UCB)\n\n-   **Type**: Non-Contextual\n-   **Description**: The UCB algorithm selects arms based on upper confidence bounds of the estimated rewards, without considering any context. It is suitable when no contextual information is available.\n\n#### How It Works\n\n-   **Model**: Estimates the average reward for each arm.\n-   **Exploration**: Adds a confidence term to the average reward to explore less-tried arms.\n-   **Exploitation**: Chooses the arm with the highest upper confidence bound.\n\n\n\n### 4. Thompson Sampling\n\n-   **Type**: Non-Contextual\n-   **Description**: Thompson Sampling is a Bayesian algorithm that selects arms based on samples drawn from the posterior distributions of the arm's reward probabilities.\n\n#### How It Works\n\n-   **Model**: Assumes Bernoulli-distributed rewards for each arm.\n-   **Exploration \u0026 Exploitation**: Balances both by sampling from the posterior distributions.\n-   **Updates**: Updates the posterior distributions based on observed rewards.\n\n\n\n### 5. KernelUCB\n\n-   **Type**: Contextual\n-   **Description**: KernelUCB uses kernel methods to capture non-linear relationships between contexts and rewards. It extends the UCB algorithm to a kernelized context space.\n\n#### How It Works\n\n-   **Model**: Uses a kernel function (e.g., RBF kernel) to compute similarity between contexts.\n-   **Exploration**: Adds an exploration term based on the uncertainty in the kernel space.\n-   **Exploitation**: Predicts the expected reward using kernel regression.\n\n\n\n### 6. NeuralLinearBandit\n\n-   **Type**: Contextual\n-   **Description**: NeuralLinearBandit uses a neural network to learn a representation of the context and then applies linear regression on the learned features.\n\n#### How It Works\n\n-   **Model**: Combines a neural network for feature extraction with a linear model for reward prediction.\n-   **Exploration**: Adds an exploration bonus based on the uncertainty of the linear model.\n-   **Exploitation**: Uses the predicted rewards from the linear model.\n\n\n### 7. DecisionTreeBandit\n\n-   **Type**: Contextual\n-   **Description**: The DecisionTreeBandit algorithm uses decision trees to model the relationship between context and rewards, allowing it to capture non-linear patterns.\n\n#### How It Works\n\n-   **Model**: Fits a decision tree regressor for each arm based on the observed contexts and rewards.\n-   **Exploration**: Relies on the decision tree's predictions; may require additional mechanisms for exploration.\n-   **Exploitation**: Selects the arm with the highest predicted reward.\n\n\u003ch2 align=\"center\"\u003e🌟⭐✨STAR ME✨⭐🌟\u003c/h2\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cb\u003eYou can give me a small 🤓 dopmaine 🤝 support by ⭐STARRING⭐ this project\u003c/b\u003e\n  \n\u003cimg src=\"https://api.star-history.com/svg?repos=singhsidhukuldeep/contextual-bandits-algos\u0026type=Date\" width=\"70%\" alt=\"🌟⭐✨STAR ME✨⭐🌟\"\u003e\n\u003c/p\u003e\n\n## Credits\n\n### Maintained by\n\n***Kuldeep Singh Sidhu*** \n\nGithub: [github/singhsidhukuldeep](https://github.com/singhsidhukuldeep)\n`https://github.com/singhsidhukuldeep`\n\nWebsite: [Kuldeep Singh Sidhu (Website)](http://kuldeepsinghsidhu.com)\n`http://kuldeepsinghsidhu.com`\n\nLinkedIn: [Kuldeep Singh Sidhu (LinkedIn)](https://www.linkedin.com/in/singhsidhukuldeep/)\n`https://www.linkedin.com/in/singhsidhukuldeep/`","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsinghsidhukuldeep%2Fcontextual-bandits","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsinghsidhukuldeep%2Fcontextual-bandits","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsinghsidhukuldeep%2Fcontextual-bandits/lists"}