Projects in Awesome Lists tagged with thompson-sampling
A curated list of projects in awesome lists tagged with thompson-sampling .
https://github.com/sail-sg/oat
🌾 OAT: A research-friendly framework for LLM online alignment, including preference learning, reinforcement learning, etc.
alignment distributed-rl distributed-training dpo dueling-bandits grpo llm llm-aligment llm-exploration online-alignment online-rl ppo r1-zero reasoning rlhf thompson-sampling
Last synced: 08 May 2025
https://github.com/eric-bradford/ts-emo
This repository contains the source code for “Thompson sampling efficient multiobjective optimization” (TSEMO).
bayesian-optimization black-box-optimization expensive-to-evaluate-functions gaussian-processes genetic-algorithms kriging machine-learning matlab multi-objective-optimization spectral-sampling surrogate-based-optimization thompson-sampling
Last synced: 03 Sep 2025
https://github.com/stitchfix/mab
Library for multi-armed bandit selection strategies, including efficient deterministic implementations of Thompson sampling and epsilon-greedy.
data-science experimentation go golang multi-armed-bandit multi-armed-bandits multiarmed-bandits reinforcement-learning thompson thompson-sampling
Last synced: 16 Jul 2025
https://github.com/playtikaoss/pybandits
Python library for Multi-Armed Bandits
bayesian-neural-networks contextual-bandit-algorithms contextual-bandits multi-armed-bandit multi-armed-bandits multiarmed-bandits offline-policy-evaluation reinforcement-learning stochastic-bandit stochastic-bandit-algorithms thompson-sampling
Last synced: 12 Mar 2026
https://github.com/michaelosthege/pyrff
pyrff: Python implementation of random fourier feature approximations for gaussian processes
bayesian-optimization gaussian-processes thompson-sampling
Last synced: 21 Jun 2025
https://github.com/IgorGanapolsky/ThumbGate
Agent governance for ThumbGate: 👍/👎 become Pre-Action Checks that block repeat mistakes before code, money, or customer systems change.
agent-reliability ai-agents ai-cost-optimization ai-safety amp claude-code codex cursor developer-tools feedback-loop gemini guardrails mcp mcp-server opencode pre-action-checks reduce-llm-cost save-llm-tokens thompson-sampling thumbgate
Last synced: 12 Jun 2026
https://github.com/v-i-s-h/mab.jl
A Julia Package for providing Multi Armed Bandit Experiments
bandit-experiments exp julia julia-language julia-package julialang mab multi-arm-bandits reinforcement-learning reinforcement-learning-algorithms thompson-sampling ucb
Last synced: 01 May 2025
https://github.com/igorganapolsky/thumbgate
Agent governance for ThumbGate: 👍/👎 become Pre-Action Checks that block repeat mistakes before code, money, or customer systems change.
agent-reliability ai-agents ai-cost-optimization ai-safety amp claude-code codex cursor developer-tools feedback-loop gemini guardrails mcp mcp-server opencode pre-action-checks reduce-llm-cost save-llm-tokens thompson-sampling thumbgate
Last synced: 30 May 2026
https://github.com/lucko515/ads-strategy-reinforcement-learning
The example of using reinforcement learning algorithms in the business, specifically finding what ads to use in our campaign.
machine-learning reinforcement-learning thompson-sampling upper-confidence-bounds
Last synced: 24 Oct 2025
https://github.com/gjjvdburg/thompsonsampling
Source code for blog post on Thompson Sampling
bandit-algorithms multi-armed-bandit multiarmed-bandits thompson-sampling
Last synced: 27 Jun 2025
https://github.com/rueian/gobandit
A golang library for solving multi armed bandit problem which can optimize your business choice on the fly without A/B testing
enforcement golang multi-armed-bandit thompson-sampling
Last synced: 10 Apr 2025
https://github.com/volvo-cars/eene-nav-bandit-sim
EENE Navigation Bandit Simulator
bayes-ucb combinatorial-bandit machine-learning machine-learning-algorithms multi-armed-bandit navigation-algorithms python thompson-sampling
Last synced: 06 Oct 2025
https://github.com/mykeels/multi-armed-bandit-problem
An implementation of solvers for the multi-armed-bandit-problem in JavaScript.
epsilon-greedy multi-armed-bandit thompson-sampling ucb1
Last synced: 19 Jul 2026
https://github.com/vmarchaud/ts-mab
Typescript implementation of a multi-armed bandit
mab thompson-sampling typescript
Last synced: 13 Jun 2026
https://github.com/weill-labs/hgm
Clean-room reimplementation of the Huxley-Gödel Machine (arXiv:2510.21614): a self-improving coding agent with CMP + Thompson-Sampling tree search, validated in $0 simulation and live on SWE-bench (mini-swe-agent + gpt-5.4).
agentic-ai ai-agents coding-agent huxley-godel-machine llm mini-swe-agent research-reproduction self-improving-ai swe-bench thompson-sampling
Last synced: 29 Jun 2026
https://github.com/mefedursun/dynamic-pricing-engine
Research simulation comparing Thompson Sampling vs. Epsilon-Greedy algorithms in non-stationary hyper-inflationary markets. Reveals "Bayesian Inertia" phenomenon.
dynamic-pricing reinforcement-learning research thompson-sampling
Last synced: 25 Apr 2026
https://github.com/paramrathour/intelligent-and-learning-agents
My programs during CS747 (Foundations of Intelligent and Learning Agents) Autumn 2021-22
epsilon-greedy kl-ucb linear-programming markov-decision-processes mountain-car multi-armed-bandit policy-control policy-iteration sarsa thompson-sampling tile-coding ucb value-iteration
Last synced: 23 Mar 2025
https://github.com/howardleegeek/growth-os
Growth 2.0: an autonomous evolutionary loop that hacks Twitter's post-December-2025 ranking algorithm. Karpathy autoresearch × EvoHarness tree search × offline Twitter simulator × Grok opponent study. Encodes lessons from Oyster Labs' Growth 1.0 bootstrap — a hunger game of content strategies inside a simulation arena.
autonomous-agent autoresearch bootstrapping distribution-engineering evoharness evolutionary-algorithm growth-2-0 growth-engineering growth-hacking harness-engineering hunger-games karpathy open-source python simulation-arena thompson-sampling tree-search twitter-algorithm vibe-coding
Last synced: 05 Jul 2026
https://github.com/preferred-pictures/php
A PHP client for the PreferredPictures API.
ab-testing optimization php-client php-library thompson-sampling
Last synced: 17 Mar 2026
https://github.com/preferred-pictures/node
A Node.js client for PreferredPictures API.
ab-testing nodejs optimization thompson-sampling typescript
Last synced: 17 Mar 2026
https://github.com/thatguychandan/adoptimization
This project implements an ad optimization system using a hybrid approach combining Thompson Sampling and Upper Confidence Bound (UCB) algorithms. The system learns to select the most effective ads based on user context and historical performance.
numpy pandas plotly python pytorch reinforcement-learning scikit-learn streamlit thompson-sampling upper-confidence-bound
Last synced: 10 Apr 2026
https://github.com/aashish22bansal/best-ads-predictor
Predicting the best Ad from the given Ads.
reinforcement-learning thompson-sampling upper-confidence-bound
Last synced: 01 May 2026
https://github.com/nathanaelcheramlak/agent-psi
OpenPsi Experimentation Repository
Last synced: 17 Feb 2026
https://github.com/mohammadi-hadi/dynamic-pricing-dashboard
Interactive in-browser simulator: dynamic pricing with demand learning, forward-looking customers, and advertising
demand-learning dynamic-pricing revenue-management simulation streamlit thompson-sampling webassembly
Last synced: 17 Aug 2026
https://github.com/miningstore/vibe-x-agent
Self-improving X (Twitter) promotion agent for your VPS. A Claude-CLI content engine posts about your product across content angles; an engagement-driven Thompson-sampling bandit learns which angles land and tunes the mix. No database, runs on a $5 box. Modeled on vibe-seo-agent.
agent claude growth marketing-automation multi-armed-bandit thompson-sampling twitter-bot vps x-twitter
Last synced: 13 Jul 2026
https://github.com/abailey81/implicit-interaction-intelligence
Adaptive AI companion that builds a model of each user from implicit interaction signals — keystroke dynamics, linguistic complexity, temporal patterns — and continuously adapts its responses. Custom TCN + transformer + contextual bandit, built from scratch in PyTorch.
adaptive-systems contextual-bandit deep-learning edge-ai fastapi from-scratch huawei human-machine-interaction language-model machine-learning personalization poetry privacy-by-design python pytorch temporal-convolutional-network thompson-sampling transformer user-modeling websocket
Last synced: 18 Apr 2026
https://github.com/posgnu/bayesian-active-learning-on-multi-armed-bandit
Bayesian active learning algorithm with Thompson sampling on multi-armed bandit with Numpy
bayesian-active-learning multi-armed-bandits thompson-sampling
Last synced: 14 Apr 2025
https://github.com/featmate/thompsonsampling-orderrpc
汤普森采样的通用服务,用于从redis中获得目标物品的alpha,beta值,然后过beta分布随机出一个数值后做排序
microservice thompson-sampling
Last synced: 14 Jan 2026