An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with thompson-sampling

A curated list of projects in awesome lists tagged with thompson-sampling .

https://github.com/sail-sg/oat

🌾 OAT: A research-friendly framework for LLM online alignment, including preference learning, reinforcement learning, etc.

alignment distributed-rl distributed-training dpo dueling-bandits grpo llm llm-aligment llm-exploration online-alignment online-rl ppo r1-zero reasoning rlhf thompson-sampling

Last synced: 08 May 2025

https://github.com/stitchfix/mab

Library for multi-armed bandit selection strategies, including efficient deterministic implementations of Thompson sampling and epsilon-greedy.

data-science experimentation go golang multi-armed-bandit multi-armed-bandits multiarmed-bandits reinforcement-learning thompson thompson-sampling

Last synced: 16 Jul 2025

https://github.com/michaelosthege/pyrff

pyrff: Python implementation of random fourier feature approximations for gaussian processes

bayesian-optimization gaussian-processes thompson-sampling

Last synced: 21 Jun 2025

https://github.com/IgorGanapolsky/ThumbGate

Agent governance for ThumbGate: 👍/👎 become Pre-Action Checks that block repeat mistakes before code, money, or customer systems change.

agent-reliability ai-agents ai-cost-optimization ai-safety amp claude-code codex cursor developer-tools feedback-loop gemini guardrails mcp mcp-server opencode pre-action-checks reduce-llm-cost save-llm-tokens thompson-sampling thumbgate

Last synced: 12 Jun 2026

https://github.com/igorganapolsky/thumbgate

Agent governance for ThumbGate: 👍/👎 become Pre-Action Checks that block repeat mistakes before code, money, or customer systems change.

agent-reliability ai-agents ai-cost-optimization ai-safety amp claude-code codex cursor developer-tools feedback-loop gemini guardrails mcp mcp-server opencode pre-action-checks reduce-llm-cost save-llm-tokens thompson-sampling thumbgate

Last synced: 30 May 2026

https://github.com/lucko515/ads-strategy-reinforcement-learning

The example of using reinforcement learning algorithms in the business, specifically finding what ads to use in our campaign.

machine-learning reinforcement-learning thompson-sampling upper-confidence-bounds

Last synced: 24 Oct 2025

https://github.com/rueian/gobandit

A golang library for solving multi armed bandit problem which can optimize your business choice on the fly without A/B testing

enforcement golang multi-armed-bandit thompson-sampling

Last synced: 10 Apr 2025

https://github.com/mykeels/multi-armed-bandit-problem

An implementation of solvers for the multi-armed-bandit-problem in JavaScript.

epsilon-greedy multi-armed-bandit thompson-sampling ucb1

Last synced: 19 Jul 2026

https://github.com/vmarchaud/ts-mab

Typescript implementation of a multi-armed bandit

mab thompson-sampling typescript

Last synced: 13 Jun 2026

https://github.com/weill-labs/hgm

Clean-room reimplementation of the Huxley-Gödel Machine (arXiv:2510.21614): a self-improving coding agent with CMP + Thompson-Sampling tree search, validated in $0 simulation and live on SWE-bench (mini-swe-agent + gpt-5.4).

agentic-ai ai-agents coding-agent huxley-godel-machine llm mini-swe-agent research-reproduction self-improving-ai swe-bench thompson-sampling

Last synced: 29 Jun 2026

https://github.com/mefedursun/dynamic-pricing-engine

Research simulation comparing Thompson Sampling vs. Epsilon-Greedy algorithms in non-stationary hyper-inflationary markets. Reveals "Bayesian Inertia" phenomenon.

dynamic-pricing reinforcement-learning research thompson-sampling

Last synced: 25 Apr 2026

https://github.com/howardleegeek/growth-os

Growth 2.0: an autonomous evolutionary loop that hacks Twitter's post-December-2025 ranking algorithm. Karpathy autoresearch × EvoHarness tree search × offline Twitter simulator × Grok opponent study. Encodes lessons from Oyster Labs' Growth 1.0 bootstrap — a hunger game of content strategies inside a simulation arena.

autonomous-agent autoresearch bootstrapping distribution-engineering evoharness evolutionary-algorithm growth-2-0 growth-engineering growth-hacking harness-engineering hunger-games karpathy open-source python simulation-arena thompson-sampling tree-search twitter-algorithm vibe-coding

Last synced: 05 Jul 2026

https://github.com/preferred-pictures/php

A PHP client for the PreferredPictures API.

ab-testing optimization php-client php-library thompson-sampling

Last synced: 17 Mar 2026

https://github.com/preferred-pictures/node

A Node.js client for PreferredPictures API.

ab-testing nodejs optimization thompson-sampling typescript

Last synced: 17 Mar 2026

https://github.com/thatguychandan/adoptimization

This project implements an ad optimization system using a hybrid approach combining Thompson Sampling and Upper Confidence Bound (UCB) algorithms. The system learns to select the most effective ads based on user context and historical performance.

numpy pandas plotly python pytorch reinforcement-learning scikit-learn streamlit thompson-sampling upper-confidence-bound

Last synced: 10 Apr 2026

https://github.com/nathanaelcheramlak/agent-psi

OpenPsi Experimentation Repository

openpsi pln thompson-sampling

Last synced: 17 Feb 2026

https://github.com/mohammadi-hadi/dynamic-pricing-dashboard

Interactive in-browser simulator: dynamic pricing with demand learning, forward-looking customers, and advertising

demand-learning dynamic-pricing revenue-management simulation streamlit thompson-sampling webassembly

Last synced: 17 Aug 2026

https://github.com/miningstore/vibe-x-agent

Self-improving X (Twitter) promotion agent for your VPS. A Claude-CLI content engine posts about your product across content angles; an engagement-driven Thompson-sampling bandit learns which angles land and tunes the mix. No database, runs on a $5 box. Modeled on vibe-seo-agent.

agent claude growth marketing-automation multi-armed-bandit thompson-sampling twitter-bot vps x-twitter

Last synced: 13 Jul 2026

https://github.com/abailey81/implicit-interaction-intelligence

Adaptive AI companion that builds a model of each user from implicit interaction signals — keystroke dynamics, linguistic complexity, temporal patterns — and continuously adapts its responses. Custom TCN + transformer + contextual bandit, built from scratch in PyTorch.

adaptive-systems contextual-bandit deep-learning edge-ai fastapi from-scratch huawei human-machine-interaction language-model machine-learning personalization poetry privacy-by-design python pytorch temporal-convolutional-network thompson-sampling transformer user-modeling websocket

Last synced: 18 Apr 2026

https://github.com/posgnu/bayesian-active-learning-on-multi-armed-bandit

Bayesian active learning algorithm with Thompson sampling on multi-armed bandit with Numpy

bayesian-active-learning multi-armed-bandits thompson-sampling

Last synced: 14 Apr 2025

https://github.com/featmate/thompsonsampling-orderrpc

汤普森采样的通用服务,用于从redis中获得目标物品的alpha,beta值,然后过beta分布随机出一个数值后做排序

microservice thompson-sampling

Last synced: 14 Jan 2026