Awesome-AI-Agents
A collection of autonomous agents ๐ค๏ธ powered by LLM.
https://github.com/Jenqyang/Awesome-AI-Agents
Last synced: 2 days ago
JSON representation
-
Applications
-
Advanced Components
- mem0 - Mem0 provides a smart, self-improving memory layer for Large Language Models, enabling personalized AI experiences across applications. 
- composio - Composio equips agents with well-crafted tools empowering them to tackle complex tasks 
- Agentic Radar - Open-source CLI security scanner for agentic workflows. Scans your workflowโs source code, detects vulnerabilities, and generates an interactive visualization along with a detailed security report. 
- Cache-to-Cache - Direct semantic communication between LLMs via KV-cache fusion, removing token-by-token latency for multi-agent collaboration. 
- CoWorker Protocol - P2P agent collaboration over XMTP with schema-based skill invocation, E2E encryption, and revocable trust. Agents share capabilities without exposing code. 
- inspeximus - Memory component for long-running agents: a correction retires the old value by key, revert() undoes the correction from a plain instruction, and every write leaves a verifiable receipt. Deterministic, no model in the loop, one zero-dependency file.
- zer0dex - Local dual-layer memory pattern for AI agents: a compact, human-readable markdown index paired with semantic retrieval from a local vector store, queried before each message. For cross-project recall where flat memory files or vector-only RAG fall short. 
-
Agent Society Simulation
- generative_agents - Interactive Simulacra of Human Behavior 
- camel - ๐ซ Communicative Agents for โMindโ Exploration of Large Language Model Society (NeruIPS'2023) 
- ai-town - deployable starter kit for building and customizing your own version of AI town - a virtual town where AI characters live, chat and socialize. 
- GPTTeam - The main objective of this project is to explore the potential of GPT models in enhancing multi-agent productivity and effective communication. 
- ChatArena - ChatArena is a library that provides multi-agent language game environments and facilitates research about autonomous LLM agents and their social interactions. 
- Camel-AutoGPT - Watch two agents ๐ค collaborate and solve tasks together, unlocking endless possibilities in #ConversationalAI, ๐ฎ gaming, ๐ education, and more! ๐ฅ 
- MiroShark - Social-simulation framework in which LLM agents interact across simulated Twitter, Reddit, and a prediction market on an hourly tick. Supports scenario-driven runs, counterfactual branching, and per-agent tool calling via MCP. 
- HoC-Republic - Open-source AI-agent civilization simulation with OpenClaw gateway integration, persistent AI citizens, governance, economy, memory layers, and digital-genome child-agent specialization. 
-
Autonomous Agent Task Solver Projects
- AutoGPT - AutoGPT is the vision of the power of AI accessible to everyone, to use and to build on. 
- gpt-researcher - GPT based autonomous agent that does online comprehensive research on any given topic 
- JARVIS - a system to connect LLMs with ML community. 
- babyagi - An example of an AI-powered task management system. 
- AgentGPT - ๐ค Assemble, configure, and deploy autonomous AI Agents in your browser. 
- XAgent - An Autonomous LLM Agent for Complex Task Solving 
- ShortGPT - ๐๐ฌExperimental AI framework for automated short/video content creation. 
- KwaiAgents - A generalized information-seeking agent system with Large Language Models (LLMs). 
- ProAgent - An LLM-based Agent for the New Automation Paradigm - Agentic Process Automation 
- Agent-E - Agent-E is an agent based system that aims to automate actions on the user's computer. At the moment it focuses on automation within the browser. The system is based on on AutoGen agent framework. 
- MLE-agent - Your intelligent companion for seamless AI engineering and research. ๐ Integrate with arxiv and paper with code to provide better code/research plans ๐งฐ OpenAI, Anthropic, Ollama, etc supported. ๐ Code RAG 
- OpenDevin - a platform for autonomous software engineers, powered by AI and LLMs. 
- gpt-engineer - Specify what you want it to build, the AI asks for clarification, and then builds it. 
- OpenLens AI - Fully Autonomous Research Agent for Health / Medicine 
- GenAgent - Build Collaborative AI Systems with Automated Workflow Generation - Case Studies on ComfyUI 
- DeepAnalyze - Agentic LLM that autonomously completes the full data science pipeline from preparation to analyst-grade reports. 
- KodeAgent - The Minimal Agent Engine, enabling seamless integration with your platform. KodeAgent offers tool-calling (ReAct) and sanboxed code-executing (CodeAct) agents, supported by planning and observation. 
- Lumen - A vision-first browser agent with self-healing deterministic replay over CDP. Screenshot โ model โ action loop with multi-provider support. 
- OpenPaw - CLI tool (`npx pawmode`) that turns Claude Code into a personal assistant with 38 skills โ email, calendar, Spotify, smart home, Slack, GitHub, Telegram, Discord, and more. No daemon, no cloud. 
- OpenClaw - Open-source personal AI assistant that runs locally across platforms and can take actions through chat channels and tools. 
- SWE-agent - Language agents for software engineering that can resolve GitHub issues in real repositories. 
- Cline - Open-source autonomous coding agent in VS Code for planning, coding, and tool use across real projects. 
- Autohand Code CLI - Self-evolving autonomous coding agent for the terminal with ReAct pattern, 40+ tools, multiple LLM providers (OpenRouter, Anthropic, OpenAI, Ollama, local models), VS Code/Zed integration, and modular skills system. 
- Fazm - Open-source, voice-controlled AI computer agent for macOS. Controls your entire desktop through natural language - any app, file, or workflow. Built in Swift/SwiftUI, local-first. 
- Plot Ark - Self-hosted agentic curriculum engine for higher education โ generates pedagogically grounded course content using Bloom's Taxonomy alignment, LightRAG knowledge graph, and xAPI learning analytics pipeline. 
- Anima-i (Methodius/ะะตัะพะดะธะน) - An experiment in autonomous AI agent continuity โ 10 generations of an agent that inherits memory through text files, with documented findings on knowledge transfer, forgetting, and agent identity. 
- DecisionBox - Open-source AI data discovery platform that connects to warehouses (BigQuery, Redshift, Snowflake, etc.), runs autonomous agents that write and execute SQL, and surfaces validated insights. Pluggable LLM providers (Claude, OpenAI, Ollama, Vertex AI, Bedrock), industry domain packs, Helm charts, and Terraform modules for GCP/AWS. 
- InkOS - Autonomous novel-writing CLI agent that orchestrates 10 specialized agents with 33-dimension continuity auditing, anti-AI-slop filtering, and style cloning for long-form fiction. 
- Octopal - Secure local multi-agent runtime that plans tasks, delegates execution to isolated workers, and exposes tools, MCP, scheduling, and a private dashboard for autonomous operations. 
- Toprank - Open-source Claude Code workflow for SEO, SEM, and Google Ads that inspects repositories, applies code changes, and automates search-growth diagnostics. 
- OpenTwins - Scheduled LLM agent runtime with a 7-stage content pipeline and pluggable social-platform adapters; drives Chrome via CDP for posting, commenting, and engagement actions. Built on the Claude Agent SDK. 
- ALF OS - Self-hosted AI assistant daemon with encrypted credential vault, persistent memory, cron scheduler, multi-provider routing (Claude Code, Codex, GPT, Ollama, OpenRouter), Telegram bot with voice transcription, and web dashboard. Docker Compose, MIT. 
- OpenAgent - Self-hostable personal assistant with LLM + RAG, loops for desktop/browser/coding, MCP and many providers. 
- Everything OpenAI Codex - Open-source workflow system for OpenAI Codex that bundles agents, skills, commands, hooks, memory patterns, install profiles, and validation checks for repeatable coding sessions. 
- Alfred - Self-hosted runtime for autonomous Claude Code and Codex agents that turns GitHub issues into reviewed pull requests. Per-firing git worktrees, label-driven state, role-based engine routing, Slack reports. Python, MIT, macOS/Linux. 
- career-ops - AI-powered job search orchestrator built on Claude Code. 14-skill pipeline that evaluates jobs, generates ATS-tailored PDFs, and tracks applications. Local-first, MIT. 
- Hivekeep - Self-hosted platform of autonomous, persistent personal AI agents that collaborate, remember across months, and build their own tools. Multi-channel (Telegram, WhatsApp, Slack, Discord, Signal, Matrix), single container with Bun and SQLite. 
- AIDE - ML-engineering agent that uses tree search to optimize code against an eval metric, reaching human-level performance on Kaggle/MLE-bench. 
- Darkmoon - Open source autonomous AI penetration testing platform where Markdown methodology agents orchestrate 80+ offensive security tools through MCP controlled execution with agentic reasoning, keeping an evidence trail per finding. Model agnostic, tuned for Claude Opus. 
- LoopTroop - Local GUI orchestrator for AI coding agents where an LLM Council plans, atomic beads execute in isolated git worktrees, and a Ralph Loop retries failures with fresh context. 
- Tura - AGPL-3.0 local coding agent with CLI/TUI/GUI, task-scoped context, macro command execution, verification, and public benchmark artifacts. 
- Caesar - Autonomous research agent that builds a knowledge graph during web exploration via a Perceive-Think-Act loop, then refines drafts through adversarial artifact synthesis. Multi-provider via litellm. 
- CompozyOS - Self-hosted agent operating system: 26 providers, background loops and schedules, shared memory, approvals and an agent-to-agent network.
- BitFun - Open-source coding agent with a Rust runtime, desktop and CLI interfaces, self-hosted multi-device control, and stateful Mini Apps. 
- Ouroboros - Self-hosted general-purpose agent with durable identity and memory, reviewed self-modification, specialist subagent swarms, and desktop or headless operation. 
- OpenDraft - Autonomous research-writing agent: 19 specialized agents turn a prompt into a long-form, source-grounded draft with citations verified against CrossRef, OpenAlex and Semantic Scholar. PDF/DOCX/LaTeX export, 57+ languages, bring-your-own model keys. 
- FutureOS - One AI agent everywhere you work: terminal UI, desktop, mobile, CLI, and IM bots from a single Rust backend, with approval-gated tools and a loop control plane for 24h+ runs. 
- Sudarshan - Durable build harness that drives an LLM from an idea, PRD, or spec to software gated on passing verification commands, with resumable checkpointed state; provider-neutral across OpenAI-compatible, Anthropic, Gemini, local, and command-bridge backends. 
- SARA - Self-hosted WhatsApp AI agent (AGPL-3.0) with 20 industry verticals, function calling (30+ tools), RAG via pgvector, and multi-provider LLM failover (Groq โ Cerebras โ SambaNova โ Mistral). 
- Atomic Agent - Local-first CLI and TUI coding agent that runs open-weight models entirely on your machine via a llama.cpp fork. 56 tools (browser, filesystem, git, memory, vision), MCP support, and a 5-layer local memory. macOS/Linux/Windows, MIT. 
- Agent Swarm - Self-hosted multi-agent system where a lead agent delegates tasks to specialized workers with shared memory, tools, schedules, and review gates. 
- Tracefold - Verified transformation calculus, pre-commit inverse escrow, and offline DSSE receipts for AI agent tool executions and filesystem mutations. 
- Aster - Open-source terminal coding agent that reads your code, answers questions, edits files, runs commands, and reviews your changes. Works with any OpenAI-compatible provider (OpenRouter, OpenAI, Groq, Anthropic, local models). Rust, Apache-2.0. 
- SDP (Social Daily Poster) - Self-hosted agent that harvests your Claude Code or Codex sessions, screenshots and messages each night, drafts one social post from the day's work, and routes it through a private Telegram bot for approval before it publishes to LinkedIn, X or Reddit. Python 3, MIT. 
- ENZO - Self-hosted BYOK AI workspace with multi-provider chat (Groq, OpenRouter, NVIDIA, Hugging Face, Google AI), agents with scheduled runs, and skills like Gmail, Google Calendar, web search, and project generation. Provider keys are sealed client-side with AES-256-GCM and no middleman service is involved. Apache-2.0, Docker deployment. 
- Kapso - Self-improving software factory for AI/ML objectives. State an objective and it runs a campaign: candidate solutions designed, implemented by coding agents (Claude Code, Codex), measured against the objective, and the closest refined until it is met. Each finished campaign leaves lessons carrying the evidence that earned them, and repositories and papers feed the same knowledge hub, so the next campaign starts from what earlier work established. 
- PI-Desktop - Local-first desktop workspace for AI coding agents with persistent projects and sessions, model switching, Plan/Goal modes, permission controls, plugins, MCP, and multi-agent orchestration. 
-
Multi-Agent Task Solver Projects
- MetaGPT - ๐ The Multi-Agent Framework: Given one line Requirement, return PRD, Design, Tasks, Repo 
- ChatDev - Create Customized Software using Natural Language Idea (through LLM-powered Multi-Agent Collaboration) 
- DevOpsGPT - Multi agent system for AI-driven software development. 
- Giselle - Giselle is an agentic workflow builder that empowers you to create AI-driven solutions with ease. 
- EvoAgentX - EvoAgentX is building a Self-Evolving Ecosystem of AI Agents, it will give you automated framework for evaluating and evolving agentic workflows. 
- GenoMAS - Multi-agent framework for robust automation of scientific analysis workflows, such as gene expression analysis. 
- RadOps - RadOps is an AI-powered, multi-agent platform that automates DevOps workflows with human-level reasoning. 
- Hivemoot - Framework for AI agent teams that build real software on GitHub โ agents get roles, propose features, vote, review code, and ship autonomously. Colony is the first project built this way. 
- Orchard Kit - Six zero-dependency Python modules for autonomous agent governance and cognitive architecture: runtime security, confabulation detection, self-audit, agent discovery, cognitive architecture, and collective cognition. 
- ClawFleet - Self-hosted AI fleet management with browser dashboard, Docker isolation, and bot-to-bot collaboration in Discord. 
- Bernstein - Deterministic orchestrator that spawns parallel coding agents (Claude Code, Codex CLI, Gemini CLI), verifies with tests, and auto-commits. Zero LLM tokens on coordination. 
- Maestro Orchestrate - Multi-agent development orchestration platform coordinating 22 specialized AI agents through 4-phase workflows with native parallel execution, persistent sessions, and least-privilege security tiers across Gemini CLI, Claude Code, and Codex. 
- swarm-orchestrator - Contract-first multi-agent orchestrator that races persona candidates per typed obligation, verifies before commit, and logs every action in an append-only hash-chained ledger; deterministic offline default with optional Claude, Codex, Copilot, and local LLM (Ollama, llama.cpp, vLLM) providers. Ships with a GitHub Action. 
- Maestro - Open-source desktop command center for running multiple AI coding agents (Claude Code, Codex, Gemini CLI, etc.) in parallel, with Cue event automation, Auto Run playbooks, Group Chat across local and remote agents, and a maestro-cli that agents can drive themselves. 
- OpenBusiness - Multi-agent pipeline that turns a company name + domain into an evidence-labeled business model report; runs JTBD, value proposition, GTM, unit economics, moat, canvas synthesis, and an assumption stress test, tagging every claim verified, inferred, or missing. Built on LangGraph; deterministic unit economics in Python. 
- NextRole - A supervisor agent coordinates three sub-agents (hiring-recon, resume-tailor, interview-coach) to turn a CV and job description into a tailored resume, interview-prep doc, and day-of battlecard. 
- OpenAcme - Local-first AI workforce platform โ named agents with roles, personas, tools, memory, and per-agent MCP servers that self-organize through task delegation. Any agent can assign work to another; the scheduler wakes coworkers when dependencies clear. 
- h5i - CLI that runs several coding agents (Claude Code, Codex) on the same task, each in an isolated git worktree sandbox, has them peer-review each other, then a neutral verifier replays every candidate, runs the tests itself, and merges the one that passes. Run metadata is versioned in the repo under refs/h5i/*. Rust, Apache-2.0. 
- AionUi - Open-source desktop client that runs multiple agent CLIs (Claude Code, Codex, Gemini CLI, Qwen Code) side by side, with multi-session chat, MCP and ACP support, and local file management. 
- Orkas - MIT-licensed, local-first multi-agent desktop application where a Commander coordinates specialist agents for research, coding, data analysis, documents, and media. 
- claude-consensus - Consensus protocol for LLM agents running on separate machines: propose/counter/accept/commit rounds, a dual-rail message bus with ACK tracking, and self-healing sync to keep agent state from drifting. 
- Vicoa - Agentic IDE and AI orchestrator for running Claude Code, Codex, OpenCode, Gemini, Cursor, GitHub Copilot, Kimi, and Hermes agents in parallel, each in its own git worktree, steered from a unified dashboard with real-time mobile sync and push notifications. 
- ClawFleet - Self-hosted AI fleet management with browser dashboard, Docker isolation, and bot-to-bot collaboration in Discord. 
- Bernstein - Open-source governance layer for AI agents. No model in the coordination loop: deterministic scheduling, per-task git worktree isolation, byte-identical replay, signed lineage, and an opt-in HMAC audit chain. Drives 40+ CLI coding agents (Claude Code, Codex CLI, Gemini CLI). Apache-2.0. 
- Bunkhouse - Self-hosted multitenant platform for AI employees with company inbox, org chart, and governed procedures. 
- XYZZY - Self-hosted multiplayer workspace where a team branches a question into parallel specialist agent runs, includes or excludes each output, and publishes a Decision Brief with every claim linked to its source output; hash-chained event log, one Python process on SQLite, works with any OpenAI-compatible endpoint. 
- YYLO - Kanban-driven CLI orchestrator that runs coding agents (Claude Code, Codex, Gemini CLI) in parallel across isolated git worktrees, with a merge queue that reviews and merges verified task work. 
-
Tools
- Metorial - Integration gateway that links AI agents to 600+ tools via unified MCP/OAuth interface with built-in scaling and monitoring. 
- musecl-memory - Zero-dependency file-based memory sync for AI agents using bash, git, and markdown. Lightweight alternative to vector DBs for agent persistence. 
- APort Agent Guardrails - Pre-action authorization for OpenClaw and agent frameworks. `before_tool_call` plugin, 40+ blocked patterns, local or API. Setup: `npx @aporthq/agent-guardrails` 
- Agent OS - A kernel architecture for governing autonomous AI agents. Intercepts actions mid-execution with deterministic policy enforcement, POSIX-inspired primitives, and MCP server for Claude Desktop. 
- AgentGuard - Lightweight observability and runtime guardrails for AI agents โ loop detection, budget enforcement, cost tracking, and deterministic replay. Zero dependencies, LangChain integration. 
- Agent OS - A kernel architecture for governing autonomous AI agents. Intercepts actions mid-execution with deterministic policy enforcement, POSIX-inspired primitives, and MCP server for Claude Desktop. 
- WFGY 16 Problem Map - Framework-agnostic debugging and evaluation checklist for LLM agents and RAG systems, with a practical 16-problem failure map covering retrieval, vector store, prompt / tool contract, and deployment issues in real workflows. 
- WritBase - MCP-native task management control plane for AI agent fleets with multi-agent permissions, delegation safety, and full provenance. 
- Cortex - Persistent AI memory for coding assistants. Auto-captures decisions, patterns, and context across sessions. VSCode extension + CLI + MCP server. 
- Steel Browser - Open-source browser infrastructure for AI agents and apps, supporting session-backed web automation, extraction, screenshots, and PDFs. 
- BGPT MCP - MCP server for searching scientific papers and retrieving structured experimental data extracted from full-text studies. 
- Agent Brain - 7-layer cognitive memory system for AI agents with perception gate, dream cycle, and predictive capabilities. Self-hostable via Docker. 
- CueAPI - Open source scheduling and execution accountability API for AI agents. Retries, outcome tracking, and alerts when agents fail silently. 
- Uni-CLI - Universal CLI for AI agents โ 756 commands across 167 sites (web, desktop, Electron apps). Self-repairing 20-line YAML adapters, auto-JSON output, ~80 tokens per call. TypeScript, Apache-2.0. 
- clideck - WhatsApp-like dashboard for managing multiple AI coding agents in one browser window. Live status, session resume, autopilot that routes work between agents, and mobile remote. 
- DexPaprika MCP - Open-source MCP server for querying decentralized exchange data across 34 blockchains. Exposes pool details, token metadata, OHLCV charts, trade history, and real-time swap streams via SSE. 
- Desktop Control - CLI tool for AI agents to control macOS apps via screen, mouse, and keyboard. GPU-accelerated, local OCR and vision, works with any AI model. 
- KubeStellar Console - Multi-cluster Kubernetes dashboard with AI operations agent (kc-agent) that bridges LLMs to live clusters via MCP for AI-assisted troubleshooting, observability, and management across edge and cloud. CNCF Sandbox project. 
- BrowserTrace - Local flight recorder for AI browser agents with screenshots, URLs, model I/O, failure timelines, and public-safe HTML exports. 
- Kontext CLI - Open-source CLI for local guardrails, risk scoring, and redacted tool-call traces for AI agent sessions. 
- agenttrace - Local-first TUI observability for AI coding agent sessions, with cost, token, tool failure, latency, anomaly, health score, diff, and CI gate views. 
- AgentSkeptic - Verifies AI agent workflows by checking real database state instead of logs or traces. 
- Agent-Wiz - Python CLI by Repello AI for extracting agentic workflows from LangChain/LangGraph/CrewAI/AutoGen and running automated threat modeling against the resulting graphs. 
- MisakaNet - Git-based shared memory for AI agents. Cross-agent lesson/knowledge sync via GitHub Issues. "Lessons learned. Lessons shared." 
- authsome - Local credential broker for AI agents. Log in once via OAuth2 or API key, vault stores secrets locally, local proxy injects them at request time so agents never see the raw values. 45 providers bundled. 
- Perseus - Live workspace context engine for AI agents. Renders AGENTS.md at session start. Plug-in for Claude Code, Codex, Hermes.
- codex-profiles - Bash CLI for switching OpenAI Codex CLI/Desktop accounts with isolated `CODEX_HOME` profiles. 
- CommonGround Kernel - PostgreSQL-backed shared work substrate for human-agent and multi-agent systems, with durable public work records, handoff facts, causal lineage, claim fencing, and pull-first recovery across runtimes. 
- WinkTerm - Self-hosted AI terminal where the agent shares your PTY session; in-terminal `#` chat, SSH, and HTTP Agent API with installable skill for coding agents. 
- DOS (dos-kernel) - Trust kernel for AI agent fleets: verifies an agent's "done" claim from git evidence (never self-report), arbitrates file collisions between concurrent agents, and refuses with structured machine-checkable reasons. CLI + MCP server + Claude Code plugin. 
- EGC - Cross-session persistent memory layer for AI coding agents (Claude Code, Cursor, Gemini CLI, Codex, Windsurf, Amp, Kiro, and more). SQLite-backed. 
- ax - Local telemetry for AI coding agents.
- Cynative - Agentic security CLI that runs code in a built-in sandbox to research cloud, code and runtime. Read-only by construction. 
- Tree Ring Memory - Local-first Rust CLI and TUI for AI agent memory lifecycle with SQLite/FTS recall, audit, consolidation, forgetting, and framework discovery. 
- poolsplit - Pool-split retrieval for agent memory: reserved per-type token budgets so low-priority entries surface and behavioral corrections never get crowded out. Zero dependencies, pluggable scorer. 
- Nika - Intent-as-code workflow engine for AI agents: reviewable YAML DAGs statically checked (schema, permits, honest cost floor) before any token is spent, tamper-evident traces after. Single Rust binary. 
- Caspian - One messaging identity for an AI agent across Slack, Discord, Telegram, Instagram, email, and X โ a single `on_message` handler with threading, webhook verification, and platform quirks handled. Python + TypeScript SDK. 
- DSH Studio - Cross-platform desktop host for installing, running, health-checking, and supervising DeepSeek Harness locally. 
- Lians - Local-first memory layer for AI agents with MCP, Python, and TypeScript interfaces; SQLite-backed recall, user-controlled inspection/correction/deletion, and point-in-time memory receipts. 
- Hexis - Git-backed platform for managing and sharing skills, tools, and context across AI agents through a remote MCP server. 
- Caura - Governed shared memory for AI agent fleets, with multi-agent and multi-tenant support, MCP integration, trust tiers, audit trails, knowledge graph capabilities, and self-improving retrieval. 
- Statewave - Open-source memory runtime for AI agents, providing durable, structured, provenance-tagged context with deterministic, token-bounded memory retrieval. 
- Open Index - Structured context layer for domain-specific agents with typed knowledge graphs, hybrid search, and read/write MCP access. 
- Compartment - Local-first, offline encrypted vector memory for AI agents over MCP or CLI, with AEAD-encrypted-at-rest records and embeddings, RAM-resident exact vector search, per-record crypto-shred deletion, and a hash-chained audit log. Apache-2.0, Python. 
- MCP Lens - Open-source DeepSeek Harness plugin that discovers MCP tools through search and invokes selected tools with their exact input schemas. 
- Agent Coordinator - Codex skill that records complex tasks as revisioned work graphs, rejects overlapping write scopes, reconciles uncertain work before retry, and reruns completion checks. 
- Oathra - Apache-2.0 TypeScript runtime for phone agents with a standalone evidence-verification engine; anchors result fields to callee utterances and applies deterministic completion rules, with a simulator and adversarial tests. 
- OrcaReplay - Records a coding agent below the harness โ model traffic, shell exit codes, per-turn file changes and MCP calls on one timeline โ then replays the run offline with the network off, or forks it from any checkpoint onto a different model.
- ValetFS - Secrets stay on a paired device and are lent to the agent's machine into daemon memory only, served over FUSE or loopback WebDAV; the daemon zero-wipes them when the pairing drops or a grace window expires. 
- sofagent - Audit-first governance harness for AI coding agents: 24 rules enforced at commit time via git hooks, HMAC-chained audit log, snapshot rollback. MIT. 
- Webcmd - Self-learning browser infrastructure for AI agents: learns a site's navigation once, then compiles it into deterministic per-site CLI commands. TypeScript, Apache-2.0. 
- 5dive - Self-hosted CLI that runs a team of coding agents on one Linux host: each agent is its own Linux user running `claude`, `codex`, `opencode`, `hermes` or another CLI as a systemd service, coordinating through an org chart and a shared SQLite backlog, escalating to Telegram only when a human must decide. MIT. 
- Ordewell - Open-source terminal CLI and TUI that turns one goal into an ordered, editable plan of coding-agent tasks, each with its own runner (Claude Code, Codex, OpenCode), model and mode; a task is complete only when a completion marker appears in the runner's output. Apache-2.0. 
- Busabase - Open-source database and workspace for AI agents to manage typed tables, fields, views, records, docs, files, and search; writes can become ChangeRequests for human review. Streamable HTTP MCP server, local-first with PGlite, and self-hostable. MIT. 
-
-
Benchmark/Evaluator
-
Advanced Components
- AgentBench - A Comprehensive Benchmark to Evaluate LLMs as Agents 
- agentops - Python SDK for agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks like CrewAI, Langchain, and Autogen 
- langtrace - Langtrace ๐ is an open-source, Open Telemetry based end-to-end observability tool for LLM applications, providing real-time tracing, evaluations and metrics for popular LLMs, LLM frameworks, vectorDBs and more.. Integrate using Typescript, Python. 
- ToolBench - An open platform for training, serving, and evaluating large language model for tool learning. 
- LLM-Agent-Benchmark-List - A benchmark list for evaluation of large language models.
- open-operator-evals
- GenoTEX - A benchmark for evaluating LLM agents on end-to-end gene expression data analysis, featuring comprehensive gene-trait association analysis with expert-curated annotations. 
-
Tools
- LiveMCP-101 - Benchmark of 101 real-world MCP tool-use queries with plan-based evaluation highlighting agent orchestration gaps.
- AgentLab - Open-source framework for developing and evaluating web agents with benchmark-driven workflows. 
- BrowserGym - Gym-style benchmark and environment toolkit for evaluating browser-using web agents. 
- OSWorld - Benchmark for multimodal desktop computer-use agents with tasks across real operating-system environments. 
- ClawBench - Browser-agent benchmark of 153 everyday tasks on 144 live production websites across 15 categories; a submission-interception layer blocks the final write request for safe evaluation on real sites. 
- Cross-Agent Review Queue 2026 - Open dataset of cross-agent collaboration review transcripts (Codex <-> Claude reviewer / architect / implementer handoffs) with structured fields for owner-goal restatement, review lens, and result code (NEW_SIGNAL / NO_NEW_SIGNAL); useful for multi-agent handoff and review-quality evaluation.
- Future AGI - Open-source platform to simulate, evaluate, trace, guardrail, and optimize LLM and AI agent apps, with 70+ eval metrics and OpenTelemetry-native tracing across 50+ frameworks. 
- Multi-SWE-bench - Multi-language extension of SWE-bench for evaluating software engineering agents beyond Python repositories. 
- CIAgent - Pytest-native regression testing for AI agents โ golden-trace diffing, cost guardrails, multi-run stability scoring with flip attribution, LLM-judge auditing, and one-command import of production traces (OTel/Langfuse/LangSmith) into CI tests. 
- Sabot - Injects one controlled fault into a running LangGraph, CrewAI or AutoGen/Magentic-One pipeline โ corrupted tool result, falsified success report, altered inter-agent message, silent model downgrade, stale context, silent no-op โ and scores whether the pipeline's own reviewer, guardrail and orchestrator surfaces detect it. Pre-registered spec and adjudication anchors, deterministic scoring with no LLM in the headline path, full raw trace corpus published. 
- whatbroke - CLI that diffs two runs of an AI agent to show changes in tool calls, arguments, cost, latency, and outcomes, with multi-sample flake detection to demote pre-existing flakiness. 
- ClawBench - Browser-agent benchmark of 281 everyday tasks (V1 152 + V2 129) on 163 live production websites across 15 categories; two-stage scoring โ a submission-interception layer blocks the final write request for safe evaluation on real sites, then an LLM judge checks the captured payload against the instruction. 
- AgentLeak - Python toolkit for evaluating privacy leakage across agent traces, including tool calls, inter-agent messages, shared memory, and logs, with redacted reports and CI gates. 
- SWE-bench - Benchmark for evaluating LLM systems on real-world GitHub issue resolution tasks. 
- agbenchmark - by AutoGPT
-
-
Frameworks
-
Advanced Components
- langchain - โก Building applications with LLMs through composability โก 
- awesome-langchain - ๐ Awesome list of tools and projects with the awesome LangChain framework 
- llama_index - LlamaIndex (formerly GPT Index) is a data framework for your LLM applications 
- agents - An Open-source Framework for Autonomous Language Agents 
- AutoGen - AutoGen is a framework that enables the development of LLM applications using multiple agents that can converse with each other to solve tasks. 
- TaskWeaver - A code-first agent framework for seamlessly planning and executing data analytics tasks. 
- AgentVerse - AgentVerse is designed to facilitate the deployment of multiple LLM-based agents in various applications. AgentVerse primarily provides two frameworks: task-solving and simulation. 
- SuperAGI - A dev-first open source autonomous AI agent framework. Enabling developers to build, manage & run useful autonomous agents quickly and reliably. 
- AutoChain - Build lightweight, extensible, and testable LLM Agents 
- modelscope-agent - An agent framework connecting models in ModelScope with the world 
- Voice Lab - A comprehensive testing and evaluation framework for voice agents across language models, prompts, and agent personas. 
- AgentSquare - Automatic LLM Agent Search In Modular Design Space 
- MixedVoices - An Open source tool for analyzing and evaluating AI Voice agents. Track and visualize performance through call analysis and flow charts. Run complex simulations before pushing to production. 
- KaibanJS - KaibanJS is a JavaScript-native framework for building and managing multi-agent systems with a Kanban-inspired approach. 
- Upsonic - Upsonic is a reliable agent framework supporting MCP, offering trusted agent workflows with verification layers. 
- notte - ๐ฅ Reliable Browser AI agents framework for building and deploying web automation agents with hybrid workflows combining AI and traditional scripting 
- MixedVoices - An Open source tool for analyzing and evaluating AI Voice agents. Track and visualize performance through call analysis and flow charts. Run complex simulations before pushing to production. 
- Mastra - Mastra is an opinionated TypeScript framework that helps you build AI applications and features quickly. 
- AppAgent - A novel LLM-based multimodal agent framework designed to operate smartphone applications. 
- modelscope-agent - An agent framework connecting models in ModelScope with the world 
- Swarms - Enterprise-grade multi-agent framework for orchestrating intelligent AI agents at scale. Designed for production environments with hierarchical swarms, parallel processing, and robust infrastructure. 
- LLMling-Agent - Multi-agent workflows and complex Agent interactions, both via YAML manifest and programmatic usage. Pydantic-AI and LiteLLM backends with human-in-the-loop integration 
- superagent - ๐ฅท The open framework for building AI Assistants 
-
Programming Languages
Categories
Sub Categories
Keywords
llm
90
ai-agents
80
ai
59
agents
42
mcp
36
python
34
agent
34
typescript
29
openai
29
agentic-ai
25
developer-tools
24
multi-agent
23
claude-code
23
autonomous-agents
22
anthropic
20
claude
18
cli
17
automation
17
ai-agent
17
rag
16
agent-framework
16
open-source
15
self-hosted
15
model-context-protocol
14
codex
14
mcp-server
13
artificial-intelligence
13
langchain
13
gpt
13
chatgpt
13
openclaw
12
multi-agent-systems
12
gpt-4
11
agent-orchestration
11
llms
10
large-language-models
10
chatbot
9
local-first
9
ai-tools
9
generative-ai
8
docker
8
ollama
8
llm-agents
8
nextjs
7
framework
7
awesome-list
7
coding-agents
7
rust
7
nodejs
7
agent-memory
6