{"id":51546316,"url":"https://github.com/adithyan-ak/claude-conclave","last_synced_at":"2026-07-09T18:31:59.651Z","repository":{"id":364868872,"uuid":"1269523129","full_name":"adithyan-ak/claude-conclave","owner":"adithyan-ak","description":"Convene a Claude Conclave: a Constitutional Tournament of sequestered agents that arbitrates hard, open-ended engineering decisions and grounds the verdict in executed verification — not confident prose. A Claude Code skill.","archived":false,"fork":false,"pushed_at":"2026-06-14T20:35:32.000Z","size":37,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-06-14T22:21:21.919Z","etag":null,"topics":["agentic-ai","ai-agents","ai-decision-making","anthropic","claude","claude-code","decision-support","llm-agents","llm-as-judge","multi-agent"],"latest_commit_sha":null,"homepage":null,"language":"JavaScript","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/adithyan-ak.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-06-14T20:21:06.000Z","updated_at":"2026-06-14T20:35:35.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/adithyan-ak/claude-conclave","commit_stats":null,"previous_names":["adithyan-ak/claude-conclave"],"tags_count":null,"template":false,"template_full_name":null,"purl":"pkg:github/adithyan-ak/claude-conclave","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/adithyan-ak%2Fclaude-conclave","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/adithyan-ak%2Fclaude-conclave/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/adithyan-ak%2Fclaude-conclave/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/adithyan-ak%2Fclaude-conclave/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/adithyan-ak","download_url":"https://codeload.github.com/adithyan-ak/claude-conclave/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/adithyan-ak%2Fclaude-conclave/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35309827,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-09T02:00:07.329Z","response_time":57,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["agentic-ai","ai-agents","ai-decision-making","anthropic","claude","claude-code","decision-support","llm-agents","llm-as-judge","multi-agent"],"created_at":"2026-07-09T18:31:59.507Z","updated_at":"2026-07-09T18:31:59.641Z","avatar_url":"https://github.com/adithyan-ak.png","language":"JavaScript","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n\n# 🕊️ Claude Conclave\n\n### Make AI give you *one* trustworthy decision — or admit it can't.\n\n**A Claude Code skill that convenes a \"Constitutional Tournament\" of sequestered agents to arbitrate hard, open-ended engineering decisions — and grounds the verdict in *executed verification*, not confident prose.**\n\n[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)\n[![Claude Code Skill](https://img.shields.io/badge/Claude%20Code-Skill-8A2BE2)](https://docs.claude.com/en/docs/claude-code)\n[![PRs Welcome](https://img.shields.io/badge/PRs-welcome-brightgreen.svg)](CONTRIBUTING.md)\n[![Status](https://img.shields.io/badge/status-research%20preview-orange.svg)](#status--honest-limitations)\n\n\u003c/div\u003e\n\n---\n\n## The problem\n\nAsk an LLM the same architecture question three times and you get three different, **mutually exclusive** answers — each defended with flawless, authoritative, mathematically-plausible reasoning. The model never says \"I'm not sure.\" It samples a path and then justifies it to the hilt. Worse, when one model seeds the constraints, every downstream agent confidently agrees with its baked-in bias.\n\nThis is **probabilistic variance + persuasive inconsistency**, and it has three faces:\n\n| Symptom | What you see |\n| --- | --- |\n| **Non-determinism** | Same prompt, system state → entirely different paradigms across runs. |\n| **Authority-hallucination** | Not uncertainty — a *flawless, convincing* justification for whichever path was randomly sampled. |\n| **Mono-model groupthink** | One model's latent bias becomes the foundation; downstream agents confidently ratify a flawed premise. |\n\nThey all reduce to **one root cause**: open-ended design has *no objective ground-truth function*, so the model substitutes a proxy — its own fluent self-certainty — and reports it as if it were truth.\n\n## The idea\n\n\u003e **Stop trying to make the *model* deterministic or trustworthy. Make the *decision procedure around it* deterministic, and let executed verification — not model confidence — decide. When verification can't decide, refuse to fake a winner.**\n\nClaude Conclave borrows the one human institution built to extract a single trusted decision from a roomful of biased, fallible voters: the **conclave**.\n\n| Conclave | Claude Conclave |\n| --- | --- |\n| Diverse electors, **sequestered** so they can't sway each other | N agents generate criteria \u0026 candidates **independently** (no shared context, authorship stripped) |\n| A fixed **constitution** governs the vote | A **frozen, hashed weighted objective** — no agent may edit it mid-run |\n| Voting in **rounds** | Deterministic **single-elimination tournament** |\n| **Two-thirds supermajority** required (literally `S = 2/3` here) | A criterion enters the objective only on a **2/3 supermajority** of electors |\n| No supermajority → **keep deliberating**, don't force a result | No grounded majority → **abstain** to an honest Pareto shortlist |\n| White smoke = one verdict | `SINGLE` verdict = one winner, with its grounding stated |\n\nThe result you can trust **100% on procedure** (same input + same frozen constitution always replays to the same ranking, every elimination traceable to a runnable check) — and a **measured \"grounding dial\"** that tells you exactly how much of that verdict is backed by *executed reality* versus model estimate, and **abstains** rather than over-claim when the dial is low.\n\n---\n\n## How it works\n\nFive phases. The LLM is demoted from *decider* to *hypothesis generator*; the deciding is done by frozen math + executed checks.\n\n```\n  ┌─ 0. CONSTITUTION ──────────────────────────────────────────────┐\n  │  D diverse electors draft criteria independently               │\n  │  → cluster by meaning → keep only 2/3-supermajority criteria    │\n  │  → FORGE a runnable check per criterion (with negative control) │\n  │  → freeze weighted objective + SHA, never edited again          │\n  └────────────────────────────────────────────────────────────────┘\n  ┌─ 1. GENERATE ──────────────────────────────────────────────────┐\n  │  N candidate solutions across decorrelated persona×principle×   │\n  │  framing axes; provenance stripped, live tallies hidden         │\n  └────────────────────────────────────────────────────────────────┘\n  ┌─ 2. VERIFY ────────────────────────────────────────────────────┐\n  │  Score each candidate per criterion. Where a check is runnable, │\n  │  AGENTS ACTUALLY RUN CODE and record EXECUTED vs ESTIMATED.     │\n  │  Unverifiable criteria → weight 0 (describe, never decide).     │\n  └────────────────────────────────────────────────────────────────┘\n  ┌─ 3. TOURNAMENT ────────────────────────────────────────────────┐\n  │  Deterministic single-elimination. A match is a pure score      │\n  │  comparison; only in an ε-near-tie may an agent break it — and  │\n  │  ONLY with an executed, criterion-citing artifact, not rhetoric.│\n  └────────────────────────────────────────────────────────────────┘\n  ┌─ 4. VERDICT ───────────────────────────────────────────────────┐\n  │  Pure-JS honesty gate computes grounding ratio + effective      │\n  │  independence (n_eff). Emits SINGLE winner iff grounded —       │\n  │  else an honest, ranked PARETO shortlist + named assumptions.   │\n  └────────────────────────────────────────────────────────────────┘\n```\n\n### Where trust actually comes from\n\n- **Semantic ground-truth (the only thing allowed to decide):** external verifiers run in phase 2 — a test's PASS/FAIL, a benchmark's measured number, a static-analysis result. These are deterministic functions of the artifact, independent of any model's prose.\n- **Procedural determinism:** the frozen, hashed constitution + the fixed gate remove the degrees of freedom a model would use to rationalize.\n- **Categorically barred from deciding:** LLM consensus, debate convergence, self-certainty, verbalized confidence. Each is correlated error or uncalibrated confidence.\n\n### The grounding dial \u0026 honest abstention\n\nEvery verdict reports a **grounding ratio** ∈ [0,1] — the fraction of the *winning margin* backed by an executed check vs. a model estimate — and an **effective-independence** number `n_eff` (are your N candidates really N opinions, or one opinion wearing N hats?).\n\nThe conclave returns a **single winner only if**:\n\n```\nnot a near-tie (Δscore \u003e ε)\nAND ( grounded_margin ≥ τ_g     # the win is backed by real execution\n      OR (n_eff ≥ τ_n AND families ≥ 2) )   # OR genuinely independent consensus\n```\n\nOtherwise it **abstains** and hands you a ranked Pareto shortlist with the exact assumptions you'd need to accept to choose. That refusal is the feature: *a system that honestly says \"I can't ground this\" is strictly better than one that confidently hands you the wrong answer.*\n\n---\n\n## Install\n\n\u003e Requires [Claude Code](https://docs.claude.com/en/docs/claude-code) with the **Workflow** tool available (the deterministic multi-agent orchestration runtime). The engine is plain JavaScript executed by that runtime — no `npm install`, no dependencies.\n\n```bash\ngit clone https://github.com/adithyan-ak/claude-conclave.git\n# Install as a user-level skill:\nmkdir -p ~/.claude/skills/conclave\ncp claude-conclave/skill/SKILL.md     ~/.claude/skills/conclave/\ncp claude-conclave/skill/conclave.js  ~/.claude/skills/conclave/\n```\n\nThen in Claude Code:\n\n```\n/conclave \u003cyour decision, with its hard constraints\u003e\n```\n\nThat's it. `/conclave` is user-invocable only (it won't auto-trigger mid-conversation).\n\n---\n\n## Usage\n\n```\n/conclave For a service ingesting ~5k webhook events/sec with bursty spikes,\nchoose the queueing/backpressure approach: Kafka, Redis Streams, a managed\nqueue (SQS/PubSub), or something else. Optimize for operational simplicity,\ncost, and not losing events under burst.\n```\n\n**Tuning knobs** (optional, safe defaults):\n\n| Arg | Default | Meaning |\n| --- | --- | --- |\n| `D` | 6 | electors who draft the constitution |\n| `N` | 6 | candidate solutions generated |\n| `tau_g` | 0.6 | grounded-margin bar to crown a single winner |\n| `tau_n` | 3 | effective-independence bar for a consensus claim |\n| `eps` | 0.04 | near-tie band on the aggregate score |\n| `tiers` | false | mix opus/sonnet/haiku for extra error-decorrelation (else uniform Opus) |\n\nPass them in the skill invocation, e.g. *\"…use D=8, N=8, tiers=true.\"*\n\n---\n\n## Test runs (real, reproducible)\n\nTwo live runs are included verbatim in [`docs/examples/`](docs/examples). Both were executed by the engine in this repo; nothing is mocked.\n\n### 1. The blind test — a problem with a known-correct answer ([full transcript](docs/examples/01-streaming-variance.md))\n\nWe handed the conclave a streaming-statistics problem whose *optimal* answer (Welford's online algorithm) and *seductive-but-wrong* answer (naive sum-of-squares, which suffers catastrophic cancellation on a large baseline) were **never named in the input** — only constraints. The author sealed a prediction beforehand.\n\n**What happened:** the conclave forged a numerical-accuracy criterion, **wrote Python that computed exact ground-truth variance via rational arithmetic, and ran every candidate against it**:\n\n```\nnaive sum-of-squares :  relerr 1.5e0   → FAIL (variance even went NEGATIVE: -26850)\nWelford              :  relerr 3.3e-11 → PASS\nshifted/assumed-mean :  relerr 3.5e-14 → PASS\n```\n\nIt converged on **Welford** (5 of 6 surviving candidates) with shifted-mean as the legitimate runner-up — exactly the sealed prediction — and **proved the popular wrong answer wrong by executing code**, not by arguing. Grounding ratio: **100%**.\n\n\u003e It returned `PARETO` rather than crowning one Welford variant, because *all* survivors passed *all* executed checks — there was no grounded margin between equally-correct answers. Honest, if conservative. (See [issue #1](docs/examples/01-streaming-variance.md#the-finding) — the gate should treat \"executed elimination of all distinct rivals\" as a SINGLE condition.)\n\n### 2. A genuinely contested decision ([full transcript](docs/examples/02-webhook-queue.md))\n\nThe Kafka-vs-Redis-vs-SQS webhook problem above. The conclave forged fault-injection and load-saturation checks (and ran in-process simulations of each: *\"leader-only-ack run lost 36000 acked writes while durable run lost 0\"*), but the candidates **tied on the executed durability/burst checks and differed only on *estimated* cost/ops dimensions**. So it correctly returned `PARETO` — grounded margin 0%, `n_eff` 1.38 — with a ranked shortlist and the named cost/ops assumptions you'd have to accept to pick. **It refused to fake a winner.**\n\n---\n\n## Why not just… ?\n\n| Approach | Why it doesn't solve this |\n| --- | --- |\n| **Lower the temperature** | Controls diversity, not correctness; semantic variance lives in the distribution, not the sampler. |\n| **Self-consistency / majority vote** | Assumes one extractable correct answer. Open-ended design has none — and a confident wrong mode just wins the vote. |\n| **Ask 5 models, take the majority** | Same-lineage models share blind spots; you get one opinion wearing five hats (the `n_eff` problem). |\n| **Multi-agent debate** | Debate is a belief-martingale: it converges to *consensus*, not *truth*, and rewards persuasiveness. |\n| **LLM-as-judge** | The judge is another probabilistic, persuadable model. Conclave's \"judge\" has no authority — only re-executed artifacts decide. |\n\nConclave's bet: don't add another confident model layer — **route the decision to executed reality where it exists, and abstain honestly where it doesn't.**\n\n---\n\n## Status \u0026 honest limitations\n\n**This is a research preview, not a validated product.** Limitations that survive adversarial review (the same honesty the tool enforces, applied to itself):\n\n- **The verifiable slice is small.** Structural/perf/cost checks are decisive; the *most consequential* calls (bounded-context boundaries, build-vs-buy, 3-year evolvability) resist executable checks. There the conclave correctly **abstains to a human** — it degrades to an honest, well-instrumented escalation router, not an oracle.\n- **The constitution's completeness is unprovable.** It's drafted by Claude-family models and supermajority-filtered, so a blind spot shared across *all* electors survives as \"consensus.\" The tool makes this foundation **inspectable** (it lists excluded criteria and reports constitution-level independence) rather than hiding it. Re-run with a human-ratified constitution to harden it.\n- **Claude-only independence is limited.** In a Claude-only environment cross-model-family diversity is unavailable; `n_eff` measures *commitment* divergence, not true *error* independence, so it can look healthier than it is. The verdict says so, and tells you to trust the grounding dial over agreement.\n- **\"Wrong-but-green verifier\" risk is reduced, not eliminated.** Forging negative controls catches the obvious cases; a subtly mis-specified check can still pass and lend false authority.\n- **Benchmark-to-reality transfer is unproven.** The composed techniques were validated on tasks *with* ground truth; their transfer to no-oracle design is plausible and hedged (only verifiable parts ever decide) but not proven.\n\nSee [`docs/DESIGN.md`](docs/DESIGN.md) for the full grounded design rationale and citations.\n\n## Roadmap\n\n- [ ] Gate rule: treat *executed elimination of all distinct rivals* as a `SINGLE`-winner condition even at zero margin (the finding from test run #1).\n- [ ] Human-ratification checkpoint on the frozen constitution (accept/veto, no silent edit).\n- [ ] Pluggable cross-family elector hook (use non-Anthropic models when credentials exist; degrade gracefully when not).\n- [ ] Persist `constitution.\u003chash\u003e.json` artifacts for audit \u0026 replay.\n- [ ] A small benchmark of decisions-with-known-answers to calibrate `τ_g` / `τ_n`.\n\n## Contributing\n\nPRs welcome — see [CONTRIBUTING.md](CONTRIBUTING.md). Good first issues are tagged. The cardinal rule: **don't add a feature that lets a model's *opinion* decide something — only executed artifacts or frozen rules may decide.**\n\n## How it was built\n\nFittingly, Claude Conclave was designed *using its own method*: a multi-agent workflow that ran parallel grounded literature research, generated competing solution architectures from incompatible philosophies, adversarially red-teamed each, and synthesized the survivors — with every load-bearing citation independently verified. The design doc is [`docs/DESIGN.md`](docs/DESIGN.md).\n\n## License\n\n[MIT](LICENSE).\n\n## Acknowledgements\n\nGrounded in published work on semantic entropy \u0026 uncertainty (Kuhn et al.; Farquhar et al.), LLM-as-judge reliability and panels (Zheng et al.; Verga et al.), multi-agent debate dynamics (Du et al.; Khan et al.), correlated-judge effective sample size (the \"nine judges, two effective votes\" result), sycophancy (Sharma et al.), and classical architecture-decision rigor (SEI ATAM; evolutionary-architecture fitness functions). Full citations in [`docs/DESIGN.md`](docs/DESIGN.md).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fadithyan-ak%2Fclaude-conclave","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fadithyan-ak%2Fclaude-conclave","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fadithyan-ak%2Fclaude-conclave/lists"}