{"id":51173539,"url":"https://github.com/kazi-org/kazi","last_synced_at":"2026-07-19T02:01:14.387Z","repository":{"id":366710790,"uuid":"921901248","full_name":"kazi-org/kazi","owner":"kazi-org","description":"Make your coding agent actually finish the job. Install one skill and Claude Code keeps working — planning, fixing, testing, deploying — until your goal is objectively true (tests pass, endpoint live, deployed), or it tells you why.","archived":false,"fork":false,"pushed_at":"2026-07-14T01:06:45.000Z","size":5964,"stargazers_count":2,"open_issues_count":23,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-07-14T02:17:58.498Z","etag":null,"topics":["agent-orchestration","agentic","ai-agents","automation","claude-code","codex","coding-agent","developer-tools","elixir","llm","mcp","reconciliation"],"latest_commit_sha":null,"homepage":"https://kazi.sire.run","language":"Elixir","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/kazi-org.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":"NOTICE","maintainers":null,"copyright":null,"agents":"AGENTS.md","dco":null,"cla":null}},"created_at":"2025-01-24T20:45:20.000Z","updated_at":"2026-07-14T01:06:37.000Z","dependencies_parsed_at":null,"dependency_job_id":"13474a1b-e729-4072-9881-c91dfff8b3bc","html_url":"https://github.com/kazi-org/kazi","commit_stats":null,"previous_names":["kazi-org/kazi"],"tags_count":195,"template":false,"template_full_name":null,"purl":"pkg:github/kazi-org/kazi","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kazi-org%2Fkazi","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kazi-org%2Fkazi/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kazi-org%2Fkazi/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kazi-org%2Fkazi/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/kazi-org","download_url":"https://codeload.github.com/kazi-org/kazi/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kazi-org%2Fkazi/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35637763,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-19T02:00:06.923Z","response_time":112,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["agent-orchestration","agentic","ai-agents","automation","claude-code","codex","coding-agent","developer-tools","elixir","llm","mcp","reconciliation"],"created_at":"2026-06-27T02:04:55.036Z","updated_at":"2026-07-19T02:01:14.376Z","avatar_url":"https://github.com/kazi-org.png","language":"Elixir","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n  \u003cpicture\u003e\n    \u003csource media=\"(prefers-color-scheme: dark)\" srcset=\"assets/logo/kazi-wordmark-dark.svg\"\u003e\n    \u003cimg alt=\"kazi\" src=\"assets/logo/kazi-wordmark.svg\" width=\"260\"\u003e\n  \u003c/picture\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://kazi.sire.run\"\u003e\u003cb\u003eWebsite\u003c/b\u003e\u003c/a\u003e \u0026nbsp;\u0026middot;\u0026nbsp;\n  \u003ca href=\"https://kazi.sire.run/proof\"\u003eProof\u003c/a\u003e \u0026nbsp;\u0026middot;\u0026nbsp;\n  \u003ca href=\"https://kazi.sire.run/blog\"\u003eBlog\u003c/a\u003e \u0026nbsp;\u0026middot;\u0026nbsp;\n  \u003ca href=\"docs/concept.md\"\u003eConcept\u003c/a\u003e \u0026nbsp;\u0026middot;\u0026nbsp;\n  \u003ca href=\"https://github.com/kazi-org/kazi/releases\"\u003eReleases\u003c/a\u003e \u0026nbsp;\u0026middot;\u0026nbsp;\n  \u003ca href=\"https://github.com/kazi-org/homebrew-tap\"\u003eHomebrew tap\u003c/a\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://img.shields.io/github/stars/kazi-org/kazi?style=flat-square\u0026logo=github\u0026color=f1c40f\u0026labelColor=555555\" alt=\"GitHub Stars\"/\u003e\n  \u003ca href=\"https://github.com/kazi-org/kazi/releases/latest\"\u003e\u003cimg src=\"https://img.shields.io/github/v/release/kazi-org/kazi?style=flat-square\" alt=\"Latest Release\"/\u003e\u003c/a\u003e\n  \u003cimg src=\"https://img.shields.io/badge/license-Apache--2.0-blue?style=flat-square\" alt=\"License: Apache-2.0\"/\u003e\n  \u003ca href=\"https://github.com/kazi-org/kazi/actions/workflows/ci.yml\"\u003e\u003cimg src=\"https://img.shields.io/github/actions/workflow/status/kazi-org/kazi/ci.yml?branch=main\u0026style=flat-square\u0026label=CI\" alt=\"CI Status\"/\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n# kazi\n\n**Your coding agent says \"done.\" kazi proves it.**\n\nGive **Claude Code** the power to actually *finish*. You chat with Claude the way you\nalready do; kazi works in the background to make \"done\" **objective** — looping your agent\nuntil every check passes (tests green, the endpoint live, the change deployed), or stopping\nto tell you why (`stuck`, or out of budget) instead of pretending it's finished.\n**You never run kazi yourself — Claude does.**\n\n## Try it in 10 seconds\n\n```sh\nbrew install kazi-org/tap/kazi\nkazi install-skill        # teaches Claude Code the kazi skill (writes ~/.claude/skills/kazi:\n                          # SKILL.md + AUTHORING.md + RECIPES.md; your own LOCAL.md is never touched)\n```\n\n**Or install the Claude Code plugin** — one marketplace install bundles the skill, the\nkazi MCP server, and the session-bus hooks together, and refreshes them on the release\ncadence instead of on-demand re-runs of the explicit commands\n([ADR-0077](docs/adr/0077-claude-code-plugin-distribution.md)):\n\n```text\n/plugin marketplace add kazi-org/claude-plugins\n/plugin install kazi@kazi\n```\n\nBoth paths are fully supported and neither is \"the\" way — pick the explicit `kazi\ninstall-skill` commands or the plugin. See [Install via the Claude Code\nplugin](#install-via-the-claude-code-plugin) for the full comparison. You still need the\n`kazi` binary on your `PATH` (`brew install kazi-org/tap/kazi`) either way.\n\nThen drive it from Claude Code with the kazi skill — author the checks, then converge:\n\n```text\n/kazi plan \"add a /healthz endpoint that returns 200 ok, with a test, deployed\"\n# Claude drafts the acceptance predicates; glance at them, then:\n/kazi apply\n```\n\nClaude loops — editing, testing, deploying — and reports back only when every predicate\nis *objectively* true (or it is genuinely `stuck`). You never leave your chat with Claude.\n\nPrefer plain English? Just say **have kazi drive this until done** — the skill runs the\nsame `/kazi plan` → `/kazi apply` for you.\n\n## How it works\n\nUnder the hood, *kazi* is **the outer/reconciliation loop for coding agents** — your agent\nruns it, not you. (*kazi* is Swahili for *work / a job*.) It *drives* the coding agent you\nalready use (Claude Code, Codex, …) in a reconcile loop: observe every failing check,\ndispatch a fix, integrate, re-check.\n\nThink of it like **Kubernetes for coding goals**: you declare desired state, kazi\nwatches actual state, and it keeps closing the gap until the two match.\n\n```mermaid\nflowchart TD\n    U([\"You: 'build a URL-shortener web service and ship it live in production'\"]) --\u003e K\n\n    subgraph kazi [kazi reconcile loop]\n        O[\"Observe\u003cbr/\u003eWhat's failing?\"] --\u003e D[\"Dispatch\u003cbr/\u003ean agent to fix it\"]\n        D --\u003e C{\"Every check passes?\"}\n        C -- No --\u003e O\n        C -- Yes --\u003e I[\"Integrate\u003cbr/\u003ePR / Merge\"]\n        I --\u003e Dep[\"Deploy and Verify Live\"]\n    end\n\n    style U fill:#1e293b,stroke:#cbd5e1,color:#f8fafc\n    style kazi fill:#0f172a,stroke:#334155,color:#f8fafc\n    style O fill:#334155,stroke:#475569,color:#f8fafc\n    style D fill:#334155,stroke:#475569,color:#f8fafc\n    style C fill:#334155,stroke:#475569,color:#f8fafc\n    style I fill:#334155,stroke:#475569,color:#f8fafc\n    style Dep fill:#334155,stroke:#475569,color:#f8fafc\n```\n\nThat loop drives the failing predicates to zero and only then reports done. The\nrecording below is a **real** `kazi apply` run (not a mockup): a goal whose one\nacceptance predicate — *\"`go test` passes\"* — is false at t0, driven by the\n`claude` harness until it is objectively true. Reproduce it with\n[`priv/examples/hero_cast_demo`](priv/examples/hero_cast_demo) (the committed\nasciicast is [`assets/proof-loop.cast`](assets/proof-loop.cast)):\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"assets/proof-loop.gif\" alt=\"A real kazi apply run: kazi.loop reports iter=1 failing=[\u0026quot;tests-pass\u0026quot;], then iter=2 failing=[], then CONVERGED — every predicate is satisfied (predicate vector: [pass] tests-pass).\" width=\"760\"\u003e\n\u003c/p\u003e\n\nIt is **not** another coding agent, terminal, or IDE. kazi *drives* the agent you\nalready use. As that agent gets better, kazi gets better for free.\n\n---\n\n## Why kazi?\n\nTwo problems nobody else owns:\n\n1. **\"Done\" is the agent's opinion.** A coding agent stops when it *thinks* it's\n   finished — even when the work is merely plausible. kazi makes \"done\" objective:\n   the loop can only succeed when *every* check (kazi calls them **predicates**)\n   evaluates true, with stored evidence. Truth lives in the controller, not the agent.\n2. **Parallel agents collide.** Locking a *task* doesn't stop two agents editing the\n   *same files*. kazi coordinates on **resources** — an agent leases its \"blast\n   radius\" before touching code — so concurrent runs converge instead of conflict.\n3. **Bring Your Own Model (BYOM) \u0026 Privacy.** Use cloud models or run entirely locally. Wire kazi to a local model (e.g., Llama 3, Qwen) via `opencode` for zero data leaks. Your code and context never leave your hardware.\n\n\u003e **Why now?** Coding agents are finally good enough to do real work — ship\n\u003e features, fix bugs, wire up tests. That is exactly *why* kazi exists: once\n\u003e agents act autonomously, you need a **controller above them** to decide when\n\u003e they are truly done. That layer didn't exist until now, and it's precisely what\n\u003e kazi is.\n\n---\n\n## The 60-second mental model\n\nA **goal** is just a list of checkable statements plus a budget:\n\n- **predicates** — the checks that define \"done\": `the unit tests pass`, `GET /health\n  returns 200 ok`, `the production error rate is 0 over 30m`, …\n- a **budget** — a hard ceiling (iterations / wall-clock / tokens) so it can never\n  run forever or burn money.\n- a **scope** — the repo + paths agents are allowed to touch.\n\nkazi loops: **observe** every predicate → the failing ones *are* the to-do list →\n**dispatch** an agent to fix them → **integrate** (open a PR, rebase-merge) →\n**deploy** → **re-check**. It stops only when all predicates are true (`converged`),\nthe same checks keep failing (`stuck` → escalate to you), or the budget runs out.\n\n---\n\n## With kazi vs. without\n\n| Without kazi | With kazi |\n|---|---|\n| *\"The agent says it's done.\"* You trust it on faith. | **Every predicate verified true**, with stored evidence. Truth lives in the controller, not the agent. |\n| Two parallel agents edit the same files → merge conflicts. | Agents **lease their blast radius** first — disjoint work runs free, overlapping work serializes. |\n| Green tests on a laptop, broken in production. | A **live predicate** probes the *deployed* endpoint. Green-on-my-machine is never enough. |\n| It stops when it *feels* finished. | It stops only on `converged`, `stuck`, or `over-budget` — and tells you which. |\n\n---\n\n## Proof — real goals kazi converged\n\nGoals a naive pipeline leaves *subtly broken* — the file looks created, the answer looks\nright, the parallel split looks done — that kazi drove to **objective convergence**. Every\nnumber below is copied verbatim from a real `kazi apply --json` run recorded in\n[`docs/devlog.md`](docs/devlog.md); the full reproduction steps (goal-file, command, version,\nand the metered `cost_usd` per run) are in\n[`docs/dogfood-methodology.md`](docs/dogfood-methodology.md). No unverifiable claims —\nif a number can't be traced to a captured run, it isn't here.\n\n| Goal | The subtle break | Result | Source |\n|---|---|---|---|\n| **Exact-content file** — `VERSION.txt` must be exactly `1.0.0` | \"The file exists\" looks done, but the bytes must be exact | `converged`, 2 iters / 18.5 s / 39,712 tokens (v1.64.2) | T26.6 |\n| **Self-correcting on an opaque oracle** — `solution.py` graded by a one-way sha256 | The first plausible attempt is *wrong*; nothing grades it | `converged`, 2 iters / 39.3 s, Haiku self-corrected on iter 2 (v1.64.1) | T30.4 |\n| **A real cross-group dependency, parallelized** — streaming consumes a contract type | A naive parallel split compiles against a type that doesn't exist yet | `collective: converged`, 2 disjoint groups concurrent + 1 gated, single-node, NATS-free (v1.64.2) | T21.12 / T23.9 |\n\nFull gallery + before/after evidence + reproducible method (incl. per-run cost):\n**\u003chttps://kazi.sire.run/proof\u003e** (see also the [methodology doc](docs/dogfood-methodology.md)).\n\n---\n\n## How the skill routes\n\n`install-skill` adds a trigger to the kazi skill, so `/kazi plan` / `/kazi apply` — and the\nplain-English phrase **have kazi drive this until done** — only route to kazi once you have\ninstalled it. From there Claude authors the acceptance predicates with `kazi plan` and runs\n`kazi apply` until they are *objectively* true. You do not operate kazi directly; your\nagent does.\n\nThe skill ships as three files (ADR-0074): `SKILL.md` (the router), `AUTHORING.md`\n(predicate authoring quality — dense briefs, capability-vs-guard, the red-at-t0 rule), and\n`RECIPES.md` (escalation ladder, streaming, the check-only gate variant, the session bus).\nIt is fully self-contained — it never assumes any other skill exists. To wire kazi into\nyour own local workflow (routing conventions, model policy), put them in\n`~/.claude/skills/kazi/LOCAL.md`: the skill reads it first when present, and\n`install-skill` never overwrites it. That is a **stable path**, deliberately decoupled\nfrom wherever the skill content is installed (ADR-0077): a Claude Code plugin update\nreplaces the skill directory wholesale, so keeping your customization at the stable path\nmeans it survives plugin updates, re-installs, and any future relocation of the skill\ndirectory. If you already have a `LOCAL.md` inside an old skill directory, `install-skill`\nmigrates it to the stable path (or warns if one exists at both — it never silently\ndiscards your wiring).\n\n---\n\n## Token economy without local models\n\nThe cheapest way to run an agent loop isn't a local GPU — it's spending frontier\nreasoning **once**, then grinding on a cheap model that the predicates keep honest.\nkazi makes that an **in-family Claude** move, so any Claude Code user gets it with\n**no local model and no local GPU host**\n([ADR-0033](docs/adr/0033-cheaper-via-in-family-claude-tiering.md)):\n\n- You **chat with Claude Code**; it drives kazi.\n- **The grind** runs on the **default grind tier** (**Sonnet 5**, `claude-sonnet-5`) —\n  cheaper than the frontier, reliable enough that the loop converges instead of\n  burning iterations (re-derive the tiering call from your own fleet any time:\n  `kazi economy --json`).\n- **Hard reasoning** runs on a **frontier model** (e.g. **Opus 4.8**, `claude-opus-4-8`).\n- The **predicates keep the cheap model honest** — it cannot declare a false \"done,\"\n  because convergence is gated on objective checks, not the model's say-so.\n\nYou pay frontier rates only for the judgment (authoring the predicates) and cheap\nrates for the bulk of the work.\n\n### Worked example — author on a frontier model, grind on a cheap one\n\n```sh\n# 1. Author the acceptance predicates ONCE — let your session's frontier model\n#    (e.g. Opus 4.8) draft them with `kazi plan`. It PRINTS a proposal-ref:\nkazi plan \"add a /healthz endpoint that returns 200 ok\" --workspace ./svc\n# review the draft (or run `kazi list-proposed` to see pending refs), then approve\n# it by passing the ref `kazi plan` printed:\nkazi approve \u003cproposal-ref\u003e\n\n# 2. Drive the N-iteration grind on the default grind tier — no local model needed:\nkazi apply my-goal.toml --workspace ./svc --harness claude --model claude-sonnet-5\n```\n\n`--harness claude --model \u003cid\u003e` selects which Claude model runs that call\n([ADR-0033](docs/adr/0033-cheaper-via-in-family-claude-tiering.md)); `claude` is\nalready the default harness, so the only new thing is naming a cheaper model for\nthe grind.\n\n### Start on the default tier, escalate on stuck (the smart default)\n\nStatic tiering always grinds on one model. The adaptive default — **start on the\ndefault grind tier, step UP only when kazi reports the same slice is stuck** —\npays frontier rates only for the slices that actually need them\n([ADR-0035](docs/adr/0035-skill-driven-adaptive-model-tiering.md), amended on\nfleet data — `claude-haiku-4-5` is an explicit opt-down for known-trivial\nslices, not a ladder rung):\n\n```\nclaude-sonnet-5  -\u003e  claude-opus-4-8   (cap — do not escalate past Opus)\n```\n\nThe policy lives in the orchestrating skill, **never in kazi**: kazi reports\nper-iteration state (`converged` / `stuck` / `over_budget`) via `kazi apply --json`,\nand the skill owns the ladder and the per-slice rung counter. The full\ncopy-paste recipe is in [`AGENTS.md`](AGENTS.md) (\"Escalate-on-stuck\") and the\ninstalled `kazi` skill.\n\n\u003e **Designed-for, not yet measured.** The cost win is the *intended* economics —\n\u003e frontier judgment once, cheap iterations gated by predicates. The headline dollar\n\u003e figure is being measured by the multi-iteration benchmark; until it lands, we state\n\u003e the *shape* of the saving, not an unproven number.\n\n\u003e **Want full privacy instead?** Local / bring-your-own-model is the **secondary**\n\u003e option: point kazi at a local model via `opencode` so your code and context never\n\u003e leave your hardware (see [Use a different coding harness](#use-a-different-coding-harness)).\n\u003e It trades the in-family convenience for on-prem privacy.\n\n---\n\n## What a coding agent says\n\n\u003e *\"Left to myself, I'll tell you a task is done the moment the code looks\n\u003e right. kazi won't let me — it holds the predicates and re-checks them against\n\u003e reality, so I stop claiming 'done' when it isn't. I end up shipping the thing\n\u003e you actually asked for, not the thing I hoped was finished.\"*\n\u003e\n\u003e — Claude (Anthropic), describing kazi in its own words. Agent-authored, kept\n\u003e verbatim and labelled as such — not a human testimonial.\n\n---\n\n## Who it's for\n\n- **Builders who ship fast but need reliability** — if an agent has ever\n  \"finished\" something that wasn't actually done, objective termination is the\n  guardrail against plausible-but-broken output.\n- **Teams running parallel coding agents** — resource leases coordinate who edits\n  what *before* any file changes, so concurrent runs converge instead of collide.\n- **Engineers who refuse \"works on my machine\"** — predicates can verify the live,\n  deployed system, not just the local checkout.\n\n**Not for you (yet) if:** you want an agent to decide *what* to build — that's your\ncall; kazi only drives toward an outcome you declare. It also needs a coding\nharness (`claude`, `opencode`, …) on your `PATH`; kazi drives one, it isn't one.\n\n---\n\n## Install\n\nThe fastest way — a single self-contained binary via Homebrew (no Erlang\nprerequisite; ERTS and the SQLite NIF are bundled, so you get the full read-model):\n\n```sh\nbrew install kazi-org/tap/kazi\nkazi --help\n```\n\nPrebuilt binaries are published for **Apple Silicon macOS**, **x86_64 Linux**, and\n**ARM Linux** (`aarch64`) on each [GitHub Release](https://github.com/kazi-org/kazi/releases)\n(Intel macOS is not yet built — build from source, below). The binary is a\nBurrito wrap of a `mix release` ([ADR-0014](docs/adr/0014-binary-distribution-burrito-homebrew.md)),\nso unlike the escript it carries the native `exqlite` NIF and persists every\niteration.\n\n\u003e **Runtime requirement:** kazi DRIVES a coding agent ([ADR-0001](docs/adr/0001-positioning-outer-loop-reconciler.md));\n\u003e it does not bundle one. A harness binary — `claude` (default) or `opencode` —\n\u003e must be on your `PATH` to actually run a goal.\n\nTo build from source instead, you need **Elixir / Erlang** (OTP 26+) and `mix`,\nplus **git** (kazi commits/opens PRs in your target repo) and *(optional, for live\ndeploys)* **gcloud** / a deploy command and `gh`:\n\n```sh\ngit clone https://github.com/kazi-org/kazi \u0026\u0026 cd kazi\nmix setup         # deps + SQLite read-model + git hooks (core.hooksPath -\u003e .githooks)\nmix test          # hermetic test suite, should be green\n```\n\nContributing? `mix setup` also wires the repo's committed git hooks\n(`.githooks/`), including a pre-push guard that rejects direct pushes to\n`main` — `main` auto-releases via release-please, so all changes go through a\nbranch + PR + rebase-merge. If you skip `mix setup`, wire the hooks by hand\nwith `git config core.hooksPath .githooks`.\n\nTwo ways to invoke kazi (same behavior):\n\n```sh\n# Mix task — recommended. Boots the full app and persists every iteration to a\n# local SQLite read-model (created + migrated automatically on first run).\nmix kazi.apply \u003cgoal-file\u003e --workspace \u003cpath-to-your-project\u003e\n\n# Or build a standalone binary:\nmix escript.build          # produces ./kazi\n./kazi apply \u003cgoal-file\u003e --workspace \u003cpath-to-your-project\u003e\n./kazi --help\n```\n\n### Use a different coding harness\n\n`claude` (Claude Code) is the **default** harness — if you do nothing, kazi\nshells out to `claude` exactly as before. But kazi drives whatever CLI coding\nagent you already have installed and configured; it does not reimplement provider\nplumbing ([ADR-0016](docs/adr/0016-generic-harness-profiles.md)). Pick another\nharness per-run with a flag:\n\n```sh\nkazi apply \u003cgoal-file\u003e --workspace \u003cpath\u003e \\\n  --harness opencode --model local-ollama/qwen3.6:35b-a3b\n```\n\n`--harness \u003cid\u003e` selects the harness (`claude`, `opencode`, `codex`,\n`antigravity`, `claw`, or `gemini_cli` today — see the [tier table](#tiered-harness-support-adr-0022)\nbelow); `--model \u003cprovider/model\u003e` selects the model that harness should use.\n\n**Point opencode at a local model (e.g. a locally-hosted Qwen3.6).** If you run\n[`opencode`](https://opencode.ai) wired to a local model (for example a\n**Qwen3.6 35B-A3B** on a local GPU host), **opencode's own provider config\nis the source of truth** for the endpoint and credentials. `--model` is\nopencode's `provider/model` string — the provider (`local-ollama` above) and its\nbase URL live in your opencode config, not in kazi. kazi can also forward\nprovider/endpoint environment variables to the harness subprocess when a local\nsetup expects them — declare them as the harness `:env` and kazi passes them\nstraight through to the underlying call.\n\n**Workspace isolation.** `opencode run` resolves its own project root rather\nthan honoring the directory it is launched from, so kazi passes the goal's\n`--workspace` explicitly as `opencode run … --dir \u003cworkspace\u003e` — the inner\nagent's edits always land in the goal's workspace, where the acceptance\npredicates are evaluated.\n\n#### A per-goal or global default\n\nYou don't have to pass `--harness`/`--model` on every run. A goal-file can carry\nits own preferred harness in an optional `[harness]` table:\n\n```toml\n[harness]\nid = \"opencode\"                            # a KNOWN harness id (claude / opencode / codex / antigravity / claw)\nmodel = \"local-ollama/qwen3.6:35b-a3b\"       # optional provider/model override\ncommand = \"opencode\"                       # optional binary override\neffort = \"high\"                            # optional Claude-only reasoning-effort level\npermission_mode = \"acceptEdits\"            # optional Claude-only permission mode\nallowed_tools = [\"Write\", \"Bash\", \"Edit\"]  # optional Claude-only tool allow-list\n```\n\n`permission_mode` / `allowed_tools` are Claude-only (silently dropped by every\nother harness) and forward `claude --permission-mode \u003cmode\u003e` /\n`claude --allowed-tools \u003ct\u003e …`. They're also settable per run as CLI flags —\n`--permission-mode` / `--allowed-tools` (each overrides the matching\n`[harness]` field) — which matters for a **headless dispatch against a\nworkspace that has not been through Claude Code's interactive trust dialog**,\nwhere every tool call would otherwise be silently denied.\n\nOr set a machine-wide default in app config:\n\n```elixir\n# config/config.exs\nconfig :kazi, :harness, :opencode\n```\n\nkazi resolves the harness with a fixed **precedence** (highest first):\n\n1. the **`--harness` / `--model` CLI flags**;\n2. the goal-file **`[harness]` table**;\n3. the **app config** `config :kazi, :harness`;\n4. the default, **`claude`**.\n\nSo a CLI flag always wins; absent every layer, kazi drives `claude`.\n\n#### Add a harness = declare a profile\n\nThere is no new adapter module per harness. A harness is a **profile** — a value\nin [`Kazi.Harness.Registry`](lib/kazi/harness/registry.ex) built from\n[`Kazi.Harness.Profile`](lib/kazi/harness/profile.ex): a `command`, an argv\nrenderer (`build_args`), and a stdout parser (`parse`), plus the set of optional\nflags the harness understands. One generic adapter (`Kazi.Harness.CliAdapter`)\nruns every profile. Adding `codex`, `gemini-cli`, etc. is profile DATA — often\nreusing an existing parser — not a new module; a fully custom harness can be\ndeclared in config without touching kazi.\n\n\u003e **Runtime requirement.** The chosen harness binary (`claude`, `opencode`, …)\n\u003e must be installed and on your `PATH`. kazi shells out to it as a subprocess; it\n\u003e does not bundle or install harnesses.\n\n#### Tiered harness support (ADR-0022)\n\nNot every CLI agent clears the same bar. kazi drives every harness as a\nnon-interactive subprocess and parses its stdout, so a harness is **first-class**\nonly when it runs from a single prompt AND emits machine-parseable output\n(JSON/JSONL) correctly under a non-TTY subprocess\n([ADR-0022](docs/adr/0022-harness-onboarding-conformance.md)). Some tools are\nadded with a documented workaround, and one is **best-effort only**:\n\n| Harness (`--harness`) | Tier | Notes |\n| --- | --- | --- |\n| `claude` (default) | First-class | single JSON envelope; full cost/token parse. |\n| `opencode` | First-class | NDJSON event stream; point it at a local model. kazi passes `--dir \u003cworkspace\u003e` (opencode ignores the launch cwd), so edits land in the goal's workspace. |\n| `codex` | First-class | `codex exec … --json` JSONL stream; auth `OPENAI_API_KEY` / `codex login`. |\n| `antigravity` | Conformant **with a workaround** | non-TTY stdout bug (`antigravity-cli#76`) handled via `--prompt-file --output json`; auth `GEMINI_API_KEY` / `ANTIGRAVITY_API_KEY`. |\n| `claw` | **Best-effort / demo-grade** | claw-code emits **no** structured output and has no model flag — kazi surfaces its raw stdout as the result with **no cost/token extraction**. It runs, but fidelity is degraded; treat it as a demo (\"an agent-managed museum exhibit, not a production tool\"), not a budgeted production run. Auth is via env API keys (`ANTHROPIC_API_KEY` / `OPENAI_API_KEY`). |\n| `gemini_cli` | First-class | `gemini -p … -o json` single JSON envelope; `--approval-mode yolo` runs non-interactively; full result + best-effort token parse; auth `GEMINI_API_KEY` (or Google OAuth / Vertex `GOOGLE_API_KEY`). |\n\n### Build a self-contained release (full read-model)\n\nThe escript can't bundle the native SQLite NIF, so it runs **without** the\nread-model. A `mix release` bundles ERTS *and* the compiled NIFs, so the released\nbinary has the **full read-model** (and is the foundation the per-platform binary\nis built from — see [ADR-0014](docs/adr/0014-binary-distribution-burrito-homebrew.md)):\n\n```sh\nMIX_ENV=prod mix release --overwrite     # builds _build/prod/rel/kazi\n\n# The CLI is invoked through the release's `eval` command, which propagates the\n# CLI's exit code (0 on convergence / a recorded proposal / approval, non-zero\n# otherwise) — so the release composes in scripts and CI like the escript:\n_build/prod/rel/kazi/bin/kazi eval 'Kazi.Release.cli([\"--help\"])'\n_build/prod/rel/kazi/bin/kazi eval \\\n  'Kazi.Release.cli([\"apply\", \"\u003cgoal-file\u003e\", \"--workspace\", \"\u003cpath\u003e\"])'\n_build/prod/rel/kazi/bin/kazi eval 'Kazi.Release.cli([\"list-proposed\"])'\n```\n\n`Kazi.Release.cli/1` dispatches to the same `Kazi.CLI` core as the escript and\n`mix kazi.apply`, so every subcommand (`apply` / `plan` / `list-proposed` /\n`approve` / `reject` / `--help`) behaves identically.\n\n### Build a single-file native binary (Burrito)\n\n[Burrito](https://github.com/burrito-elixir/burrito) wraps the `mix release`\nabove into one self-contained per-platform executable that bundles ERTS **and**\nthe compiled exqlite NIF — so the binary has the **full SQLite read-model** with\nno Erlang prerequisite on the user's machine (T6.2, [ADR-0014](docs/adr/0014-binary-distribution-burrito-homebrew.md)).\nThe `kazi` release declares four targets: macOS `aarch64`/`x86_64` and Linux\n`aarch64`/`x86_64`.\n\nBuilding requires [Zig](https://ziglang.org) **0.15.2** (Burrito's pinned\nversion) and `xz` on `PATH`; cross-target builds also need `7z` for Windows\n(kazi ships no Windows target). Build the host target and run it:\n\n```sh\n# Build the binary for the current host platform (set BURRITO_TARGET to one of\n# macos_aarch64 / macos_x86_64 / linux_aarch64 / linux_x86_64; omit it to build\n# every declared target). Output lands in ./burrito_out/.\nBURRITO_TARGET=macos_aarch64 MIX_ENV=prod mix release --overwrite\n\n# The wrapped binary takes the CLI args directly — no `eval`. It reads them via\n# Burrito's argv and dispatches through the same Kazi.CLI core:\n./burrito_out/kazi_macos_aarch64 --help\n./burrito_out/kazi_macos_aarch64 apply \u003cgoal-file\u003e --workspace \u003cpath\u003e\n./burrito_out/kazi_macos_aarch64 list-proposed\n```\n\nThe binary persists its read-model to `$KAZI_DB` if set, otherwise\n`~/.kazi/kazi.db` (created on first run; see `config/runtime.exs`). Unlike the\nescript, every iteration and proposal is persisted — the NIF is bundled.\n\nIf the VM crashes (e.g. the installed binary was replaced mid-run by an\nin-place upgrade), `erl_crash.dump` is written under `$KAZI_STATE_DIR/crash` if\nset, otherwise `~/.kazi/crash` — never the workspace CWD (`Kazi.CrashDump`,\nissue #856). And if the crash is specifically caused by the on-disk release\nhaving changed underneath the running VM, every entry point (escript, `mix\nkazi.run`, the release `eval` path, the Burrito binary) reports one clear line\n— `installed release changed under a running VM -- re-run under the new\nbinary` — instead of a harness stack trace plus Logger formatter-crash spam\n(`Kazi.SwapDiagnosis`).\n\nIf a CLI invocation ever hangs at startup (issue #1255), a diagnostic watchdog\n(`Kazi.StartupWatchdog`) fires after a deadline and prints to STDERR *where* the\nprocess is stuck — its current stacktrace, run-queue lengths, and open ports/fds\n— so a hang is diagnosable in seconds instead of a from-scratch native-stack\ninvestigation. It defaults to dump-and-CONTINUE (a slow-but-healthy startup is\nnever turned into a failure). Tune with `KAZI_STARTUP_WATCHDOG_MS` (deadline in\nms, default `30000`; `0` disables) and opt into a hard exit on timeout with\n`KAZI_STARTUP_WATCHDOG_HALT=1` (exit code `124`). The Burrito extraction step runs\nbefore the BEAM, so its time is not counted against the deadline.\n\n\u003e **macOS 26 + Zig note.** Burrito 1.5.0 pins Zig **0.15.2**, which cannot link\n\u003e native binaries against the macOS 26 SDK (Xcode 26); Zig 0.16 links it but is\n\u003e API-incompatible with Burrito's `build.zig`. On a macOS 26 host the wrap step\n\u003e fails at the Zig link; build the macOS binaries on a macOS 15 (or earlier)\n\u003e runner — which is what the release CI matrix (T6.3) targets.\n\n### Install via the Claude Code plugin\n\nkazi also ships as a [Claude Code plugin](https://code.claude.com/docs/en/plugins-reference):\none install bundles the kazi skill, the kazi MCP server registration, and the\nsession-bus hooks together, and marketplace updates refresh them with the binary\nrelease cadence instead of on-demand re-runs of the explicit installers\n([ADR-0077](docs/adr/0077-claude-code-plugin-distribution.md)). The explicit\n`install-skill` / `init --with-mcp` / `install-hooks` commands are unchanged — the\nplugin is an *additional* channel rendered from the SAME sources, never a fork.\n\n**Install from the marketplace.** The release pipeline publishes the bundle to the\n[kazi-org/claude-plugins](https://github.com/kazi-org/claude-plugins) marketplace on\nevery release, at the release version (lockstep — the plugin version always equals\nthe binary release it was cut from, so marketplace content can never lag the binary):\n\n```\n/plugin marketplace add kazi-org/claude-plugins\n/plugin install kazi@kazi\n```\n\nBecause the binary (brew/direct install) and the plugin (marketplace) can be\nupgraded independently, kazi warns you if they drift out of lockstep: on\n`SessionStart` the session-bus hook (T61.5, ADR-0077) compares the local `kazi\nversion` against the installed plugin's declared version and, on a mismatch,\nemits a single line naming both versions and which channel to update. It is\nsilent when the versions match or when no plugin is installed, and never blocks\nthe session.\n\n`mix kazi.plugin` renders that bundle from those single sources of truth (the\ngenerator adds no new teaching content — it only re-renders what the installers\nalready produce):\n\n```sh\nmix kazi.plugin --out ./dist/plugin              # plugin version = the kazi version\nmix kazi.plugin --out ./dist/plugin --version 1.246.0   # pin the version (release pipeline)\n```\n\nIt writes a self-contained plugin directory:\n\n```\ndist/plugin/\n├── .claude-plugin/plugin.json   # metadata + inline MCP server + hook declarations\n└── skills/kazi/\n    ├── SKILL.md                 # the router (same content install-skill writes)\n    ├── AUTHORING.md\n    └── RECIPES.md\n```\n\nThe render is deterministic — the same version yields a byte-identical bundle\n(no timestamps, no randomness) — so the release pipeline (T61.4) can publish it\nreproducibly. `LOCAL.md` is deliberately never bundled: a plugin update replaces\nthe skill directory wholesale, so operator customization stays at the stable\n`~/.claude/skills/kazi/LOCAL.md` path outside the bundle\n([ADR-0077](docs/adr/0077-claude-code-plugin-distribution.md)).\n\n---\n\n## Quickstart 1 — describe what you want in plain English\n\nYou don't have to write a goal-file by hand, and you don't have to break the work\ndown yourself. Tell kazi the *app* (or feature) you want — as high-level as\n\"build an X\" — and it drafts the machine-checkable predicates that define \"done\"\nfor you (using your coding agent), then holds them for your review. **Nothing runs\nuntil you approve**, and you can trim or edit what it drafted:\n\n```sh\n# 1. Describe the app you want. In a terminal, kazi asks a few sharp clarifying\n#    questions FIRST (so \"done\" is precise — especially the live-verification\n#    target), then drafts the acceptance predicates and an inline rationale:\nkazi plan \"create a URL-shortener web service\" --workspace ./shortener\n#\n#   A few questions to make the goal precise (press Enter for the default):\n#   What is the live-verification target for this goal?\n#     1) A deployed URL probed over HTTP *\n#     2) Production logs / a runtime signal\n#     3) None for now — green tests are enough\n#   \u003e 1\n#   ...\n#   PROPOSED  proposal=prop-url-shortener-3f9c1a2b  goal=url-shortener\n#     • go test ./... passes\n#     • POST /shorten returns 201 with a short code for a submitted URL\n#     • GET /\u003ccode\u003e redirects (302) to the original URL\n#     • GET / renders a form to submit a URL\n#   rationale: probe the deployed shortener over HTTP; auth is out of scope for v1\n\n# 2. Review what it drafted (you're the approver — agents propose, humans dispose).\n#    Too much? Too little? Refine inline with a sharper sentence when prompted.\nkazi list-proposed\n#   prop-url-shortener-3f9c1a2b   proposed   url-shortener   (4 predicates)\n\n# 3. Approve the goal you want kazi to pursue:\nkazi approve prop-url-shortener-3f9c1a2b\n#   APPROVED   proposal=prop-url-shortener-3f9c1a2b  goal=url-shortener\n#   The goal is now runnable: kazi apply \u003cgoal-file\u003e --workspace \u003cpath\u003e\n```\n\nThe clarify phase is a HYBRID (ADR-0019): a deterministic floor of gap-checks kazi\nalways runs (it insists on a live-verification target and a scope boundary) plus\nquestions your coding agent drafts for the specific idea. Scripting it? `--yes` (or\nany non-TTY pipe) skips the questions and drafts best-effort; `--strict` refuses an\nunderspecified idea instead of guessing; `--adr` also writes an ADR-lite rationale\ndoc under `docs/adr/`.\n\n`plan` / `approve` are the natural-language **front door** (an agent drafts,\na human approves — the only write path the dashboard shares too).\nThe higher-level the idea, the more predicates kazi drafts — and the more you'll\nwant to curate them before approving, because every predicate becomes a wall kazi\nwon't declare \"done\" until it's objectively true. Approving blesses the goal; to\ndrive it, hand `kazi apply` a goal-file (next section) — the same predicates, captured\nas a file you can version and re-run.\n\n\u003e More \"build an app for X\" ideas kazi can draft predicates for:\n\u003e - `kazi plan \"create a paste-bin app with a create-paste API and a raw view\"`\n\u003e - `kazi plan \"build a webhook receiver that validates signatures and stores events\"`\n\u003e - `kazi plan \"create a REST API for a to-do list with the usual CRUD endpoints\"`\n\n---\n\n## Quickstart 2 — write a tiny goal-file and ship it\n\nA goal-file is a few lines of TOML. Here's one that says *\"the unit tests pass AND\nthe deployed `/livez` endpoint returns `ok`\"* — code **and** live production, in one\ndeclaration:\n\n```toml\n# my-goal.toml\nid = \"health-green-and-live\"\nname = \"health endpoint returns ok, tests green and live\"\n\n[budget]\nmax_iterations = 8        # hard ceilings — kazi can never loop forever\nmax_tokens = 500000\n\n[scope]\nworkspace = \".\"           # the repo kazi may edit\npaths = [\"main.go\"]\n\n# A CODE check: the project's tests must pass.\n[[predicate]]\nid = \"tests\"\nprovider = \"test_runner\"\ndescription = \"unit tests pass\"\ncmd = \"go\"\nargs = [\"test\", \"./...\"]\n\n# A LIVE check: the *deployed* service must answer correctly. This is what makes\n# convergence real — green-on-my-laptop is not enough.\n[[predicate]]\nid = \"livez-live\"\nprovider = \"http_probe\"\ndescription = \"deployed GET /livez returns 200 body \\\"ok\\\"\"\nurl = \"https://your-service.run.app/livez\"\nexpect_status = 200\nexpect_body = \"ok\"\nbody_match = \"exact\"      # exact, not substring — \"ok\" is a substring of \"not-ok\"!\n```\n\nPicking `[budget]` numbers by hand? `kazi plan` and `kazi init` will suggest\none LEARNED from your own run history once you have some (p95 x 1.5\nheadroom, with provenance — advisory only, never applied silently; see\n[docs/economy.md](docs/economy.md#learned-budget-proposals-kazi-plan--kazi-init-t489)).\n\nRun it:\n\n```sh\nmix kazi.apply my-goal.toml --workspace ./my-service\n```\n\nkazi prints each iteration and a final verdict, and exits `0` only on convergence:\n\n```\nkazi.loop goal=health-green-and-live iter=1 failing=[\"tests\",\"livez-live\"]   → dispatch agent\nkazi.loop goal=health-green-and-live iter=2 failing=[\"livez-live\"]           → integrate (PR #42)\nkazi.loop goal=health-green-and-live iter=3 failing=[\"livez-live\"]           → deploy\nkazi.loop goal=health-green-and-live iter=4 failing=[]                       → CONVERGED ✓\nOUTCOME: :converged   (tests pass · live /livez = \"ok\")\n```\n\n**Predicate providers** you can use today:\n\n| `provider`     | checks… | key config |\n|----------------|---------|------------|\n| `test_runner`  | a command's exit code (unit/integration tests) | `cmd`, `args` |\n| `http_probe`   | a live URL's status + body, optionally **sustained** over N samples | `url` (**required**), `expect_status`, `expect_body`, `body_match`, `samples`, `interval_ms` |\n| `browser`      | a real browser flow (Playwright), optionally a **journey** over N runs | `url` (**required**), per-flow config, `samples` |\n| `prod_log`     | a production-log condition (e.g. 5xx rate) — a coarse safety net | per-check config |\n| `metrics`      | a live **RED/SLO** signal (PromQL): windowed quantile, error-rate, or **burn-rate** gate | `query_url`, `query`, `pass_when`, `quantile`, `burn_rate` |\n| `custom_script`| ANY CLI checker (scanner, mutation tester, contract check) | `cmd`, `args`, `verdict`, `path`, `pass_when` |\n| `ratchet`      | a metric may not regress vs a baseline (coverage, perf, size) | `metric`, `baseline`, `direction`, `allowed_regression` |\n| `static`       | static analysis / type-check / lint (Dialyzer-led, SARIF-general) | `cmd`, `args`, `format`, `baseline`, `allowed_regression` |\n| `no_stubs`     | the diff-vs-base adds no stub/placeholder marker to a production (non-test) file | `patterns`, `base`, `exclude` |\n| `docs_updated` | a user-facing surface change ships with a docs update or a `[no-docs]` marker | `base`, `surface_patterns`, `doc_patterns` |\n\n`http_probe` and `browser` are the **live** predicates: `url` is VALIDATED at\ngoal-load time (`Kazi.Goal.Loader`, T48.1, ADR-0058) — a missing or blank `url`\nfails loudly, naming the predicate and the key, instead of silently erroring\n`missing_url` on every observation until the loop exhausts its budget. Neither\nprovider resolves any other key (e.g. a relative `path`) into a url at dispatch\ntime, so `url` must be the full target.\n\n`custom_script` is the **escape hatch**: it turns any command-line tool into a\npredicate without a kazi release. Crucially the **verdict is declared, not\nassumed** — a SARIF/JSON scanner that exits `0` *with* findings is gated on its\nparsed output, not its exit code (the class of \"the gate silently passed\" bug,\ndesigned out). See [`docs/custom-script-provider.md`](docs/custom-script-provider.md),\n`kazi schema custom_script`, and the recipes in\n[`priv/examples/`](priv/examples/) (`custom_script_sarif.toml`,\n`custom_script_junit.toml`, `custom_script_mutation.toml`). A fuller\noff-the-shelf catalog — contract/schema compat, perf/size, secret scanning, a11y,\nIaC/container scan, visual regression — plus the two evidence tiers and the\nper-tool exit-code gotchas is in\n[`docs/custom-script-recipes.md`](docs/custom-script-recipes.md).\n\n`ratchet` is the **no-regression** mode: a metric passes only while it stays\nwithin `allowed_regression` of a `baseline`, read through `direction`\n(`higher_better` for coverage/mutation score, `lower_better` for size/latency).\nThe baseline is a fixed number, the metric's own stored prior value (`\"stored\"` —\nseeded on the first run, tightened on every pass), or a **git ref** (`\"main\"` —\nthe metric recomputed at that ref). Coverage, perf, and size are configs of this\none mode. With `allowed_regression = 0` a metric \"may only improve.\" See\n[`docs/ratchet-predicate.md`](docs/ratchet-predicate.md), `kazi schema ratchet`,\nand the recipes in [`priv/examples/`](priv/examples/) (`ratchet_coverage.toml`,\n`ratchet_size.toml`).\n\n`static` is the **analysis / type-check / lint** mode: the cheapest, most\ndeterministic check, run every iteration to catch defects on paths the tests\nnever execute. It **leads with Dialyzer** (kazi-native, zero false positives) and\ngeneralizes to the polyglot SARIF tools (`tsc`, `mypy`, `golangci-lint`, Semgrep)\nvia `format`. The verdict is gated on the **parsed findings, not the exit code**,\nand a `baseline` turns it into a ratchet that **fails only on NEW findings** (so\npre-existing debt can only shrink, never block). Findings surface as localized\n`file:line:col` evidence. See [`docs/static-predicate.md`](docs/static-predicate.md),\n`kazi schema static`, and the recipes in [`priv/examples/`](priv/examples/)\n(`static_dialyzer.toml`, `static_sarif.toml`).\n\nThe **live providers** (`http_probe` sustained-health, `browser` journeys,\n`metrics`, `prod_log`) verify a *deployed* service. The discipline they enforce:\n**never converge on a single sample** — `http_probe` and `browser` require N\n*consecutive* healthy samples (the Kubernetes `failureThreshold` model), and a\n`metrics` burn-rate gate fires only when both a long and a short window breach.\nAbsent a metrics endpoint, `metrics` degrades to *not applicable* (never a false\npass). See [`docs/live-providers.md`](docs/live-providers.md) and\n`kazi schema http_probe` / `kazi schema browser` / `kazi schema metrics`.\n\nAdd `guard = true` to a predicate to make it an **invariant** (e.g. \"coverage must\nnot drop\") — kazi blocks the \"delete the failing test\" shortcut.\n\n---\n\n## A real worked example: failing test → live production\n\nThis is kazi's own end-to-end proof (the **T0.12 dogfood**), and you can read it in\n[`docs/devlog.md`](docs/devlog.md). The fixture in\n[`fixtures/deploy-target/`](fixtures/deploy-target/) is a tiny Go web service whose\n`/livez` endpoint returns `\"not-ok\"` and whose unit test therefore **fails on\npurpose**. Given the goal *\"tests pass AND deployed `/livez` returns ok\"*, kazi:\n\n1. **Observed** both checks failing — and refused to call it done.\n2. **Dispatched** a `claude -p` agent, which made the one-line fix.\n3. **Integrated** it — opened a PR and rebase-merged it to `main`.\n4. **Deployed** the new build to Cloud Run.\n5. **Re-checked** the live endpoint → `200 \"ok\"` → **converged**.\n\nCrucially, through steps 1–4 the live check kept failing and kazi **stayed\nnon-converged** — it only reported success once the real, deployed endpoint was\ncorrect. That's the whole point: done is observed, not asserted.\n\n\u003e Live deploys need a deploy target configured (the service / project / region and a\n\u003e deploy command). The fixture's setup — GCP roles, the Cloud Run quirks kazi\n\u003e discovered, and the goal-file — is documented in\n\u003e [`fixtures/deploy-target/README.md`](fixtures/deploy-target/README.md) and\n\u003e [`docs/lore.md`](docs/lore.md).\n\n---\n\n## Adopt an existing project\n\nAlready have a working repo? `kazi init \u003crepo-dir\u003e` reverse-engineers a starter\ngoal-file — the equivalent of `terraform import` for \"what already works\"\n([ADR-0013](docs/adr/0013-adopt-reverse-engineer-goals.md)):\n\n```sh\nkazi init ./my-service --out my-service.goal.toml\n```\n\nIt detects the stack from marker files (`go.mod` → `go test ./...`, `mix.exs` →\n`mix test`, `package.json`'s test script, `pyproject.toml`/`setup.cfg` →\n`pytest`) and writes one baseline goal-file with:\n\n- a **`test_runner` predicate** naming the detected test command, so kazi holds\n  the suite green;\n- conservative **guards** (e.g. a coverage ratchet when a coverage tool is\n  configured) — never a guard it cannot evaluate;\n- a **commented live-predicate TODO** — an `http_probe` scaffold for you to point\n  at the real deployed endpoint. Live predicates are scaffolded, never guessed.\n\nDetection is deterministic: the same repo always produces the same goal-file.\nPass `--enrich` (off by default) to have your coding agent propose live\npredicates from discovered endpoints; the deterministic detection always stands.\nReview the goal-file, fill in the live TODO, then `kazi apply` it.\n\nPass `--with-gist` to opt **this repo** into the Gist context store\n([ADR-0045](docs/adr/0045-context-store-layer-gist-provider.md)) — a budget-fitted\ntext-artifact memory that keeps each agent prompt small. It verifies `gist doctor`,\nwrites the project-local `.kazi/context.toml` naming the provider, registers the\n`gist serve` MCP server in the repo's `.mcp.json`, and recommends setting\n`KAZI_GIST_DSN` to a PostgreSQL DSN for cross-iteration persistence. It is\n**project-local only** — it never mutates a global agent config — and requires the\n[`gist`](https://github.com/sirerun/gist) binary on `PATH` (absent it, the command\nreports the missing dep and the goal-file is still written):\n\n```sh\nkazi init ./my-service --with-gist\nexport KAZI_GIST_DSN=\"postgres://USER:PASS@HOST:5432/gist\"   # cross-call persistence\n```\n\n### Worked example\n\nRun it against the Go fixture that ships with this repo:\n\n```sh\nkazi init fixtures/deploy-target --out my.goal.toml\n```\n\nIt detects the `go.mod` and writes\n[`priv/examples/adopt_deploy_target.goal.toml`](priv/examples/adopt_deploy_target.goal.toml):\n\n```toml\nid = \"adopt-deploy-target\"\nname = \"Adopted baseline for deploy-target\"\n[scope]\nworkspace = \"fixtures/deploy-target\"\n\n[[predicate]]\nid = \"tests-pass\"\nprovider = \"test_runner\"\ndescription = \"project test suite passes\"\nargs = [\"test\", \"./...\"]\ncmd = \"go\"\n\n# ... a `tests-pass-baseline` guard, then a COMMENTED live-predicate TODO\n# you uncomment and point at the real deployed endpoint.\n```\n\nThe acceptance predicate names the detected `go test ./...`; the live predicate\nis left as a commented scaffold for you to fill in. A hermetic end-to-end test\npins this output, so the example never drifts from what the tool produces.\n\n---\n\n## Watch it work\n\n- **LiveView dashboard** — a goal board, live agent presence, the lease map, a\n  live dependency-DAG \"wave\" view (`/dag`: groups by running / ready / blocked /\n  converged, the `needs` edges, per-group convergence), per-goal convergence\n  history, and a drill-in convergence heatmap with an iteration scrubber\n  (`/goals/:id/drillin`: predicates x iterations, regression flips visually\n  distinct). Read-only inspection, decoupled from the loop\n  ([ADR-0011](docs/adr/0011-slice3-operator-surfaces.md)).\n- **`kazi dashboard` — Mission Control** — several concurrent `kazi apply`\n  runs on one machine are a black box by default; `kazi dashboard` boots a\n  standalone, read-only web endpoint over the shared run registry. Mission\n  Control is the landing page (`/`, alias `/starmap`): an ops-center card grid —\n  every registered run as a card with its state (running / converged / stale /\n  stuck / over-budget), inline predicate DNA, iteration burn, and a convergence\n  sparkline; a NEEDS ATTENTION row and a fleet-wide event river. With\n  `--roadmap`, the grid groups into topological wave sections. See\n  [docs/dashboard.md](docs/dashboard.md)\n  ([ADR-0070](docs/adr/0070-mission-control-dashboard.md), superseding the\n  ADR-0057 starmap home view).\n\n---\n\n## CLI reference\n\n```\nkazi init \u003crepo-dir\u003e [--out \u003cfile\u003e] [--enrich] [--with-mcp] [--with-gist]  # adopt a repo -\u003e a goal-file (+ .mcp.json / context store)\nkazi plan \"\u003cidea\u003e\" [--workspace \u003cpath\u003e]   # draft predicates from plain English\nkazi list-proposed [--status \u003cstate\u003e]        # review drafts (proposed/approved/rejected)\nkazi approve \u003cproposal-ref\u003e                  # bless a drafted goal\nkazi reject  \u003cproposal-ref\u003e                  # discard a draft\nkazi apply \u003cgoal-file\u003e --workspace \u003cpath\u003e      # drive a goal to convergence\n        [--env \u003cname\u003e]                       #   target a deploy environment (staging/prod)\n        [--standing]                         #   run continuously (re-converge on drift)\n        [--check]                            #   observe-only: evaluate the vector once, dispatch nothing (issue #805)\n        [--explain]                          #   pure planning: print the computed wave schedule + per-partition worktree isolation, dispatch nothing\n        [--in-place] [--base \u003cref\u003e]          #   edit the workspace directly / pick the task-worktree base ref (ADR-0065)\n        [--parallel [N]]                     #   native parallel scheduler over the partitioned goal-set; one git worktree per partition (ADR-0027, #937)\n        [--pause-between-waves]              #   pause at each wave boundary with a resume_token (issue #936)\n        [--resume \u003ctoken\u003e]                   #   continue a paused run from its checkpoint\nkazi apply --fleet \u003cdir|manifest\u003e --workspace \u003cpath\u003e [--fleet-concurrency N]  # a DAG of goal-files, one worktree per member (ADR-0065)\nkazi status \u003cref\u003e                            # report a run's (or proposal's) current state\nkazi status                                  # list every currently LIVE run (pre-upgrade check, issue #971)\nkazi portfolio [--full]                      # sitrep \"where are we / how is it going?\": headline % across done/in-progress/blocked/todo/planned, bounded per-bucket summaries (blocked entries name their blocker), honest predicates-green rate — never a projected date (ADR-0046); --full restores the complete ledger (E64, #1427)\nkazi orphans [--reap]                         # list runs whose harness child process is still alive (#1073/#857); --reap sends TERM then KILL to each\nkazi install-hooks [--local] [--uninstall] # opt-in: register session-bus delivery hooks (SessionStart + UserPromptSubmit -\u003e `kazi bus hook \u003cevent\u003e`, ADR-0076); --uninstall reverts exactly\nkazi economy [--goal \u003cref\u003e]                  # run-economics history: p50/p95 by goal-shape/model/harness (ADR-0058)\nkazi context index \u003clabel\u003e \u003cfile\u003e            # context store: index a heavy artifact\nkazi context search \"\u003cquery\u003e\" [--budget N]   #   budget-fitted recall (--provider gist)\nkazi context stats                           #   byte accounting (indexed/returned/saved)\nkazi memory recall \"\u003cquery\u003e\" [--budget N]    # budgeted FTS recall over the git-native corpus (ADR-0062)\nkazi spec import \u003cfeature-file\u003e... --into \u003cgoal-ref\u003e  # derive predicates from a reviewed behavior spec (ADR-0050)\nkazi spec import ... --lower scenario         #   lower @interface:web/@interface:cli Scenarios to runtime scenario predicates (ADR-0054 d3)\nkazi export \u003cgoal-file\u003e --obsidian \u003cdir\u003e     # write an Obsidian vault of the goal tree\nkazi lint \u003cgoal-file\u003e                        # advisory near-duplicate group-name warnings\nkazi economy --rediscovery \u003cgoal\u003e            # ranked rediscovery-pressure report (report-only)\nkazi mcp                                      # start the MCP server over stdio (ADR-0044)\nkazi dashboard [--port \u003cn\u003e] [--bind \u003cip\u003e]    # Mission Control: every registered run, read-only (ADR-0070)\n        [--roadmap \u003cgoal-file\u003e]              #   group the fleet grid into the roadmap's needs-DAG wave sections (T47.2)\nkazi help [--json]                           # the command/flag surface (--json for machines)\nkazi version                                 # print the kazi version and exit\n```\n\n`kazi apply` exits `0` on convergence, non-zero otherwise — so it composes in CI/scripts.\n\n\u003e **Drive kazi over MCP (preferred).** An MCP-speaking harness wires kazi as an MCP\n\u003e server and drives its self-describing `kazi_plan` / `kazi_approve` / `kazi_apply` /\n\u003e `kazi_status` / `kazi_list_proposed` tools — no JSON-CLI shell-out. The same server\n\u003e also exposes the session-bus verbs (ADR-0067) as `kazi_bus_post` / `kazi_bus_read` /\n\u003e `kazi_bus_who` / `kazi_bus_tell`, mirroring `kazi bus post|read|who|tell` — each\n\u003e requires a running `kazi daemon` and reports a structured `no_daemon` tool error\n\u003e otherwise. The canonical client config references the installed binary verb\n\u003e (`kazi init --with-mcp` writes exactly this `.mcp.json`):\n\u003e\n\u003e ```json\n\u003e { \"mcpServers\": { \"kazi\": { \"command\": \"kazi\", \"args\": [\"mcp\"] } } }\n\u003e ```\n\n\u003e **Read-model note.** The Mix task (`mix kazi.apply`) creates and migrates the SQLite\n\u003e read-model on startup, so every iteration is persisted. The standalone escript\n\u003e can't bundle the native SQLite NIF, so it runs without persistence (it still\n\u003e converges; it just won't record history).\n\n\u003e **`kazi economy --rediscovery \u003cgoal\u003e` (T48.10, ADR-0058 decision 3).** Folds a\n\u003e goal's recorded per-iteration tool counters (file reads / search calls / code-graph\n\u003e queries) into a RANKED, report-only candidate list: a tool category that keeps\n\u003e recurring past the first dispatch instead of falling off is a candidate for a\n\u003e stronger orientation pack / retrieval cache. A goal with no recorded tool-use\n\u003e stream reports `unknown`, never a fabricated empty ranking (ADR-0046\n\u003e honest-unknown). This is advisory only — it feeds nothing back into a dispatch\n\u003e prompt; a candidate ships as an actual prompt/context change only once the\n\u003e benchmark gate (T48.12) measures a real reduction.\n\n---\n\n## How it works (under the hood)\n\n- **Positioning** — a harness-agnostic outer loop, never a harness ([ADR-0001](docs/adr/0001-positioning-outer-loop-reconciler.md)).\n- **Goals** — machine-checkable predicate sets, evidence-backed ([ADR-0002](docs/adr/0002-goals-as-predicates.md)).\n- **Runtime** — Elixir / OTP + Phoenix LiveView ([ADR-0003](docs/adr/0003-language-elixir-otp.md)); one supervised process per active goal.\n- **Coordination** — NATS JetStream KV leases (revision-CAS + TTL) and graph-aware blast-radius partitioning ([ADR-0004](docs/adr/0004-coordination-substrate-nats-jetstream.md), [ADR-0006](docs/adr/0006-coordination-leases-and-graph-partitioning.md)).\n- **Data split** — Git (code) · JetStream (coordination) · ETS (live state) · SQLite (read-model) ([ADR-0005](docs/adr/0005-data-layer-split.md)).\n- **Harness \u0026 context** — stateless per iteration; kazi owns context as a thin, deterministic evidence projection plus a blast-radius orientation pack — never conversation memory ([ADR-0008](docs/adr/0008-harness-invocation-and-context.md), [ADR-0009](docs/adr/0009-prompt-construction-thin-evidence-projection.md), [ADR-0010](docs/adr/0010-context-injection-reexploration-mitigation.md)).\n- **Intended vs. actual** — kazi imports intent from standard specs and prose, and surfaces drift / dead code via a surface-coverage meta-predicate ([ADR-0021](docs/adr/0021-intended-vs-actual-reconciliation.md)).\n- **Any harness, self-taught** — kazi onboards any CLI coding harness through a profile conformance contract, and is a harness-friendly, agent-drivable CLI that teaches itself to harnesses via a skill, an MCP server, and machine-readable help ([ADR-0022](docs/adr/0022-harness-onboarding-conformance.md), [ADR-0023](docs/adr/0023-harness-friendly-agent-drivable-cli.md), [ADR-0024](docs/adr/0024-kazi-self-teaching-to-harnesses.md)).\n- **Native scheduler \u0026 predicate-graph waves** — kazi owns parallelization with a native scheduler over a partitioned goal-set, running dependency-aware predicate-graph waves instead of leaning on an external pool ([ADR-0026](docs/adr/0026-kazi-under-apply-pool.md), [ADR-0027](docs/adr/0027-kazi-owns-parallelization-native-scheduler.md), [ADR-0028](docs/adr/0028-dependency-aware-partitioning-predicate-graph-waves.md)).\n- **Agent-native surfaces** — the docs and website lead with the agent-driven on-ramp, the `kazi` skill is a router for code goals, and your coding agent (not a separate chat bridge) is the mobile interface ([ADR-0025](docs/adr/0025-docs-lead-with-agent-driven-onramp.md), [ADR-0030](docs/adr/0030-content-marketing-agent-native-positioning.md), [ADR-0031](docs/adr/0031-kazi-skill-router-subsumes-loop-apply-qualify.md)).\n- **One verb set** — the CLI verbs are `kazi plan` (draft predicates from an idea) and `kazi apply` (drive a goal to convergence), unifying the human, skill, and CLI surfaces ([ADR-0032](docs/adr/0032-rename-cli-verbs-run-apply-propose-plan.md)).\n\nFull narrative: [`docs/concept.md`](docs/concept.md). Decisions: [`docs/adr/`](docs/adr/).\nBuild plan: [`docs/plan.md`](docs/plan.md).\n\n---\n\n## Status\n\nSlices 0–3 are implemented and green (Elixir/OTP; ~700 hermetic ExUnit tests), and\nthe live idea → production loop is proven end-to-end (the T0.12 dogfood above). What\nworks today:\n\n- **Convergence core** — the reconcile loop drives predicates to truth via a\n  stateless agent harness plus integrate (branch → PR → rebase-merge) and deploy\n  actions; evidence persisted to SQLite.\n- **Trustworthy loops** — regression detection, flake quarantine, hard budgets,\n  stuck-escalation, and a production-log predicate.\n- **Creation mode** — kazi builds *new* features from failing acceptance predicates,\n  not only repairs. From here on, kazi builds kazi.\n- **Coordination \u0026 surfaces** — NATS leases + presence, graph partitioning,\n  natural-language authoring, and a LiveView dashboard.\n- **Safe concurrent work** ([ADR-0065](docs/adr/0065-safe-concurrent-work-serial-worktree-fleet.md)) —\n  a serial `kazi apply` never edits your checkout: it works in its own task\n  worktree off the workspace's HEAD and lands converged commits back on the base\n  by rebase-merge (`--in-place` opts out). `kazi apply --fleet \u003cdir|manifest\u003e`\n  runs a whole DAG of goal-files the same way — one worktree per member goal,\n  `depends_on`/scope-overlap edges ordering them, and\n  `--pause-between-waves`/`--resume` for supervised frontier-by-frontier runs.\n- **Context injection** — every stateless iteration starts *oriented* (a\n  deterministic blast-radius pack + an optional, off-by-default retrieval adapter),\n  without reintroducing conversation memory.\n\n**By design, kazi will never**: become a coding agent/harness; decide *what* to\nbuild (that's your judgment); or put a vector DB in the core loop (the retrieval\nadapter is an optional augmentation, never the foundation).\n\n## Learn more\n\nNew to driving coding agents to an objective \"done\"? The blog walks the same\nladder we climbed, one rung per post — from prompting by feel up to a\nreconciliation workflow. Each post is independently useful, even if you never\nadopt kazi.\n\n- **[From Vibe Coding to Reconciliation](https://kazi.sire.run/blog/from-vibe-coding-to-reconciliation)** — the twelve-part series.\n- **[All posts](https://kazi.sire.run/blog)** — the kazi blog index.\n\n## Community \u0026 help\n\nQuestions? Start a [GitHub Discussion](https://github.com/kazi-org/kazi/discussions) | Read [`concept.md`](docs/concept.md) for the architecture.\n\n## License\n\nLicensed under the [Apache License, Version 2.0](LICENSE). See the [NOTICE](NOTICE)\nfile for attribution. Copyright 2026 Sire Run, Inc.\n\n---\n\n\u003cp align=\"center\"\u003e\n  Built by the team behind \u003ca href=\"https://sire.run\"\u003e\u003cb\u003eSire\u003c/b\u003e\u003c/a\u003e.\n\u003c/p\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkazi-org%2Fkazi","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fkazi-org%2Fkazi","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkazi-org%2Fkazi/lists"}