{"id":51169274,"url":"https://github.com/ggwhite/4x","last_synced_at":"2026-07-13T08:01:14.054Z","repository":{"id":363985545,"uuid":"1265260318","full_name":"ggwhite/4x","owner":"ggwhite","description":"Agentic AI development loop that splits Design, Code, Review, and Test into isolated roles with deterministic guardrails","archived":false,"fork":false,"pushed_at":"2026-07-12T18:04:42.000Z","size":65052,"stargazers_count":33,"open_issues_count":0,"forks_count":9,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-07-12T19:08:47.122Z","etag":null,"topics":["agentic-coding","ai","ai-agents","ai-code-review","automation","claude","cli","code-review","developer-tools","golang","llm","multi-agent","software-engineering"],"latest_commit_sha":null,"homepage":"https://ggwhite.github.io/4x/","language":"Go","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/ggwhite.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":"THREAT_MODEL.md","audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":"AGENTS.md","dco":null,"cla":null}},"created_at":"2026-06-10T15:59:03.000Z","updated_at":"2026-07-12T18:05:20.000Z","dependencies_parsed_at":null,"dependency_job_id":"a5035106-4d70-4367-9e83-a18a5bc09a85","html_url":"https://github.com/ggwhite/4x","commit_stats":null,"previous_names":["ggwhite/4x"],"tags_count":46,"template":false,"template_full_name":null,"purl":"pkg:github/ggwhite/4x","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ggwhite%2F4x","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ggwhite%2F4x/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ggwhite%2F4x/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ggwhite%2F4x/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/ggwhite","download_url":"https://codeload.github.com/ggwhite/4x/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ggwhite%2F4x/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35414732,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-13T02:00:06.543Z","response_time":119,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["agentic-coding","ai","ai-agents","ai-code-review","automation","claude","cli","code-review","developer-tools","golang","llm","multi-agent","software-engineering"],"created_at":"2026-06-26T23:02:33.906Z","updated_at":"2026-07-13T08:01:14.047Z","avatar_url":"https://github.com/ggwhite.png","language":"Go","funding_links":[],"categories":["Cross-Agent References"],"sub_categories":[],"readme":"**English** | [繁體中文](docs/translate/README.zh-TW.md) | [简体中文](docs/translate/README.zh-CN.md) | [日本語](docs/translate/README.ja.md) | [한국어](docs/translate/README.ko.md) | [Español](docs/translate/README.es.md)\n\n[![Go Reference](https://pkg.go.dev/badge/github.com/ggwhite/4x.svg)](https://pkg.go.dev/github.com/ggwhite/4x)\n[![Go Report Card](https://goreportcard.com/badge/github.com/ggwhite/4x)](https://goreportcard.com/report/github.com/ggwhite/4x)\n[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)\n[![CI](https://github.com/ggwhite/4x/actions/workflows/ci.yml/badge.svg)](https://github.com/ggwhite/4x/actions/workflows/ci.yml)\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/assets/4x-banner.svg\" alt=\"4X — Design. Code. Review. Test.\" width=\"480\"\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/assets/demo.gif\" alt=\"4x demo\" width=\"720\"\u003e\n\u003c/p\u003e\n\n**4x** is an open-source CLI that orchestrates AI coding agents into a multi-role development loop — each role (Design, Code, Review, Test) runs in isolation with deterministic guardrails, so features survive contact with production. Like 4X strategy games (eXplore, eXpand, eXploit, eXterminate), the name reflects a system where distinct roles with distinct strengths converge to conquer complexity.\n\n## Key Features\n\n| Category | Highlights |\n|---|---|\n| **Multi-Role Loop** | Design → Code → Review → Test → Deep Review → Accept, with role isolation. Adaptive pipeline selects profile (full / mini / quick) by feature complexity. |\n| **6 AI Runners** | Claude Code · Codex · Gemini CLI · Antigravity · Copilot · Cursor — same `.4x/` file protocol, mix and match per role. |\n| **Dashboard (4x Live)** | macOS native (Swift) + Windows / Linux (Tauri). Real-time SSE monitoring, dependency graph, runner log streaming, screenshot gallery, settings UI, batch monitoring. 6-language i18n, system notifications, menu bar integration. |\n| **Deterministic Guardrails** | State machine, scope lock, baseline snapshots, evidence-based testing gate, dependency gate — enforced by the Go CLI, not by prompting an LLM. |\n| **Crash Recovery** | Runner crash → auto-resume from last saved state. Transient API errors (network, rate limits) → automatic backoff retry. |\n| **Batch Mode** | Dependency-aware DAG scheduling, auto-merge on completion, batch reports, graceful stop. Queue dozens of features and review in the morning. |\n| **MCP Server** | Model Context Protocol server for integration with MCP-compatible clients. |\n| **Issue-First MR Flow** | Optional `issue_tracker` mode: `4x new` creates or links an issue, `4x done` pushes the branch and opens a PR/MR instead of merging locally. Auto-detects GitHub vs. GitLab (including self-hosted) from each repo's remote — no per-repo config needed. |\n| **20+ CLI Commands** | `run`, `batch`, `live`, `doctor`, `clean`, `verify`, `mcp`, phase hooks, health checks, structured logging, and more. |\n| **Self-Evolution** | History mining from past runs, auto-discovered feature enrichment, evolution value gate with anti-hack, self-modification scope guard, and continuous improvement driver (`4x evolve`). 4x learns from its own failures and iterates itself. |\n\n---\n\n## Why 4x?\n\nSingle-agent coding is fast but fragile. You ask one AI to design, implement, review, and test — all in the same breath, with the same biases. It works for small tasks. It falls apart on real features.\n\n4x splits the loop. Each role has a focused job, limited scope, and no access to the others' reasoning. The Designer doesn't write code. The Coder doesn't judge its own work. The Reviewer is adversarial by design. The Tester validates against criteria written before implementation.\n\nThe result: features that survive contact with production.\n\n## Trade-offs\n\nChoosing 4x means trading speed and cost for structure and correctness. Be honest about whether your project needs that trade.\n\n### Strengths\n\n- **Role isolation eliminates self-review bias.** The Coder never judges its own work. The Reviewer is adversarial by design. Single-agent workflows let the same model write and approve code — 4x doesn't.\n- **Deterministic guardrails don't depend on AI judgment.** Scope lock, state machine, evidence requirements — these are enforced by the CLI in Go, not by prompting an LLM to \"please stay in scope.\"\n- **File-based protocol makes it LLM-agnostic.** Switch between Claude, Gemini, Codex, or mix them per role. No vendor lock-in, no SDK dependency.\n- **Crash-resistant state.** Everything lives in `.4x/` files. Session dies, machine reboots — `4x run` picks up exactly where it stopped.\n- **Human stays in the loop.** The `pending-review` gate ensures a human always reviews AI work before it's marked done. The AI proposes, you dispose.\n- **Tames large-scale refactoring.** Changes too big for a single AI session — splitting god objects, extracting packages, migrating APIs — can be broken into dependent features with appropriate profiles. 4x handles the sequencing, review, and verification across multiple phases that would overwhelm a single context window.\n- **Batch mode scales.** Dependency-aware scheduling lets you queue dozens of features overnight and review them in the morning.\n\n### Weaknesses\n\n- **Significantly higher token cost.** Every feature runs through 4+ separate LLM calls at minimum. A review failure doubles that. Expect 3-10x the token cost of a single-agent approach for the same task. See [Usage Tips](docs/guide/usage-tips.md) for cost estimates.\n- **Slower for simple tasks.** A one-line bug fix doesn't need a Designer, Reviewer, and Tester. The overhead of the full loop is wasted on trivial changes. Use single-agent tools for quick fixes.\n- **Setup cost.** `4x init`, feature YAML, settings configuration — there's ceremony before you start. Not worth it for a throwaway script.\n- **Rigid loop structure.** The Design → Code → Review → Test sequence is fixed. If your workflow doesn't fit four roles, you'll fight the framework instead of using it.\n- **Quality depends on prompt quality.** Vague feature descriptions produce vague specs, which produce wrong code. 4x adds structure, but garbage in still means garbage out — just with more steps.\n\n### When to use 4x\n\n- Features that need to be correct (payments, auth, data pipelines)\n- Work that benefits from adversarial review (security-sensitive code)\n- Batch processing of a feature backlog\n- Teams that want audit trails of AI-generated code\n\n### When NOT to use 4x\n\n- Quick one-off fixes or exploratory prototyping\n- Tasks where speed matters more than correctness\n- Projects where token budget is tight\n- Solo hacking sessions where you'd review the code yourself anyway\n\n## Architecture\n\n```\n You\n  |\n  v\n+--------------------------------------------------+\n|  4x CLI (Go)                                     |\n|  Deterministic guardrails. No LLM calls.         |\n|  Scope checks, protocol, state machine, batch    |\n+--------+-----------------------------------------+\n         |  .4x/ directory (file-based protocol)\n         v\n+--------------------------------------------------+\n|  Runners                                         |\n|  Claude Code | Codex | Gemini | Antigravity      |\n|  Copilot | Cursor                                |\n|  Each uses native platform capabilities          |\n+--------+-----------------------------------------+\n         |  SSE events\n         v\n+--------------------------------------------------+\n|  4x Live (Dashboard)                             |\n|  Multi-project real-time monitoring              |\n+--------------------------------------------------+\n```\n\n**Layer 1 — CLI** handles everything deterministic: scope validation, state transitions, baseline snapshots, evidence collection. It never calls an LLM. Guardrails don't depend on AI judgment.\n\n**Layer 2 — Runners** bridge the CLI protocol to your AI tool of choice. Claude Code, Codex, Gemini, Antigravity, Copilot, Cursor — each speaks the same `.4x/` file protocol but uses native platform capabilities.\n\n**Layer 3 — Live** is the multi-project dashboard. Watch your AI agents work in real-time, see phase transitions, stream logs. REST + SSE API.\n\n## Installation\n\n### Homebrew (macOS / Linux)\n\n```bash\nbrew install ggwhite/tap/fourx\n```\n\n### Go Install\n\n```bash\ngo install github.com/ggwhite/4x/cmd/4x@latest\n```\n\n### Shell Script\n\n```bash\ncurl -sSfL https://raw.githubusercontent.com/ggwhite/4x/main/install.sh | sh\n```\n\n### Verify Checksums\n\nThe shell script verifies the download's checksum automatically and aborts if it fails. When downloading a release archive or binary manually from the [Releases](https://github.com/ggwhite/4x/releases) page, verify it yourself before extracting: download that release's `checksums.txt` into the same directory, then run one of:\n\n```bash\n# Linux (and macOS with coreutils)\nsha256sum --check --ignore-missing checksums.txt\n\n# macOS (default)\nshasum -a 256 --ignore-missing --check checksums.txt\n```\n\nOnly extract and run the binary once the checksum reports `OK`.\n\n### Download Binary\n\nPre-built binaries for macOS, Linux, and Windows (amd64 / arm64) are available on the [Releases](https://github.com/ggwhite/4x/releases) page.\n\n## Quick Start\n\n```bash\n# Initialize in your project\ncd my-project\n4x init\n\n# Create a feature\n4x new \"User authentication with OAuth2\"\n# =\u003e Created: F001-user-authentication-w\n\n# Run the full loop\n4x run F001 --runner claude\n\n# Check status\n4x status\n\n# Review and complete\n4x done F001\n\n# Or watch it live\n4x live -w\n```\n\n`4x run` drives the Design-Code-Review-Test loop automatically. If Review finds issues, Code gets another pass. If Test fails, the loop iterates. You stay in control with `--max-rounds` and `--timeout` flags.\n\n## The Four Roles\n\n| Role | Job | Outputs |\n|---|---|---|\n| **Designer** | Analyze requirements, produce spec + acceptance criteria | `task-brief.md`, `acceptance-criteria.md` |\n| **Coder** | Implement exactly what the spec says | Source code, `coder-report.md` |\n| **Reviewer** | Catch bugs and spec violations (checklist + adversarial) | `review-report.md` with verdict |\n| **Tester** | Validate against acceptance criteria with evidence | `test-report.md`, `verify.json` |\n\nEach role is **isolated**. The Coder never sees the Reviewer's prior feedback. The Tester validates against criteria written by the Designer, not the Coder. This separation prevents the blind spots that plague single-agent workflows.\n\n## How the Loop Works\n\n```\nDesigner → Coder → Reviewer → Tester → Accept → Pending Review → Done\n                      ↓           ↓                                 ↑\n                   amending ←─────┘                          human sign-off\n```\n\n- **Review failure** (verdict FAIL or CRITICAL findings) sends code back for amending\n- **Test failure** (verify not passed) sends code back for amending\n- **Escalation** (spec mismatch, criteria wrong) routes back to Designer\n- **Pending review** gate ensures a human always reviews before marking done\n- **Round budget** (default 5) prevents infinite loops\n\n## Deterministic Guardrails\n\nEnforced by the CLI, not AI judgment:\n\n| Guardrail | What it does |\n|---|---|\n| **Scope check** | Changed files must be within declared repos |\n| **Baseline snapshot** | Pre-coding state captured for safe rollback |\n| **State machine** | Phases must proceed in legal order |\n| **Evidence requirement** | Tester must provide verify.json with command output |\n| **Testing gate** | verify.json + test-report + final-report required |\n| **Dependency gate** | Features with unmet dependencies cannot start |\n\n## Batch Mode\n\n```bash\n4x batch plan            # generate dependency-aware execution plan\n4x batch run --runner claude  # run all eligible features in order\n4x batch stop            # graceful shutdown after current feature\n```\n\n## MCP Server\n\nStart the Model Context Protocol (MCP) server:\n\n```bash\n4x mcp\n```\n\n## Claude Code Skills\n\n4x ships two optional [Claude Code Skills](https://docs.anthropic.com/en/docs/claude-code) for driving the pipeline from a Claude Code session. Install them with [`npx skills`](https://github.com/vercel-labs/skills) — no need to clone this repo:\n\n```bash\nnpx skills add ggwhite/4x --skill 4x-audit       # scan past run artifacts, report + optionally file gap features\nnpx skills add ggwhite/4x --skill 4x-autopilot   # drive the full pick → run → merge → next loop unattended\nnpx skills add ggwhite/4x --all                  # install both\n```\n\n- **`4x-audit`** — scans discovered-feature-gaps, escalation/review reports, and recurring learnings patterns; produces a categorized HTML report and can optionally batch-create feature YAMLs.\n- **`4x-autopilot`** — polls 4x state and drives features through their full lifecycle (design → code → review → test → merge) without pausing for confirmation, including merging. **Owner-only** — read the skill's warning before using it on a shared repo.\n\n## Permission Model\n\n**4x runs AI agents in non-interactive mode.** During `4x init`, runners are configured with flags that skip permission prompts (`--dangerously-skip-permissions`, `-y`, `approval: full-auto`) so the loop runs autonomously.\n\nThe CLI's deterministic guardrails (scope lock, baseline snapshots, state machine) provide the safety boundary.\n\n**Run 4x only in projects where you are comfortable with autonomous AI agent execution.**\n\n## Documentation\n\n| Document | Description |\n|---|---|\n| **[User Guide](docs/guide/)** | Complete usage documentation |\n| [Getting Started](docs/guide/getting-started.md) | Installation and first run |\n| [CLI Reference](docs/guide/cli.md) | All commands and flags |\n| [Core Concepts](docs/guide/concepts.md) | Roles, state machine, protocol, guardrails |\n| [Configuration](docs/guide/configuration.md) | Settings, models, locale, runners |\n| [Runners \u0026 Plugins](docs/guide/runners.md) | Supported runners and plugin contract |\n| [Dashboard](docs/guide/dashboard.md) | 4x Live multi-project dashboard |\n| [Batch Mode](docs/guide/batch.md) | Dependency-aware batch execution |\n\n## Project Structure\n\n```\n4x/\n  cmd/4x/              CLI entry point (Cobra)\n  internal/\n    protocol/           .4x/ file format, workspace, types\n    state/              State machine (phase transitions)\n    guard/              Guardrail checks (scope, baseline, evidence)\n    batch/              Dependency DAG, batch scheduler\n    runner/             Subprocess runner interface\n    server/             SSE + REST server for Live dashboard\n  plugins/\n    claude-code/        Claude Code skill + workflow\n    codex/              Codex runner instructions\n    gemini/             Gemini runner instructions\n    agy/                Antigravity runner instructions\n    copilot/            Copilot runner instructions + workflow\n    cursor/             Cursor rules\n    embed.go            go:embed plugin files into binary\n  dashboard/\n    macos/              Swift native app (planned)\n  docs/\n    guide/              User documentation\n    architecture/       System-level design docs\n    design/             Mechanism design docs\n    reference/          Plugin contract\n```\n\n## FAQ\n\n**Q: Does 4x call any LLM APIs directly?**\nNo. The CLI is pure Go with zero LLM dependencies. Runners handle all AI interaction using their native platform capabilities.\n\n**Q: Can I use different LLMs for different roles?**\nYes. Configure per-role models in `.4x/settings.json`. Use Claude for Design, Gemini for Code — each reads the same `.4x/` files.\n\n**Q: How is this different from Devin / SWE-agent / OpenHands?**\nThose are autonomous agents that do everything in one shot. 4x is a *framework* that structures multi-role collaboration with deterministic guardrails. It's closer to a CI pipeline for AI than a single autonomous agent.\n\n## Origin Story\n\n4x was born inside a production system called DCT (Designer-Coder-Tester) that shipped 60+ features for a large-scale platform rewrite. The patterns that survived — role isolation, file-based protocol, deterministic scope checking, evidence-based testing — became 4x. The parts that didn't survive — LLM-specific hacks, shared context assumptions, trust-based guardrails — were deliberately left out.\n\n## Contributing\n\n```bash\ngit clone https://github.com/ggwhite/4x.git\ncd 4x\ngo build ./cmd/4x\ngo test ./...\n```\n\n## License\n\n[MIT](LICENSE)\n\n---\n\n\u003cp align=\"center\"\u003e\n  \u003cstrong\u003eStop hoping your AI writes correct code. Start verifying it.\u003c/strong\u003e\n\u003c/p\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fggwhite%2F4x","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fggwhite%2F4x","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fggwhite%2F4x/lists"}