{"id":51795878,"url":"https://github.com/stunspot/testforge","last_synced_at":"2026-07-23T05:01:16.109Z","repository":{"id":372101793,"uuid":"1302637373","full_name":"Stunspot/TestForge","owner":"Stunspot","description":"Free Collaborative Dynamics Augment for risk-driven software verification, skeptical review, and isolated Agent behavioral evals.","archived":false,"fork":false,"pushed_at":"2026-07-22T04:20:36.000Z","size":3994,"stargazers_count":6,"open_issues_count":0,"forks_count":1,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-07-22T04:29:41.254Z","etag":null,"topics":["agent-skills","ai-agents","claude-code","codex","llm-evaluation","ollama","python","software-testing","software-verification"],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Stunspot.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE.md","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":"SECURITY.md","support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":"NOTICE.md","maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-07-16T10:18:43.000Z","updated_at":"2026-07-22T04:19:21.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/Stunspot/TestForge","commit_stats":null,"previous_names":["stunspot/testforge"],"tags_count":7,"template":false,"template_full_name":null,"purl":"pkg:github/Stunspot/TestForge","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Stunspot%2FTestForge","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Stunspot%2FTestForge/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Stunspot%2FTestForge/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Stunspot%2FTestForge/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Stunspot","download_url":"https://codeload.github.com/Stunspot/TestForge/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Stunspot%2FTestForge/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35789247,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-23T02:00:06.683Z","response_time":57,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["agent-skills","ai-agents","claude-code","codex","llm-evaluation","ollama","python","software-testing","software-verification"],"created_at":"2026-07-21T03:34:59.017Z","updated_at":"2026-07-23T05:01:16.088Z","avatar_url":"https://github.com/Stunspot.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"![TestForge - software verification that argues back](assets/testforge-social-preview.png)\n\n# TestForge\n\nTestForge is a free Collaborative Dynamics Augment that gives inexpensive local coding Agents a verification discipline they do not reliably improvise on their own. It turns software changes, repositories, defects, and release candidates into risk-ranked evidence instead of a comforting pile of green checkmarks.\n\nThe bundled Augment testbed generalizes the same discipline beyond code. Every capability you build can carry behavioral evals, run isolated trials, expose the exact failed dimensions, guide reengineering, seal the evidence, promote a reviewed passing baseline, and detect regression later. TestForge makes “I should check this” an operative capability and “it worked before” a durable record.\n\nThis is not one-shot benchmarking. **TestForge is a quality ratchet for Agent competence:** failures drive reengineering, reviewed success becomes the new baseline, and regression gates resist backward motion. One way, upwards.\n\nThe operating loop is: **build → test → diagnose → reengineer → rerun → review → promote → regression-check**.\n\n## Built with Codex and GPT-5.6 during OpenAI Build Week\n\nTestForge was conceived and built during the July 13-21, 2026 submission period. Stun supplied the product intention, approved capability map, promptcraft system, evaluation philosophy, and release judgment. Codex with GPT-5.6 Sol turned that design into the complete Augment, repaired defects found during execution, packaged it, and then helped use TestForge to build the behavioral testbed included here.\n\nThe primary Codex build task completed roughly 35 minutes of active construction. The result is not a prompt folder wearing a fake moustache: it is an installable two-SKILL verification system with deterministic tools, schemas, examples, evals, independent review, host adapters, CI, and an evidence-recording testbed. Its practical wager is simple: externalized method and memory can buy more competent work from cheaper cognition.\n\nRead the [Build Week collaboration and architecture record](BUILD-WEEK.md), or take the [five-minute fictional judge path](JUDGE-QUICKSTART.md).\n\nIt includes two Agent SKILLs:\n\n- `$software-verification` reconstructs impact, ranks failure risk, designs meaningful oracles, creates stack-compatible tests, interprets execution evidence and issues a traceable release assessment.\n- `$verification-reviewer` independently attacks the evidence chain for catastrophic omissions, weak tests, unsupported claims and verdicts that outrun the proof.\n\nThis repository also includes the CD Augment evaluation testbed used to run isolated behavioral cases, keep answer keys away from the evaluated model, record criterion-level judgments, preserve hard gates outside the average and promote reviewed regression baselines.\n\n## What you can do with it\n\n- Give a limited local coding Agent a reusable risk model, oracle discipline, evidence vocabulary, and skeptical second pass.\n- Hand an Agent a bug, diff, feature or failing test and get a risk-driven verification plan.\n- Generate tests that try to expose consequential failure rather than merely exercise edited lines.\n- Distinguish product defects, test defects, environment failures, flaky behavior and insufficient evidence.\n- Produce `READY`, `READY_WITH_RESIDUAL_RISK`, `NOT_READY`, `INSUFFICIENT_EVIDENCE` or `BLOCKED_BY_ENVIRONMENT` with a reproducible evidence trail.\n- Challenge that conclusion with a separate skeptical reviewer.\n- Run and track behavioral evals for other Augments with Codex or local Ollama models.\n- Preserve reviewed baselines so future builds can compare behavior instead of relying on remembered impressions.\n\nTestForge is advisory verification machinery. It does not prove defect freedom, certify compliance, grant production access, authorize release, or turn model confidence into evidence.\n\n## Repository map\n\n- [`testforge/`](testforge/) - the complete portable TestForge Augment v1.1.1.\n- [`testforge/docs/QUICK-START.md`](testforge/docs/QUICK-START.md) - install and first-use guide.\n- [`RELEASE-NOTES-v1.1.1.md`](RELEASE-NOTES-v1.1.1.md) - plugin-publication changes and exact untested boundary.\n- [`ARCHIVE-CUSTODY.md`](ARCHIVE-CUSTODY.md) - canonical Augment, plugin, standalone-skill, Claude, GitHub, and backup custody.\n- [`PLUGIN-DIRECTORY-SUBMISSION-v1.1.2.md`](PLUGIN-DIRECTORY-SUBMISSION-v1.1.2.md) - exact OpenAI draft listing, portal-specific upload custody, reviewer cases, and owner-only submission gate.\n- [`testforge/docs/SALES-DEMO.md`](testforge/docs/SALES-DEMO.md) - a compact proof-of-value scenario.\n- [`tools/augment-evals/`](tools/augment-evals/) - isolated Augment behavioral evaluation harness.\n- [`tools/augment-evals/README.md`](tools/augment-evals/README.md) - testbed setup, run, review, seal, promote and regression workflow.\n\n## Quick start: install the Codex plugin\n\n```text\ncodex plugin marketplace add Stunspot/TestForge\ncodex plugin add testforge@cd-testforge\n```\n\nStart a new Codex task, then invoke `$software-verification` or `$verification-reviewer`. The plugin bundles the two self-contained TestForge v1.1.1 skills so their doctrine, tools, examples, and status vocabulary stay aligned. The separate Augment behavioral-evaluation harness remains in this repository rather than the skills-only plugin. Its marketplace namespace is product-specific, so TestForge can coexist with other Collaborative Dynamics plugin repositories.\n\n## Quick start: use the standalone Agent SKILLs\n\nDownload the latest release, unzip it and keep the `testforge/` tree together. Expose both directories under `testforge/skills/` through your Agent host's skill mechanism. Host-specific notes are included for [Codex](testforge/adapters/codex.md), [Claude Code](testforge/adapters/claude-code.md), [GitHub](testforge/adapters/github.md), [local shell](testforge/adapters/local-shell.md) and [copy-paste chat](testforge/adapters/copy-paste-chat.md).\n\nThe GitHub release also preserves the complete Augment, the installable Codex plugin, its distinct OpenAI skills-only portal upload, and each bundled skill as separately named archives. This keeps `$software-verification` and `$verification-reviewer` independently recoverable without losing the complete two-skill product. The OpenAI draft exists and both bundled skills passed automated scanning; accountable-owner attestations and submission for review remain pending.\n\nThen start with:\n\n```text\n$software-verification Verify this change. Reconstruct what could break, create the smallest credible evidence set, run only safe authorized checks, and give me an evidence-backed release assessment.\n```\n\nAfter the evidence package exists, use a fresh context when practical:\n\n```text\n$verification-reviewer Challenge this verification package and tell me whether its release status is actually supported.\n```\n\nPython 3.10+ is needed only for deterministic tools. The TestForge package itself has no mandatory third-party Python dependency.\n\n## Quick start: run Augment behavioral evals\n\nFrom the repository root:\n\n```powershell\npy -m pip install -r tools\\augment-evals\\requirements.txt\npy tools\\augment-evals\\augment_eval.py validate testforge\n```\n\nThen choose a supplied adapter or create one using the documented adapter contract. Local Ollama and signed-in Codex CLI examples are included. Raw runs remain local under `evaluation-results/`; reviewed compact baselines can be promoted into Git-tracked records.\n\n## Trust and evidence\n\nThe release contains synthetic planted-defect examples, deterministic package tooling, canonical eval cases and reviewed local baselines. Read the exact exercised and unexercised boundaries in [`testforge/docs/LIMITATIONS.md`](testforge/docs/LIMITATIONS.md), [`testforge/SECURITY.md`](testforge/SECURITY.md) and the testbed README.\n\nDo not send secrets or unnecessary proprietary code in an issue. Active security testing, destructive operations, production access, dependency installation, production-code changes, CI changes and external publication remain explicitly authorized human decisions.\n\n## License\n\nTestForge uses a split license: MIT for Python software and machine-readable schemas, and CC BY-ND 4.0 for the authored Augment materials. You may redistribute the authentic, unmodified branded Augment, including inside a larger commercial product. See [`LICENSE.md`](LICENSE.md), [`ATTRIBUTION.md`](ATTRIBUTION.md) and [`TRADEMARKS.md`](TRADEMARKS.md).\n\n## Publisher\n\nTestForge is a Collaborative Dynamics Augment. Issues and contributions are welcome under the boundaries in [`CONTRIBUTING.md`](CONTRIBUTING.md) and [`SECURITY.md`](SECURITY.md).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fstunspot%2Ftestforge","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fstunspot%2Ftestforge","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fstunspot%2Ftestforge/lists"}