{"id":45193224,"url":"https://github.com/greynewell/evaldriven.org","last_synced_at":"2026-02-22T14:01:37.633Z","repository":{"id":338664335,"uuid":"1158651890","full_name":"greynewell/evaldriven.org","owner":"greynewell","description":"Ship evals before you ship features.","archived":false,"fork":false,"pushed_at":"2026-02-17T21:32:20.000Z","size":65,"stargazers_count":12,"open_issues_count":0,"forks_count":4,"subscribers_count":3,"default_branch":"main","last_synced_at":"2026-02-20T14:48:06.995Z","etag":null,"topics":["ai-engineering","ai-evaluation","ai-quality","ai-safety","ai-testing","automation","benchmarking","best-practices","ci-cd","continuous-evaluation","devops","eval-driven-development","evaluation","llm-evaluation","machine-learning","manifesto","methodology","quality-assurance","software-engineering","testing"],"latest_commit_sha":null,"homepage":"https://evaldriven.org","language":"Nunjucks","has_issues":false,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"cc0-1.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/greynewell.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-02-15T18:01:37.000Z","updated_at":"2026-02-19T06:20:07.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/greynewell/evaldriven.org","commit_stats":null,"previous_names":["greynewell/evaldriven.org"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/greynewell/evaldriven.org","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/greynewell%2Fevaldriven.org","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/greynewell%2Fevaldriven.org/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/greynewell%2Fevaldriven.org/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/greynewell%2Fevaldriven.org/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/greynewell","download_url":"https://codeload.github.com/greynewell/evaldriven.org/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/greynewell%2Fevaldriven.org/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":29681468,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-02-21T12:30:22.644Z","status":"ssl_error","status_checked_at":"2026-02-21T12:29:55.402Z","response_time":107,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai-engineering","ai-evaluation","ai-quality","ai-safety","ai-testing","automation","benchmarking","best-practices","ci-cd","continuous-evaluation","devops","eval-driven-development","evaluation","llm-evaluation","machine-learning","manifesto","methodology","quality-assurance","software-engineering","testing"],"created_at":"2026-02-20T12:34:31.941Z","updated_at":"2026-02-22T14:01:37.624Z","avatar_url":"https://github.com/greynewell.png","language":"Nunjucks","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Eval-Driven Development\n\nWhat we can build matters less than what we can prove.\n\nAI writes code. The engineer defines \"working,\" measures it, enforces it. **Eval-Driven Development**: every probabilistic system starts with a correctness spec. Nothing ships without automated proof it passes.\n\n## Principles\n\n### 1. Evaluation is the product\n\nBuild evals first. Code is generated. Evals are engineered.\n\n### 2. Define correctness before you write a prompt\n\nCan't express \"correct\" as a deterministic function? Not ready to build. Every task needs an eval. Every eval needs a threshold. Every threshold needs a justification.\n\n### 3. Probabilistic systems require statistical proof\n\nOne passing test proves nothing about a stochastic system. Sample sizes, confidence intervals, regression baselines. Distributions, not anecdotes.\n\n### 4. Evals run in CI\n\nEvals that don't run on every change don't exist. Next to lint, type-check, build.\n\n### 5. Evaluation drives architecture\n\nCan't independently evaluate a component? Can't independently trust it. Design for measurability.\n\n### 6. Cost is a metric\n\nToken spend, latency, compute. Correct but unaffordable is a failed eval.\n\n### 7. Human judgment doesn't scale\n\nEvery manual review is a missing eval. Extract judgment into a rubric, automate it, evaluate the evaluator.\n\n### 8. Ship the eval, not the demo\n\nDemos prove something works once. Evals prove it works under distribution shift.\n\n### 9. Version your evals\n\nDefinitions, datasets, thresholds, results. Version control. Changelogs. Document why.\n\n### 10. The eval gap is the opportunity\n\n\"Works on my machine\" vs. \"passes at p \u003c 0.05.\" That gap is where defensible products get built.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fgreynewell%2Fevaldriven.org","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fgreynewell%2Fevaldriven.org","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fgreynewell%2Fevaldriven.org/lists"}