{"id":51551342,"url":"https://github.com/Zandereins/schliff","last_synced_at":"2026-07-12T14:01:16.192Z","repository":{"id":345602238,"uuid":"1186557480","full_name":"Zandereins/schliff","owner":"Zandereins","description":"Deterministic quality scorer for AI agent instruction files — 8-dimension scoring with security, multi-format (SKILL.md, CLAUDE.md, .cursorrules, AGENTS.md), anti-gaming detection, zero dependencies","archived":false,"fork":false,"pushed_at":"2026-07-03T13:21:50.000Z","size":7098,"stargazers_count":4,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-07-03T13:28:49.630Z","etag":null,"topics":["agents-md","ai-agents","autoresearch","claude","claude-code","cli","code-quality","cursorrules","deterministic","developer-tools","linter","multi-format","python","security-scoring"],"latest_commit_sha":null,"homepage":"https://pypi.org/project/schliff/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Zandereins.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":".github/FUNDING.yml","license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":"SECURITY.md","support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null},"funding":{"github":"Zandereins"}},"created_at":"2026-03-19T18:50:02.000Z","updated_at":"2026-07-03T13:21:55.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/Zandereins/schliff","commit_stats":null,"previous_names":["zandereins/skillforge","zandereins/schliff"],"tags_count":22,"template":false,"template_full_name":null,"purl":"pkg:github/Zandereins/schliff","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Zandereins%2Fschliff","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Zandereins%2Fschliff/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Zandereins%2Fschliff/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Zandereins%2Fschliff/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Zandereins","download_url":"https://codeload.github.com/Zandereins/schliff/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Zandereins%2Fschliff/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35393398,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-12T02:00:06.386Z","response_time":87,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["agents-md","ai-agents","autoresearch","claude","claude-code","cli","code-quality","cursorrules","deterministic","developer-tools","linter","multi-format","python","security-scoring"],"created_at":"2026-07-10T00:00:52.212Z","updated_at":"2026-07-12T14:01:16.186Z","avatar_url":"https://github.com/Zandereins.png","language":"Python","funding_links":["https://github.com/sponsors/Zandereins"],"categories":["Linting","Harnesses \u0026 orchestration"],"sub_categories":["Agent infrastructure"],"readme":"\u003cdiv align=\"center\"\u003e\n  \u003cpicture\u003e\n    \u003csource media=\"(prefers-color-scheme: dark)\" srcset=\"https://raw.githubusercontent.com/Zandereins/schliff/main/docs/assets/hero-dark.svg\"\u003e\n    \u003cimg src=\"https://raw.githubusercontent.com/Zandereins/schliff/main/docs/assets/hero-light.svg\" alt=\"Schliff — deterministic quality scores for AGENTS.md\" width=\"840\"\u003e\n  \u003c/picture\u003e\n\u003c/div\u003e\n\n**The Ruff for `AGENTS.md` — deterministic quality scores for the instruction files that drive your AI. Same input, same score, on every machine.**\n\n[![PyPI](https://img.shields.io/pypi/v/schliff?color=blue\u0026label=PyPI\u0026v=8.5.0)](https://pypi.org/project/schliff/)\n[![Python](https://img.shields.io/pypi/pyversions/schliff)](https://pypi.org/project/schliff/)\n[![Tests](https://github.com/Zandereins/schliff/actions/workflows/test.yml/badge.svg)](https://github.com/Zandereins/schliff/actions/workflows/test.yml)\n[![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](LICENSE)\n[![AGENTS.md quality](https://img.shields.io/endpoint?url=https%3A%2F%2Fschliff-playground.vercel.app%2Fapi%2Fbadge%3Frepo%3DZandereins%2Fschliff)](https://schliff-playground.vercel.app)\n\n*That last badge is Schliff scoring this repo's own `AGENTS.md` — live, right now: **91.6 · A**.*\n\n**Your AI instruction files silently degrade — and nothing catches it.** `AGENTS.md` is read by Cursor, Codex, Copilot, and Claude Code — one rotting file now quietly degrades four tools. A trigger phrase rots, an edge case slips, the file balloons past its token budget. No error, no red test — just agents that quietly get worse.\n\nSchliff scores `AGENTS.md` — and the rest of the family (`SKILL.md`, `CLAUDE.md`, `.cursorrules`, system prompts) — against an explicit, versioned rubric. No LLM judge in the critical path. No network. No randomness. A rule engine you can read, pin, and gate CI on.\n\n```bash\npip install schliff\nschliff score AGENTS.md   # or any SKILL.md / CLAUDE.md / .cursorrules\nschliff demo              # no instruction file handy? score a built-in bad one\n```\n\nThis is the real, current output of `schliff score AGENTS.md` on this repo's own\n[`AGENTS.md`](AGENTS.md) — clone and run it yourself:\n\n```text\nschliff v8.5.0\n\n  structure             █████████░   90/100  great\n  operational_coverage  ██████████  100/100  perfect\n  efficiency            ████████░░   78/100  good\n  composability         ████░░░░░░   45/100  poor\n  clarity               ██████████  100/100  perfect\n\n  Structural Score  ██████████████████░░  91.6/100  [A]\n\n  Tokens: 837 / 3,000 (ok)\n  Format: agents.md (normalized)\n```\n\nNo model produced that number. Run it on another laptop and you get 91.6 again. **A score you can't reproduce isn't a measurement — it's a vibe.**\n\n*Every number in this README comes from released `schliff==8.5.0` (`pip install schliff==8.5.0` to reproduce byte-for-byte). No install? Paste your file into the [playground](https://schliff-playground.vercel.app) — same engine, with AGENTS.md and SKILL.md tabs.*\n\n---\n\n## A real catch\n\nThe SKILL.md for ShieldClaw — a real prompt-injection-defense skill, now archived — is Schliff's reproducible before/after. The fixtures ship in [`docs/case-studies/shieldclaw/`](docs/case-studies/shieldclaw/); every number below is the current engine's output, reproducible with `schliff score`. (For the 27.9 row, score a copy of `SKILL-before.md` outside that directory — in place, the engine auto-discovers the sibling eval suite.)\n\n| | Score | Grade | Dimensions measured |\n| --- | --- | --- | --- |\n| Before, scored in isolation | 27.9 | F | 4/7 — no eval suite; explicit ceiling warning |\n| Before, with its eval suite | 83.7 | B | 7/7 |\n| After fixes | 93.8 | A | 7/7 |\n\nTwo separate effects, and Schliff refuses to conflate them. Adding the eval suite lifted the **measurement ceiling** (27.9 → 83.7) — that is coverage, not quality. The **quality** delta is 83.7 → 93.8, driven by composability **20 → 86** and efficiency **60 → 83** after adding scope boundaries, an I/O contract, and handoffs — structural gaps a linter can't see, caught as a number that was too low.\n\nA second field run, against an external repo ([hydra](docs/case-studies/hydra/), measured on released `schliff==8.4.0`, 2026-07-03): **71.0 [C] → 76.5 [B]** — edges 82→100, composability 56→81, fix merged upstream ([Zandereins/hydra#34](https://github.com/Zandereins/hydra/pull/34)). Efficiency deliberately stayed at 47: a ~14k-token file against a 1,000-token budget was an **informed decline**, not a blind chase of the number. (External repo — not re-runnable from these fixtures.)\n\n---\n\n## Why deterministic?\n\nMost \"AI quality\" tools ask another LLM how good your prompt *feels* — a different answer every run. That makes the score **non-reproducible** (re-run it, get a different number), **un-auditable** (the rubric lives in a hidden prompt), and **trivially gameable** (write for the judge, not the user). You can't gate a release on a number that drifts. Schliff computes how good the file *measurably is* — the same answer every run.\n\nDeterministic means reproducible and auditable — it does not automatically mean the number is right. Schliff's claim is narrower and checkable: the rubric is open source, every scorer is readable, the weights are a dict, and the case studies above show the score moving with real fixes. If you disagree with the rubric, you can read it and file an issue — you can't do that with a judge prompt.\n\nConfig linters tell you whether the file is *valid* — a list of pass/fail rules. Schliff tells you how *good* it is — one graded 0–100 score you can gate, diff across commits, and rank.\n\n- **Reproducible.** The headline composite is computed from a canonical, versioned weight registry. Calibration is **off by default**, so `verify`, `badge`, and the leaderboard return the same score on your laptop and in CI.\n- **Auditable.** Every dimension is a readable scorer in [`scripts/scoring/`](skills/schliff/scripts/scoring/). The weights are a dict you can open. There is no hidden judge prompt.\n- **Anti-gaming, precisely scoped.** A dedicated guard layer ([`guards.py`](skills/schliff/scripts/scoring/guards.py)) detects and floors padding, junk fences, platitude farms, and keyword stuffing — worthless text cannot outrank operational text. A *plausible lie* about your repo is out of reach of any static scorer (see [What the score does not measure](#what-the-score-does-not-measure)).\n- **Zero core dependencies.** Core Schliff is stdlib-only and runs on **Python ≥ 3.10**. (Optional `[evolve]` / `[judge]` extras pull in LLM clients for an opt-in smoke-test only — never for scoring.)\n\nBecause the number is stable, it does real work: gate pull requests on it [in CI](#use-it-in-ci), or diff and compare it across commits with the [CLI](#cli).\n\nAn optional LLM judge exists for exploratory work, but it is never part of the deterministic score. The number you gate on is rule-based, end to end.\n\n---\n\n## The scoring model\n\nFull methodology: [`docs/SCORING.md`](docs/SCORING.md).\n\nFor the `SKILL.md` family, Schliff runs **8 scorers** per file. **7 of them form the headline composite**; `security` (always computed) and `runtime` (opt-in) are reported as **separate signals** so a security warning never silently inflates or deflates your quality grade.\n\n| Dimension | Weight | In headline? |\n| --- | --- | --- |\n| `structure` | 0.15 | ✅ |\n| `triggers` | 0.20 | ✅ |\n| `quality` | 0.20 | ✅ |\n| `edges` | 0.15 | ✅ |\n| `efficiency` | 0.10 | ✅ |\n| `composability` | 0.10 | ✅ |\n| `clarity` | 0.05 | ✅ |\n| `security` | 0.05 | Separate signal (gate threshold 70) |\n| `runtime` | — | Separate signal (no profile weight) |\n\nThe seven headline weights are renormalized to sum to **1.0** — that is the canonical basis.\n\n\u003e [!NOTE]\n\u003e `security` is a side signal for the `SKILL.md` / `CLAUDE.md` / `.cursorrules` / `AGENTS.md` family, but a **core 0.15 headline dimension for the `system_prompt` format**, which uses its own scorer set. Only `runtime` is excluded everywhere.\n\n### The composite: a full-denominator model\n\nSchliff does **not** quietly renormalize across whatever you happened to measure. Unmeasured dimensions **contribute 0 and stay in the denominator** — so coverage gaps lower your ceiling instead of quietly disappearing. Your score ceiling equals your measurement coverage. Measure 4 of the 7 headline dimensions and your maximum possible score is capped accordingly, with an explicit warning:\n\n```text\nℹ Scored 4/7 dimensions — the score can't exceed 42% until the rest\n  are measured. Run /schliff:init to add an eval suite and score:\n  triggers, quality, edges.\n```\n\n\u003e [!IMPORTANT]\n\u003e This is deliberate. A partial measurement is an honest partial score, never a flattering one. Unmeasured work is missing points, not invisible. To lift the ceiling, measure more — don't hide the gap.\n\n\u003e [!NOTE]\n\u003e **Structural score** = the composite renormalized over the dimensions Schliff can measure deterministically without an eval suite (structure, efficiency, composability, clarity). The full 7-dimension composite additionally folds in triggers, quality, and edges — which require an eval suite, generated with the `/schliff:init` Claude Code slash command. AGENTS.md needs no eval suite: its full 3-dimension headline (structure, operational_coverage, efficiency) is always measurable.\n\n\u003e [!NOTE]\n\u003e Calibration is strictly opt-in: ambient auto-calibrated weights apply **only** when `SCHLIFF_CALIBRATED_WEIGHTS` is set and **only** for the interactive `score` command, and Schliff emits a `weight_source=calibrated` warning flagging that such scores are **not** comparable to the canonical scale. Everything that gates a release stays canonical.\n\n### Grade scale\n\n`S` ≥ 95 · `A` ≥ 85 · `B` ≥ 75 · `C` ≥ 65 · `D` ≥ 50 · `E` ≥ 35 · `F` \u003c 35\n\n---\n\n## What the score does not measure\n\n- **Structure, not truth.** Schliff cannot verify that a documented command exists or runs. A syntactically plausible fabrication — invented-but-real-looking commands in well-formed sections — scores in the S range. This is a documented, test-pinned limit ([`test_known_limit_plausible_fabrication_scores_high`](skills/schliff/tests/unit/test_operational_coverage.py), spec §11).\n- **Not agent behavior.** A high score doesn't prove your agent gets better — validity evidence today is case-study-level (the two dated before/afters above), not benchmark-level.\n- **Token counts are estimates.** stdlib `len//4`, not a tokenizer.\n- **Coverage is on you.** `triggers`/`quality`/`edges` need an eval suite (`/schliff:init`) or the ceiling warning caps the score. AGENTS.md has no such gap.\n- **Informed declines are valid.** A low dimension can be a deliberate tradeoff (hydra left efficiency at 47 rather than gut a 14k-token file).\n\nA high Schliff score is necessary, not sufficient.\n\n---\n\n## Multi-format support\n\nOne engine, five instruction-file formats — each with its own token budget and scorer set:\n\n| Format | Token budget | Scorers |\n| --- | --- | --- |\n| `SKILL.md` | 1,000 | shared 8-scorer registry |\n| `CLAUDE.md` | 2,000 | shared 8-scorer registry |\n| `.cursorrules` | 500 | shared 8-scorer registry |\n| `AGENTS.md` | 3,000 | shared 8 scorers + `operational_coverage` (own 3-dim headline) |\n| system prompts | 1,500 | dedicated set (`structure_prompt`, `output_contract`, `efficiency`, `clarity`, `security`, `composability`, `completeness`) |\n\nFormat is auto-detected; override with `--format` (`skill`, `claude`, `cursor`, `agents`, `system-prompt`).\n\n---\n\n## Use it in CI\n\nPublished on the GitHub Marketplace as [AGENTS.md Lint (Schliff)](https://github.com/marketplace/actions/agents-md-lint-schliff).\n\n### GitHub Action\n\nGate pull requests on instruction-file quality. The action defaults to your\nrepo-root `AGENTS.md` and posts a scored comment on every PR:\n\n```yaml\n# .github/workflows/agents-lint.yml\nname: AGENTS.md Lint\non: [pull_request]\njobs:\n  score:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - uses: Zandereins/schliff@v1\n        with:\n          minimum-score: '75'   # optional: fail the PR below this score\n```\n\nBy default it scores `AGENTS.md` at the repo root; set `skill-path:` to lint a\n`SKILL.md`, `CLAUDE.md`, or `.cursorrules` instead. One caveat: the Action\ninstalls the latest released engine from PyPI, so after a release its scores can\nlead an older pinned install; pin it with `schliff-version: '8.5.0'` in the\n`with:` block if you need byte-stable gates.\n\n### CI gate without the Action\n\nPrefer not to depend on a third-party action? The dependency-light equivalent:\n\n```yaml\n      - run: pip install schliff\n      - run: schliff verify AGENTS.md --min-score 75\n```\n\n`schliff verify` exits non-zero when the score falls short and works for every\nsupported format — the minimum is scaled by measurement coverage, so `SKILL.md`\nfiles without an eval suite aren't auto-failed. Requires schliff ≥ 8.5.0 for\n`AGENTS.md`: older engines scored it under the SKILL profile\n([#101](https://github.com/Zandereins/schliff/issues/101)).\n\n### README badge\n\nShow your repo's AGENTS.md quality — no setup, no CI, no account:\n\n```markdown\n![AGENTS.md quality](https://img.shields.io/endpoint?url=https%3A%2F%2Fschliff-playground.vercel.app%2Fapi%2Fbadge%3Frepo%3DOWNER%2FREPO)\n```\n\nReplace `OWNER/REPO` with your repository. The badge is scored live from your\n`AGENTS.md` at `HEAD`; GitHub's image cache (camo) may delay refreshes. Public\nrepos only.\n\n### pre-commit\n\n```yaml\n# .pre-commit-config.yaml\nrepos:\n  - repo: https://github.com/Zandereins/schliff\n    rev: v8.5.0\n    hooks:\n      - id: schliff-verify\n        args: ['--min-score', '75']\n```\n\nThe hook fires on `SKILL.md` files (its `files` filter); gate `AGENTS.md` with the Action or `schliff verify AGENTS.md` in CI.\n\n---\n\n## CLI\n\n```text\nschliff \u003ccommand\u003e [path] [options]\n```\n\n| Command | What it does |\n| --- | --- |\n| `score` | Score a file and print the grade bar |\n| `verify` | CI gate — exit 0/1 based on a minimum score |\n| `doctor` | Scan and grade every installed skill |\n| `badge` | Generate a Markdown score badge |\n| `diff` | Explain score changes between two git commits |\n| `compare` | Compare two files side by side |\n| `suggest` | Rank fixes by estimated score impact |\n| `report` | Generate a Markdown score report |\n| `demo` | Score a built-in bad skill to see Schliff in action |\n| `evolve` | Improve an instruction file's score |\n| `version` | Print the version |\n\n---\n\n## Optional: closing the loop\n\nBeyond grading, Schliff can apply fixes. The improvement engine **measures first, then fixes** (not the other way around):\n\n1. **Score** the file across all dimensions.\n2. **Generate** deterministic patch gradients for the weakest dimensions.\n3. **Apply** the safe, rule-based patches automatically — **~32% of suggested fixes** apply deterministically through the apply gate (confidence=high, single-edit), as measured by the canonical script [`measure_patch_ratio.py`](skills/schliff/scripts/measure_patch_ratio.py) — re-run it to check. The rest are handed to an optional LLM.\n4. **Re-score** and keep the change only if the score improved — otherwise revert.\n5. **Stop** on plateau detection or when the target is reached.\n\nIt also carries **cross-session episodic memory** ([`episodic_store.py`](skills/schliff/scripts/episodic_store.py)), so improvement runs learn from prior attempts instead of repeating them. Drive it from Claude Code with `/schliff:auto`, or use `schliff evolve` directly. This is an optional convenience layer — the deterministic score is the product.\n\n```bash\npip install \"schliff[evolve,judge]\"  # optional LLM extras for this layer only\n```\n\nLLM extras power this optional layer only; they are never used for scoring.\n\n| Install | Pulls in | When you need it |\n| --- | --- | --- |\n| `schliff` | stdlib only | Scoring, verify, badge, CI — everything that gates a release |\n| `schliff[judge]` | LLM client | Opt-in exploratory LLM-judge smoke-test (never scoring) |\n| `schliff[evolve]` | LLM client | Opt-in autonomous-improvement extras |\n\n---\n\n## Under the hood\n\nThe full methodology — scorer internals, the full-denominator composite, the anti-gaming guards, and the calibration model — lives in [`docs/SCORING.md`](docs/SCORING.md).\n\n```text\nscripts/\n├── cli.py                  # CLI entrypoint\n├── scoring/\n│   ├── registry.py         # canonical weights, scorer lists, headline exclusions\n│   ├── composite.py        # full-denominator composite model\n│   ├── formats.py          # format detection + token budgets\n│   ├── guards.py           # anti-gaming detection\n│   └── structure.py · triggers.py · quality.py · edges.py · …\n├── text_gradient.py        # deterministic patch gradients (apply gate)\n├── episodic_store.py       # cross-session episodic memory\n└── measure_patch_ratio.py  # canonical source for the patch-ratio claim\n```\n\n---\n\n## Links \u0026 docs\n\n- **Docs:** [`docs/SCORING.md`](docs/SCORING.md)\n- **Playground:** [schliff-playground.vercel.app](https://schliff-playground.vercel.app) — paste a SKILL.md or AGENTS.md, get a live score (or `schliff demo` in the CLI). The playground reports the structural score for SKILL.md — the CLI's full 7-dim composite can differ; AGENTS.md has no such gap.\n- **Leaderboard:** [schliff-leaderboard.vercel.app](https://schliff-leaderboard.vercel.app)\n- **Case studies:** [`docs/case-studies/`](docs/case-studies/)\n\n## License\n\nMIT © Franz Paul\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FZandereins%2Fschliff","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FZandereins%2Fschliff","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FZandereins%2Fschliff/lists"}