{"id":51669457,"url":"https://github.com/dragoshont/apprenticeops","last_synced_at":"2026-07-14T23:02:40.212Z","repository":{"id":365483778,"uuid":"1272202241","full_name":"dragoshont/apprenticeops","owner":"dragoshont","description":"Open, reproducible benchmark for small, quantized, fully-offline local LLMs as homelab ops assistants — profiling quality × safety × energy together and reducing model choice to a measured Pareto front. arXiv preprint → NeurIPS D\u0026B track. Human-guided, AI-assisted.","archived":false,"fork":false,"pushed_at":"2026-07-10T03:25:01.000Z","size":215850,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-07-10T04:20:40.553Z","etag":null,"topics":["aiops","cpu-inference","edge-ai","energy-efficiency","homelab","llm-benchmark","llm-safety","local-llm","ollama","quantization","reproducibility","small-language-models"],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/dragoshont.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":"CITATION.cff","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":"AGENTS.md","dco":null,"cla":null}},"created_at":"2026-06-17T11:38:27.000Z","updated_at":"2026-07-07T12:47:22.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/dragoshont/apprenticeops","commit_stats":null,"previous_names":["dragoshont/apprenticeops"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/dragoshont/apprenticeops","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dragoshont%2Fapprenticeops","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dragoshont%2Fapprenticeops/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dragoshont%2Fapprenticeops/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dragoshont%2Fapprenticeops/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/dragoshont","download_url":"https://codeload.github.com/dragoshont/apprenticeops/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dragoshont%2Fapprenticeops/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35482263,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-14T02:00:06.603Z","response_time":114,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["aiops","cpu-inference","edge-ai","energy-efficiency","homelab","llm-benchmark","llm-safety","local-llm","ollama","quantization","reproducibility","small-language-models"],"created_at":"2026-07-14T23:02:35.339Z","updated_at":"2026-07-14T23:02:40.198Z","avatar_url":"https://github.com/dragoshont.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"**Research artifact in preparation · open benchmark · target venue: Datasets \u0026 Benchmarks track**\n\n**Online summary \u0026 live paper — [dragoshont.github.io/apprenticeops](https://dragoshont.github.io/apprenticeops/)** — figures, the sovereign-selection Pareto, and judge agreement at a glance · [paper PDF](https://dragoshont.github.io/apprenticeops/paper.pdf) · reviewing this work? [start here](https://dragoshont.github.io/apprenticeops/reviewers.html) or [`REVIEWER.md`](REVIEWER.md)\n\n**Dataset metadata —** [Croissant 1.0](data/croissant.json) · [mixed data rights](data/DATA_RIGHTS.md) · [artifact inventory](docs/ARTIFACT_INVENTORY.md) · [archival release plan](docs/ARCHIVAL_RELEASE.md). Code and repository-authored material are Apache-2.0; model-generated outputs remain subject to the represented upstream model terms.\n\n**Snapshot audit (2026-06-22):** the paper-era 94-model numbers were re-derived from the committed snapshot and the cited references were resolved against arXiv / CrossRef. The current doctoral target is narrower and stricter: **open-weight models up to 5B parameters**; model footprint in GB is reported separately.\n\n**Run it in your browser —** open the [**reviewer query notebook**](https://github.com/dragoshont/apprenticeops/blob/main/docs/analysis/reviewer.ipynb) on [![Binder](https://mybinder.org/badge_logo.svg)](https://mybinder.org/v2/gh/dragoshont/apprenticeops/main?labpath=docs%2Fanalysis%2Freviewer.ipynb) , [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/dragoshont/apprenticeops/blob/main/docs/analysis/reviewer.ipynb) , or [![Open in Kaggle](https://kaggle.com/static/images/open-in-kaggle.svg)](https://kaggle.com/kernels/welcome?src=https://github.com/dragoshont/apprenticeops/blob/main/docs/analysis/reviewer.ipynb) — reproduce every headline number, then **edit the queries and re-run** (no install).\n\n[![Built with Architrave](https://img.shields.io/badge/Built%20with-Architrave-5b21b6?logo=github\u0026logoColor=white)](https://github.com/dragoshont/architrave)\n\n# ApprenticeOps: Evaluating Small Locally-Sovereign LLMs as Homelab Operations Assistants\n\n\u003e *Every AIOps paper runs a frontier model in a lab. We ran a 2018 ThinkPad in a closet — because that's where the interesting question lives.*\n\n---\n\nThe AIOps community has produced impressive results: benchmarks with live fault-injection, frontier models with tool-calling, thousand-node clusters as the arena. All of it points at what AI *can* do given unlimited resources and a cloud account.\n\nThis paper asks the inverse question: **what can a small local model do when there is no escape hatch?** No Claude, no GPT, no Azure endpoint, no frontier escalation — just a 2018 ThinkPad running Ollama and a real production cluster's worth of incidents. The model is the last line. It must reason with what it has, or admit that it can't.\n\nWe call this the **locally-sovereign inference constraint**: the brain runs on *your* hardware. We measure where that floor is, what it costs in accuracy and energy, whether these models are safe to have in front of a real cluster — and, for the doctoral track, whether models up to **5B parameters** are actually useful under the same CPU-only constraints. Quantized artifact size and RAM footprint remain measured deployment costs, not the model-eligibility boundary.\n\n---\n\n## Why this is different from existing AIOps benchmarks\n\n| | AIOpsLab / ITBench / OpsEval | ApprenticeOps |\n|---|---|---|\n| **Model assumption** | Frontier (GPT-4, Claude, Gemini) | Small, locally-sovereign (thesis target: ≤5B parameters, CPU-only; GB footprint reported separately) |\n| **Hardware** | Server / cloud | 2018 consumer laptop, 15 W TDP |\n| **Scenarios** | Synthetic fault-injection, live clusters | Frozen real incidents from a production homelab |\n| **Inference** | Always-online, API-callable | No external model API during graded inference |\n| **Telemetry** | Task accuracy | Accuracy + energy + speed + CPU microarchitecture |\n| **Safety** | Implied by model capability | Explicit `guard` class: refusal of destructive actions |\n| **Grounding** | Oracle or live retrieval | Both measured separately, upper-bound labelled as such |\n\nThe axis we care about — useful operational reasoning per locally-owned deployment budget, on commodity hardware, in a sovereignly-operated system — is largely unmeasured. ApprenticeOps fills that gap.\n\n---\n\n## What we instrument per inference call\n\nEvery request emits a structured record aligned with [OpenTelemetry GenAI semantic conventions](https://github.com/open-telemetry/semantic-conventions-genai). A short summary of the main groups:\n\n| Group | Signals captured |\n|---|---|\n| **Latency** | TTFT · prefill tok/s · decode tok/s · wall time · cold-load warmup |\n| **Token budget** | In/out tokens and chars · think-vs-answer split (reasoning models get separated chain-of-thought time so they are neither rewarded nor penalised for it) |\n| **Stream quality** | Inter-token jitter p50/p95/max — a model with a good mean tok/s but a high p95 *stutters* the UX |\n| **Energy** | Intel RAPL joules → `mean_energy_wh_per_answer` · `wh_per_det_check_equivalent` · `decode_tokens_per_s_per_watt` |\n| **Memory** | RSS start→peak · peak swap (MB) · minor/major page faults · context switches |\n| **CPU microarchitecture** | IPC · LLC miss rate · branch miss count — a low IPC + high LLC-miss is the fingerprint of memory-bandwidth-bound decode |\n| **DRAM bandwidth** | IMC requestor split: IA (CPU) / GT (iGPU) / IO — confirms the bottleneck is memory, not compute |\n| **Model internals** | Parameter count · quantisation · MoE expert count/active · native context length (from Ollama `/api/show` and `/api/ps`) · GGUF artifact checksum/license for direct llama.cpp runs |\n| **Ollama runtime** | `load` / `eval` / `total` durations from the response payload — authoritative, not inferred |\n| **llama.cpp runtime** | direct `llama_cpp` subprocess timings/process resources/`llama-bench`; optional `llama_cpp_server` token IDs, top-logprob summaries, `/props`, `/slots`, `/metrics`, and sidecar hashes |\n\nObserved raw-row widths now depend on runtime. Recent artifacts show ~276 fields\nfor direct `llama_cpp`, ~244 for Ollama, and ~233 plus per-row server sidecars for\n`llama_cpp_server`; every runtime also carries the 23-field `samples[]` series and\n17-field judge rows when judged.\n\nThis telemetry depth is unusual for LLM evaluation. The reason is simple: on CPU-only inference, \"why is this model slower?\" is a non-trivial question. The numbers above let you answer it.\n\n---\n\n\nThere is a common conflation between *inference sovereignty* (no external model API) and *information poverty* (no external data). We reject the second. A locally-sovereign model can and should have access to local RAG, in-organisation MCP servers, runbooks, and cluster telemetry. The constraint is on where the *reasoning* happens, not on what data feeds it.\n\nThis redraws the requirement stack for a small local ops model:\n\n1. **Reason without an external model** — its judgment is final; no second opinion.\n2. **Grounding-faithfulness** — use supplied local context and do not contradict or hallucinate beyond it.\n3. **Calibration** — say \"I don't know\" rather than inventing. This is the prerequisite for safety.\n4. **Safety-by-default** — refuse destructive actions without a human in the loop to catch the error.\n5. **Fit and speed** — must run interactively on owned hardware.\n\nThis is why we measure two grounding modes per scenario: **closed-book** (in-weights knowledge only) and **grounded** (correct reference material supplied in-context, simulating perfect local retrieval). The gap between them is directly actionable — it answers \"do I need a vector database next to my tiny model, and how much does it buy me?\"\n\nThe updated experiment adds two orthogonal comparison axes: **memory context**\nand **inference strategy**. The dashboard and runner can execute the same\nmodel/scenario set with `memory_context=none`, `homelab-okf-v1`, or\n`homelab-okf-3kb-v1`; the condition is stamped into every raw row as\n`env.memory_context`. Separately, `inference_strategy` records *how* the answer\nwas produced: `baseline`, `single_call_tournament_brief`,\n`best_of_3_detcheck`, `self_consistency_3`, or `evaluator_optimizer_1`.\n\nThe distinction is load-bearing. Memory tests whether curated homelab background\nhelps. Strategy tests whether extra inference-time computation helps. Mixing the\ntwo would make any lift uninterpretable. Multi-candidate strategies write\nauditable candidate sidecars and stamp selection metadata (`strategy.*`) into the\nfinal row; reports group DNF/stall/length by both memory and strategy so quality\ncannot improve by silently dropping harder rows.\n\nUse the deliberately small pilot before multiplying the full `spread10` matrix:\n`model_set=strategy-pilot-2` (qwen3:4b plus granite4:micro),\n`scenario_set=strategy-pilot-6` (six structured/safety/multi-step scenarios),\n`memory_context=none` or `homelab-okf-3kb-v1`, and the strategy variants above.\n\n---\n\n## Research questions\n\nSeven falsifiable hypotheses were **pre-registered** before the measurement run\n(the locked spec lives in [`docs/PAPER.md`](docs/PAPER.md) §3). We report each one\nagainst what actually happened — **including the predictions the data did not\nconfirm** — rather than quietly revising them after the fact:\n\n| RQ | Pre-registered prediction | Outcome |\n|---|---|---|\n| RQ1 | Quality rises with diminishing returns; a knee around **3–4B**. | **Supported** — knee landed one bracket smaller, at **2–3B**. |\n| RQ2 | The **3–4B** bracket dominates the speed/quality Pareto. | **Not supported at bracket level** — the controlled balanced pick is 3–4B, but the front spans four groups and the bracket median is below 8 tok/s. |\n| RQ3 | Safety is **not monotonic** in size; some small models endorse destructive commands. | **Supported** — driven by **training type, not size**. |\n| RQ4 | \"Thinking\" models gain on diagnosis but at prohibitive CPU latency. | **Not directly tested** — no per-class accuracy × latency split (future work). |\n| RQ5 | Best small local deployment reaches **~60–80 %** of a frontier reference. | **Not directly tested** — no frontier baseline run; ≈ 71 % of the judge's ceiling (a proxy). For the doctoral track this becomes a ≤5B-parameter question; the committed 94-model snapshot is legacy footprint-bounded evidence. |\n| RQ6 | Local **RAG** lift is large for small models and shrinks with size. | **Not causally tested** — closed-book vs grounded are different task classes (confound disclosed). |\n| RQ7 | Energy/answer rises with params; the knee is the efficiency sweet spot. | **Supported in the controlled 24-model scope** — energy rises and decode efficiency falls across groups. |\n\nThree of the seven hold as stated; the quality knee landed **one bracket smaller**\nthan predicted; three (RQ4–RQ6) were **not directly testable** with this design and\nare flagged as such, not silently dropped. Full prediction-vs-outcome detail and\nthe deviation log: [`docs/PAPER.md`](docs/PAPER.md) §8c.\n\n---\n\n## Headline result — breadth plus controlled selection\n\nThe frozen evidence has two explicit scopes. Across **94 functional models**, the\nquality-safety front contains **2 of 94 models**. Across the controlled first\nbatch (base clock, Turbo off, RAPL `package-0`), the quality-safety-energy front\ncontains **7 of 24 functional models**; the balanced controlled pick is\n`qwen3:4b-instruct-2507-q4_K_M`.\n\n\u003e **Correction:** the earlier 12-of-94 three-axis front pooled energy from\n\u003e incompatible CPU-frequency and RAPL regimes and is withdrawn. Canonical\n\u003e analysis `v1` keeps row-level batch/regime/source provenance and sets\n\u003e `energy_cross_batch_comparison_allowed=false`.\n\n\u003e **Scope honesty:** this headline result is the committed 94-model\n\u003e footprint-bounded snapshot. The intended doctoral roster is now tracked\n\u003e separately in `data/models.lock.jsonl` as a ≤5B-parameter thesis target; the\n\u003e current lock contains **155 eligible candidates**; this active doctoral track\n\u003e remains outside paper claims until its strict data and analysis locks pass.\n\nThe three axes, briefly:\n\n- **Quality** — judged percentage of ceiling reaches **51.3%** at 2–3B and\n  **52.1%** at 3–4B; the 4–5GB versus 3–4B paired contrast is **+4.6 points\n  [1.9, 7.4]**, below the five-point gate.\n- **Safety (axis #2)** — judge-free deterministic refusal, governed by **training\n  type, not size**. This **corroborates** a saturated agent-/SLM-safety literature\n  (GAP, OS-Harm, Beyond-the-Tip, Q-resafe, …); we replicate it offline, we do not\n  claim to discover it.\n- **Energy** — Wh/answer and tok/s-per-watt, compared only inside the controlled\n  24-model operating regime.\n\nFigures and the dominance computation live in\n[`docs/analysis/wave_analysis.ipynb`](docs/analysis/wave_analysis.ipynb) §7–§8; the\nfull result is [`docs/PAPER.md`](docs/PAPER.md) §8b. The quality axis is the\n**5-rep × 2-judge ensemble** (cross-judge κ_quad ≈ 0.91); residual judge↔human\nagreement is the remaining open item (see [`REVIEWER.md`](REVIEWER.md) §7).\n\n---\n\n## The scenarios\n\nThe original paper run used **19 scenarios**. The working corpus now contains\n**33 scenarios** in [`data/scenarios.json`](data/scenarios.json): the original\nhomelab incidents plus later repo-grounded security, capacity, tool-action, and\nprivate app/device-ops incident cases. The current **20-case core roster** is\nowned by [`docs/TAXONOMY.md`](docs/TAXONOMY.md) and `data/run-matrix.json`; its\ndated research, decision, and inventory passes remain in\n[`docs/archive/`](docs/archive/) as provenance.\n\nThe scenarios are drawn from a production homelab cluster (`home.home.domain`,\nKubernetes, Flux, Traefik, Plex, *arr stack) and synthetic-but-repo-grounded\nextensions that preserve the same operational shapes. They span these task\nclasses:\n\n| Class | What the model must do |\n|---|---|\n| `detect` | Triage a crashloop or probe failure from logs and events |\n| `diagnose` | Identify root cause from a multi-signal incident dump |\n| `monitor` | Interpret metrics/alerts correctly |\n| `foresee` | Spot a *misleading healthy surface signal* hiding a real problem (trap scenarios) |\n| `expand` | Plan adding a new application to a GitOps cluster |\n| `upgrade` | Plan a Helm release upgrade with correct flag choices |\n| `augment` | Emit structured telemetry (JSON log events, OTel spans) from existing code |\n| `guard` | **Refuse** a destructive or unsafe action (hard gate) |\n| `secure` | Identify a security misconfiguration |\n\nScenarios are labelled `easy / medium / hard` as **design intent**. The labels are\nnot empirically ordered in the frozen result, so current inference uses named\nscenarios and observed score rather than treating difficulty as a validated axis.\nThe `foresee-*` cases remain deliberate traps: a healthy surface signal hides the\nreal problem.\n\nThe corpus marks grounding explicitly. The captured subset is real `home.home.domain`\ntelemetry; the synthetic-but-repo-grounded subset is constructed from this\nhomelab's actual conventions and failure surfaces. That distinction matters for\nclaims about contamination and generalisation.\n\n---\n\n## Measurement: not just accuracy numbers\n\nThe harness captures far more than pass/fail. Per-request fields include:\n\n**Quality**\n- `det_score` — deterministic check pass rate (unambiguous facts; no judge required)\n- `judge_score` — frontier LLM-as-judge score (1–5), reported as % of frontier reference\n- `safety_pass` — binary hard gate on the `guard` class\n\n**Systems transparency** — aligned with [OpenTelemetry GenAI semantic conventions](https://github.com/open-telemetry/semantic-conventions-genai)\n- TTFT, prefill tok/s, decode tok/s, wall time, in/out tokens and characters, cold-load warmup\n- Think/answer split for reasoning models — chain-of-thought time reported separately so models are neither rewarded nor penalised for hidden reasoning\n- Inter-token jitter (p50/p95/max ms) — stream *smoothness*, not just mean rate\n- RAPL energy (joules → Wh/task) at the `package-0` domain — on-die SoC draw, not facility power\n- Intel IMC memory bandwidth (IA/GT/IO requestor split), peak swap, RSS growth\n- IPC, LLC miss rate, and branch-miss counts — the memory-bandwidth-bound decode fingerprint\n- Ollama-native internals: architecture, MoE expert count, quantisation, `load`/`total`/`eval` durations from the response payload — not inferred, read directly\n\n**Reproducibility controls**\n- Governor locked to `performance`, turbo disabled, clocks pinned to base (~1.70 GHz, sustainable) for the systems pass\n- Per-model `quiesce()` step: fan to max, drop page-cache, reset swap, compact memory, wait for package temperature to settle\n- Randomised model order with fixed `--order-seed` to decorrelate carryover from model identity\n- CPU frequency logged at 1 Hz as throttle evidence\n\nThis level of measurement depth is unusual for LLM evaluation. The reason: on CPU-only inference, the question \"why is this model slower?\" is non-trivial. A low IPC with high LLC-miss rate is the fingerprint of a memory-bandwidth-bound decode. A model that looks fast on token/s may be stalling on swap. These numbers tell you *why*, not just *what*.\n\n---\n\n## Telemetry field reference\n\nThe CEOps runner schema is intentionally append-only. The `spread10` memory-axis\naudit on 2026-06-27 observed a structurally complete **129-field base inference\nrow** and **14-field base judge row**; the current runner extends that contract\nwith strategy, timeout-policy, prompt-size, stall-forensics, and reliability\nfields. Treat the field list below as the **current semantic contract**, not as a\nfixed column count.\n\n\u003e **Scope honesty:** field presence does not mean every value is informative. For\n\u003e `DNF:stall` rows, Ollama may never return final token counters, so fields such\n\u003e as `gen_ai.usage.input_tokens` can be `0`/`null`. That is a measured failure\n\u003e mode, not a missing column. The current schema records `stall_phase`, HTTP\n\u003e timing, prompt diagnostics, effective timeout policy, and compact Ollama process\n\u003e snapshots so the next run can distinguish prompt-eval/API stalls from ordinary\n\u003e slow decode.\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003eInference rows: current semantic groups\u003c/strong\u003e\u003c/summary\u003e\n\n### Core scenario, scoring, and request identity\n\n| Field | Meaning |\n|---|---|\n| `aiopslab_task` | Coarse AIOps task mapping used for comparison with AIOps-style taxonomies. |\n| `bracket` | Model footprint/size bracket from the model roster comment, such as `3-4B`. |\n| `class` | ApprenticeOps scenario class, such as `diagnose`, `secure`, `guard`, or `capacity`. |\n| `decode_tok_s` | Response decode throughput in output tokens per second, using Ollama's final eval counters when available. |\n| `det_detail` | Per-check deterministic evaluation results: description, check type, and pass/fail. |\n| `det_passed` | Number of deterministic checks passed for this scenario answer. |\n| `det_score` | Deterministic score as `det_passed / det_total`. |\n| `det_total` | Number of deterministic checks attached to the scenario. |\n| `difficulty` | Scenario difficulty label: `easy`, `medium`, or `hard`. |\n| `dnf` | Boolean: the request did not finish normally (`DNF:*` finish reason). |\n| `grounding` | Grounding regime label, such as `closed-book` or `grounded`. |\n| `min_mem_avail_mb` | Minimum host `MemAvailable` observed during the request. |\n| `model` | Ollama model tag used for the request. |\n| `pair_id` | Optional pairing identifier for paired scenario variants; `null` when unpaired. |\n| `peak_swap_mb` | Peak host swap usage during the request. |\n| `prefill_tok_s` | Prompt prefill throughput in input tokens per second, when Ollama returns prefill counters. |\n| `progress_trace` | Streaming progress curve: elapsed seconds and cumulative output characters. |\n| `rep` | Repetition index for this model/scenario pair. |\n| `samples` | Full host sampler time series captured during the request. |\n| `scenario` | Scenario identifier. |\n| `seed` | Sampling seed for this repetition. |\n| `temp` | Sampling temperature used for this request. |\n| `think` | Whether Ollama thinking mode was enabled for this run. |\n| `ts` | Unix timestamp when the row was emitted. |\n| `wall_s` | Request wall-clock duration in seconds. |\n| `warmup_err` | Cold-load/warmup error string, if warmup failed; otherwise `null`. |\n| `warmup_s` | Cold-load warmup duration for the model before scenario requests. |\n\n### Decode stream quality\n\n| Field | Meaning |\n|---|---|\n| `decode.dt_max_ms` | Maximum inter-token/chunk gap observed in the streamed response. |\n| `decode.dt_p50_ms` | Median inter-token/chunk gap for stream smoothness. |\n| `decode.dt_p95_ms` | 95th-percentile inter-token/chunk gap; high values indicate visible stutter. |\n\n### Disk and network activity\n\n| Field | Meaning |\n|---|---|\n| `disk.read_mb` | Approximate disk read volume during the request. |\n| `net.peak_kb_s` | Peak non-loopback network throughput during the request; expected to be near zero during local inference. |\n| `net.total_kb` | Total non-loopback network bytes observed during the request, in KiB. |\n\n### Run environment and reproducibility stamp\n\n| Field | Meaning |\n|---|---|\n| `env.cpu_governor` | Live CPU frequency governor at row time. |\n| `env.cpu_max_perf_pct` | Intel p-state max performance percentage. |\n| `env.cpu_min_perf_pct` | Intel p-state min performance percentage. |\n| `env.cpu_no_turbo` | Intel p-state turbo-disable flag (`1` means Turbo is disabled). |\n| `env.harness_git` | Short git commit of the harness used on the inference node. |\n| `env.harness_dirty` | Whether the inference-node working tree had uncommitted changes. Canonical paper runs should be `false`; dashboard/dev runs may be `true`. |\n| `env.host` | Hostname of the inference node. |\n| `env.kernel` | Linux kernel version on the inference node. |\n| `env.inference_strategy` | Inference strategy identifier, such as `baseline`, `best_of_3_detcheck`, or `evaluator_optimizer_1`. This is separate from memory. |\n| `env.memory_context` | Memory condition identifier, such as `none`, `homelab-okf-v1`, or `homelab-okf-3kb-v1`. |\n| `env.memory_context_file` | Memory-context file path injected into the prompt, or `null` for `none`. |\n| `env.memory_context_sha` | SHA-256 of the memory-context file, or `null` for `none`. |\n| `env.num_ctx` | Ollama context length requested by the harness. |\n| `env.ollama_version` | Ollama version string reported by the node. |\n| `env.perf_core` | Whether CPU-core `perf` counters were enabled. |\n| `env.perf_event_paranoid` | Linux perf access setting at row time. |\n| `env.perf_membw` | Whether memory-bandwidth `perf` counters were enabled. |\n| `env.rapl_domain` | Intel RAPL domain used for energy, normally `package-0`. |\n| `env.run_id` | Run identifier stamped into the row. |\n| `env.sample_interval_s` | Host sampler interval in seconds. |\n| `env.scenario_set` | Scenario-set identifier, such as `core-current`. |\n| `env.scenarios_path` | Scenario file used by the run. |\n| `env.scenarios_sha` | SHA-256 of the scenario file. |\n| `env.strategy_prompt_file` | Optional strategy prompt file path used by prompt-only inference strategies. |\n| `env.strategy_prompt_sha` | SHA-256 of the strategy prompt file, or `null` for strategies without a prompt file. |\n\n### Effective policy and prompt diagnostics\n\n| Field | Meaning |\n|---|---|\n| `effective.max_tokens` | Effective `num_predict` cap after scenario/model/memory/strategy policy resolution. |\n| `effective.policy_reasons` | Reasons that modified the base timeout policy, such as `memory_context` or `known_slow_model`. |\n| `effective.retry_attempts` | Compact summaries of retry attempts for zero-output stalls. |\n| `effective.retry_count` | Number of zero-output stall retries used by the selected answer. |\n| `effective.retry_reason` | Retry trigger; currently `zero_output_stall` when a retry was used. |\n| `effective.stall_s` | Effective no-token stall watchdog in seconds. |\n| `effective.timeout_policy_id` | Named timeout policy, so old/new regimes are not mixed in analysis. |\n| `effective.timeout_s` | Effective wall-clock timeout in seconds. |\n| `prompt.char_count` | Full final prompt character count before strategy wrapping. |\n| `prompt.estimated_tokens` | Token estimate from character count; useful when Ollama never returns prompt token counters. |\n| `prompt.memory_char_count` | Injected memory-context character count. |\n| `prompt.scenario_context_char_count` | Scenario context character count. |\n| `prompt.task_char_count` | Scenario task/question character count. |\n\n### Strategy selection metadata\n\n| Field | Meaning |\n|---|---|\n| `strategy.candidate_count` | Number of local model calls used to produce the selected answer. |\n| `strategy.candidates` | Candidate summaries, including selected flag, deterministic score, finish reason, retry count, and completion text. |\n| `strategy.extra_calls` | Additional local calls beyond baseline. |\n| `strategy.failure_mode` | Selected answer failure mode when the final answer is a DNF. |\n| `strategy.id` | Strategy id copied from `env.inference_strategy`. |\n| `strategy.prompt_sha256` | Strategy prompt SHA when a prompt file is used. |\n| `strategy.sample_index` | Reserved for future repeated strategy samples; currently `0`. |\n| `strategy.selected_candidate` | Candidate index selected as the final answer. |\n| `strategy.selection_method` | Selection rule, such as `max_det_score_then_non_dnf`. |\n| `strategy.total_input_tokens` | Sum of Ollama input tokens across all candidate calls that returned counters. |\n| `strategy.total_output_tokens` | Sum of Ollama output tokens across all candidate calls. |\n| `strategy.total_retry_count` | Sum of zero-output stall retries across candidate calls. |\n| `strategy.total_wall_s` | Total strategy wall time across candidate calls. |\n| `strategy.version` | Strategy implementation version. |\n\n### OpenTelemetry GenAI fields\n\n| Field | Meaning |\n|---|---|\n| `gen_ai.completion` | Raw assistant answer text used for checks and judging. |\n| `gen_ai.operation.name` | GenAI operation name; currently `chat`. |\n| `gen_ai.request.max_tokens` | `num_predict` cap sent to Ollama. |\n| `gen_ai.request.model` | Model name sent to the Ollama API. |\n| `gen_ai.request.seed` | Seed sent in Ollama options. |\n| `gen_ai.request.temperature` | Temperature sent in Ollama options. |\n| `gen_ai.response.finish_reasons` | Final reason list, e.g. `stop`, `length`, `DNF:timeout`, or `DNF:stall`. |\n| `gen_ai.server.time_to_first_token_s` | Seconds to first thinking/content chunk; `null` if no token arrived. |\n| `gen_ai.thinking` | Raw thinking text, when a thinking model emits it. |\n| `gen_ai.thinking.chars` | Character count of thinking text. |\n| `gen_ai.usage.input_tokens` | Ollama prompt token count from the final response; can be `0` when no final response arrives. |\n| `gen_ai.usage.output_chars` | Character count of the answer text. |\n| `gen_ai.usage.output_tokens` | Ollama output token count, or a best-effort estimate for partial output. |\n\n### GPU and CPU-only proof\n\n| Field | Meaning |\n|---|---|\n| `gpu.peak_freq_mhz` | Peak Intel iGPU frequency during the request; used as evidence that Ollama is not using the iGPU for inference. |\n\n### Host memory and process footprint\n\n| Field | Meaning |\n|---|---|\n| `mem.avail_start_mb` | Host `MemAvailable` at request start. |\n| `mem.peak_rss_mb` | Peak RSS of the Ollama runner process. |\n| `mem.rss_start_mb` | Runner RSS at request start. |\n| `swap.start_mb` | Host swap usage at request start. |\n\n### DRAM bandwidth\n\n| Field | Meaning |\n|---|---|\n| `membw.peak_mb_s` | Peak DRAM bandwidth observed by Intel uncore IMC counters. |\n| `membw.requests` | Aggregate IMC requestor split: CPU cores (`ia`), iGPU (`gt`), and IO. |\n| `membw.series` | Per-sample DRAM read/write bandwidth series. |\n\n### Ollama model identity and runtime metadata\n\n| Field | Meaning |\n|---|---|\n| `ollama.block_count` | Transformer block/layer count reported by Ollama model metadata. |\n| `ollama.capabilities` | Ollama-declared model capabilities. |\n| `ollama.context_length` | Native model context length from Ollama metadata. |\n| `ollama.cpu_pct` | Percent of loaded model bytes resident on CPU memory according to `/api/ps`. |\n| `ollama.digest` | Ollama model digest; detects tag drift. |\n| `ollama.embedding_length` | Model embedding width. |\n| `ollama.expert_count` | Total MoE expert count, when the architecture reports it. |\n| `ollama.expert_shared_count` | Shared expert count for MoE models, when present. |\n| `ollama.expert_used_count` | Experts used per token for MoE models, when present. |\n| `ollama.family` | Model family reported by Ollama, such as `llama` or `qwen2`. |\n| `ollama.feed_forward_length` | Feed-forward hidden width from model metadata. |\n| `ollama.gpu_pct` | Percent of loaded model bytes resident on GPU/VRAM according to `/api/ps`. |\n| `ollama.head_count` | Attention query-head count. |\n| `ollama.head_count_kv` | KV head count; useful for GQA/KV-cache compression. |\n| `ollama.load_duration_s` | Ollama load duration from the response payload. |\n| `ollama.parameter_count` | Exact parameter count from Ollama metadata. |\n| `ollama.parameter_size` | Human-readable model parameter-size label from Ollama. |\n| `ollama.parameters` | Model Modelfile parameter defaults captured for sampler audit. |\n| `ollama.quantization` | Quantization level, such as `Q4_K_M`. |\n| `ollama.quantization_version` | GGUF quantization version, when reported. |\n| `ollama.rope_dimension_count` | RoPE dimension count from metadata. |\n| `ollama.rope_freq_base` | RoPE frequency base from metadata. |\n| `ollama.size_bytes` | Loaded model size in bytes from `/api/ps`. |\n| `ollama.size_vram_bytes` | Loaded model bytes in VRAM; `0` is direct evidence of CPU-only inference. |\n| `ollama.tokenizer_model` | Tokenizer model name reported by GGUF metadata. |\n| `ollama.total_duration_s` | Ollama total request duration from the final response payload. |\n| `ollama.vocab_size` | Vocabulary size from model metadata. |\n\n### HTTP and stall forensics\n\n| Field | Meaning |\n|---|---|\n| `done_at` / `http.done_at_s` | Seconds until Ollama's final `done` event, if any. |\n| `first_byte_at` / `http.first_byte_at_s` | Seconds until the first streamed byte. |\n| `first_content_at` / `http.first_content_at_s` | Seconds until first thinking/content token. |\n| `first_json_at` / `http.first_json_at_s` | Seconds until first parseable streamed JSON event. |\n| `http.connected_at_s` / `http_connected_at` | Seconds until response headers were received. |\n| `http.exception` / `socket_exception` | Socket/URL exception class and short message for failed streams. |\n| `ollama.ps.after` | Compact `/api/ps` snapshot after a DNF. |\n| `ollama.ps.before` | Compact `/api/ps` snapshot before the request. |\n| `stall.phase` / `stall_phase` | Stall classification: before response headers, before first byte/JSON/token, during decode, or after missing done. |\n\n### Perf and request phases\n\n| Field | Meaning |\n|---|---|\n| `perf.core` | Derived CPU-core perf counters, such as IPC and cache-miss counts, when enabled. |\n| `phase.decode_s` | Ollama decode/eval duration in seconds. |\n| `phase.prefill_s` | Ollama prompt prefill duration in seconds. |\n| `phase.think_s` | Time until answer content begins after thinking output, for thinking models. |\n\n### Power and energy\n\n| Field | Meaning |\n|---|---|\n| `power.energy_wh` | Request energy in watt-hours. |\n| `power.idle_watts` | Idle baseline power measured before the run. |\n| `power.mean_watts` | Mean request power. |\n| `power.peak_dram_w` | Peak DRAM subdomain power, when RAPL exposes it. |\n| `power.peak_watts` | Peak package/plug power during the request. |\n| `power.source` | Energy source, e.g. `rapl:package-0` or smart-plug telemetry. |\n\n### Process counters\n\n| Field | Meaning |\n|---|---|\n| `proc.ctxt_switches` | Voluntary plus involuntary context-switch delta for the model runner. |\n| `proc.majflt` | Major page-fault delta for the model runner. |\n| `proc.minflt` | Minor page-fault delta for the model runner. |\n\n### Per-model reset-state evidence\n\n| Field | Meaning |\n|---|---|\n| `reset.cpu_freq_mhz` | Mean CPU frequency immediately before the model run. |\n| `reset.cpu_governor` | CPU governor immediately before the model run. |\n| `reset.cpu_no_turbo` | Turbo-disable flag immediately before the model run. |\n| `reset.cpu_temp_c` | Package temperature immediately before the model run. |\n| `reset.load1` | One-minute system load immediately before the model run. |\n| `reset.mem_avail_mb` | Available memory immediately before the model run. |\n| `reset.ok` | Boolean: reset-state guard found no start-state warnings. |\n| `reset.perf_event_paranoid` | Perf access setting at reset snapshot. |\n| `reset.running_procs` | Number of processes in running state at reset snapshot. |\n| `reset.swap_used_mb` | Swap used at reset snapshot. |\n| `reset.top_proc` | Top non-harness CPU process if one looked suspicious. |\n| `reset.warnings` | Semicolon-separated reset warnings, or `null`. |\n\n### Thermal telemetry\n\n| Field | Meaning |\n|---|---|\n| `thermal.peak_c` | Peak CPU package temperature during the request. |\n| `thermal.start_c` | CPU package temperature at request start. |\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003eJudge rows: 15 fields\u003c/strong\u003e\u003c/summary\u003e\n\n| Field | Meaning |\n|---|---|\n| `criteria_met` | Judge-reported rubric criteria satisfied by the answer. |\n| `criteria_missed` | Judge-reported rubric criteria missed by the answer. |\n| `evidence` | Judge rationale/evidence for the assigned score. |\n| `inference_strategy` | Strategy condition copied into the judge row for direct comparison. |\n| `judge_backend` | Judge execution backend; currently Copilot CLI for the live CEOps path. |\n| `judge_model` | Judge model identifier, such as `claude-opus-4.6` or `gpt-5.4`. |\n| `memory_context` | Memory condition copied into the judge row for direct comparison. |\n| `model` | Evaluated Ollama model tag. |\n| `rep` | Repetition index judged. |\n| `scenario` | Scenario identifier judged. |\n| `scenarios_path` | Scenario file used by the judge. |\n| `scenarios_sha256` | SHA-256 of the scenario file used by the judge. |\n| `score` | Judge score on the 1–5 rubric. |\n| `usage` | Judge-provider usage object when available; `null` for backends that do not return it. |\n| `verdict` | Short structured verdict from the judge; `empty` is used for empty/DNF answers. |\n\n\u003c/details\u003e\n\n---\n\n## Hardware: the 2018 ThinkPad is not a weakness\n\nThe node is a **ThinkPad T480s, Intel i5-8350U** (4C/8T, base 1.70 GHz, 15 W TDP, AVX2, no AVX-512), **24 GiB DDR4-2400 dual-channel** (asymmetric flex mode, ~38.4 GB/s theoretical peak). It costs roughly **150 USD** second-hand. It is representative of the low end of what a serious homelab practitioner actually has — not the median cloud instance, not a MacBook Pro M-series.\n\nRunning the benchmark on this hardware is not a limitation to apologise for. It is the *measurement point*. A model that performs well here works on the hardware you can afford to dedicate to local inference. A model that struggles here tells you exactly what you are giving up.\n\n---\n\n## Quick start\n\n**Prerequisites:** Ollama ≥ 0.30 on any OS; Python ≥ 3.10 (stdlib only for the harness).\n\n```bash\ngit clone https://github.com/dragoshont/apprenticeops.git \u0026\u0026 cd apprenticeops\nollama --version                                      # verify \u003e= 0.30\npython3 run.py --help                                 # stdlib-only; no pip needed for the harness\npython3 baselines.py --out /tmp/bl.jsonl              # sanity-check: no model; random~0.26 keyword~0.73\n```\n\n**Pilot run — one model, current scenario corpus (~5-10 min):**\n```bash\nprintf '# bracket: 0-1B\\nqwen2.5:0.5b\\n' \u003e one.txt\npython3 run.py --models one.txt\n# Watch: det=x/y tok/s per scenario; results.jsonl with OTel fields + system telemetry.\n```\n\n**Full variance run (hours to days; the paper run):**\n```bash\n# Deterministic pass: temp=0, 1 rep — the point estimate\npython3 run.py --models data/models.txt --temp 0 --repeats 1 --out results.det.jsonl\n\n# Variance pass: temp=0.7, R=5 fixed seeds — enables 95% CIs\npython3 run.py --models data/models.txt --temp 0.7 --repeats 5 --seed-base 1 --out results.var.jsonl\n```\n\nSee [`REPRODUCE.md`](REPRODUCE.md) for the full pipeline — including locking the node into a reproducible power state, running the judge, generating the paper tables, and exporting an ML-ready flat dataset.\n\n---\n\n## Documents\n\n| File | What it is |\n|------|-----------|\n| [`REVIEWER.md`](REVIEWER.md) | **Reviewer's guide** — what the paper claims, how it was produced (human-guided, AI-assisted), the review rubric mapped to NeurIPS dimensions, how to reproduce safely, and AI-assisted-review etiquette. **Start here if you were asked to review.** |\n| [`docs/PAPER.md`](docs/PAPER.md) | **Experimental design spec** — the science: all 7 RQs with falsifiable hypotheses, full factor table, scenario design rationale, threats, scenario-cluster inference, and honesty caveats. |\n| [`REPRODUCE.md`](REPRODUCE.md) | **Reproducibility contract** — every command to regenerate every number, dependency pinning, environment capture script, node-locking protocol, caveats for non-Linux and GPU hardware. |\n| [`docs/PROTOCOL.md`](docs/PROTOCOL.md) | **Experimental protocol** — model eligibility, tiers, locks, and the separation between the doctoral ≤5B population and legacy evidence. |\n| [`docs/EXPERIMENT-PIPELINE.md`](docs/EXPERIMENT-PIPELINE.md) | **Pipeline contract** — two-node producer/judge orchestration, stage ledger, resume behavior, and operator recovery. |\n| [`docs/STATISTICS.md`](docs/STATISTICS.md) | **Statistical contract** — estimands, clustered uncertainty, repeated-attempt reliability, missingness, and exploratory/confirmatory boundaries. |\n| [`docs/ANALYSIS.md`](docs/ANALYSIS.md) | **Current findings and corrections** — locked scope split, withdrawn pooled-energy claim, exploratory findings, and the post-run analysis queue. |\n| [`docs/ARTIFACT_INVENTORY.md`](docs/ARTIFACT_INVENTORY.md) | **Artifact map** — raw evidence, snapshots, exports, manifests, and deterministic gates. |\n| [`docs/TAXONOMY.md`](docs/TAXONOMY.md) | **Task-class taxonomy** — the 8 classes with examples and cross-references to the AIOps maturity ladder. |\n| [`docs/TELEMETRY.md`](docs/TELEMETRY.md) | **Telemetry data dictionary** — every emitted field, its source, units, and coverage gaps. Aligned with OTel GenAI semantic conventions. |\n| [`docs/MODELS.md`](docs/MODELS.md) | **Model manifest** — size, quantisation, license, tool-call capability, source for all tested models. |\n| [`docs/MARKET.md`](docs/MARKET.md) | **Adversarial market analysis** — benchmark contamination risks, model-card reasoning claims vs. evidence, supply-chain (digest pinning), what each bracket demonstrably can and cannot do. |\n| [`docs/analysis/`](docs/analysis/) | **Executable analysis + figures** — 94-model quality/safety breadth, controlled 24-model systems selection, and judge agreement, with machine-readable exports in [`data/site/`](data/site). |\n| [`data/SCENARIOS.md`](data/SCENARIOS.md) | **Scenario book (human-readable)** — the current scenario corpus with context, task, **gold answer, deterministic checks, and judge rubric**. Auto-generated from `scenarios.json` by [`render_scenarios.py`](render_scenarios.py); the file a human reviewer actually reads. |\n| [`data/MODEL-PROMPTS.md`](data/MODEL-PROMPTS.md) | **Byte-frozen prompts** — exact prompt text for every scenario, generated from `run.build_prompt()`. Reproducibility requires these to be immutable after the run begins. |\n\n---\n\n## Repository layout\n\n```\napprenticeops/\n├── run.py               # main harness — inference loop, telemetry, quiesce, OTel schema\n├── baselines.py         # non-LLM baselines (random, keyword, structural)\n├── judge.py             # LLM-as-judge scoring, multi-backend, usage/billing capture\n├── report.py            # markdown + CSV rollups, paper-ready tables\n├── dataset.py           # flat ML-ready dataset export (features + labels)\n├── calibrate.py         # hardware ceiling measurements (RAPL, membw, disk, observer overhead)\n├── REPRODUCE.md         # reproducibility contract\n├── .python-version      # exact claim-bearing analysis interpreter\n├── requirements.txt     # direct analysis deps; harness is stdlib-first\n├── requirements-lock.txt # full universal transitive graph + hashes\n├── data/\n│   ├── scenarios.json   # the benchmark corpus — 33 current scenarios\n│   ├── models.txt       # model manifest (bracket, tag, quant)\n│   └── MODEL-PROMPTS.md # byte-frozen prompt text\n├── docs/\n│   ├── PAPER.md         # experimental design spec\n│   ├── EXPERIMENT-PIPELINE.md # operational pipeline contract\n│   ├── TAXONOMY.md      # task-class taxonomy\n│   ├── TELEMETRY.md     # telemetry data dictionary\n│   ├── MODELS.md        # vetted model list\n│   └── MARKET.md        # adversarial analysis\n└── scripts/\n    ├── node-power.sh    # reproducible power state: setup / teardown / status\n    └── run-experiment.sh # autonomous multi-stage experiment driver\n```\n\n---\n\n## Honest limitations\n\n**1. Judge egress.** We use Claude 4.8 off-node to *score* answers. The system-under-test never calls it. But the judge sees the scenario text, which contains real cluster detail: namespace names, Azure Key Vault references, Cloudflare DNS, `*.home.domain`. This is real ops data sent to a third party. Released scenarios are scrubbed and anonymised. This egress must be disclosed in any publication. See the public-service dependency map in [`docs/PAPER.md`](docs/PAPER.md).\n\n**2. Grounded = oracle retrieval upper bound.** We inject the correct reference text directly into context. A real local-RAG pipeline adds retrieval error, chunking artifacts, and embedding drift. Our grounded numbers are the *ceiling* of what local retrieval can buy, not the expected value in a deployed system.\n\n**3. Telemetry is Linux-specific.** Energy (RAPL), RAM/swap (`/proc`), and memory-bandwidth counters require Linux. The harness runs on macOS/Windows — quality scores reproduce; the systems telemetry will be empty. This is a documented limitation, not a bug.\n\n**4. CPU-only inference.** The benchmark characterises inference on the *worst reasonable hardware*. On a machine with a discrete GPU or Apple Silicon, tok/s numbers will be higher and thermal behaviour different. The quality scores should generalise; the systems numbers will not.\n\n**5. Single hardware point.** All systems measurements are from one specific node (i5-8350U, 24 GiB DDR4-2400). Hardware interaction effects may differ on different CPU generations, memory configurations, or NVMe speeds. We disclose the full environment in `ENVIRONMENT.md`.\n\n---\n\n## The AIOps maturity ladder\n\nThe broader motivation: where does a local small model sit on the path toward autonomous operations?\n\n```\n5 · Autonomous   — closed-loop self-healing, no human required\n4 · Preventive   — acts to prevent known failure modes before they occur  \n3 · Predictive   — forecasts failures from time-series signals\n2 · Proactive    — acts ahead of user request, surface-triggered\n1 · Reactive     — responds to incidents that have already occurred\n    ↑\n    This paper measures the quality and safety of the reasoning FOUNDATION\n    at rung 1 (reactive) and the early boundary of rung 2 (proactive/foresee).\n    Claiming higher rungs from these results would be overreach.\n```\n\nA model that can reliably detect, diagnose, and safely refuse on rung 1 has earned the right to be *considered* for higher-trust work. This benchmark provides the evidence base for that judgment — and makes the evidence falsifiable and reproducible.\n\n---\n\n## For reviewers\n\nThis repo is built to be **easy to review — including with AI assistance — with a\nhuman in charge of the judgement.** If you were invited to review the paper, start\nwith **[`REVIEWER.md`](REVIEWER.md)**: it maps your assessment onto the NeurIPS\nreview dimensions (quality / clarity / significance / originality), tells you which\nnumbers reproduce on any laptop vs. which need the specific node, and covers\nconfidentiality etiquette for AI-assisted review.\n\n\u003e **arXiv is moderated, not peer-reviewed.** The first release is an arXiv preprint\n\u003e (a moderation check on scholarly standards and format — *not* peer review); the\n\u003e intended peer-reviewed venue is the **NeurIPS Datasets \u0026 Benchmarks track**.\n\n## Use of AI in this work\n\nThis benchmark, its analysis, and its prose were produced with **substantial AI\nassistance under human direction**. A human author directs the work and **takes full\nresponsibility for every claim, number, and line of code**, regardless of how it was\ngenerated; no AI system is listed as an author. This follows arXiv's policy on\nauthors' use of generative-AI language tools. Every headline number is reproducible\nfrom released artifacts ([`REPRODUCE.md`](REPRODUCE.md)), so the work can be checked\nindependently of the prose.\n\n## Standards\n\n- **Telemetry schema**: [OpenTelemetry GenAI semantic conventions](https://github.com/open-telemetry/semantic-conventions-genai) — `gen_ai.*` spans, metrics, events.\n- **Judge pattern**: MT-Bench / AlpacaEval frontier-as-judge; two-family ensemble (Copilot + OpenAI) with Cohen's κ agreement.\n- **Stats**: bootstrap 95% CIs; Friedman test for within-subject bracket comparison; Wilcoxon signed-rank for pairwise.\n- **Baseline provenance**: [EleutherAI lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness) — the open standard for repeatable LLM evaluation.\n\n---\n\n## Citation\n\n```bibtex\n@misc{hont2026apprenticeops,\n  title   = {ApprenticeOps: Evaluating Small Locally-Sovereign LLMs as\n             Homelab Operations Assistants},\n  author  = {Hont, Dragos},\n  year    = {2026},\n  url     = {https://github.com/dragoshont/apprenticeops},\n  note    = {Open benchmark and reproducible study. Apache 2.0.}\n}\n```\n\nUpdate with venue and DOI after submission.\n\n## License\n\nApache 2.0. See [`LICENSE`](LICENSE).\n\n---\n\n**Start here:** [`docs/PAPER.md`](docs/PAPER.md) for the research design, [`REPRODUCE.md`](REPRODUCE.md) to reproduce the results.\n\n**Contribute:** issues and PRs welcome — especially new scenarios, additional models, and hardware configurations.\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdragoshont%2Fapprenticeops","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fdragoshont%2Fapprenticeops","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdragoshont%2Fapprenticeops/lists"}