awesome-cpu-first-ai
Curated, evidence-backed list of runtimes, formats & tools for running AI inference on CPU — start with CPU, justify the GPU.
https://github.com/ranjithrajv/awesome-cpu-first-ai
Last synced: 9 days ago
JSON representation
-
Benchmarks and Evidence
- llama.cpp performance tracking - Community-maintained thread with tokens/second figures for various models across CPU and GPU hardware; useful as a real-world comparison baseline.
- LLM Inference Benchmarking Cheat-Sheet (llm-tracker.info) - Canonical reference explaining llama.cpp benchmark metrics (pp512/tg128), quantization naming conventions, and how to correctly interpret and compare community-reported figures across hardware platforms.
- MyAIHardware — llama.cpp benchmarks - Aggregated llama.cpp benchmark scoreboard across CPUs, GPUs, and NPUs under standardized test conditions; useful for hardware selection and cross-platform throughput comparison.
- MLPerf Inference — edge CPU submissions - Industry-audited inference benchmark with CPU-only submissions in the edge category; provides verified latency/throughput figures under defined, reproducible test conditions.
- MLPerf Inference v5.0 — datacenter CPU submissions (MLCommons, Apr 2025) - Industry-audited inference benchmark with CPU-only datacenter submissions on Intel Xeon 6 Granite Rapids; reports GPT-J at 316 tok/s (INT4), Llama-3.1-8B at 450 tok/s (server) and 1,196 tok/s (offline). Intel remains the only vendor submitting server CPU results, holding through the v5.1 round (Sept 2025). *(last verified: 2026-07)* ([Dell 2S-GNR results](https://github.com/mlcommons/inference_results_v5.0/tree/main/closed/Dell/results/1-node-2S-GNR_86C), [Supermicro results](https://github.com/mlcommons/inference_results_v5.0/tree/main/closed/Supermicro/results/1-node-2S-GNR_128C))
- ONNX Runtime GenAI CPU benchmark (ISE Developer Blog, May 2025) - Production benchmark comparing ONNX Runtime GenAI, llama.cpp, and Hugging Face Optimum for Phi-3 on CPU; ONNX Runtime achieved 137.6 tok/s vs 109.5 for llama.cpp and 108.3 for Optimum — 1.2–1.6× higher throughput across prompt lengths with equivalent latency.
- OpenVINO Model Hub benchmarks (Intel, 2025) - Intel's centralized benchmark catalog for LLMs on Xeon CPUs with INT4/INT8/FP16 precision; reports DeepSeek-R1-Distill-Llama-8B at 155.4 tok/s and Llama-3-8B at 376 tok/s (OpenVINO Model Server) on Intel Xeon Platinum. ([OpenVINO LLM benchmark tool](https://github.com/openvinotoolkit/openvino.genai/tree/master/tools/llm_bench), [white paper](https://www.intel.com/content/dam/develop/public/us/en/documents/llm-with-model-server-white-paper.pdf))
- Simon Willison's llama.cpp experiments - Practitioner write-ups with real-world timing data across diverse CPU hardware; useful for calibrating expectations before purchasing cloud instances.
- Model size distribution on Hugging Face - Independent analyses of the Hugging Face ecosystem find 40–51% of models are sub-7B and 55–65% are sub-13B parameters, with 92% of all downloads going to models under 1B params and the median downloaded model at 406M params. ([MoClaw, Apr 2026](https://moclaw.ai/blog/huggingface-hub-state-2026), [HF model stats, Oct 2025](https://huggingface.co/blog/lbourdois/huggingface-models-stats))
- MyAIHardware — llama.cpp benchmarks - Aggregated llama.cpp benchmark scoreboard across CPUs, GPUs, and NPUs under standardized test conditions; useful for hardware selection and cross-platform throughput comparison.
-
Cloud ARM Servers
-
Summary Dashboard
- OCI Ampere Altra A1 instances - Oracle Cloud shapes based on Ampere Altra (Neoverse N1); benchmarked at 119 tok/s aggregate throughput for Llama-2 7B with 16 concurrent users using an optimized llama.cpp stack, with up to 152% improvement over upstream llama.cpp reported. *(last verified: 2026-07)*
- Azure Cobalt 100 (Neoverse N2) - Microsoft's 128-core Neoverse N2 processor; Arm-optimized ONNX Runtime (KleidiAI kernels) delivers 1.9× higher token-generation throughput and 2.8× better price/performance compared to AMD Genoa-based instances for LLM inference. *(last verified: 2026-07)*
- Azure Cobalt 200 (Neoverse V3) - Microsoft's 132-core Neoverse V3 processor on TSMC 3 nm; delivers up to 50% better CPU performance over Cobalt 100 and is positioned explicitly for agentic AI inference workloads; early-access VMs available as of Build 2026. *(last verified: 2026-06)*
- AWS Graviton4 — c8g instances - Amazon's Neoverse V2-based fourth-generation Graviton; up to 30% better performance and up to 3× more vCPUs than Graviton3 (c7g) at the largest sizes; llama.cpp MMLA kernels are supported and distributed multi-node inference is documented in the [Arm Learning Paths guide](https://learn.arm.com/learning-paths/servers-and-cloud-computing/distributed-inference-with-llama-cpp/). *(last verified: 2026-06)*
- Google Axion (Neoverse V2) - Google Cloud's custom Neoverse V2 processor (C4A instances); benchmarked with llama.cpp on Llama-3.1 8B and reports up to 2× better prompt-processing and token-generation performance vs current-generation x86 instances. *(last verified: 2026-07)*
- aarch64.cloud — Graviton vs Axion vs Cobalt benchmark - Independent benchmark comparing AWS Graviton3, Google Axion, and Azure Cobalt 100 on llama.cpp with Llama-3.1 8B and Llama-3.2 1B; documents tokens/s and price/performance ratios. *(Note: predates Graviton4 and Cobalt 200; use as a Graviton3-generation baseline.)*
-
-
Contributing
-
Project & community
-
-
Cost and Deployment Economics
-
Summary Dashboard
- AWS Lambda pricing - Serverless compute priced per GB-second on CPU; viable for low-throughput embedding and small-model inference without a persistent GPU instance.
- Hetzner dedicated servers - Example of high-core-count x86 servers at commodity pricing; a useful reference point when constructing cost-per-token calculations to compare against GPU instances.
- Fly.io CPU machines - Container-level CPU VMs with per-second billing; commonly used for llama.cpp-backed inference APIs at low traffic volumes where a persistent GPU instance would be idle most of the time.
- Modal — GPU selection guide - Serverless platform that makes the CPU/GPU choice explicit at the function level; useful for hybrid deployments where embeddings run CPU-side and generation runs GPU-side.
-
-
CPU AI Gap Map
- AI Potluck Open Source AI Gap Map - nativeness** — the degree to which a tool is designed for CPU inference rather than treating it as a secondary fallback.
-
CPU Fine-Tuning
-
Summary Dashboard
- Unsloth - GPU-accelerated LoRA/QLoRA fine-tuning library; included here because it produces GGUF-compatible LoRA adapters that can be merged and deployed on CPU via llama.cpp. Fine-tune on GPU (or free Colab T4), export adapter, run inference on CPU.
- llama.cpp fine-tuning - Built-in fine-tuning example in llama.cpp supporting LoRA-style adapter training on CPU; produces `.lora` files loadable by `llama-cli` at inference time. Suited for small-scale domain adaptation (classification heads, instruction tuning on 1K–10K examples). Runs entirely on CPU with no GPU dependency at any stage.
- LoRAX - Multi-LoRA inference server that serves thousands of fine-tuned adapters from a single base model; designed for GPU by default but the LoRA-weight merging pattern applies to CPU deployments — pre-merge adapters into a single GGUF with `llama.cpp`'s `export-lora` for CPU serving.
- PEFT - Hugging Face's parameter-efficient fine-tuning library (LoRA, IA³, Prefix Tuning, AdaLoRA); runs on CPU for small models when `device="cpu"` is set, though throughput is 10–50× slower than a single GPU. Practical for models ≤ 3B where training data is small (< 5K examples) and iteration time is not critical.
- LlamaFactory - Unified fine-tuning framework supporting LoRA, QLoRA, full-parameter, and DoRA; CPU mode works for small-scale adapter training (batch size 1–2, ≤ 3B base models) with `CUDA_VISIBLE_DEVICES=""` to force CPU execution.
- Axolotl - Flexible fine-tuning toolkit supporting QLoRA and multi-GPU training; provides CPU-offloaded optimizer states via `deepspeed` ZeRO-3 CPU offload, keeping activations on GPU while optimizer states reside in system RAM — a hybrid approach that reduces GPU VRAM requirements for larger models.
-
-
Mixture-of-Experts on CPU
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (DeepSeek, Jan 2025) - Introduces DeepSeek-R1 (671B total, 37B activated per token) and distilled dense variants from 1.5B to 70B; the dense distillations run on any CPU with llama.cpp at Q4, and the full MoE model with IQ1_S quantization fits within ~10 GB RAM on CPU.
- Deploy DeepSeek-R1 on Arm Servers with llama.cpp (Arm Learning Paths, Apr 2026) - Walkthrough for DeepSeek-R1-Distill-Qwen-7B Q4_K_M on AWS Graviton4; benchmarks 18–22 tok/s generation and ~420 tok/s prompt processing on 24 vCPU, 192 GB RAM with ~5.8 GB model RAM use.
- DeepSeek-R1 7B on OCI Ampere A1: Full CPU Inference Guide (asknikhil.com, May 2026) - Practitioner guide deploying DeepSeek-R1-Distill-Qwen-7B Q4_K_M on OCI Ampere A1 free-tier ARM instances; reports ~18–22 tok/s generation, ~420 tok/s prompt processing, and ~5.8 GB RAM utilisation with no CUDA/driver setup required.
- BigMoeOnEdge - On-device MoE engine built on llama.cpp's public API that streams only the experts each token routes to from flash storage, running models up to 7× larger than RAM on a phone's CPU (DeepSeek V4 Flash 0731 at ~91 GB on a 12 GB phone, ~1 tok/s) with byte-identical output and a capped expert cache. Apache 2.0. *(Also listed under On-Device, Edge, ARM, and SBCs.)*
-
Mobile Phone CPUs
-
Summary Dashboard
- Bonsai 27B - class model on an iPhone 17 Pro CPU at ~11 tok/s — see [Ternary / 1-bit models](#quantization-and-model-formats). *(vendor-reported; last verified: 2026-07)*
- Apple A19 Pro - core ANE, ~35 TOPS), 🧠 [Snapdragon 8 Elite](https://www.qualcomm.com/products/mobile/snapdragon/smartphones/snapdragon-8-series-mobile-platforms/snapdragon-8-elite) (Hexagon NPU, ~60 TOPS), 🧠 [Exynos 2500](https://semiconductor.samsung.com/processor/mobile-processor/exynos-2500/) (59 TOPS NPU), 🧠 [Dimensity 9500](https://i.mediatek.com/mediatek-dimensity-ai) (NPU 890, ~50 TOPS), 🧠 [Tensor G5](https://store.google.com/pixel_10) (TPU, NPU path experimental).
- PocketPal AI - source (MIT, 7.6k★) cross-platform chat app running GGUF models on CPU/GPU/NPU with on-device TTS, tool use, and Hugging Face integration; [PocketLFM](https://github.com/Jeevav62/pocketlfm) — Android app for Liquid AI's LFM2.5 models on CPU via llama.cpp; [PrivateFoundationModels](https://github.com/john-rocky/PrivateFoundationModels) — Swift package unifying Apple FoundationModels, Core ML, and MLX for on-device LLMs on iOS/macOS.
- Off Grid AI (OGAM) - licensed cross-platform offline AI suite for Android, iOS, and macOS (React Native) running GGUF LLMs via llama.cpp on CPU with optional OpenCL/Metal GPU and experimental Snapdragon NPU acceleration, plus on-device vision (SmolVLM/Qwen3-VL), Whisper STT, Stable Diffusion image generation, tool calling, MCP, and local-network OpenAI-compatible servers; fully offline with per-model RAM management; [PocketPal AI](https://github.com/a-ghorbani/pocketpal-ai) — open-source (MIT, 7.6k★) cross-platform chat app running GGUF models on CPU/GPU/NPU with on-device TTS, tool use, and Hugging Face integration; [PocketLFM](https://github.com/Jeevav62/pocketlfm) — Android app for Liquid AI's LFM2.5 models on CPU via llama.cpp; [PrivateFoundationModels](https://github.com/john-rocky/PrivateFoundationModels) — Swift package unifying Apple FoundationModels, Core ML, and MLX for on-device LLMs on iOS/macOS.
-
-
Model Selection and Hardware Fit
-
Summary Dashboard
- llmfit - Terminal tool (MIT) that detects your RAM, CPU, GPU, and backend, then ranks hundreds of models with a single 0–100 Fit score across memory fit, estimated speed, quality, and context length. CPU-aware rather than VRAM-only, launches matched models directly through Ollama or llama.cpp, and includes a simulation mode (`S`) to override RAM/VRAM/core count for upgrade planning.
- whichllm - CLI (MIT) that auto-detects CPU, RAM, and GPU and ranks the Hugging Face models that actually run on your machine — ordered by real, recency-aware benchmarks rather than parameter count, in one command with no project setup.
- Local AI Master — Model Recommender - Browser-based recommender: pick a task (chat, coding, reasoning, RAG, vision, audio) and your available RAM/VRAM to get ranked models with quality scores, Q4 memory requirements, tokens/second estimates, and one-line install commands; explicitly covers CPU-only and Apple Silicon unified-memory cases.
-
-
Multimodal CPU Workloads
-
Summary Dashboard
- faster-whisper - computer/transcribe.cpp), [Piper](https://github.com/OHF-Voice/piper1-gpl), [PocketTTS](https://github.com/kyutai-labs/pocket-tts), [PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR), [Tesseract](https://github.com/tesseract-ocr/tesseract), [MobileSAM](https://github.com/ChaoningZhang/MobileSAM), [rembg](https://github.com/danielgatis/rembg), [InsightFace](https://github.com/deepinsight/insightface).
-
-
On-Device, Edge, ARM, and SBCs
-
Summary Dashboard
- Intel Core Ultra NPU + OpenVINO - Intel Meteor Lake/Arrow Lake/Lunar Lake CPUs include a dedicated NPU for on-device AI inference. OpenVINO's NPU plugin deploys INT8/FP16 models compiled for the Intel AI engine via the standard OpenVINO API, with automatic CPU fallback. Phi-2 INT4 runs entirely on the NPU at competitive latencies. ([NPU device docs](https://docs.openvino.ai/2024/openvino-workflow/running-inference/inference-devices-and-modes/npu-device.html))
- XNNPACK - Google's accelerated neural network inference library for ARM and x86; the shared CPU kernel backend behind [TFLite](https://www.tensorflow.org/lite) / [LiteRT](https://github.com/google-ai-edge/LiteRT) (including [LiteRT.js](https://ai.google.dev/edge/litert/web) WebAssembly), [ExecuTorch](https://github.com/pytorch/executorch), and ONNX Runtime's mobile path — hand-tuned for NEON, SSE/AVX, and WASM SIMD.
- TensorFlow Lite - Google's inference runtime for mobile and embedded; the default execution path is CPU (ARM/x86), with delegate APIs for optional hardware accelerators. *(Note: [LiteRT](https://github.com/google-ai-edge/LiteRT) is the official successor to TensorFlow Lite, keeping the same `.tflite` format and XNNPACK CPU backend; the browser target [LiteRT.js](https://ai.google.dev/edge/litert/web) is listed under Runtimes and Inference Engines.)*
- MLC LLM (WebAssembly/CPU target) - Compiles LLMs to native CPU code or WebAssembly via TVM; the browser/WebAssembly target is inherently CPU-only. *(Note: also targets GPU; relevant here specifically for its WebAssembly/CPU compilation path.)*
- llama.cpp Android build - Official docs for cross-compiling llama.cpp for Android ARM; runs on-device without network access or cloud inference costs.
- V-Seek — LLM inference on RISC-V server CPUs (arxiv:2503.17422) - Paper documenting LLM inference optimizations on the Sophon SG2042, the first commercially available many-core RISC-V server CPU (64 RVV-capable cores); achieves 13 tok/s for 7B models and 5.5× throughput over baseline llama.cpp by exploiting RISC-V Vector (RVV) extensions with vectorized GEMM kernels.
- MediaTek Genio 720 / 520 (MediaTek, 2025) - Edge AI IoT platforms (6 nm) with octa-core Arm CPU (2× Cortex-A78 + 6× Cortex-A55) and 10 TOPS NPU; supports LLMs (Llama, Phi, DeepSeek) on-device via LiteRT and ONNX Runtime with CPU fallback. ([Genio AI Developer Guide](https://genio.mediatek.com/doc/iot-yocto/latest/sw/yocto/iot-ai-hub.html))
- turbo-fieldfare - Model-specific Swift + Metal runtime that runs Gemma 4 26B-A4B (26B total, ~3.9B active per token) in ~2 GB of RAM on any Apple Silicon Mac, even 8 GB ones, by keeping the shared core and KV cache resident and streaming only the routed experts needed per token from SSD. Apache 2.0. *(Note: executes on Apple's Metal GPU, not the CPU; included for its expert-streaming memory-efficiency approach on on-device ARM hardware.)*
- mlx-od-moe - On-demand MoE inference for Apple Silicon that memory-maps experts as `.npy` files on NVMe, keeps an LRU cache of hot experts in RAM, and uses a lightweight shadow model to predict and asynchronously prefetch the next top-K experts; runs 375 GB models (e.g. Kimi-K2.5) in 192 GB RAM. MIT. *(Note: executes on Apple's Metal/MLX stack, not the CPU; included for its on-demand expert-loading and learned-prefetch approach on on-device ARM hardware.)*
- moe-stream - SSD-streaming MoE inference engine (Rust + Metal) that runs 80B-parameter models on a 24 GB Apple Silicon Mac, auto-selecting between GPU-resident, GPU-hybrid, and SSD-streaming modes, with fused MXFP4 matvec kernels and OpenAI-compatible + MCP servers. Apache 2.0. *(Note: executes on Apple's Metal GPU, not the CPU; included for its SSD-streaming memory-efficiency approach on on-device ARM hardware.)*
- s-moe - Model-agnostic inference engine for fine-grained MoE LLMs on Apple Silicon that streams experts from NVMe with a ring-buffer LRU and a Direct I/O prefetch pump, running Qwen3-235B on a standard 48 GB MacBook. MIT. *(Note: executes on Apple's Metal GPU, not the CPU; included for its NVMe-streaming memory-efficiency approach on on-device ARM hardware.)*
- sparsify - MLX-based runtime that treats the SSD as a first-class memory tier, paging router-selected experts into a bounded RAM cache with byte-identical output; Mixtral 8x7B (26.3 GB stored) runs in 3.33 GB RSS on a 16 GB MacBook Air. MIT. *(Note: executes on Apple's Metal/MLX stack, not the CPU; included for its expert-paging memory-efficiency approach on on-device ARM hardware.)*
- IREE Compiler - V edge inference through its LLVM-based CPU backend, incorporating MLIR-native microkernels (ukernels) optimized for RVV to handle heavy operations like matrix multiplication efficiently.
- RunAnywhere SDKs - Production SDK toolkit (Android, iOS, React Native, Flutter, web, C++, server) over one C++ core that runs LLM chat, VLM, speech-to-text, TTS, voice agents, embeddings, RAG, and image generation locally on phones, browsers, desktops, and servers; offline by design, with a capability registry routing each call to the best engine the device actually has (llama.cpp-based LLM path). *(License: non-OSI source-available — free for individuals, small orgs under $1M funding, education, and nonprofits; commercial license required beyond that.)*
- cactus - Low-latency hybrid edge-cloud AI engine for mobile devices and wearables (Apple, Samsung, Pixel SoCs) with CPU/GPU kernels, a custom rotation-based quantization scheme, zero-copy computation graph, built-in RAG, and OpenAI-compatible APIs for text, speech, and vision; brew-installable with a local-run mode and cloud handoff. *(License: non-OSI source-available — free for individuals, small orgs under $2M funding/revenue, education, and nonprofits; commercial license required beyond that.)*
-
-
Performance Tuning
- llama.cpp token generation performance tips - Official guidance on setting `--threads`, `--threads-batch`, CPU affinity masks, and NUMA-aware memory allocation for multi-socket servers.
- OpenBLAS - Optimized BLAS implementation with auto-tuned kernels for x86 (SSE/AVX/AVX-512) and ARM; a drop-in dependency for frameworks that delegate GEMM to BLAS.
- Intel MKL / oneMKL - Intel's math kernel library with AVX-512 and AMX-optimized GEMM; free to use and typically the fastest BLAS on recent Xeon hardware.
- Intel AMX (Advanced Matrix Extensions) - Hardware matrix multiplication tiles in Sapphire Rapids and later Xeon CPUs; AMX delivers 2,048 INT8 operations per cycle vs 256 for AVX-512 VNNI — an 8× arithmetic throughput improvement for quantized inference on the same silicon; llama.cpp and ONNX Runtime both expose AMX code paths. [(Intel AMX solution brief)](https://www.intel.com/content/dam/www/central-libraries/us/en/documents/2022-12/accelerate-ai-with-amx-sb.pdf)
- numactl - Linux utility to bind a process to specific NUMA nodes and CPU cores; essential for avoiding cross-socket memory latency on multi-socket inference servers.
- perf + Linux PMU - Standard Linux profiling tool; useful for measuring LLC miss rates and memory bandwidth saturation during inference, which are the dominant bottlenecks on CPU.
- likwid - Hardware performance counter tool suite for x86; provides memory bandwidth and FLOP/s measurements useful for diagnosing inference throughput limits on specific µarchs.
-
Quantization and Model Formats
- Ternary / 1-bit models (BitNet b1.58 lineage) - Weights constrained to {−1, 0, +1} (~1.58 bits) or binary {−1, +1}, turning weight multiplications into additions — a natural fit for CPU, where memory bandwidth rather than FLOPs is the bottleneck. Served via [bitnet.cpp](https://github.com/microsoft/BitNet) or llama.cpp. The first open-source native 1-bit LLM is [BitNet b1.58 2B4T](https://huggingface.co/microsoft/bitnet-b1.58-2B-4T) (Microsoft, Apr 2025, MIT), a 2B-parameter model trained from scratch with ternary weights on 4T tokens — its non-embedding weights use only ~0.4 GB. A recent multimodal flagship is [Bonsai 27B (PrismML, Jul 2026)](https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf), a 27B model in 1-bit (~3.9 GB) and 1.58-bit ternary (~5.9 GB) GGUF form running on llama.cpp CPU builds under Apache 2.0 (~11 tok/s on iPhone 17 Pro CPU). See [docs/cpu-native-models.md](docs/cpu-native-models.md) for the full ternary model catalog, a cost-advantage analysis of ≤3B/ternary models, and a coming-soon tracker for in-development ternary projects (BitNet v2, larger native ternary models, NPU backends, TWLA). *(vendor-reported benchmarks; last verified: 2026-07)*
- Intel Neural Compressor - Framework-agnostic post-training quantization and pruning toolkit targeting CPU inference; supports ONNX, PyTorch, and TensorFlow backends.
- Optimum - Hugging Face's optimization toolkit; the `optimum[onnxruntime]` and `optimum-intel` paths export and quantize models for CPU inference via ONNX Runtime and OpenVINO respectively.
- AutoGPTQ - GPTQ quantization library. *(Caveat: primarily targets GPU inference; include only when the produced GPTQ checkpoints are subsequently converted to GGUF for CPU use. Do not assume CPU parity.)*
- llama.cpp quantize tool - Built-in `llama-quantize` binary converting Hugging Face checkpoints to GGUF; covers k-quants (Q4_K_M, Q5_K_M, Q6_K) and importance-matrix–guided i-quants (IQ3_XS, IQ4_XS) that route more bits to high-impact weights — IQ4_XS saves ~400 MB vs Q4_K_M on a 7B model at comparable accuracy. Pair with `--imatrix` for any format below Q5_K_M.
- llama.cpp imatrix tool - Calibration pass that runs a small corpus through the unquantized model and records per-layer weight importance; the resulting `.imatrix` file is passed to `llama-quantize` and significantly improves output quality at aggressive compression ratios (IQ3_XS, IQ4_XS, Q3_K_S).
-
Runtimes and Inference Engines
- ArcLight - Many-core CPU inference framework designed for NUMA systems with tensor parallelism across multiple CPU sockets; claims 46% higher throughput than llama.cpp through optimized NUMA-aware scheduling and parallel decomposition on multi-socket x86 servers. MIT. ([arXiv:2603.07770](https://arxiv.org/abs/2603.07770)) *(vendor-reported benchmarks; last verified: 2026-07)*
- bitnet.cpp - Microsoft's official inference framework for 1-bit and 1.58-bit ternary LLMs; ternary weights ({−1, 0, +1}) replace multiplications with additions, and its CPU-optimized kernels (x86 and ARM) report 2.4–6.2× speedup with 71.9–82.2% energy reduction on x86 (1.37–5.07× / 55.4–70.0% on ARM) versus llama.cpp, running a 100B b1.58 model on a single CPU at 5–7 tok/s. *(CPU and GPU kernels; NPU support planned.)*
- candle - Hugging Face's Rust ML framework; CPU execution is the primary target, with optional CUDA support compiled in separately.
- colibri - Pure C inference engine for large MoE models (GLM-5.2 744B) that treats VRAM/RAM/disk as one memory hierarchy, streaming routed experts from disk on demand; runs on 25 GB RAM with zero dependencies and no GPU required. Apache 2.0.
- ctransformers - Python bindings for GGUF models; lets Python callers run quantized models on CPU without touching C++.
- cpubrrr - From-scratch NEON and SME SIMD kernels written in Rust for MoE model inference on Apple M4; claims 110 tok/s for gpt-oss:20b (7.5× llama.cpp) with speculative decoding and hardware-aware scheduling. Apache 2.0. *(vendor-reported benchmarks; last verified: 2026-07)*
- distributed-llama - Tensor-parallel inference that splits a model's compute and RAM across a cluster of ARM/x86 AVX2 CPU nodes (power-of-2 node counts), letting commodity or Raspberry Pi devices jointly run models too large for a single machine; MIT-licensed and actively maintained, with experimental Vulkan GPU support. <!-- TODO: cite CPU throughput benchmark -->
- eLLM - Rust-based LLM inference framework built specifically for CPU servers (Intel Xeon 4th Gen+ with AMX). Uses a static computation graph with dimension-first tensor layout and head-by-head attention to reduce runtime scheduling overhead; targets long-context prefill-heavy workloads where the CPU's larger memory capacity lets it compete with multi-GPU systems. Reports ~1.6× decode speedup vs SGLang CPU baseline (random-parameter benchmarks; alpha). Apache 2.0. <!-- TODO: verify with real model weights once available -->
- Ferrite - CPU-native Rust inference engine with pure Rust SIMD kernels for x86 (AVX2/AVX-512) and ARM (NEON); no GPU code paths. Supports GGUF models and targets CPU-only deployment; reports ~85 tok/s for TinyLlama-1.1B Q4 on an i7-13700K. Apache 2.0. *(vendor-reported benchmarks; last verified: 2026-07)*
- ExecuTorch - PyTorch's on-device inference runtime; designed for mobile and embedded, with CPU kernels for ARM (XNNPACK backend) as the primary deployment target and cross-compilation support for Android, iOS, and Linux. *(Also listed under On-Device, Edge, ARM, and SBCs.)*
- ggml - The tensor library underlying llama.cpp; hand-optimized CPU kernels using SIMD intrinsics for AVX2, AVX-512, NEON, and SVE.
- Intel Extension for Transformers - Drop-in optimization layer for Hugging Face Transformers that applies CPU-specific INT4/INT8 kernels, AMX acceleration, and weight-only quantization.
- KTransformers - CPU-GPU heterogeneous LLM inference framework designed for large MoE models; Intel AMX/AVX-512/AVX2-optimized CPU kernels for INT4/INT8 quantized inference with NUMA-aware expert scheduling and SGLang integration. Apache 2.0. *(Also listed under Mixture-of-Experts on CPU.)*
- LiteRT.js - Google AI Edge's in-browser ML runtime; the web target of [LiteRT](https://github.com/google-ai-edge/LiteRT) (the successor to TensorFlow Lite), running `.tflite` models via WebAssembly (CPU, XNNPACK backend) or WebGPU with zero server dependency, supporting INT8 quantized models on the CPU path.
- llamafile - Distributable single-file LLM executables (built on llama.cpp + Cosmopolitan libc) that run on CPU across Linux, macOS, and Windows with no install.
- llama.cpp - C/C++ LLM inference engine designed from day one for CPU; optional GPU offload of individual layers rather than GPU-first design.
- llama2.c - Andrej Karpathy's minimal C implementation of LLaMA 2 inference; a pedagogical reference showing that CPU inference requires no ML framework, only a few hundred lines of C.
- MNN - Alibaba's inference engine for mobile and edge; CPU is the primary target, with quantization-aware kernels for ARM NEON and x86 SSE/AVX.
- MojoLlama - High-throughput CPU inference engine built on Modular MAX with a pure Mojo backend; optimized for Mixture-of-Experts architectures and claims 1.3× throughput over llama.cpp on CPU with INT4 quantization. Apache 2.0. *(vendor-reported benchmarks; last verified: 2026-07)*
- ncnn - Mobile and embedded neural network inference framework optimized for ARM and x86 CPUs; no dependencies, builds for Raspberry Pi, Jetson (CPU-only mode), and RISC-V with no OS-level GPU driver requirement.
- Qualcomm AI Engine Direct (QNN) - Qualcomm's NPU SDK for Snapdragon platforms; runs INT8/INT4 models on the Hexagon NPU with CPU fallback. Achieves 12–25 tok/s for 7B INT4 on Snapdragon 8 Elite. Snapdragon-only; CPU fallback for cross-device. ([Qualcomm AI Hub](https://aihub.qualcomm.com/))
- Apple Core ML (ANE path) - Apple's on-device ML framework with explicit Neural Engine deployment path; models compiled via `coremltools` target the ANE for vision and small LLMs. CPU/GPU fallback for unsupported ops. ([Core ML ANE deployment](https://developer.apple.com/videos/play/wwdc2024/10160))
- Intel OpenVINO NPU plugin - OpenVINO's NPU execution provider targeting Intel NPU (Meteor Lake, Arrow Lake, Lunar Lake). Supports INT8/FP16 models compiled for the Intel AI engine; integrated into the standard OpenVINO API with automatic CPU fallback for unsupported operations.
- AMD Ryzen AI (Vitis AI / XDNA) - AMD's NPU inference stack for Ryzen AI PC processors; deploys ONNX models via Vitis AI Execution Provider on the XDNA NPU. Supports INT4/INT8/BF16 with automatic CPU fallback for unsupported operators. ([ONNX Runtime VitisAI EP](https://onnxruntime.ai/docs/execution-providers/Vitis-AI-ExecutionProvider.html))
- ONNX Runtime (CPU EP) - The CPU Execution Provider in ONNX Runtime; production-grade, supports operator fusion and quantized INT8 models natively on x86 and ARM.
- ollama - Local model runner that falls back to full CPU execution when no GPU is present; convenient for development and low-traffic deployments. *(Note: GPU is used when available; included here for its CPU fallback path and single-binary packaging story.)*
- OpenVINO - Intel's model optimization and inference toolkit; targets x86 CPU as first-class hardware with graph optimization passes specific to Intel µarchs.
- Project Zero - Pure C inference engine for BitNet-style 1.58-bit models requiring no SIMD extensions; reports ~1,000 tok/s for 1.58-bit models and 1.56× speedup over llama.cpp on INT1 workloads. MIT. *(vendor-reported benchmarks; last verified: 2026-07)*
- rwkv.cpp - CPU inference library for RWKV v4–v7 language models (INT4/INT5/INT8 and FP16); RWKV's recurrent state requires O(1) memory per token at inference time with no growing KV cache, making it especially suited to CPU inference under long context lengths where transformer KV-cache memory becomes prohibitive.
- Reame - CPU-first inference server built on llama.cpp with disk-backed KV cache and self-regulating speculative decoding; targets CPU-only deployments and reports 25–40 tok/s on a 7B Q4 model using 18 GB RAM on a laptop. MIT. *(vendor-reported benchmarks; last verified: 2026-07)*
- Transformers.js - Hugging Face's in-browser transformer inference library; runs ONNX models via WebAssembly (CPU) or WebGPU, supporting 200+ architectures across NLP, vision, and audio with zero server dependency. CPU execution uses ONNX Runtime Web's WebAssembly backend with INT8 quantization.
- WebLLM - In-browser LLM inference engine built on MLC LLM and Apache TVM; uses WebGPU when available and falls back to WebAssembly for CPU-only execution, delivering an OpenAI-compatible API callable from browser JavaScript with no server required.
- whisper.cpp - Port of OpenAI Whisper to ggml; runs speech-to-text inference entirely on CPU with explicit ARM NEON and AVX paths; on Raspberry Pi 5, the base model achieves 3–5× real-time throughput and the JFK benchmark completes in approximately 9 s with the float32 tiny model.
- hummingbird - Zero-dependency C17 runtime that unifies SSD, RAM, and VRAM into a single memory hierarchy, streaming expert weights from disk on demand (with io_uring-backed async reads) for large MoE models such as GPT-OSS 120B, GLM, DeepSeek, and Qwen; model-agnostic through a flexible adapter interface rather than per-model engines. Apache 2.0 (placeholder LICENSE). *(Also listed under Mixture-of-Experts on CPU.)*
- kimi-k3-in-c - Portable C99 inference engine that runs the 2.78T-parameter Kimi K3 MoE model on a single CPU in 8.24 GB RAM by streaming the 1.45 TB of routed experts directly out of their packed mxfp4 form; no BLAS, no framework, no GPU, and byte-identical output at any memory budget from 8 GB to 224 GB. Apache 2.0. *(Also listed under Mixture-of-Experts on CPU.)*
- midge - Spec-driven C engine that runs 100B+ MoE LLMs (gpt-oss, Mixtral, Qwen3-MoE) on ordinary CPU machines by keeping the dense trunk resident and streaming routed experts from disk in packed 4-bit form; ships an OpenAI-compatible server with tool calling. Apache 2.0. *(Also listed under Mixture-of-Experts on CPU.)*
- LARQL - Rust engine that runs transformer models entirely on CPU by decompiling their weights into a queryable "vindex" graph format, letting you browse, edit, and recompile model knowledge with the LQL query language and serve it over HTTP/gRPC with token streaming; no GPU required. Apache 2.0.
- TinyChatEngine - MIT Han Lab's on-device LLM/VLM inference library (MLSys 2024 Best Paper) implementing the AWQ low-precision weights; from-scratch C/C++ with no library dependency, running W4A16 quantized models on x86 (Intel/AMD), ARM (Apple M1/M2, Raspberry Pi), and CUDA. MIT. *(last updated 2024; research-grade reference for AWQ-based on-device inference.)*
- Qualcomm GenieX - Qualcomm's on-device GenAI runtime (community version of GENIE) that runs almost any GGUF model from Hugging Face across the Hexagon NPU, Adreno GPU, or CPU with a few lines of code; one C SDK exposed through CLI, Python, Kotlin/Java, Docker, and an OpenAI-compatible server. Snapdragon-only (Windows ARM64, Android, Linux IoT). BSD-3-Clause. *(Also listed under On-Device, Edge, ARM, and SBCs.)*
-
Talks, Papers, and Articles
-
Summary Dashboard
- llama.cpp — initial commit thread (Georgi Gerganov, Jan 2023) - The discussion that kicked off serious community interest in CPU-first LLM inference; the commit and issue thread document the initial benchmark results that made the case.
- "Efficient LLM Inference on CPUs" (Intel Labs, 2023) - Describes weight-only quantization for CPU inference on Sapphire Rapids: weights stored in INT4/INT8 while activations compute in BF16, reducing memory bandwidth pressure without sacrificing activation precision; a key technique behind Intel Extension for Transformers.
- "Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs" (Arm, Dec 2024) - Proposes Q4_0_8_8 and related kernels exploiting SVE/NEON MMLA instructions on Neoverse V1/V2; reports 3–3.2× prefill and up to 2× decode throughput improvement on Graviton3 vs standard Q4_0, with fine-grained codebooks that recover accuracy at the same bit-width.
- "Benchmarking AI Workloads on Intel Xeon with AMX" (itsabout.ai, 2024) - Practical benchmark guide for AMX on Sapphire Rapids; measures sustained GEMM throughput across batch sizes and establishes the 8× INT8 ops/cycle advantage of AMX over AVX-512 VNNI in the context of quantized LLM inference.
- "CPUs are finally having their AI moment" (TechZine, Jul 2026) - Analysis arguing CPUs have been underappreciated as AI infrastructure, especially for agentic workloads where tool calling, orchestration, and deterministic execution are inherently CPU-bound; surveys AMD EPYC Venice, Arm AGI CPU, and Nvidia Vera positioning.
- "LLM Optimization and Deployment on SiFive RISC-V Intelligence Processors" (SiFive, Jan 2026) - End-to-end deployment of TinyLlama and Llama-2-7B-Q4 on RISC-V using the IREE compiler with RVV vectorization; includes MLPerf accuracy validation and demonstrates readiness of open-hardware RISC-V for LLM inference.
- "Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone" (Microsoft, Apr 2024) - Reports Phi-3-mini (3.8B) quantized to 4-bit fits in 1.8 GB and achieves 12+ tok/s on iPhone 14 CPU fully offline, establishing small models as viable for on-device CPU inference.
- "Llama 3.2 vs Phi-3 Mini vs Gemma 2: iPhone Bake-Off" (PocketLLM, Apr 2026) - Head-to-head benchmark of Llama 3.2 3B, Phi-3.5 Mini 3.8B, and Gemma 2 2B on iPhone 15 Pro (Q4); reports 25–38 tok/s generation and 300–500 ms first-token latency on CPU-only inference, with Gemma 2 fastest and Llama 3.2 most versatile.
- "The Business Value of CPU-Based Inference" (Lenovo Press, 2025) - Enterprise TCO analysis of on-prem CPU inference (Intel Xeon 6) vs cloud GPU; reports $1.02M net savings over 5 years and 4.5-month breakeven vs cloud rental for sustained LLM workloads.
- "ML Model Server Resource Saving – Transition From GPUs to Intel CPUs" (NAVER / PyTorch Blog, 2024) - Production case study migrating 15 GPU instances to Intel CPUs across Korea/Japan data centers; achieved equivalent serving quality with 3× scale-out and ~$340K annual cost savings using IPEX and oneDNN optimizations.
- "WebLLM: A High-Performance In-Browser LLM Inference Engine" (MLCo/CMU, Dec 2024) - Demonstrates LLM inference running entirely in-browser via WebGPU and WebAssembly; the WebAssembly path handles CPU workloads including grammar-constrained generation and KV-cache management, retaining up to 80% of native device performance; ships an OpenAI-compatible browser-side API.
- "Why CPUs Also Make Sense for AI Inference" (Scaleway / Ampere CTO Jeff Wittich) - Interview with Ampere Computing CTO arguing that GPUs are compute overkill for inference; includes a concrete power-efficiency figure: the 128-core Ampere Altra consumes 3.6× less power per inference than an Nvidia A10 and 5.6× less than a Tesla T4, measured on OpenAI Whisper — a strong argument for CPU in energy-constrained or cost-per-watt–sensitive deployments.
- Tim Dettmers — "Which GPU for Deep Learning" - GPU-centric guide that nonetheless clearly articulates when GPU matters and when it does not; useful as a foil for the CPU-first argument.
- "FairyFuse: Multiplication-Free Ternary Inference on CPU" (arXiv, Jul 2026) - Proposes removing all multiplications from ternary (1.58-bit) model inference on CPU by fusing operations into lookup tables and bitwise logic; reports 3.6× speedup over GGML on CPU with 97.4% MMLU retention on Phi-3 Mini.
- "AdaptiveSD: Adaptive Speculative Decoding for CPU-Only Inference" (arXiv, Mar 2026) - Speculative decoding framework designed specifically for CPU-only inference with adaptive draft model selection; reports 1.9× speedup on CPU with no GPU dependency.
- "SMEPilot: Optimizing ARM SME Instructions for CPU Inference" (arXiv, Jul 2026) - Optimization guide for ARM Scalable Matrix Extension (SME) instructions targeting CPU inference workloads; provides auto-tuned SME kernels with performance analysis across ARM Cortex-X/A cores.
- "Kimi K3 runs on 8GB CPU, No GPU" (Data Science in Your Pocket, Aug 2026) - Practitioner deep-dive into kimi-k3-in-c's four core ideas — bounded resident working set, streaming from storage, on-demand expert loading, and expert LRU caching — that run the 2.78T Kimi K3 MoE on a single CPU in 8.24 GB RAM, summarised in [CPU-First Inference R&D Ideas](docs/cpu-first-rd-ideas.md).
- "PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU" (SOSP 2024) - The seminal hot/cold locality result: LLM neuron activation follows a skewed power-law, so hot neurons can stay resident while cold ones are computed on demand (offloaded to the CPU in a GPU-CPU hybrid). The neuron-level ancestor of today's expert-level disk-streaming engines; traced in [CPU-First Inference R&D Ideas](docs/cpu-first-rd-ideas.md).
-
-
Vision on CPU
-
Summary Dashboard
- YOLOv8 with OpenVINO (Ultralytics, 2025) - Official Ultralytics integration exporting YOLOv8–YOLO26 models to OpenVINO IR; benchmarks on Intel Xeon show the largest FP32 models exceeding 360 fps in async mode and smallest models approaching 5,000 fps with up to 14× throughput improvement over native PyTorch CPU. ([Lenovo Press: YOLO on Xeon 6](https://lenovopress.lenovo.com/lp2345-accelerating-real-time-object-detection-yolo-models-intel-xeon-6-openvino))
- Ultralytics OpenVINO on CPU — Production Guide - Practical guide comparing PyTorch CPU, ONNX, and OpenVINO for YOLO inference; reports OpenVINO delivers 2–3× speedup over ONNX Runtime and 5× over PyTorch CPU on Intel hardware, with INT8 quantization halving latency.
- CLIP-ONNX — CPU benchmarks - Benchmark comparing ONNX Runtime and PyTorch for CLIP ViT-B/32 on CPU (Xeon 2.3 GHz); ONNX achieves ~2.5 img/s at batch=2 for image encoding with 3× improvement over PyTorch at larger batch sizes.
- clip.cpp — CLIP inference in GGML - Dependency-free CLIP model inference using ggml with 4/5/8-bit quantization; supports text-only and vision-only modes, short startup time suitable for serverless deployments.
- DFN5B-CLIP-ViT-H-14-378 — INT8 ONNX on CPU - Large CLIP model (~5B params) quantized to INT8 ONNX; runs 2.3× faster on CPU with cosine similarity 0.985 vs FP32; benchmarks show 405 ms/image and ~20 text seq/s on current-gen Intel i7.
-
-
When You Actually Do Want a GPU
-
Summary Dashboard
-
Programming Languages
Categories
Runtimes and Inference Engines
39
Talks, Papers, and Articles
18
On-Device, Edge, ARM, and SBCs
15
Benchmarks and Evidence
10
Performance Tuning
7
CPU Fine-Tuning
6
Cloud ARM Servers
6
Quantization and Model Formats
6
Vision on CPU
5
Cost and Deployment Economics
4
Mixture-of-Experts on CPU
4
Mobile Phone CPUs
4
Model Selection and Hardware Fit
3
CPU AI Gap Map
1
Contributing
1
When You Actually Do Want a GPU
1
Multimodal CPU Workloads
1
Sub Categories
Keywords
llm
18
inference
11
transformers
9
deep-learning
9
quantization
8
ai
6
pytorch
5
llm-inference
5
gemma
5
qwen
5
llms
5
moe
4
deepseek
4
llama
4
simd
4
neural-network
4
large-language-models
4
transformer
4
c
3
llama3
3
gpt-oss
3
on-device-ai
3
gguf
3
android
3
python
3
language-model
3
lora
3
machine-learning
3
fine-tuning
3
nlp
3
cpu
3
openai
2
javascript
2
gpt
2
arm
2
vulkan
2
agent
2
cpp
2
minimax
2
ollama
2
tvm
2
chatgpt
2
webml
2
habana
2
llama-cpp
2
mixture-of-experts
2
cpu-inference
2
linux
2
ios
2
inference-engine
2