https://github.com/laude-institute/terminal-bench
A benchmark for LLMs on complicated tasks in the terminal
https://github.com/laude-institute/terminal-bench
Last synced: about 1 year ago
JSON representation
A benchmark for LLMs on complicated tasks in the terminal
- Host: GitHub
- URL: https://github.com/laude-institute/terminal-bench
- Owner: laude-institute
- License: apache-2.0
- Created: 2025-01-17T22:34:26.000Z (over 1 year ago)
- Default Branch: main
- Last Pushed: 2025-08-14T16:48:25.000Z (about 1 year ago)
- Last Synced: 2025-08-14T18:33:25.606Z (about 1 year ago)
- Language: JetBrains MPS
- Homepage: https://www.tbench.ai
- Size: 39.2 MB
- Stars: 393
- Watchers: 7
- Forks: 108
- Open Issues: 82
-
Metadata Files:
- Readme: README.md
- License: LICENSE
Awesome Lists containing this project
- Awesome-LLMOps - terminal-bench - institute/terminal-bench.svg?style=flat&color=green)   (Training / Benchmark)
- awesome-prompts - **Terminal-Bench** - terminal agent benchmark (Stanford/Laude) — compile code, train models, set up servers in Docker-sandboxed environments; the de facto benchmark for agentic coding (2026). | (Frameworks / Eval & Testing)
- awesome-ai-agent-benchmarks - Terminal-Bench - to-end tasks in a real com… | 🔴 heavy | 2.4k | (The index / Tier 1 — frontier model-card standard (27))