https://github.com/harbor-framework/terminal-bench
A benchmark for LLMs on complicated tasks in the terminal
https://github.com/harbor-framework/terminal-bench
Last synced: 28 days ago
JSON representation
A benchmark for LLMs on complicated tasks in the terminal
- Host: GitHub
- URL: https://github.com/harbor-framework/terminal-bench
- Owner: harbor-framework
- License: apache-2.0
- Created: 2025-01-17T22:34:26.000Z (over 1 year ago)
- Default Branch: main
- Last Pushed: 2026-01-22T05:27:03.000Z (6 months ago)
- Last Synced: 2026-06-17T15:06:48.617Z (about 1 month ago)
- Language: Python
- Homepage: https://www.tbench.ai
- Size: 138 MB
- Stars: 2,366
- Watchers: 14
- Forks: 543
- Open Issues: 302
-
Metadata Files:
- Readme: README.md
- License: LICENSE
Awesome Lists containing this project
- awesome-evals - Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces - framework/terminal-bench> ยท *benchmark* โ A concrete instance of the thesis: each task ships a Docker environment + programmatic verification test suite + oracle โ i.e. a benchmark that IS an RL environment (and is used as one). 2.4k stars, active. ๐ (2 ยท "If you can eval it, you have built it" โ eval โ capability โ RL environment)
- awesome-agent-harness - GitHub - 2373-f4b400?style=flat-square)](https://github.com/harbor-framework/terminal-bench) | terminal, benchmark, long-horizon | Terminal-native benchmark suite for long-horizon, verification-heavy agent tasks. | (Catalog / Evaluation Harnesses & Benchmarks)
- awesome-harness-engineering - harbor-framework/terminal-bench