Projects in Awesome Lists tagged with humaneval
A curated list of projects in awesome lists tagged with humaneval .
https://github.com/bin123apple/autocoder
We introduced a new model designed for the Code generation task. Its test accuracy on the HumanEval base dataset surpasses that of GPT-4 Turbo (April 2024) and GPT-4o.
code-generation code-interpreter humaneval llm nlp nlp-machine-learning text-generation
Last synced: 21 Apr 2025
https://github.com/bin123apple/AutoCoder
We introduced a new model designed for the Code generation task. Its test accuracy on the HumanEval base dataset surpasses that of GPT-4 Turbo (April 2024) and GPT-4o.
code-generation code-interpreter humaneval llm nlp nlp-machine-learning text-generation
Last synced: 05 Apr 2025
https://github.com/the-crypt-keeper/can-ai-code
Self-evaluating interview for AI coders
ai ggml humaneval langchain llama-cpp llm transformers
Last synced: 05 Apr 2025
https://github.com/abacaj/code-eval
Run evaluation on LLMs using human-eval benchmark
Last synced: 06 Apr 2025
https://github.com/zorse-project/coboleval
Evaluate LLM-generated COBOL
cobol evaluation humaneval llm
Last synced: 13 Apr 2025
https://github.com/declare-lab/llm-reasoningtest
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
Last synced: 18 Jul 2025
https://github.com/abhaymundhara/llm-benchmark-suite
Benchmark suite for evaluating LLMs and SLMs on coding and SE tasks. Features HumanEval, MBPP, SWE-bench, and BigCodeBench with an interactive Streamlit UI. Supports cloud APIs (OpenAI, Anthropic, Google) and local models via Ollama. Tracks pass rates, latency, token usage, and costs.
benchmark bigcodebench claude code-generation evaluation gemini humaneval llm mbpp ollama openai python streamlit swe-bench
Last synced: 29 Apr 2026
https://github.com/viplismism/benchmark_suite
Universal LLM evaluation framework — HumanEval, SWE-bench, GPQA, BigCodeBench, τ-Bench, Terminal-Bench via a single make command
benchmarks evaluation gpqa humaneval llm swe-bench
Last synced: 31 May 2026
https://github.com/jcartu/qwen-bench-2026-05-11-v2-followup
Study #4: FP8+MTP{3,5} speed on repne/vllm:v2 + max_tokens=8192 quality re-runs for BF16+DFlash n=8 and FP8+MTP=3. Follow-up to studies #2 and #3.
benchmark blackwell dflash humaneval inference mbpp mtp qwen-bench qwen3 speculative-decoding vllm
Last synced: 12 Jun 2026
https://github.com/north-shore-ai/crucible_datasets
Dataset management and caching for AI research benchmarks
ai ai-benchmarks beam benchmark-datasets benchmarking data-loading dataset-management datasets elixir ensemble-methods gsm8k humaneval llm machine-learning ml-datasets mmlu otp reliability research statistical-testing
Last synced: 21 Oct 2025
https://github.com/jcartu/llm-stress-harness
Diagnostic toolkit for self-hosted LLM inference: failure-taxonomic stress harness + 4-phase orchestrator + parametric vLLM launchers
benchmarking humaneval inference llm mbpp python sglang speculative-decoding stress-testing vllm
Last synced: 12 Jun 2026