An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with humaneval

A curated list of projects in awesome lists tagged with humaneval .

https://github.com/bin123apple/autocoder

We introduced a new model designed for the Code generation task. Its test accuracy on the HumanEval base dataset surpasses that of GPT-4 Turbo (April 2024) and GPT-4o.

code-generation code-interpreter humaneval llm nlp nlp-machine-learning text-generation

Last synced: 21 Apr 2025

https://github.com/bin123apple/AutoCoder

We introduced a new model designed for the Code generation task. Its test accuracy on the HumanEval base dataset surpasses that of GPT-4 Turbo (April 2024) and GPT-4o.

code-generation code-interpreter humaneval llm nlp nlp-machine-learning text-generation

Last synced: 05 Apr 2025

https://github.com/the-crypt-keeper/can-ai-code

Self-evaluating interview for AI coders

ai ggml humaneval langchain llama-cpp llm transformers

Last synced: 05 Apr 2025

https://github.com/abacaj/code-eval

Run evaluation on LLMs using human-eval benchmark

humaneval wizardcoder

Last synced: 06 Apr 2025

https://github.com/zorse-project/coboleval

Evaluate LLM-generated COBOL

cobol evaluation humaneval llm

Last synced: 13 Apr 2025

https://github.com/declare-lab/llm-reasoningtest

Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions

gsm8k humaneval reasoning

Last synced: 18 Jul 2025

https://github.com/abhaymundhara/llm-benchmark-suite

Benchmark suite for evaluating LLMs and SLMs on coding and SE tasks. Features HumanEval, MBPP, SWE-bench, and BigCodeBench with an interactive Streamlit UI. Supports cloud APIs (OpenAI, Anthropic, Google) and local models via Ollama. Tracks pass rates, latency, token usage, and costs.

benchmark bigcodebench claude code-generation evaluation gemini humaneval llm mbpp ollama openai python streamlit swe-bench

Last synced: 29 Apr 2026

https://github.com/viplismism/benchmark_suite

Universal LLM evaluation framework — HumanEval, SWE-bench, GPQA, BigCodeBench, τ-Bench, Terminal-Bench via a single make command

benchmarks evaluation gpqa humaneval llm swe-bench

Last synced: 31 May 2026

https://github.com/jcartu/qwen-bench-2026-05-11-v2-followup

Study #4: FP8+MTP{3,5} speed on repne/vllm:v2 + max_tokens=8192 quality re-runs for BF16+DFlash n=8 and FP8+MTP=3. Follow-up to studies #2 and #3.

benchmark blackwell dflash humaneval inference mbpp mtp qwen-bench qwen3 speculative-decoding vllm

Last synced: 12 Jun 2026

https://github.com/jcartu/llm-stress-harness

Diagnostic toolkit for self-hosted LLM inference: failure-taxonomic stress harness + 4-phase orchestrator + parametric vLLM launchers

benchmarking humaneval inference llm mbpp python sglang speculative-decoding stress-testing vllm

Last synced: 12 Jun 2026