Projects in Awesome Lists by open-compass
A curated list of projects in awesome lists by open-compass .
https://github.com/open-compass/opencompass
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
benchmark chatgpt evaluation large-language-model llama2 llama3 llm openai
Last synced: 17 Nov 2025
https://github.com/open-compass/OpenCompass
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
benchmark chatgpt evaluation large-language-model llama2 llama3 llm openai
Last synced: 30 Jul 2025
https://github.com/open-compass/vlmevalkit
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
chatgpt claude clip computer-vision evaluation gemini gpt gpt-4v gpt4 large-language-models llava llm multi-modal openai openai-api pytorch qwen vit vqa
Last synced: 13 May 2025
https://github.com/open-compass/VLMEvalKit
Open-source evaluation toolkit of large vision-language models (LVLMs), support ~100 VLMs, 40+ benchmarks
chatgpt claude clip computer-vision evaluation gemini gpt gpt-4v gpt4 large-language-models llava llm multi-modal openai openai-api pytorch qwen vit vqa
Last synced: 20 Jul 2025
https://github.com/open-compass/mixtralkit
A toolkit for inference and evaluation of 'mixtral-8x7b-32kseqlen' from Mistral AI
Last synced: 12 Apr 2025
https://github.com/open-compass/MixtralKit
A toolkit for inference and evaluation of 'mixtral-8x7b-32kseqlen' from Mistral AI
Last synced: 12 Apr 2025
https://github.com/open-compass/lawbench
Benchmarking Legal Knowledge of Large Language Models
Last synced: 06 Apr 2025
https://github.com/open-compass/t-eval
[ACL2024] T-Eval: Evaluating Tool Utilization Capability of Large Language Models Step by Step
Last synced: 07 Apr 2025
https://github.com/open-compass/T-Eval
[ACL2024] T-Eval: Evaluating Tool Utilization Capability of Large Language Models Step by Step
Last synced: 09 May 2025
https://github.com/open-compass/LawBench
Benchmarking Legal Knowledge of Large Language Models
Last synced: 01 Apr 2025
https://github.com/open-compass/mmbench
Official Repo of "MMBench: Is Your Multi-modal Model an All-around Player?"
Last synced: 25 Jan 2026
https://github.com/open-compass/botchat
Evaluating LLMs' multi-round chatting capability via assessing conversations generated by two LLM instances.
Last synced: 30 Jan 2026
https://github.com/open-compass/mathbench
[ACL 2024 Findings] MathBench: A Comprehensive Multi-Level Difficulty Mathematics Evaluation Dataset
Last synced: 28 Feb 2026
https://github.com/open-compass/DevEval
A Comprehensive Benchmark for Software Development.
Last synced: 09 May 2025
https://github.com/open-compass/deveval
A Comprehensive Benchmark for Software Development.
Last synced: 14 Apr 2025
https://github.com/open-compass/compassjudger
The All-in-one Judge Models introduced by Opencompass
Last synced: 12 Nov 2025
https://github.com/open-compass/GTA
[NeurIPS 2024 D&B Track] GTA: A Benchmark for General Tool Agents
Last synced: 09 Jul 2025
https://github.com/open-compass/gta
[NeurIPS 2024 D&B Track] GTA: A Benchmark for General Tool Agents
Last synced: 07 Apr 2025
https://github.com/open-compass/mmbench-gui
Official repo of "MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents". It can be used to evaluate a GUI agent with a hierarchical manner across multiple platforms, including Windows, Linux, macOS, iOS, Android and Web.
benchmark-framework computer-use gui-agent vision-language-model
Last synced: 15 Sep 2025
https://github.com/open-compass/MathBench
[ACL 2024 Findings] MathBench: A Comprehensive Multi-Level Difficulty Mathematics Evaluation Dataset
Last synced: 16 Mar 2025
https://github.com/open-compass/ada-leval
The official implementation of "Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks"
Last synced: 14 Aug 2025
https://github.com/open-compass/criticeval
[NeurIPS 2024] A comprehensive benchmark for evaluating critique ability of LLMs
Last synced: 04 Mar 2026
https://github.com/open-compass/anah
[ACL 2024] ANAH & [NeurIPS 2024] ANAH-v2 & [ICLR 2025] Mask-DPO
acl alignment gpt hallucination-detection hallucination-mitigation iclr llms neurips
Last synced: 24 Apr 2025
https://open-compass.github.io/CriticEval/
[NeurIPS 2024] A comprehensive benchmark for evaluating critique ability of LLMs
Last synced: 29 Mar 2025
https://github.com/open-compass/prosa
[EMNLP 2024 Findings] ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
Last synced: 24 Apr 2025
https://github.com/open-compass/code-evaluator
A multi-language code evaluation tool.
Last synced: 02 Apr 2026
https://github.com/open-compass/GPassK
Official Repository of Are Your LLMs Capable of Stable Reasoning?
Last synced: 05 Mar 2025
https://github.com/open-compass/Creation-MMBench
Assessing Context-Aware Creative Intelligence in MLLMs
Last synced: 08 May 2025
https://github.com/open-compass/cibench
Official Repo of "CIBench: Evaluation of LLMs as Code Interpreter "
Last synced: 06 Oct 2025