An open API service indexing awesome lists of open source software.

Projects in Awesome Lists by open-compass

A curated list of projects in awesome lists by open-compass .

https://github.com/open-compass/opencompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

benchmark chatgpt evaluation large-language-model llama2 llama3 llm openai

Last synced: 17 Nov 2025

https://github.com/open-compass/OpenCompass

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

benchmark chatgpt evaluation large-language-model llama2 llama3 llm openai

Last synced: 30 Jul 2025

https://github.com/open-compass/vlmevalkit

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

chatgpt claude clip computer-vision evaluation gemini gpt gpt-4v gpt4 large-language-models llava llm multi-modal openai openai-api pytorch qwen vit vqa

Last synced: 13 May 2025

https://github.com/open-compass/VLMEvalKit

Open-source evaluation toolkit of large vision-language models (LVLMs), support ~100 VLMs, 40+ benchmarks

chatgpt claude clip computer-vision evaluation gemini gpt gpt-4v gpt4 large-language-models llava llm multi-modal openai openai-api pytorch qwen vit vqa

Last synced: 20 Jul 2025

https://github.com/open-compass/mixtralkit

A toolkit for inference and evaluation of 'mixtral-8x7b-32kseqlen' from Mistral AI

llm mistral moe

Last synced: 12 Apr 2025

https://github.com/open-compass/MixtralKit

A toolkit for inference and evaluation of 'mixtral-8x7b-32kseqlen' from Mistral AI

llm mistral moe

Last synced: 12 Apr 2025

https://github.com/open-compass/lawbench

Benchmarking Legal Knowledge of Large Language Models

benchmark chatgpt law llm

Last synced: 06 Apr 2025

https://github.com/open-compass/t-eval

[ACL2024] T-Eval: Evaluating Tool Utilization Capability of Large Language Models Step by Step

Last synced: 07 Apr 2025

https://github.com/open-compass/T-Eval

[ACL2024] T-Eval: Evaluating Tool Utilization Capability of Large Language Models Step by Step

Last synced: 09 May 2025

https://github.com/open-compass/LawBench

Benchmarking Legal Knowledge of Large Language Models

benchmark chatgpt law llm

Last synced: 01 Apr 2025

https://github.com/open-compass/mmbench

Official Repo of "MMBench: Is Your Multi-modal Model an All-around Player?"

Last synced: 25 Jan 2026

https://github.com/open-compass/botchat

Evaluating LLMs' multi-round chatting capability via assessing conversations generated by two LLM instances.

Last synced: 30 Jan 2026

https://github.com/open-compass/mathbench

[ACL 2024 Findings] MathBench: A Comprehensive Multi-Level Difficulty Mathematics Evaluation Dataset

Last synced: 28 Feb 2026

https://github.com/open-compass/DevEval

A Comprehensive Benchmark for Software Development.

Last synced: 09 May 2025

https://github.com/open-compass/deveval

A Comprehensive Benchmark for Software Development.

Last synced: 14 Apr 2025

https://github.com/open-compass/compassjudger

The All-in-one Judge Models introduced by Opencompass

Last synced: 12 Nov 2025

https://github.com/open-compass/GTA

[NeurIPS 2024 D&B Track] GTA: A Benchmark for General Tool Agents

llm-agent llm-evaluation

Last synced: 09 Jul 2025

https://github.com/open-compass/gta

[NeurIPS 2024 D&B Track] GTA: A Benchmark for General Tool Agents

llm-agent llm-evaluation

Last synced: 07 Apr 2025

https://github.com/open-compass/mmbench-gui

Official repo of "MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents". It can be used to evaluate a GUI agent with a hierarchical manner across multiple platforms, including Windows, Linux, macOS, iOS, Android and Web.

benchmark-framework computer-use gui-agent vision-language-model

Last synced: 15 Sep 2025

https://github.com/open-compass/MathBench

[ACL 2024 Findings] MathBench: A Comprehensive Multi-Level Difficulty Mathematics Evaluation Dataset

Last synced: 16 Mar 2025

https://github.com/open-compass/ada-leval

The official implementation of "Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks"

gpt4 llm long-context

Last synced: 14 Aug 2025

https://github.com/open-compass/criticeval

[NeurIPS 2024] A comprehensive benchmark for evaluating critique ability of LLMs

Last synced: 04 Mar 2026

https://github.com/open-compass/anah

[ACL 2024] ANAH & [NeurIPS 2024] ANAH-v2 & [ICLR 2025] Mask-DPO

acl alignment gpt hallucination-detection hallucination-mitigation iclr llms neurips

Last synced: 24 Apr 2025

https://open-compass.github.io/CriticEval/

[NeurIPS 2024] A comprehensive benchmark for evaluating critique ability of LLMs

Last synced: 29 Mar 2025

https://github.com/open-compass/prosa

[EMNLP 2024 Findings] ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs

Last synced: 24 Apr 2025

https://github.com/open-compass/code-evaluator

A multi-language code evaluation tool.

Last synced: 02 Apr 2026

https://github.com/open-compass/GPassK

Official Repository of Are Your LLMs Capable of Stable Reasoning?

Last synced: 05 Mar 2025

https://github.com/open-compass/compassbench

Demo data of CompassBench

Last synced: 02 Mar 2026

https://github.com/open-compass/Creation-MMBench

Assessing Context-Aware Creative Intelligence in MLLMs

Last synced: 08 May 2025

https://github.com/open-compass/cibench

Official Repo of "CIBench: Evaluation of LLMs as Code Interpreter "

Last synced: 06 Oct 2025

https://github.com/open-compass/saga

Last synced: 12 Feb 2026