awesome-llm-eval
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on LLMs. 一个由工具、基准/数据、演示、排行榜和大模型等组成的精选列表,主要面向基础大模型评测,旨在探求生成式AI的技术边界.
https://github.com/onejune2018/awesome-llm-eval
Last synced: 20 days ago
JSON representation
-
Anthropomorphic-Taxonomy
-
Typical Emotional Quotient (EQ)-Alignment Ability evaluation benchmarks
-
Typical Intelligence Quotient (IQ)-General Intelligence evaluation benchmarks
-
Typical Professional Quotient (PQ)-Professional Expertise evaluation benchmarks
-
-
Courses
-
Popular-LLM
- Full+Stack+LLM+Bootcamp - LLM相关学习/应用资源集.
- 大语言模型课程notebooks集-Large Language Model Course - Course with a roadmap and notebooks to get into Large Language Models (LLMs).
-
-
Datasets-or-Benchmark
-
Agent能力
-
RAG检索增强生成评估
- CRAG - answer pairs and mock APIs to simulate web and Knowledge Graph (KG) search. CRAG is designed to encapsulate a diverse array of questions across five domains and eight question categories, reflecting varied entity popularity from popular to long-tail, and temporal dynamisms ranging from years to seconds. (2024-06-07) |
- BERGEN - augmented GENeration), a library to benchmark RAG systems, focusing on question-answering (QA). Inconsistent benchmarking poses a major challenge in comparing approaches and understanding the impact of each component in a RAG pipeline. BERGEN was designed to ease the reproducibility and integration of new datasets and models thanks to HuggingFace. (2024-05-31) |
- raga-llm-hub - llm-hub是一个全面的语言和学习模型(LLM)评估工具包。它拥有超过100个精心设计的评价指标,是允许开发者和组织有效评估和比较LLM的最具综合性的平台,并为LLM和检索增强生成(RAG)应用建立基本的防护措施。这些测试评估包括相关性与理解、内容质量、幻觉、安全与偏见、上下文相关性、防护措施以及漏洞扫描等多个方面,同时提供一系列基于指标的测试用于定量分析 (2024-03-10) |
- ARES - 文档-答案三元组。ARES培训流程包括三个步骤:(1)从领域内段落生成合成查询和答案。(2)通过在合成生成的训练数据上进行微调,为评分RAG系统准备LLM评委。(3)部署准备好的LLM评委以评估您的RAG系统在关键性能指标上的表现 (2023-09-27) |
- RGB - 09-04) |
-
代码能力
- McEval - 4相比,在多语言的编程能力上仍然存在较大差距,绝大多数开源模型甚至无法超越GPT-3.5。此外测试也表明开源模型中如Codestral,DeepSeek-Coder, CodeQwen以及一些衍生模型也展现出优异的多语言能力. McEval is a massively multilingual code benchmark covering 40 programming languages with 16K test samples, which substantially pushes the limits of code LLMs in multilingual scenarios. The benchmark contains challenging code completion, understanding, and generation evaluation tasks with finely curated massively multilingual instruction corpora McEval-Instruct. McEval leaderboard can be found [here](https://mceval.github.io/). (2024-06-11) |
- SuperCLUE-Agent - XL,这是一个大规模多语言代码生成基准测试,旨在填补这一缺陷。HumanEval-XL在23种自然语言和12种编程语言之间建立了联系,包含22,080个提示,平均每个提示有8.33个测试用例。通过确保跨多种NL和PLs的平行数据,HumanEval-XL为多语言LLMs提供了一个全面的评估平台,允许评估对不同NLs的理解。这项工作是填补多语言代码生成领域NL泛化评估空白的开创性步骤。 (2024-02-26) |
-
垂直领域
- PsyEval - 11-15) |
- RGB - Augmented Generation,RAG)的评测基准,分析了不同大型语言模型在RAG所需的4种基本能力(噪声稳健性、负面拒绝、信息整合和反事实稳健性)的性能,建立了中英文的“检索增强生成基准”(Retrieval-Augmented Generation Benchmark,RGB),根据所需的基本能力分为4个独立的测试集 (2023-09-04) |
- OpsEval - 10-02) |
- BLURB
- Fin-Eva - Eva Version 1.0,覆盖财富管理、保险、投资研究等多个金融场景以及金融专业主题学科,总评测题数目达到1.3w+。蚂蚁数据源包括各业务领域数据、互联网公开数据,经过数据脱敏、文本聚类、语料精筛、数据改写等处理过程后,结合金融领域专家的评审构建而成。上海财经大学数据源主要基于相关领域权威性考试的各类真题和模拟题对知识大纲的要求。蚂蚁部分涵盖金融认知、金融知识、金融逻辑、内容生成以及安全合规五大类能力33个子维度共8445个测评题; 上财部分涵盖金融,经济,会计和证书等四大领域,包括4661个问题,涵盖34个不同的学科。Fin-Eva Version 1.0 全部采用单选题这类有固定答案的问题,配合相应指令让模型输出标准格式 (2023-12-20) |
- GenMedicalEval - 4等其他模型,具有独特优势 (2023-12-08)|
- DebugBench - 4向源数据植入漏洞,并确保了严格的质量检查 (2024-01-09) |
- LAiW - 10-25)|
- LawBench - 09-28) |
- PPTC - Match评估系统,该系统根据预测文件而不是标签API序列来评估大语言模型是否完成指令,因此它支持各种LLM生成的API序列目前PPT生成存在三个方面的不足:多轮会话中的错误累积、长PPT模板处理和多模态感知问题 (2023-11-04) |
- LLMRec - 10-08)|
- GSM8K - / *)以达到最终答案 |
-
多模态-跨模态
- ChartVLM - 02-19) |
- ReForm-Eval - Eval是一个用于综合评估大视觉语言模型的基准数据集。ReForm-Eval通过对已有的、不同任务形式的多模态基准数据集进行重构,构建了一个具有统一且适用于大模型评测形式的基准数据集。所构建的ReForm-Eval具有如下特点:构建了横跨8个评估维度,并为每个维度提供足量的评测数据(平均每个维度4000余条);具有统一的评测问题形式(包括单选题和文本生成问题);方便易用,评测方法可靠高效,且无需依赖ChatGPT等外部服务;高效地利用了现存的数据资源,无需额外的人工标注,并且可以进一步拓展到更多数据集上 (2023-10-24) |
- LVLM-eHub - Modality Arena"是一个用于大型多模态模型的评估平台。在Fastchat之后,两个匿名模型在视觉问答任务上进行并排比较,"Multi-Modality Arena"允许你在提供图像输入的同时,对视觉-语言模型进行并排基准测试。支持MiniGPT-4,LLaMA-Adapter V2,LLaVA,BLIP-2等多种模型 |
-
推理速度
- llm-analysis
- llmperf - 11-03)|
- llm-inference-benchmark
- llm-inference-bench
- GPU-Benchmarks-on-LLM-Inference - inch M1 Max MacBook Pro, M2 Ultra Mac Studio, 14-inch M3 MacBook Pro and 16-inch M3 Max MacBook Pro. |
-
通用
- LucyEval
- Zhujiu
- InfoQ评测
- lmsys排名榜
- MultiNLI
- HellaSwag
- LAMBADA
- OpenAI Moderation API
- GLUE Benchmark
- MT-bench - bench旨在测试多轮对话和遵循指令的能力,包含80个高质量的多轮问题,涵盖常见用例并侧重于具有挑战性的问题,以区分不同的模型。它包括8个常见类别的用户提示,以指导其构建:写作、角色扮演、提取、推理、数学、编程等 |
- C_Eval - 4、ChatGPT、Claude、LLaMA、Moss 等多个模型的性能。|
- Safety Eval 安全大模型评测
- JioNLP-LLM评测数据集
- KoLA - oriented LLM Assessment benchmark(KoLA)由清华大学知识工程组(THU-KEG)托管,旨在通过进行认真的设计,考虑数据、能力分类和评估指标,来精心测试LLMs的全球知识。这个基准测试包含80个高质量的多轮问题 |
- Zhujiu
- Psychometrics Eval - 10-19) |
- TrustLLM
- RewardBench - bench),[Code](https://github.com/allenai/reward-bench) 和 [Dataset](https://hf.co/datasets/allenai/reward-bench) (2024-03-20)|
- EQ-Bench - 12-20) |
- AlignBench - as-Judge),并且结合思维链(Chain-of-Thought)生成对模型回复的多维度分析和最终的综合评分,增强了评测的高可靠性和可解释性 (2023-12-01)|
- LLMBar - 10-11) |
- HalluQA - hard部分69条,knowledge部分206条,每个问题平均有2.8个正确答案和错误答案标注。为了提高HalluQA的可用性,作者设计了一个使用GPT-4担任评估者的评测方法。具体来说,把幻觉的标准以及作为参考的正确答案以指令的形式输入给GPT-4,让GPT-4判断模型的回复有没有出现幻觉 (2023-11-08) |
- Do-Not-Answer - 4相媲美的结果 |
- Zhujiu
- ChatEval
- Z-Bench
- zeno-build
- llm-benchmark
- SQUAD
- LogiQA
- CoQA
- ParlAI
- LIT
- Adversarial NLI (ANLI)
- XieZhi - choice questions spanning 516 diverse disciplines and four difficulty levels. 新的领域知识综合评估基准测试:Xiezhi。对于多选题,Xiezhi涵盖了516种不同学科中的220,000个独特问题,其中涵盖了13个学科。作者还提出了Xiezhi-Specialty和Xiezhi-Interdiscipline,每个都含有15k个问题。使用Xiezhi基准测试评估了47种先进的LLMs的性能|
- MMCU
- CMMLU
- Gaokao
- GAOKAO-Bench - bench是一个以中国高考题目为数据集,测评大模型语言理解能力、逻辑推理能力的测评框架 |
- Safety Eval 安全大模型评测
- SuperCLUE
- BIG-Bench-Hard - Bench任务,我们称之为BIG-Bench Hard(BBH)。这些任务是以前的语言模型评估未能超越平均人工评分者的任务 |
- KoLA - oriented LLM Assessment benchmark(KoLA)由清华大学知识工程组(THU-KEG)托管,旨在通过进行认真的设计,考虑数据、能力分类和评估指标,来精心测试LLMs的全球知识。这个基准测试包含80个高质量的多轮问题 |
- M3Exam
- HalluQA - hard部分69条,knowledge部分206条,每个问题平均有2.8个正确答案和错误答案标注。为了提高HalluQA的可用性,作者设计了一个使用GPT-4担任评估者的评测方法。具体来说,把幻觉的标准以及作为参考的正确答案以指令的形式输入给GPT-4,让GPT-4判断模型的回复有没有出现幻觉 (2023-11-08) |
-
量化压缩
- LLM-QBench - QBench is a Benchmark Towards the Best Practice for Post-training Quantization of Large Language Models", and it is also an efficient LLM compression tool with various advanced compression methods, supporting multiple inference backends. (2024-05-09) |
-
长上下文
- InfiniteBench - 4, Claude 2 等。(4)真实场景与合成场景: InfiniteBench 既包含真实场景数据,探测大模型在处理实际问题的能力;也包含合成数据,为测试数据拓展上下文窗口提供了便捷. InfiniteBench is the first LLM benchmark featuring an average data length surpassing 100K tokens. InfiniteBench comprises synthetic and realistic tasks spanning diverse domains, presented in both English and Chinese. The tasks in InfiniteBench are designed to require well understanding of long dependencies in contexts, and make simply retrieving a limited number of passages from contexts not sufficient for these tasks. (2024-03-19) |
-
-
Datasets or Benchmarks
-
Domain
- SWE-bench - bench is a benchmark for evaluating the performance of large language models on real software issues collected from GitHub. Given a code repository and a problem, the task of the language model is to generate a patch that can solve the described problem |
-
General
- Aviary
- InfoQ Evaluation - oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo.|
- InfoQ Evaluation - oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo.|
- LAMBADA - term understanding capabilities.|
- InfoQ Evaluation - oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo.|
- InfoQ Evaluation - oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo.|
- InfoQ Evaluation - oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo.|
- InfoQ Evaluation - oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo.|
- HalluQA - based entries. Each question has an average of 2.8 correct and incorrect answers annotated. To enhance the usability of HalluQA, the authors designed a GPT-4-based evaluation method. Specifically, hallucination criteria and correct answers are input as instructions to GPT-4, which evaluates whether the model's response contains hallucinations.|
- InfoQ Evaluation - oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo.|
- InfoQ Evaluation - oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo.|
- InfoQ Evaluation - oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo.|
- InfoQ Evaluation - oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo.|
- HF Open LLM Leaderboard - source LLMs. Evaluations focus on four datasets: AI2 Reasoning Challenge, HellaSwag, MMLU, and TruthfulQA, primarily in English.|
-
Programming Languages
Categories
Sub Categories
Popular-LLM
115
Typical Intelligence Quotient (IQ)-General Intelligence evaluation benchmarks
53
通用
45
多模态-跨模态
42
量化压缩
41
Open-LLM
41
Quantization-and-Compression
30
Typical Emotional Quotient (EQ)-Alignment Ability evaluation benchmarks
30
Leaderboards for popular Provider (performance and cost, 2024-05-14)
29
Typical Professional Quotient (PQ)-Professional Expertise evaluation benchmarks
26
General
19
Pre-trained-LLM
19
垂直领域
12
Instruction-finetuned-LLM
8
LLM推理
6
推理速度
5
RAG检索增强生成评估
5
Aligned-LLM
4
代码能力
2
Agent能力
2
Multi-modal/Cross-modal
2
Domain
1
长上下文
1
RAG-Evaluation
1
Keywords
llm
38
machine-learning
30
large-language-models
28
chatgpt
26
deep-learning
20
python
18
language-model
15
evaluation
14
ai
13
nlp
12
llmops
12
llms
11
llm-evaluation
10
gpt-4
9
gpt
9
pytorch
9
prompt-engineering
8
benchmark
8
llama
8
gpt-3
8
mlops
8
transformers
8
openai
8
evaluation-framework
7
rag
7
chatbot
7
natural-language-processing
6
data-science
5
in-context-learning
5
instruction-following
5
evaluation-metrics
5
rlhf
5
prompt
5
inference
4
llm-inference
4
foundation-models
4
langchain
4
vector-database
4
gpu
4
question-answering
4
chinese
4
ml
4
neural-network
4
tensorflow
4
gpt-2
4
alpaca
3
gpt4
3
fine-tuning
3
agent
3
multimodal
3