{"id":29809006,"url":"https://github.com/alibaba-damo-academy/MedEvalKit","last_synced_at":"2025-07-28T16:03:49.652Z","repository":{"id":298482376,"uuid":"999941701","full_name":"alibaba-damo-academy/MedEvalKit","owner":"alibaba-damo-academy","description":"MedEvalKit: A Unified Medical Evaluation Framework","archived":false,"fork":false,"pushed_at":"2025-07-28T02:35:01.000Z","size":3490,"stargazers_count":113,"open_issues_count":8,"forks_count":9,"subscribers_count":4,"default_branch":"master","last_synced_at":"2025-07-28T04:22:43.908Z","etag":null,"topics":["evaluation-framework","llm","medicalai","multimodal"],"latest_commit_sha":null,"homepage":"https://alibaba-damo-academy.github.io/lingshu/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/alibaba-damo-academy.png","metadata":{"files":{"readme":"Readme.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-06-11T03:12:27.000Z","updated_at":"2025-07-28T02:35:04.000Z","dependencies_parsed_at":"2025-06-20T03:25:24.151Z","dependency_job_id":null,"html_url":"https://github.com/alibaba-damo-academy/MedEvalKit","commit_stats":null,"previous_names":["alibaba-damo-academy/medevalkit"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/alibaba-damo-academy/MedEvalKit","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/alibaba-damo-academy%2FMedEvalKit","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/alibaba-damo-academy%2FMedEvalKit/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/alibaba-damo-academy%2FMedEvalKit/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/alibaba-damo-academy%2FMedEvalKit/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/alibaba-damo-academy","download_url":"https://codeload.github.com/alibaba-damo-academy/MedEvalKit/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/alibaba-damo-academy%2FMedEvalKit/sbom","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":267543275,"owners_count":24104538,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-07-28T02:00:09.689Z","response_time":68,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["evaluation-framework","llm","medicalai","multimodal"],"created_at":"2025-07-28T16:01:13.809Z","updated_at":"2025-07-28T16:03:49.643Z","avatar_url":"https://github.com/alibaba-damo-academy.png","language":"Python","funding_links":[],"categories":["评估 Evaluation","4. Benchmarks"],"sub_categories":["3.2 Multimodal"],"readme":"\u003ch3 align=\"center\"\u003e\n  🩺 MedEvalKit: A Unified Medical Evaluation Framework\n\u003c/h3\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://arxiv.org/abs/2506.07044\" target=\"_blank\"\u003e📖 arXiv Paper\u003c/a\u003e •\n  \u003ca href=\"https://huggingface.co/collections/lingshu-medical-mllm/lingshu-mllms-6847974ca5b5df750f017dad\" target=\"_blank\"\u003e🤗 Lingshu Models\u003c/a\u003e •\n  \u003ca href=\"https://alibaba-damo-academy.github.io/lingshu/\" target=\"_blank\"\u003e🌐 Lingshu Project Page\u003c/a\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://opensource.org/license/apache-2-0\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/Code%20License-Apache_2.0-green.svg\" alt=\"License\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://github.com/alibaba-damo-academy\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/Institution-DAMO-red\" alt=\"Institution\"\u003e\n  \u003c/a\u003e\n  \u003ca\u003e\n    \u003cimg src=\"https://img.shields.io/badge/PRs-Welcome-red\" alt=\"PRs Welcome\"\u003e\n  \u003c/a\u003e\n\u003c/p\u003e\n\n---\n\n## 📌 Introduction\nA comprehensive evaluation framework for **Large Medical Models (LMMs/LLMs)** in the healthcare domain.  \nWe welcome contributions of new models, benchmarks, or enhanced evaluation metrics!\n\n---\n\n## Eval Results\n\u003cp align=\"center\"\u003e\n  \u003ca\u003e\n    \u003cimg src=\"assets/eval.png\"\u003e\n  \u003c/a\u003e\n\u003c/p\u003e\n\n\n\n\n## 🔥 Latest News\n* **2025-06-12** - Initial release of MedEvalKit v1.0!\n\n---\n\n## 🧪 Supported Benchmarks\n\n| Multimodal Medical Benchmarks | Text-Only Medical Benchmarks |\n|-----------------------|----------------------|\n| MMMU-Medical-test     | MedQA-USMLE          |\n| MMMU-Medical-val      | MedMCQA              |\n| PMC_VQA               | PubMedQA             |\n| OmniMedVQA            | Medbullets-op4       |\n| IU XRAY               | Medbullets-op5       |\n| MedXpertQA-Multimodal | MedXpertQA-Text      |\n| CheXpert Plus         | SuperGPQA            |\n| MIMIC-CXR             | HealthBench          |\n| VQA-RAD               | CMB                  |\n| SLAKE                 | CMExam               |\n| PATH-VQA              | CMMLU                |\n| MedFrameQA            | MedQA-MCMLE          |\n\n---\n\n## 🤖 Supported Models\n### HuggingFace Exclusive\n\u003cdiv style=\"column-count: 2;\"\u003e\n\n* BiMediX2\n* BiomedGPT\n* HealthGPT\n* Janus\n* Med_Flamingo\n* MedDr\n* MedGemma\n* NVILA\n* VILA_M3\n\n\u003c/div\u003e\n\n### HF + vLLM Compatible\n\u003cdiv style=\"column-count: 2;\"\u003e\n\n* HuatuoGPT-vision\n* InternVL\n* Llama_3.2-vision\n* LLava\n* LLava_Med\n* Qwen2_5_VL\n* Qwen2_VL\n\n\u003c/div\u003e\n\n---\n\n## 🛠️ Installation\n```bash\n# Clone repository\ngit clone https://github.com/DAMO-NLP-SG/MedEvalKit\ncd MedEvalKit\n\n# Install dependencies\npip install -r requirements.txt\npip install 'open_clip_torch[training]'\npip install flash-attn --no-build-isolation\n\n# For LLaVA-like models\ngit clone https://github.com/LLaVA-VL/LLaVA-NeXT.git\ncd LLaVA-NeXT \u0026\u0026 pip install -e .\n```\n\n---\n\n## 📂 Dataset Preparation\n### HuggingFace Datasets (Direct Access)\n```python\n# Set DATASETS_PATH='hf'\nVQA-RAD: flaviagiammarino/vqa-rad\nSuperGPQA: m-a-p/SuperGPQA\nPubMedQA: openlifescienceai/pubmedqa\nPATHVQA: flaviagiammarino/path-vqa\nMMMU: MMMU/MMMU\nMedQA-USMLE: GBaker/MedQA-USMLE-4-options\nMedQA-MCMLE: shuyuej/MedQA-MCMLE-Benchmark\nMedbullets_op4: tuenguyen/Medical-Eval-MedBullets_op4\nMedbullets_op5: LangAGI-Lab/medbullets_op5\nCMMMU: haonan-li/cmmlu\nCMExam: fzkuji/CMExam\nCMB: FreedomIntelligence/CMB\nMedFrameQA: SuhaoYu1020/MedFrameQA\n```\n\n### Local Datasets (Manual Download Required)\n| Dataset          | Source |\n|------------------|--------|\n| MedXpertQA       | [TsinghuaC3I](https://huggingface.co/datasets/TsinghuaC3I/MedXpertQA) |\n| SLAKE            | [BoKelvin](https://huggingface.co/datasets/BoKelvin/SLAKE) |\n| PMCVQA           | [RadGenome](https://huggingface.co/datasets/RadGenome/PMC-VQA) |\n| OmniMedVQA       | [foreverbeliever](https://huggingface.co/datasets/foreverbeliever/OmniMedVQA) |\n| MIMIC_CXR        | [MIMIC_CXR](https://physionet.org/content/mimic-cxr/2.1.0/) |\n| IU_Xray          | [IU_Xray](https://openi.nlm.nih.gov/faq?download=true) |\n| CheXpert Plus    | [CheXpert Plus](https://aimi.stanford.edu/datasets/chexpert-plus) |\n| HealthBench       | [Normal](https://openaipublic.blob.core.windows.net/simple-evals/healthbench/2025-05-07-06-14-12_oss_eval.jsonl),[Hard](https://openaipublic.blob.core.windows.net/simple-evals/healthbench/hard_2025-05-08-21-00-10.jsonl),[Consensus](https://openaipublic.blob.core.windows.net/simple-evals/healthbench/consensus_2025-05-09-20-00-46.jsonl) |\n\n---\n\n## 🚀 Quick Start\n### 1. Configure `eval.sh`\n```bash\n#!/bin/bash\nexport HF_ENDPOINT=https://hf-mirror.com\n# MMMU-Medical-test,MMMU-Medical-val,PMC_VQA,MedQA_USMLE,MedMCQA,PubMedQA,OmniMedVQA,Medbullets_op4,Medbullets_op5,MedXpertQA-Text,MedXpertQA-MM,SuperGPQA,HealthBench,IU_XRAY,CheXpert_Plus,MIMIC_CXR,CMB,CMExam,CMMLU,MedQA_MCMLE,VQA_RAD,SLAKE,PATH_VQA,MedFrameQA\nEVAL_DATASETS=\"Medbullets_op4\" \nDATASETS_PATH=\"hf\"\nOUTPUT_PATH=\"eval_results/{}\"\n# TestModel,Qwen2-VL,Qwen2.5-VL,BiMediX2,LLava_Med,Huatuo,InternVL,Llama-3.2,LLava,Janus,HealthGPT,BiomedGPT,Vllm_Text,MedGemma,Med_Flamingo,MedDr\nMODEL_NAME=\"Qwen2.5-VL\"\nMODEL_PATH=\"Qwen2.5-VL-7B-Instruct\"\n\n#vllm setting\nCUDA_VISIBLE_DEVICES=\"0\"\nTENSOR_PARALLEL_SIZE=\"1\"\nUSE_VLLM=\"False\"\n\n#Eval setting\nSEED=42\nREASONING=\"False\"\nTEST_TIMES=1\n\n\n# Eval LLM setting\nMAX_NEW_TOKENS=8192\nMAX_IMAGE_NUM=6\nTEMPERATURE=0\nTOP_P=0.0001\nREPETITION_PENALTY=1\n\n# LLM judge setting\nUSE_LLM_JUDGE=\"True\"\n# gpt api model name\nGPT_MODEL=\"gpt-4.1-2025-04-14\"\nOPENAI_API_KEY=\"\"\n\n\n# pass hyperparameters and run python sccript\npython eval.py \\\n    --eval_datasets \"$EVAL_DATASETS\" \\\n    --datasets_path \"$DATASETS_PATH\" \\\n    --output_path \"$OUTPUT_PATH\" \\\n    --model_name \"$MODEL_NAME\" \\\n    --model_path \"$MODEL_PATH\" \\\n    --seed $SEED \\\n    --cuda_visible_devices \"$CUDA_VISIBLE_DEVICES\" \\\n    --tensor_parallel_size \"$TENSOR_PARALLEL_SIZE\" \\\n    --use_vllm \"$USE_VLLM\" \\\n    --max_new_tokens \"$MAX_NEW_TOKENS\" \\\n    --max_image_num \"$MAX_IMAGE_NUM\" \\\n    --temperature \"$TEMPERATURE\"  \\\n    --top_p \"$TOP_P\" \\\n    --repetition_penalty \"$REPETITION_PENALTY\" \\\n    --reasoning \"$REASONING\" \\\n    --use_llm_judge \"$USE_LLM_JUDGE\" \\\n    --judge_gpt_model \"$GPT_MODEL\" \\\n    --openai_api_key \"$OPENAI_API_KEY\" \\\n    --test_times \"$TEST_TIMES\" \n```\n\n### 2. Run Evaluation\n```bash\nchmod +x eval.sh  # Add execute permission\n./eval.sh\n```\n\n---\n\n## 📜 Citation\n```bibtex\n@article{xu2025lingshu,\n  title={Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning},\n  author={Xu, Weiwen and Chan, Hou Pong and Li, Long and Aljunied, Mahani and Yuan, Ruifeng and Wang, Jianyu and Xiao, Chenghao and Chen, Guizhen and Liu, Chaoqun and Li, Zhaodonghui and others},\n  journal={arXiv preprint arXiv:2506.07044},\n  year={2025}\n}\n```\n\n\u003cdiv align=\"center\"\u003e\n  \u003csub\u003eBuilt with ❤️ by the DAMO Academy Medical AI Team\u003c/sub\u003e\n\u003c/div\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Falibaba-damo-academy%2FMedEvalKit","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Falibaba-damo-academy%2FMedEvalKit","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Falibaba-damo-academy%2FMedEvalKit/lists"}