{"id":43552,"url":"https://github.com/stardust-coder/awesome-latest-LLM","name":"awesome-latest-LLM","description":"最新LLMの一覧を作成します","projects_count":176,"last_synced_at":"2026-09-10T04:00:25.337Z","repository":{"id":211811149,"uuid":"728512687","full_name":"stardust-coder/awesome-latest-LLM","owner":"stardust-coder","description":"最新LLMの一覧を作成します","archived":false,"fork":false,"pushed_at":"2026-06-20T05:50:45.000Z","size":2019,"stargazers_count":23,"open_issues_count":0,"forks_count":0,"subscribers_count":2,"default_branch":"master","last_synced_at":"2026-08-01T22:04:44.877Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/stardust-coder.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2023-12-07T05:12:41.000Z","updated_at":"2026-07-13T09:25:07.000Z","dependencies_parsed_at":"2023-12-14T07:31:05.624Z","dependency_job_id":"120bf3f1-8708-4ff4-bd52-9879e94b63de","html_url":"https://github.com/stardust-coder/awesome-latest-LLM","commit_stats":null,"previous_names":["stardust-coder/awesome-latest-llm"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/stardust-coder/awesome-latest-LLM","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/stardust-coder%2Fawesome-latest-LLM","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/stardust-coder%2Fawesome-latest-LLM/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/stardust-coder%2Fawesome-latest-LLM/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/stardust-coder%2Fawesome-latest-LLM/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/stardust-coder","download_url":"https://codeload.github.com/stardust-coder/awesome-latest-LLM/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/stardust-coder%2Fawesome-latest-LLM/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":341189360,"owners_count":37188035,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-09-10T02:00:06.818Z","response_time":107,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2024-01-13T21:18:41.024Z","updated_at":"2026-09-10T04:00:25.338Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["English-centric","Japanese-centric","Dataset","Model","Uncategorized","Evaluation","Small language models (SLM)","Leaderboard","Table of Contents"],"sub_categories":["Only Text","Uncategorized","Image + Text","Evaluation benchmarks","Corpus"],"readme":"# 🧠 Awesome Latest LLM / LLMの最新情報のまとめ\n\nA continuously updated, curated list of the latest Large Language Models in chronological order.\nStay ahead of the rapidly evolving LLM ecosystem.\n\n\u003cp align=\"center\"\u003e \u003cimg src=\"https://img.shields.io/badge/Updates-Monthly-blue\" /\u003e \u003cimg src=\"https://img.shields.io/github/stars/stardust-coder/awesome-latest-LLM?style=social\" /\u003e \u003cimg src=\"https://img.shields.io/badge/PRs-Welcome-brightgreen\" /\u003e \u003c/p\u003e\n\n**NEWS**\n\n- 🔥2026.8 Kimi K3がリリースされました！\n- 🔥2026.6 GLM5.2がリリースされました！\n- 🔥2026.6 Gemma4に12Bが追加されました！画像、音声も入力できます\n- 🔥2026.6 MicrosoftからMAI-Thinking-1が発表されました！\n- 🔥2026.4 GLM, Gemma, DeepSeek, Qwenの最新版がリリースされました！\n\n\u003cdetails\u003e\n\n\u003csummary\u003eHistory\u003c/summary\u003e\n\n- 🔥2026.3 Rakuten 3.0がリリースされました！\n- 🔥2026.2 Qwen3.5シリーズがリリースされました！\n- 🔥2026.2 Swallowシリーズの新作がリリースされました！\n- 🔥2026.1 Kimiの最新版モデルがリリースされました！\n- 🔥2025.12 ZAIからスマホを操作するLLM（AutoGLM）やCodingに長けたLLM（GLM）がリリースされました！\n- 🔥2025.12 Mistralから最新版モデルがリリースされました！Code特化のDevestral 2も!\n- 🔥2025.11 DeepSeekの最新版モデルがリリースされました！\n- 🔥2025.11 エンプラモデルをも超えうる1TパラメタのKimi-K2がリリースされました！\n- 🔥2025.9 Qwen3-Nextがリリースされました！\n- 🔥2025.5 DeepSeek-R1の最新版が公開されました！\n- 🔥2025.5 Swallowチームから最新版Gemma Swallowがリリースされました！\n- 🔥2025.5 Ai2からOLMo2-1Bの最新版がリリースされました！\n- 🔥2025.4 Qwen3がリリースされました！\n- 🔥2025.4 Llama4がリリースされました！\n- 2025.3 🔥Gemma3がリリースされました！\n- 2025.3 🔥Llama-Swallowの最新版がリリースされました！\n- 2025.3 🔥Sarashina2.2がリリースされました.\n- 2025.2 小規模言語モデル（SLM）の特集を始めました\n- 2025.2 🔥 Grok3がxAIから発表されました！\n- 2025.2 モデルの絞り込みを行いました.\n- 2025.1 🔥 [Qwen2.5-Max](https://qwenlm.github.io/blog/qwen2.5-max/)がリリースされました. モデル非公開.\n- 2025.1 🔥Minimax-01がリリースされました！\n- 2024.11 🔥Qwenチームからreasoningに優れたとされる実験的モデルQWQがリリースされました！\n- 2024.7 🔥東工大からLlama3の日本語継続学習モデルが発表！\n- 2024.6 🔥ELYZAからLlama3の日本語継続学習モデルが発表！\n- 2024.6 🔥Googleから27BのGemma2が公開！何が強みか教えて！\n- 2024.6 🔥NVIDIAが340Bの巨大モデルを公開！publicにしては最大級\n- 2024.6 🔥Qwen2シリーズが登場！日本語も優秀！\n- 2024.5 🔥MicrosoftからPhi-3シリーズが登場！\n- 2024.5 🔥Stockmarkから100Bの日本語モデルがリリース!さすがGENIAC\n- 2024.4 🔥MetaからLlama3がリリース!まずは8Bと70B!\n- 2024.4 🔥CohereからCommand-R+がリリース!研究用に重みも公開.\n- 2024.4 🔥Databricksより132BのMoEモデルが公開されました！大きい！\n- 2024.3 Cohereからプロダクション向けCommand-Rがリリース!研究用に重みも公開.\n- 2024.3 ELYZAからLlama2の追加学習日本語モデルのデモがリリースされました！\n- 2024.3 東工大からMixtralの追加学習日本語モデル[Swallow-MX](), [Swallow-MS]()がリリースされました！👏\n- 2024.2 GoogleからGeminiで用いられているLLM [Gemma](https://blog.google/technology/developers/gemma-open-models/)をオープンにするとのお達しが出ました!\n- 2024.2 Kotoba Technologyと東工大から[日本語Mamba 2.8B](https://huggingface.co/kotoba-tech/kotomamba-2.8B-v1.0)が公開されました!\n- 2024.2 Alibabaの[QWen](https://qwenlm.github.io/blog/qwen1.5/)が1.5にアップグレードされました！！\n- 2024.2 Reka AIから21BでGemini Pro, GPT-3.5超えと発表されました.\n- 2024.2 LLM-jpのモデルが更新されました！v1.1\n- 2024.2 カラクリから70B日本語LLMが公開されました！\n- 2024.1 [リコー](https://www.nikkei.com/article/DGXZRSP667803_R30C24A1000000/)が13B日本語LLMを発表しました！\n- 2024.1 Phi-2のMoE, Phixtralが公開されました！\n- 2023.12 Phi-2のライセンスがMITに変更されました！  \n- 2023.12 ELYZAから日本語[13Bモデル](https://huggingface.co/elyza/ELYZA-japanese-Llama-2-13b)がリリースされました.  \n- 2023.12 東工大から[Swallow](https://tokyotech-llm.github.io)がリリースされました.  \n- 2023.12 MistralAIから[Mixtral-8x7B](https://github.com/open-compass/MixtralKit)がリリースされました.    \n- 2023.12 [日本語LLMの学習データを問題視する記事](https://github.com/AUGMXNT/shisa/wiki/A-Review-of-Public-Japanese-Training-Sets#analysis)が公開されました.\n\u003c/details\u003e\n\n## Table of Contents\n[My Pickup](#pickup)  \n[Omni](#omni)  \n[Computer-use](#computer-use)  \n[English-centric](#english-centric)  \n[Japanese-centric](#japanese-centric)  \n[Small Language Model](#SLM)  \n[Medical Adaptation](#medical-adaptation)  \n\n\n\u003ca id=\"pickup\"\u003e\u003c/a\u003e\n\n# My Pickup\n\nComing soon...\n\n- Personally recommended\n- Will be updated anytime\n\n# Omni\n\u003ca id=\"omni\"\u003e\u003c/a\u003e\n\n| When? | Name |  HF?  | Size(max) | License | pretraining/base | finetuning | misc.|\n|---|---|---|---|---|---|---|---|\n|2025.11| [Uni-Moe-Omni (HIT)](https://idealistxy.github.io/Uni-MoE-v2.github.io/) | [HF](https://huggingface.co/HIT-TMG/Uni-MoE-2.0-Omni) | 33B-1.5~18B | apache-2.0 | 75B token   |  | MoE, surpass Qwen2.5-Omni |\n|2025.9| [Qwen3-Omni (Alibaba)](https://github.com/QwenLM/Qwen3-Omni) | [HF](https://huggingface.co/collections/Qwen/qwen3-omni-68d100a86cd0906843ceccbe) | 30B-A3B | apache-2.0 |  text-first pretraining and mixed multimodal training |  | [demo](https://huggingface.co/spaces/Qwen/Qwen3-Omni-Demo) |\n\n# Computer Use / Tool Use / Function Calling\n\u003ca id=\"computer-use\"\u003e\u003c/a\u003e\n\n\n| When? | Name |  HF?  | Size(max) | License | pretraining/base | finetuning | misc.|\n|---|---|---|---|---|---|---|---|\n|2025.12| [FunctionGemma (Google)](https://huggingface.co/google/functiongemma-270m-it) | [HF](https://huggingface.co/zai-org/AutoGLM-Phone-9B-Multilingual) | 0.27B |  | | | function calling |\n|2025.12| [AutoGLM-Phone-9B-Multilingual (ZAI)](https://x.com/Zai_org/status/1999118106543919203?s=20) | [HF](https://huggingface.co/zai-org/AutoGLM-Phone-9B-Multilingual) | 9B | mit (for research and educational purposes only.) | | | smartphone |\n|2025.11| [Fara (Microsoft)](https://www.microsoft.com/en-us/research/blog/fara-7b-an-efficient-agentic-model-for-computer-use/) | [HF](https://huggingface.co/microsoft/Fara-7B) | 7B | mit | Qwen2.5-VL-7B |  |  |\n|2025.11| [Jan-v2]() | [HF](https://huggingface.co/collections/janhq/jan-v2-vl) | 8B | apache-2.0 | [Qwen3-VL-8B-Thinking](https://huggingface.co/Qwen/Qwen3-VL-8B-Thinking)|  |  |\n\n# LLM List\n\u003ca id=\"english-centric\"\u003e\u003c/a\u003e\n## English-centric\n\n| When? | Name |  HF?  | Size(max) | License | pretraining/base | finetuning | misc.|\n|---|---|---|---|---|---|---|---|\n|2026.7| [Kimi K3](https://huggingface.co/moonshotai/Kimi-K3) | [HF](https://huggingface.co/moonshotai/Kimi-K3) | 2.8TB | [Kimi K3 License](https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE) | ? | ? | 1T context window, moe(104Ba) |\n|2026.7| [FIM-Midtraining (TIGER AI Lab)](https://github.com/TIGER-AI-Lab/FIM-Midtraining) | [HF](https://huggingface.co/collections/TIGER-Lab/fim-midtraining) | 14B | Apache-2.0 | Qwen2.5-Coder / Qwen3 with function-aware fill-in-the-middle mid-training | R2E-Gym / SWE-Smith / SWE-Lego | [paper](https://arxiv.org/abs/2607.12463)|\n|2026.7| [Mistral-Medium-3.5]() | [HF](https://huggingface.co/mistralai/Mistral-Medium-3.5-128B) | 128B | apache-2.0  |  |  |  |\n|2026.6| [GLM 5.2](https://z.ai/blog/glm-5.1) | [HF](https://huggingface.co/zai-org/GLM-5.2) | 754B | MIT  |  |  |  |\n|2026.6| [Kimi-K2.6]() | [HF](https://huggingface.co/moonshotai/Kimi-K2.6) | 1TA32B | modifiedMIT |  ? | ? | moe, 256k context |\n|2026.4| [Deepseek V4 Pro]() | [HF](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro) | 1.6TB-A49B | MIT  |  |  |  |\n|2026.2| [Qwen3.6 (Alibaba)]() | [HF](https://huggingface.co/collections/Qwen/qwen36) | 27B | apache-2.0 |  |  |  |\n|2026.4| [Gemma 4]() | [HF](https://huggingface.co/collections/google/gemma-4) | 2.3~31B | apache-2.0  |  |  |  |\n|2026.3| [Nemotron3]() | [HF](https://huggingface.co/collections/nvidia/nvidia-nemotron-v3) | 4~235B | [license](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/)  |  |  |  |\n|2026.3| [Mistral-Small-4]() | [HF](https://huggingface.co/mistralai/Mistral-Small-4-119B-2603) | 119B | apache-2.0  |  |  | moe |\n|2026.2| [Qwen3.5 (Alibaba)]() | [HF](https://huggingface.co/collections/Qwen/qwen35) | 0.8~397B | apache-2.0 |  |  |  |\n|2025.12| [Mistral-Large-3](https://x.com/MistralAI/status/1995872766177018340?s=20) | [HF](https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512) | 675B | |  |  | |\n|2025.8| [GPT-OSS (OpenAI)]() | [HF](https://huggingface.co/openai/gpt-oss-120b) | 20B~120B |  | |  |  |\n\n\u003c!-- \n|2026.4| [GLM 5.1](https://z.ai/blog/glm-5.1) | [HF](https://huggingface.co/zai-org/GLM-5.1) | 754B | MIT  |  |  |  |\n|2026.1| [Kimi-K2.5](https://x.com/Kimi_Moonshot/status/2016024049869324599?s=20) | [HF](https://huggingface.co/moonshotai/Kimi-K2.5) | 1TA32B | modifiedMIT | Kimi-K2-Base  | 15T tokens | moe |\n|2025.12| [DeepSeek-V3.2]() | [HF](https://huggingface.co/collections/deepseek-ai/deepseek-v32) | 671B | MIT |  |  |  |\n|2025.12| [rnj-1(EssentialAI)](https://x.com/essential_ai/status/1997123628765524132?s=20) | [HF](https://huggingface.co/EssentialAI/rnj-1-instruct) | 8B | apache2.0 | 8.4T+380B tokens  | 150B tokens | code and STEM |\n|2025.11| [Olmo 3 (Allen)](https://allenai.org/blog/olmo3) | [HF](https://huggingface.co/collections/allenai/olmo-3) | 7, 32B | apache-2.0 |  |  |  |\n|2025.10| [Ling-1T (InclusionAI)]() | [HF](https://huggingface.co/collections/inclusionAI/ling-v2-68bf1dd2fc34c306c1fa6f86) | 1T-A50B | MIT | 20T+ |  | moe |\n|2025.9| [Qwen3-Next (Alibaba)]() | [HF](https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct) | 80B-A3B | apache-2.0 | 15T |  | moe |\n|2025.11| [Kimi-K2]() | [HF](https://huggingface.co/collections/moonshotai/kimi-k2) | 1T-A32B | [modified MIT](https://huggingface.co/moonshotai/Kimi-K2-Thinking/blob/main/LICENSE) |  |  | moe |\n|2025.11| [Kimi-Linear]() | [HF](https://huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct) | 48B-A3B | [modified MIT](https://huggingface.co/moonshotai/Kimi-K2-Thinking/blob/main/LICENSE) |  |  | moe |\n|2025.5| [DeepSeek-R1-0528]() | [HF](https://huggingface.co/deepseek-ai/DeepSeek-R1-0528) | 671B |  | |  | approaching o3 \u0026 Gemini 2.5 Pro |\n|2025.4| [Qwen3 (Alibaba)]()|[HF](https://huggingface.co/collections/Qwen/qwen3-67dd247413f0e2e4f653967f)| 0.6~235B |apache-2.0| | |\n|2025.4| [Llama4 (Meta)](https://ai.meta.com/blog/llama-4-multimodal-intelligence/)|[HF](https://huggingface.co/collections/meta-llama/llama-4-67f0c30d9fe03840bc9d0164)|17B|llama4|30T token||10M token|\n|2025.1| [DeepSeek-V3-0324]() | [HF](https://huggingface.co/deepseek-ai/DeepSeek-V3-0324) | 671B | [link](https://github.com/deepseek-ai/DeepSeek-V3/blob/main/LICENSE-MODEL) | 14.8T |  | MoE(37B) |\n|2025.3| [Gemma3 (Google)]() | [HF](https://huggingface.co/collections/google/gemma-3-release-67c6c6f89c4f76621268bb6d)| 27B | [gemma](https://ai.google.dev/gemma/terms) | | |\n|2025.1| [DeepSeek-R1](https://github.com/deepseek-ai/DeepSeek-R1/blob/main/DeepSeek_R1.pdf) | [HF](https://huggingface.co/deepseek-ai/DeepSeek-R1)| 671B | MIT | | |\n|2024.12| [Llama3.3 (Meta)]() | [HF](https://huggingface.co/collections/meta-llama/llama-33-67531d5c405ec5d08a852000) | 70B | [llama3.3](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct/blob/main/LICENSE) | |  |  |\n|2024.12| [Phi-4 (Microsoft)](https://www.microsoft.com/en-us/research/uploads/prod/2024/12/P4TechReport.pdf) |[HF](https://huggingface.co/NyxKrage/Microsoft_Phi-4) | 14B | msrla |  |  | small, sft, dpo |\n|2024.11| [QWQ (Alibaba)]() |[HF](https://huggingface.co/Qwen/QwQ-32B-Preview) | 32B | apache-2.0 | Qwen2.5 |  |  reasoning |\n|2024.9| [Qwen 2.5 (Alibaba)]() |[HF](https://huggingface.co/collections/Qwen/qwen25-66e81a666513e518adb90d9e) | 0.5~72B | apache2.0 |  |  | [long context available](https://huggingface.co/collections/Qwen/qwen25-1m-679325716327ec07860530ba) | \n|2024.12| [DeepSeek-V3](https://github.com/deepseek-ai/DeepSeek-V3) | [HF](https://huggingface.co/deepseek-ai/DeepSeek-V3) | 671B | [link](https://github.com/deepseek-ai/DeepSeek-V3/blob/main/LICENSE-MODEL) | 14.8T | sft, RL | MoE |\n|2024.7| [Reflection]() |[HF](https://huggingface.co/mattshumer/ref_70_e3) | 70B | llama3.1 | Llama 3.1 | synthetic data (Glaive) |  |\n|2025.1| [InternLM v3]() | [HF](https://huggingface.co/internlm/internlm3-8b-instruct) | 8B | apache-2.0 | 4T token |  | deep thinking | \n|2024.10| [Llama3.2(Meta)]() |[HF](https://huggingface.co/collections/meta-llama/llama-32-66f448ffc8c32f949b04c8cf) | 1B,3B | llama3.2 |  llama3.2 |  |  |\n|2024.7| [Llama3.1(Meta)]() |[HF]() | 70B, 405B | Llama3.1 |  |  |  |\n|2024.6| [Gemma2(Google)]() |[HF](https://huggingface.co/collections/google/gemma-2-release-667d6600fd5220e7b967f315) | 2B, 9B, 27B | gemma |  |  |  |\n|2024.6| [Nemotron(NVIDIA)]() |[HF](https://huggingface.co/nvidia/Nemotron-4-340B-Instruct) | 340B |  | - | - |  |\n|2024.6| [Qwen2(Alibaba)]() |[HF](https://huggingface.co/Qwen/Qwen2-72B) | 7~72B | tongyi-qianwen | - | - |  |\n|2024.4| [Phi-3(Microsoft)](https://arxiv.org/abs/2404.14219) |[HF](microsoft/Phi-3-medium-128k-instruct) | 3.8B, 13B | MIT |  Phi-3 datasets | - |  |\n|2024.4| [Llama 3(Meta)](https://llama.meta.com/llama3/) |[HF](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct) | 70B | [META LLAMA3](https://llama.meta.com/llama3/license/) | || [extended to 120B](https://huggingface.co/mlabonne/Meta-Llama-3-120B-Instruct) |\n|2024.4| [Wizart-8x22B(Microsoft)]() |[HF](https://huggingface.co/microsoft/WizardLM-2-8x22B) | 8x22B | apache-2.0 | [Mixtral-8x22B(Mistral)](https://mistral.ai/news/mixtral-8x22b/) | | MoE, closed now |\n|2024.4| [Mixtral-8x22B(Mistral)](https://mistral.ai/news/mixtral-8x22b/) |[HF](https://huggingface.co/mistral-community/Mixtral-8x22B-v0.1) | 8x22B | apache-2.0 | || MoE |\n|2024.4| [Command-R+(Cohere)](https://txt.cohere.com/command-r/) |[HF](https://huggingface.co/CohereForAI/c4ai-command-r-plus) | 104B | non commercial | || RAG capability |\n|2024.4| [DBRX(Databricks)]() |[HF](https://huggingface.co/databricks/dbrx-instruct) | 132B | databricks | || MoE |\n|2024.3| [Grok-1](https://github.com/xai-org/grok-1) | | 314B | | twitter | | MoE |\n|2024.3| [BTX(Meta)](https://arxiv.org/pdf/2403.07816.pdf)|||||| MoE |\n|2024.3| [Command-R(Cohere)](https://txt.cohere.com/command-r/) |[HF](https://huggingface.co/CohereForAI/c4ai-command-r-v01) | 35B | non commercial | || RAG capability |\n|2024.2| [Aya(Cohere)](https://cohere.com/research/aya?ref=txt.cohere.com) |[HF](https://huggingface.co/CohereForAI/aya-101) | 13B | apache-2.0 | || multilingual |\n|2024.2| [Gemma(Google)](https://blog.google/technology/developers/gemma-open-models/) | | 8.5B | || |application open for reseachers |\n|2024.2| [Miqu](https://twitter.com/arthurmensch/status/1752737462663684344?ref_src=twsrc%5Etfw%7Ctwcamp%5Etweetembed%7Ctwterm%5E1752737462663684344%7Ctwgr%5Ecd2e234e5fa688c1a14852aa90158cd4f59facb4%7Ctwcon%5Es1_\u0026ref_url=https%3A%2F%2Fgigazine.net%2Fnews%2F20240201-hugging-face-miqu-mistral-model%2F) | [HF](https://huggingface.co/miqudev/miqu-1-70b/tree/main) | 70B | none ||| leaked from Mistral |\n|2024.2| [Reka Flash](https://reka.ai/reka-flash-an-efficient-and-capable-multimodal-language-model/) |  | 21B | ||| not public|\n|2024.1| [Self-Rewarding(Meta)]() | [arxiv](https://arxiv.org/pdf/2401.10020.pdf) | 70B | Llama2 | Llama2| - | DPO |\n|2024.1| [Phixtral]() | [HF](https://huggingface.co/mlabonne/phixtral-4x2_8) | 2.7Bx4 | MIT |||MoE|\n|2023.12| [LongNet(Microsoft)](https://github.com/microsoft/torchscale) | [arXiv](https://arxiv.org/pdf/2307.02486.pdf) | - | apache-2.0 | [MAGNETO](https://arxiv.org/pdf/2210.06423.pdf)| input 1B token| |\n|2023.12| [Phi-2(Microsoft)]() | [HF](https://huggingface.co/microsoft/phi-2) | 2.7B | MIT |||\n|2023.12| [gigaGPT(Cerebras)](https://github.com/Cerebras/gigaGPT) | | 70B, 175B | apache-2.0 | | |\n|2023.12| [Mixtral-8x7B](https://github.com/open-compass/MixtralKit)| [HF](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1) | 8x7B | apache-2.0 |||MoE, [offloading](https://github.com/dvmazur/mixtral-offloading)|\n|2023.12| [Mamba](https://github.com/state-spaces/mamba)| [HF](https://huggingface.co/state-spaces/mamba-2.8b) | 2.8B | apache-2.0 | based on state space model| | \n|2023.11| [QWen(Alibaba)](https://github.com/QwenLM/Qwen) | [HF](https://huggingface.co/Qwen/Qwen-72B) | 72B | [license](https://github.com/QwenLM/Qwen/blob/main/Tongyi%20Qianwen%20LICENSE%20AGREEMENT)| 3T tokens | | beats Llama2 |\n|2023.9| [TinyLlama](https://github.com/jzhang38/TinyLlama) | [HF](https://huggingface.co/TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T) | apache-2.0 | 1.1B | based on Llama, 3T token |  | |\n|2023.9| [Xwin-LM](https://github.com/Xwin-LM/Xwin-LM) | [HF](https://huggingface.co/Xwin-LM/Xwin-LM-70B-V0.1)  | 70B | Llama2 |based on Llama2| also codes and math|\n|2023.7| [Llama2(Meta)](https://ai.meta.com/llama/) | [HF](https://huggingface.co/meta-llama) | 70B | Llama2 | 2T tokens| chat-hf seems the best| \n|2024.1| [LLaMa-Pro-8B(Tencent)]() | [HF](https://huggingface.co/TencentARC/LLaMA-Pro-8B) | 8B | Llama2 |||\n|2023.12| [Amber](https://www.llm360.ai) | [HF](https://huggingface.co/LLM360/Amber) | 7B | apache-2.0 | Llama|| totally open|\n|2023.11| [Orca2(Microsoft)]() | [HF](https://huggingface.co/microsoft/Orca-2-13b) | 13B | MSRA-license| based on Llama2|||\n|2023.9| [Phi-1.5(Microsoft)](https://arxiv.org/abs/2309.05463) | [HF](https://huggingface.co/microsoft/phi-1_5) | 1.3B| MSRA-license||textbooks| --\u003e\n\n- See also [Awesome-LLM](https://github.com/Hannibal046/Awesome-LLM)\n- See [OpenVLM Leaderboard](https://huggingface.co/spaces/opencompass/open_vlm_leaderboard) for VLMs.\n\n\u003ca id=\"japanese-centric\"\u003e\u003c/a\u003e\n\n## Japanese-centric\n\n| When? | Name |  HF?  | Size | License | pretraining | finetuning | misc.|\n|---|---|---|---|---|---|---|---|\n|2026.4| [LLM-jp-4（NII）]() | [HF](https://huggingface.co/collections/llm-jp/llm-jp-4-models) | 8, 32B | apache2.0 |  | | Japanese flagship |\n|2026.3| [Rakuten 3.0]() | [HF](https://huggingface.co/Rakuten/RakutenAI-3.0) | 671B |  | DeepseekV3.2 | |  |\n|2026.2| [GPTOSS-Swallow （科学大）]() | [HF](https://huggingface.co/tokyotech-llm/GPT-OSS-Swallow-120B-SFT-v0.1) | 120B |  |  | |  |\n|2025.11| [PLaMo 3（PFN）]() | [HF](https://huggingface.co/pfnet/plamo-3-nict-31b-base) | 31B |  |  | |  |\n|2025.7| [Stockmark 2（Stockmark）]() | [HF](https://huggingface.co/stockmark/Stockmark-2-100B-Instruct) | 100B |  |  | |  |\n|2025.5| [Llama3.3 Swallow （科学大）]() | [HF](https://huggingface.co/tokyotech-llm/Llama-3.3-Swallow-70B-Instruct-v0.4) | 70B | Llama3.3 | Llama3.3 | |  |\n|2025.5| [LLM-jp-3.1（NII）]() | [HF](https://huggingface.co/collections/llm-jp/llm-jp-31-fine-tuned-models-68368681b9b35de1c4ac8de4) | 1.8B, 13B, 8x13B | apache2.0 | Wikipedia etc. | | Japanese flagship |\n\u003c!-- |2025.5| [Gemma2 Swallow （科学大）]() | [HF](https://huggingface.co/collections/tokyotech-llm/gemma-2-swallow-67f2bdf95f03b9e278264241) | 2, 9, 27B |  |  | |  | --\u003e\n\u003c!-- |2025.3| [Stockmark 2]() | [HF](https://huggingface.co/stockmark/Stockmark-2-100B-Instruct-beta) | 100B |  |  | |  |\n|2025.3| [Llama-3.3-Swallow-70B-Instruct-v0.4]() | [HF](https://huggingface.co/tokyotech-llm/Llama-3.3-Swallow-70B-Instruct-v0.4) | 70B | [llama3.3](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct/blob/main/LICENSE) | Llama3.3 | | JMT-Bench 0.772 |\n|2025.2| [PlaMo 2 (PFN)](https://tech.preferred.jp/ja/blog/plamo-2-8b/) | [HF](https://huggingface.co/pfnet/plamo-2-8b) | 8B | [plamo](https://tech.preferred.jp/ja/blog/plamo-community-license/) ||| Samba |\n|2025.3| [Bakeneko (rinna)]()| [HF](https://huggingface.co/collections/rinna/qwen25-bakeneko-67aa2ef444910bbc55a21222) [HF](https://huggingface.co/rinna/qwq-bakeneko-32b) | 32B | apache-2.0 |  |  | |\n|2025.1| [DeepSeek-R1-Distil-Qwen-Japanese(CyberAgent)]()| [HF](https://huggingface.co/cyberagent/DeepSeek-R1-Distill-Qwen-32B-Japanese) | 32B | MIT | [distil](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B) | Japanese dataset | |\n|2024.12| [llm-jp-3-172b-instruct3]() | [HF](https://huggingface.co/llm-jp/llm-jp-3-172b-instruct3) | 172B | [利用規約](https://huggingface.co/llm-jp/llm-jp-3-172b-instruct3/raw/main/LICENSE_ja)  |  | |  |\n|2024.12| [Llama-3.1-Swallow-70B-Instruct-v0.3]() | [HF](https://huggingface.co/tokyotech-llm/Llama-3.1-Swallow-70B-Instruct-v0.3) | 70B | [llama3.1](https://huggingface.co/meta-llama/Llama-3.1-70B-Instruct/blob/main/LICENSE) | Llama3.1 | |  |\n|2024.7| [Llama-3.1-70B-Japanese-Instruct-2407]() | [HF](https://huggingface.co/cyberagent/Llama-3.1-70B-Japanese-Instruct-2407) | 70B | Llama3.1 | Llama3.1  | |  |\n|2024.7| [LLama3-Swallow]() | [HF](https://huggingface.co/tokyotech-llm/Llama-3-Swallow-70B-Instruct-v0.1) | 70B | Llama3 | Llama3  | |  |\n|2024.6| [LLama3ELYZA-JP-8B](https://elyza.ai/news/2024/06/26/elyza-llm-for-jp%E3%82%B7%E3%83%AA%E3%83%BC%E3%82%BA%E3%81%AE%E6%9C%80%E6%96%B0%E3%83%A2%E3%83%87%E3%83%ABllam) | [HF](https://huggingface.co/elyza/Llama-3-ELYZA-JP-8B) | 8B | Llama3 | Llama3  | | 70B not open |\n|2024.6| [KARAKURI LM 8x7B](https://karakuri.ai/seminar/news/karakuri-lm-8x7b-instruct-v0-1/) | [HF](karakuri-ai/karakuri-lm-8x7b-chat-v0.1) | 8x7B | Apache-2.0 |  | | MoE |\n|2024.5| [Stockmark-100B]() | [HF](stockmark/stockmark-100b) | 100B | MIT |  | | |\n|2024.3| [youko(rinna)]() | [HF](https://huggingface.co/rinna/llama-3-youko-8b) | 8B | Llama3 | Llama3 | | |\n|2024.3| [EvoLLM-JP]() | [HF](https://huggingface.co/SakanaAI/EvoLLM-JP-v1-7B) | 7B | MSR(non-commercial) | | | |\n|2024.3| [Swallow-MX(東工大)]() | [HF](https://huggingface.co/tokyotech-llm/Swallow-MX-8x7b-NVE-v0.1) | 8x7B | | Mixtralベース |\n|2024.2| [KARAKURI 70B](https://karakuri.ai/seminar/news/karakuri-lm/) | [HF](https://huggingface.co/karakuri-ai/karakuri-lm-70b-v0.1) | 70B | cc-by-sa-4.0 | Llama2-70Bベース | | [note](https://note.com/ngc_shj/n/n46ced665b378?sub_rt=share_h)|\n|2023.12| [ELYZA-japanese-Llama-2-13b](https://note.com/elyza/n/n5d42686b60b7) | [HF](https://huggingface.co/elyza/ELYZA-japanese-Llama-2-13b) | 13B | | Llama-2-13b-chatベース |\n|2023.12| [Swallow(東工大)](https://tokyotech-llm.github.io) | [HF](https://huggingface.co/tokyotech-llm) | 70B | | Llama2-70Bベース |\n|2023.11| [StableLM(StabilityAI)](https://ja.stability.ai/blog/japanese-stable-lm-beta) | [HF](https://huggingface.co/stabilityai/japanese-stablelm-base-beta-70b) | 70B | | Llama2-70Bベース |\n|2023.10| [LLM-jp](https://llm-jp.nii.ac.jp/blog/2024/02/09/v1.1-tuning.html) | [HF](https://huggingface.co/llm-jp) | 13B | DPO追加あり | --\u003e\n\n- See more on [awesome-japanese-llm](https://github.com/llm-jp/awesome-japanese-llm), [日本語LLMまとめ](https://llm-jp.github.io/awesome-japanese-llm/) and [日本語LLM評価](https://swallow-llm.github.io/evaluation/about.ja.html)\n- Let's go soverign AI!\n\n\u003ca id=\"SLM\"\u003e\u003c/a\u003e\n\n## Small language models (SLM)\n\n| When? | Name |  HF?  | Size | License | pretraining | finetuning | misc.|\n|---|---|---|---|---|---|---|---|\n|2026.8| [Ling-3.0-tiny (inclusionAI)](https://huggingface.co/inclusionAI/Ling-3.0-tiny) | [HF](https://huggingface.co/inclusionAI/Ling-3.0-tiny) | 7.9B | mit| ? | ? | moe 1.3Ba |\n|2026.4| [Bonsai (PrismML)](https://prismml.com/) | [HF](https://huggingface.co/collections/prism-ml/bonsai) | 1.7~8B | apache2.0| | |\n|2026.4| [LFM2.5 (LiquidAI)]() | [HF](https://huggingface.co/collections/LiquidAI/lfm25) | 0.35, 1.2B | [LFMv1](https://huggingface.co/LiquidAI/LFM2.5-350M/blob/main/LICENSE) | | also japanese |\n|2025.12| [Ministral 3]() | [HF](https://huggingface.co/collections/mistralai/ministral-3) | 3B |  | | |\n|2025.3| [Sarashina2.2](https://huggingface.co/sbintuitions/sarashina2.2-3b-instruct-v0.1) | [HF](https://huggingface.co/collections/sbintuitions/sarashina22) | 0.5B,1B,3B | mit | | | ELYZA-tasks=3.75 |\n\u003c!-- |2025.10| [Granite-4.0-H-Micro (IBM)]() | [HF](https://huggingface.co/ibm-granite/granite-4.0-h-micro) | 3B | apache-2.0 | | | --\u003e\n\u003c!-- |2025.8| [Gemma3-270M]() | [HF](https://huggingface.co/google/gemma-3-270m) | 0.27B | | | | --\u003e\n\u003c!-- |2025.7| [Qwen3]() |[HF](https://huggingface.co/collections/Qwen/qwen3) | 0.6~4B | apache2.0 |  |  |  | --\u003e\n\u003c!-- |2025.7| [Phi-4-mini-instruct]() |[HF](https://huggingface.co/microsoft/Phi-4-mini-instruct) | 3.8B | apache2.0 | 5T |  |  | --\u003e\n\u003c!-- |2025.5| [OLMo-2](https://huggingface.co/allenai) | [HF](https://huggingface.co/allenai/OLMo-2-0425-1B-Instruct) | 1B | | | | --\u003e\n\u003c!-- |2025.3| [Gemma3]() | [HF](https://huggingface.co/collections/google/gemma-3-release-67c6c6f89c4f76621268bb6d)| 1B,4B | [gemma](https://ai.google.dev/gemma/terms) | | | --\u003e\n\u003c!-- |2025.2| [SmolLM (Huggingface)](https://github.com/huggingface/smollm) | [HF](https://huggingface.co/collections/HuggingFaceTB/smollm2-6723884218bcda64b34d7db9)| 135M~1.7B| apache-2.0 | |  --\u003e\n\u003c!-- |2025.1| [TinySwallow-1.5B-Instruct]() | [HF](https://huggingface.co/SakanaAI/TinySwallow-1.5B-Instruct) | 1.5B | apache-2.0 | qwen | Japanese | TAID from Qwen2.5-32B| --\u003e\n\u003c!-- |2025.2| [Phi-4 mini](https://huggingface.co/microsoft/Phi-4-mini-instruct) |  | 3.8B | mit | | |  | --\u003e\n\u003c!-- |2025.2| [PLaMo 2 2B](https://tech.preferred.jp/ja/blog/plamo-2-2b/) | None | 2B | apache-2.0| |  | pruning, tested on HumanEval+ |\n|2025.2| [PLaMo 2 1B]() | [HF](https://huggingface.co/pfnet/plamo-2-1b) | 1B | apache-2.0| 4T (1.25T tokens Japanese)| | base model only | --\u003e\n\u003c!-- |2024.9| [Qwen2.5]() |[HF](https://huggingface.co/collections/Qwen/qwen25-66e81a666513e518adb90d9e) | 0.5,1.5,3B | apache2.0 |  |  |  | --\u003e\n\u003c!-- |2024.8| [Phi-3.5-mini-instruct]() |[HF](https://huggingface.co/microsoft/Phi-3.5-mini-instruct) | 3.8B | apache2.0 | MIT |  |  | --\u003e\n\n\n--- \n\n\u003ca id=\"medical-adaptation\"\u003e\u003c/a\u003e\n\n# Medical Domain Adaptation \n\n## Model\n|When? | Name |  HF?  | Size | License | pretraining | finetuning/continual | test | misc.|\n|---|---|---|---|---|---|---|---|---|\n|2026.6| [MeditronFO (EPFL)](https://huggingface.co/collections/EPFLiGHT/meditronfo) | [HF](https://huggingface.co/collections/EPFLiGHT/meditronfo) | 8~70B | apache-2.0, dataset is NonCommercial. | Apertus、OLMo、EuroLLM  | fullSFT with QA  |  fully open |\n|2026.6| [MedPsy (QVAC)]() | [HF](https://huggingface.co/qvac/MedPsy-1.7B) |  1.7, 4B | apache-2.0 | Qwen3 | SFT,RL  | QA, HealthBench |  |\n|2026.4| [ChatGPT for Clinicians (OpenAI)](https://chatgpt.com/plans/clinicians/) | None |  | | |   |  | ? |\n|2026.3| [SIP-jmed-llm-3-13b-OP-32k-R0.1]() | [HF](https://huggingface.co/SIP-med-LLM/SIP-jmed-llm-3-13b-OP-32k-R0.1) | 13B |  | [llm-jp-3-13b]() | [list](https://huggingface.co/SIP-med-LLM/SIP-jmed-llm-3-8x13b-AC-32k-instruct)  | - | japanese |\n|2026.3| [Med-V1]() | [HF](https://huggingface.co/ncbi/Med-V1-Q3B) | 3B | MIT | Qwen2.5/Llama3.2 |   |  |  |\n|2026.1| [ChatGPT Health (OpenAI)](https://openai.com/ja-JP/index/introducing-chatgpt-health/) | None |  | | |   |  | not a model |\n|2025.10| [SIP-jmed-llm-3-8x13b-AC-32k-instruct]() | [HF](https://huggingface.co/SIP-med-LLM/SIP-jmed-llm-3-8x13b-AC-32k-instruct) | 8x13B | CC BY-NC-SA 4.0 | [llm-jp-3-8x13b](https://huggingface.co/llm-jp/llm-jp-3-8x13b) | [list](https://huggingface.co/SIP-med-LLM/SIP-jmed-llm-3-8x13b-AC-32k-instruct)  | - | japanese |\n|2025.7| [ELYZA-Med-Base-1.0-Qwen2.5-72B](https://prtimes.jp/main/html/rd/p/000000061.000047565.html) | None | 72B | Qwen | Qwen2.5 |   | IgakuQA | japanese |\n|2025.5| [MedGemma (Google)](https://deepmind.google/models/gemma/medgemma/) |[HF](https://huggingface.co/collections/google/medgemma-release-680aade845f90bec6a3f60c4)| 1.5, 4, 27B | | Gemma3 | | | |\n|2025.4| [Med-R1 (IEEE)](https://arxiv.org/pdf/2503.13939v4) |[HF](https://huggingface.co/yuxianglai117/Med-R1)| 2B | | Qwen2-VL | | | VLM |\n|2025.4| [Med-R1 8B (IQVIA)](https://www.iqvia.com/blogs/2025/04/introducing-iqvia-medical-reasoning-med-r1-8b) | None | 8B | |  | | | reasoning |\n|2025.4| [OmniV-Med(Alibaba)](https://arxiv.org/abs/2504.14692) | | 1.5,7B |  | 252K instruction data|    | 11 benchmarks (2D/3D image and video) |  |\n|2025.4| [JPharmatron(EQUES)](https://arxiv.org/pdf/2505.16661) | [HF](https://huggingface.co/EQUES/JPharmatron-7B) | 7B | cc-by-sa-4.0 | Qwen2.5 | pharma corpus | None | Japanese, AACL2025 |\n|2025.2| [Preferred-MedLLM-Qwen-72B]() | [HF](https://huggingface.co/pfnet/Preferred-MedLLM-Qwen-72B) | 72B | Qwen | Qwen2.5 | original corpus   | IgakuQA | japanese |\n|2025.2| [OpenMeditron](https://huggingface.co/OpenMeditron)|[HF](https://huggingface.co/OpenMeditron/Meditron3-70B) | 7~70B | |||MedQA etc. | \n|2025.1| [Huatuo-o1](https://github.com/FreedomIntelligence/HuatuoGPT-o1)|[HF](https://huggingface.co/FreedomIntelligence/HuatuoGPT-o1-72B) | 72B | apache-2.0 | \n|2024.8| [LLaVA-Med++](https://github.com/UCSC-VLAA/MedTrinity-25M) | [HF](https://huggingface.co/MBZUAI/LLaVA-Meta-Llama-3-8B-Instruct-FT-S2) | 8B | ? | MedTrinity-25M | VQA-RAD etc. |   |  | |\n|2024.7| [MedLlama3-JP (EQUES)]() | [HF](https://huggingface.co/EQUES/MedLLama3-JP-v2) | 8B | Llama3 | Llama3 |   |  | japanese, merge model|\n|2024.7| [Llama3-Preferred-MedSwallow]() | [HF](https://huggingface.co/pfnet/Llama3-Preferred-MedSwallow-70B) | 70B | Llama3 | Llama3 |   |  | japanese |\n|2024.7| [Med42-v2]() | [HF](https://huggingface.co/m42-health/Llama3-Med42-70B) | 8,70B | Llama3 | llama3 |  ~1B tokens, including medical flashcards, exam questions, and open-domain dialogues. |  | |\n|2024.7| [JMedLLM-v1]() | [HF](https://huggingface.co/stardust-coder/jmedllm-7b-v1) | 7B | qwen | Qwen2 |   |  | japanese |\n|2024.6| [MedSwallow]() | [HF](https://huggingface.co/AIgroup-CVM-utokyohospital/MedSwallow-70b) | 70B | cc-by-nc-sa | Swallow |   |  | japanese |\n|2024.5| [MMed-LLama3-8B(上海交通大学)](https://github.com/MAGIC-AI4Med/MMedLM) | [HF](https://huggingface.co/Henrychur/MMed-Llama-3-8B) | 8B | cc-by-sa | Llama3 |   |  |  |\n|2024.5| [medX(JiviAI)]() | [HF](https://huggingface.co/jiviai/medX_v1) | 8B | Apache-2.0 | Llama3 |  100,000+ data, [ORPO](https://huggingface.co/blog/mlabonne/orpo-llama-3) |  |  |\n|2024.4| [UltraMedical(TsinghuaC3I)](https://arxiv.org/html/2406.03949v1) | [HF](https://huggingface.co/TsinghuaC3I) | 8B | - | Llama3 |  | | |\n|2024.4| [Meditron(EPFL)](https://www.meditron.io) | - | 8B | - | Llama3 |  | MedQA, MedMCQA, PubmedQA | SOTA |\n|2024.4| [OpenBioLLM]() | [HF](https://huggingface.co/aaditya/Llama3-OpenBioLLM-70B) | 8, 70B |  | Llama3 | |  | SOTA |\n|2024.4| [Med-Gemini(Google)](https://arxiv.org/pdf/2404.18416) | closed | ? | - | Gemini | | |multimodal|\n|2024.4| [Hippocrates](https://cyberiada.github.io/Hippocrates/) | [HF]() | 7B | |  | | |  | |\n|2024.3| [AdaptLLM(Microsoft Research)](https://github.com/microsoft/LMOps/tree/main/adaptllm) | [HF](https://huggingface.co/AdaptLLM/medicine-LLM-13B) | 7B, 13B | | reading comprehensive corpora | | |  | ICLR2024 |\n|2024.3| [Apollo](https://github.com/FreedomIntelligence/Apollo) | [HF](https://huggingface.co/FreedomIntelligence/Apollo-7B) | ~7B | | | | |  | multilingual |\n|2024.2| [BiMediX](https://arxiv.org/pdf/2402.13253) | [HF](https://huggingface.co/BiMediX) | non-commercial | 8x7B | mixtral8x7B | | | MoE |\n|2024.2| [Health-LLM(Rutgersなど)](https://arxiv.org/pdf/2402.00746.pdf) | | | | | | | RAG |\n|2024.2| [BioMistral](https://arxiv.org/pdf/2402.10373.pdf) | [HF](https://huggingface.co/BioMistral) | 7B | - |  |  |  | | \n|2024.1| [AMIE(Google)](https://arxiv.org/pdf/2401.05654.pdf) | not open | - | - | based on PaLM 2 |  |  | EHR| \n|2023.12| [Medprompt(Microsoft)]() | not open | - | - | GPT-4 | none |  |multi-modal| \n|2023.12| [JMedLoRA(UTokyo)](https://arxiv.org/abs/2310.10083) | [HF](https://huggingface.co/AIgroup-CVM-utokyohospital/llama2-jmedlora-3000) | 70B | none | none | QLoRA | IgakuQA | Japanese, insufficient quality | \n|2023.11| [Meditron(EPFL)](https://github.com/epfLLM/meditron) | [HF](https://huggingface.co/epfl-llm/meditron-70B) | 70B | Llama2 | Llama2 | GAP-Replay(48.1B) | [dataset](img/meditron-testdata.png),[score](img/meditron-eval2.png) | |\n|2023.8| [BioMedGPT(Luo et al.)](https://github.com/PharMolix/OpenBioMed) | [HF]() | 10B | |\n|2023.8| [PMC-LLaMa](https://github.com/chaoyi-wu/PMC-LLaMA)| [HF]() | 13B | |\n|2023.7| [Med-Flamingo](https://github.com/snap-stanford/med-flamingo) | [HF]() | 8.3B| ? | OpenFlamingo | MTB | Visual USMLE|based on Flamingo |\n|2023.7| [LLaVa-Med(Microsoft)](https://github.com/microsoft/LLaVA-Med) | [HF](https://huggingface.co/microsoft/llava-med-7b-delta) | 13B | - | LLaVa| medical dataset | VAQ-RAD, SLAKE, PathVQA |multi-modal| \n|2023.7| [Med-PaLM M(Google)](https://arxiv.org/abs/2307.14334) | not open | | - | PaLM2 | | |multi-modal| \n|2023.5| [Almanac(Stanford)](https://arxiv.org/pdf/2303.01229.pdf)| ? | ? | text-davinci-003 |  | | RAG |\n|2023.5| [Med-PaLM2(Google)](https://arxiv.org/abs/2305.09617) | not open | 340B | - | PaLM2 | | |\n|2022.12| [Med-PaLM(Google)](https://arxiv.org/abs/2212.13138) | not open | 540B| - | PaLM | | | |\n\n\nSee also \n- [Awesome-Healthcare-Foundation-Models](https://github.com/Jianing-Qiu/Awesome-Healthcare-Foundation-Models)\n- [Awesome-Medical-Large-Language-Models](https://github.com/burglarhobbit/Awesome-Medical-Large-Language-Models)\n- [Awesome-Medical-LLM](https://github.com/BARUDA-AI/Awesome-Medical-LLM)\n- [MedLLMsPracticalGuide](https://github.com/AI-in-Health/MedLLMsPracticalGuide).\n- [医療分野に特化したLLM紹介](https://speakerdeck.com/stardust11)\n\n\n## Leaderboard\n- [Medical LLM Leaderboard (~2024/11)](https://huggingface.co/spaces/fenglinliu/medical_llm_leaderboard)\n    - 22 LLMs in the clinic:\n    - Currently, Medical LLM Leaderboard covers 11 tasks, 7 metrics, 17 datasets, and over 20,000 test samples.\n- [Open Medical-LLM Leaderboard (down)](https://huggingface.co/spaces/openlifescienceai/open_medical_llm_leaderboard)\n- [MIRAGE Leaderboard](https://teddy-xionggz.github.io/MIRAGE/)\n- [MedHELM Leaderboard](https://crfm.stanford.edu/helm/medhelm/latest/#/leaderboard)\n- [PMC-Patients Leaderboard (2023)](https://pmc-patients.github.io/)\n- [MEDIC Leaderboard](https://huggingface.co/spaces/m42-health/MEDIC-Benchmark)\n\n## Dataset\n\nFor Japanese medical dataset, see [JMedData4LLM](https://github.com/stardust-coder/jmed-data-for-llm).\n\n\n### Corpus\n- [Healthcare Datasets (Aloe Beta)](https://huggingface.co/collections/HPAI-BSC/healthcare-datasets-aloe-beta-672374294ed56f43dc302499)\n- [PMC-Patients](https://huggingface.co/datasets/zhengyun21/PMC-Patients) from Pubmed Central, 167k patient summaries.\n\n### Evaluation benchmarks\n\n#### Pickups\n- [MMedBench](https://github.com/MAGIC-AI4Med/MMedLM)\n- [MedEval](https://arxiv.org/pdf/2310.14088)\n- [MEDIC](https://arxiv.org/pdf/2409.07314)\n- [CLIMB](https://github.com/DDVD233/CLIMB), multimodal\n- [MAST: Medical AI Superintelligence Test](https://bench.arise-ai.org/)\n- [Opencompass MedBench (Chinese)](https://medbench.opencompass.org.cn/home) and [arxiv](https://arxiv.org/pdf/2511.14439)\n\n\n#### Text-Only\n\n**classical medical benchmarks**\n- [MedQA (created from USMLE)](https://github.com/jind11/MedQA) \n- [MedMCQA](https://arxiv.org/abs/2203.14371)\n- [PubMedQA](https://arxiv.org/abs/1909.06146)\n\n\n**collections**\n\nDataset from [**FreedomIntelligence**](https://github.com/FreedomIntelligence)\n- [Medical O1 Reasoning](https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT)\n- [Medical O1 Verifiable Problem](https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem)\n- [Disease Database](https://huggingface.co/datasets/FreedomIntelligence/Disease_Database)\n- etc.\n\nDataset from [**OnDeviceMedNotes**](https://huggingface.co/OnDeviceMedNotes)\n- [synthetic-medical-conversations-deepseek-v3](https://huggingface.co/datasets/OnDeviceMedNotes/synthetic-medical-conversations-deepseek-v3)\n\n**Others**\n- [MMLU](https://github.com/hendrycks/test) : includes medicine and other related fields(clinical topics covering clinical knowledge,\ncollege biology, college medicine, medical genetics, professional medicine and anatomy)\n- [MMLUProX](https://mmluprox.github.io/)\n- [HealthsearchQA](https://huggingface.co/datasets/katielink/healthsearchqa) : 3173 samples, used in MedPaLM paper\n- [LiveQA](https://github.com/abachaa/LiveQA_MedicalTask_TREC2017) : 634+10, used in MedPaLM paper\n- [PubHealth](https://github.com/neemakot/Health-Fact-Checking)\n- [HeadQA](https://huggingface.co/datasets/dvilares/head_qa) : Spanish healthcare system\n- [K-Q\u0026A](https://github.com/Itaymanes/K-QA)\n- [MeDiSumQA](https://physionet.org/content/medisumqa/1.0.0/) : discharge summaries from the MIMIC-IV. released on Physionet.\n- [MedNLI](https://jgc128.github.io/mednli/) : MIMIC-III dataset, logical relationship between a premise and a hypothesis\n- [MeQSum](https://huggingface.co/datasets/albertvillanova/meqsum) : summarizing health queries\n- [LongHealth](https://github.com/kbressem/LongHealth) : 20 patient records, answer questions about them from a long document.\n- [Medical Eval Sphere](https://github.com/lavita-ai/medical-eval-sphere) : Long form medical questions\n- [MedCalcBench](https://github.com/ncbi-nlp/MedCalc-Bench):  [HF](https://huggingface.co/datasets/ncbi/MedCalc-Bench-v1.0)\n- [MedQA-Calc](https://huggingface.co/datasets/Nicholas-Wan/MedQA-Calc)\n- [MedS-Bench](https://huggingface.co/datasets/Henrychur/MedS-Bench)\n- [MedQuAD](https://github.com/abachaa/MedQuAD)\n- [TJH Dataset](https://github.com/HAIRLAB/Pre_Surv_COVID_19)\n- [MIMIC-IV](https://physionet.org/content/mimiciv/3.1/) : Sourced from the EHRs of the Beth Israel Deaconess Medical Center.\n- [ClinicBench](https://github.com/AI-in-Health/ClinicBench?tab=readme-ov-file) : 17 comprehensive benchmarks. \n- [EquityMedQA](https://huggingface.co/datasets/katielink/EquityMedQA) : Open-ended Q\u0026A for equity and bias mitigation.\n- [MedDistractQA](https://huggingface.co/datasets/KrithikV/MedDistractQA)\n- [AlpaCare-MedInstruct-52k](https://huggingface.co/datasets/lavita/AlpaCare-MedInstruct-52k)\n- [HealthBench (OpenAI)](https://openai.com/index/healthbench/)\n- [CLUE](https://github.com/TIO-IKIM/CLUE)\n- [CUREBench (Harvard)](https://github.com/mims-harvard/CUREBench)\n- [GlobMed](https://huggingface.co/collections/ruiyang-medinfo/globmed)\n- [OpenMed](https://huggingface.co/openmed-community)\n- [MultiMedX](https://huggingface.co/datasets/li-lab/MultiMed-X): 7 languages（BioNLI, [LiveQA](https://github.com/abachaa/LiveQA_MedicalTask_TREC2017)\n- [MedFact-Synth](https://huggingface.co/datasets/ncbi/MedFact-Synth): synthetic training set including 1.5 million instances\n\n#### Image + Text\n\n- [Clinical NLP 2023](https://clinical-nlp.github.io/2023/resources.html)\n- [PMC-15M](https://github.com/microsoft/BiomedCLIP_data_pipeline) : the largest biomedical image-text dataset\n- [LLaVA-Med Dataset](https://github.com/microsoft/LLaVA-Med/blob/main/README.md#data-download): used GPT-4 to generate diverse biomedical multimodal instruction-following data using image-text pairs from PMC-15M.\n- [PMC-OA](https://huggingface.co/datasets/axiong/pmc_oa) : 1.6M image-caption pairs\n- [MedICaT](https://github.com/allenai/medicat): image, caption, textual reference\n- [VQA-RAD](https://osf.io/89kps/) : 3515 question–answer pairs on 315 radiology images.\n- [SLAKE](https://huggingface.co/datasets/BoKelvin/SLAKE) : bilingual dataset (English\u0026Chinese) consisting of 642 images and 14,028 question-answer pairs\n- [PathVQA](https://huggingface.co/datasets/flaviagiammarino/path-vqa) : pathology image + caption\n- [MedVTE](https://github.com/ynklab/MedVTE): numeric understanding\n- [MedAlign(Stanford)](https://github.com/som-shahlab/medalign)\n- [MedEval](https://github.com/ZexueHe/MedEval)\n- [MedTrinity](https://github.com/UCSC-VLAA/MedTrinity-25M): 25M\n- [OmniMedVQA](https://huggingface.co/datasets/foreverbeliever/OmniMedVQA): 73 different medical datasets, contains 118,010 images with 127,995 QA-items, covering 12 different medical image modalities and referring to more than 20 human anatomical regions.\n- [MIMIC-ECG-IV](https://physionet.org/content/mimic-iv-ecg/) : ECG-caption dataset\n- [ECG-QA](https://github.com/Jwoo5/ecg-qa)\n- [CheXThought (Stanford, coming...)](https://aimi.stanford.edu/data)\n\n\nSee more on \n- [OpenLifeSciences (collection)](https://huggingface.co/openlifescienceai)\n- [MedLLMsPracticalGuide](https://github.com/AI-in-Health/MedLLMsPracticalGuide?tab=readme-ov-file#-practical-guide-for-medical-data)\n- [Medical datasets for LLMs (collection)](https://huggingface.co/collections/mfmezger/medical-datasets-for-llms-66bbb4e371cc405bc4d6b28a)\n- [Awesome-Medical-Dataset (~2025/1)](https://github.com/openmedlab/Awesome-Medical-Dataset).\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/stardust-coder%2Fawesome-latest-llm/projects"}