{"id":16240618,"url":"https://github.com/NexaAI/Awesome-LLMs-on-device","last_synced_at":"2025-10-25T02:30:29.332Z","repository":{"id":250054692,"uuid":"821467413","full_name":"NexaAI/Awesome-LLMs-on-device","owner":"NexaAI","description":"Awesome LLMs on Device: A Comprehensive Survey","archived":false,"fork":false,"pushed_at":"2025-01-12T21:16:04.000Z","size":1370,"stargazers_count":939,"open_issues_count":1,"forks_count":101,"subscribers_count":51,"default_branch":"main","last_synced_at":"2025-02-01T15:01:49.765Z","etag":null,"topics":["awesome","awesome-list","llm","on-device"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/NexaAI.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":".github/CODEOWNERS","security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-06-28T15:50:38.000Z","updated_at":"2025-02-01T04:04:10.000Z","dependencies_parsed_at":"2024-07-24T23:32:56.842Z","dependency_job_id":"cca04bb6-60ee-449d-a686-225c2d67d73a","html_url":"https://github.com/NexaAI/Awesome-LLMs-on-device","commit_stats":null,"previous_names":["nexaai/awesome-llms-on-device"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NexaAI%2FAwesome-LLMs-on-device","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NexaAI%2FAwesome-LLMs-on-device/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NexaAI%2FAwesome-LLMs-on-device/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NexaAI%2FAwesome-LLMs-on-device/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/NexaAI","download_url":"https://codeload.github.com/NexaAI/Awesome-LLMs-on-device/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":238059129,"owners_count":19409601,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["awesome","awesome-list","llm","on-device"],"created_at":"2024-10-10T14:00:44.478Z","updated_at":"2025-10-25T02:30:24.293Z","avatar_url":"https://github.com/NexaAI.png","language":null,"funding_links":[],"categories":["NLP","Other Lists","Related Awesome Repositories"],"sub_categories":["TeX Lists","Papers"],"readme":"# 🚀 Awesome LLMs on Device: A Must-Read Comprehensive Hub by Nexa AI\n\n\u003cdiv align=\"center\"\u003e\n\n[![Discord](https://dcbadge.limes.pink/api/server/thRu2HaK4D?style=flat\u0026compact=true)](https://discord.gg/thRu2HaK4D)\n\n[On-device Model Hub](https://model-hub.nexa4ai.com/) / [Nexa SDK Documentation](https://docs.nexaai.com/)\n\n[release-url]: https://github.com/NexaAI/nexa-sdk/releases\n[Windows-image]: https://img.shields.io/badge/windows-0078D4?logo=windows\n[MacOS-image]: https://img.shields.io/badge/-MacOS-black?logo=apple\n[Linux-image]: https://img.shields.io/badge/-Linux-333?logo=ubuntu\n\n\u003c/div\u003e\n\n\n\n\u003cdiv style=\"text-align: center;\"\u003e\n  \u003cimg src=\"resources/Summary_of_on-device_LLMs_evolution.jpeg\" alt=\"Summary of on-device LLMs’ evolution\" width=\"800\"\u003e\n  \u003cdiv style=\"font-size: 10px;\"\u003eSummary of On-device LLMs’ Evolution\u003c/div\u003e\n\u003c/div\u003e\n\n\n\n## 🌟 About This Hub\nWelcome to the ultimate hub for on-device Large Language Models (LLMs)! This repository is your go-to resource for all things related to LLMs designed for on-device deployment. Whether you're a seasoned researcher, an innovative developer, or an enthusiastic learner, this comprehensive collection of cutting-edge knowledge is your gateway to understanding, leveraging, and contributing to the exciting world of on-device LLMs.\n\n## 🚀 Why This Hub is a Must-Read\n- 📊 Comprehensive overview of on-device LLM evolution with easy-to-understand visualizations\n- 🧠 In-depth analysis of groundbreaking architectures and optimization techniques\n- 📱 Curated list of state-of-the-art models and frameworks ready for on-device deployment\n- 💡 Practical examples and case studies to inspire your next project\n- 🔄 Regular updates to keep you at the forefront of rapid advancements in the field\n- 🤝 Active community of researchers and practitioners sharing insights and experiences\n\n\n  \n# 📚 What's Inside Our Hub\n- [Awesome LLMs on Device: A Comprehensive Survey](#-awesome-llms-on-device-a-must-read-comprehensive-hub)\n- [Contents](-whats-inside-our-hub)\n  - [Foundations and Preliminaries](#foundations-and-preliminaries)\n    - [Evolution of On-Device LLMs](#evolution-of-on-device-llms)\n    - [LLM Architecture Foundations](#llm-architecture-foundations)\n    - [On-Device LLMs Training](#on-device-llms-training)\n    - [Limitations of Cloud-Based LLM Inference and Advantages of On-Device Inference](#limitations-of-cloud-based-llm-inference-and-advantages-of-on-device-inference)\n    - [The Performance Indicator of On-Device LLMs](#the-performance-indicator-of-on-device-llms)\n  - [Efficient Architectures for On-Device LLMs](#efficient-architectures-for-on-device-llms)\n    - [Model Compression and Parameter Sharing](#model-compression-and-parameter-sharing)\n    - [Collaborative and Hierarchical Model Approaches](#collaborative-and-hierarchical-model-approaches)\n    - [Memory and Computational Efficiency](#memory-and-computational-efficiency)\n    - [Mixture-of-Experts (MoE) Architectures](#mixture-of-experts-moe-architectures)\n    - [Hybrid Architectures](#hybrid-architectures)\n    - [General Efficiency and Performance Improvements](#general-efficiency-and-performance-improvements)\n  - [Model Compression and Optimization Techniques for On-Device LLMs](#model-compression-and-optimization-techniques-for-on-device-llms)\n    - [Quantization](#quantization)\n    - [Pruning](#pruning)\n    - [Knowledge Distillation](#knowledge-distillation)\n    - [Low-Rank Factorization](#low-rank-factorization)\n  - [Hardware Acceleration and Deployment Strategies](#hardware-acceleration-and-deployment-strategies)\n    - [Popular On-Device LLMs Framework](#popular-on-device-llms-framework)\n    - [Hardware Acceleration](#hardware-acceleration)\n  - [Applications](#applications)\n- [Tutorials and Learning Resources](#tutorials-and-learning-resources)\n- [Citation](#-cite-our-work)\n\n## Foundations and Preliminaries\n\n### Evolution of On-Device LLMs\n\n- Tinyllama: An open-source small language model \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2401.02385) [[Github]](https://github.com/jzhang38/TinyLlama)\n- MobileVLM V2: Faster and Stronger Baseline for Vision Language Model \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2402.03766) [[Github]](https://github.com/Meituan-AutoML/MobileVLM)\n- MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2406.10290)\n- Octopus series papers \u003cbr\u003e arXiv 2024 [[Octopus]](https://arxiv.org/abs/2404.01549) [[Octopus v2]](https://arxiv.org/abs/2404.01744) [[Octopus v3]](https://arxiv.org/abs/2404.11459) [[Octopus v4]](https://arxiv.org/abs/2404.19296) [[Github]](https://github.com/NexaAI)\n- The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2402.17764)\n- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration \u003cbr\u003e arXiv 2023 [[Paper]](https://arxiv.org/abs/2306.00978) [[Github]](https://github.com/mit-han-lab/llm-awq)\n- Small Language Models: Survey, Measurements, and Insights \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/pdf/2409.15790)\n\n\n### LLM Architecture Foundations\n\n- The case for 4-bit precision: k-bit inference scaling laws \u003cbr\u003e ICML 2023 [[Paper]](https://arxiv.org/abs/2212.09720)\n- Challenges and applications of large language models \u003cbr\u003e arXiv 2023 [[Paper]](https://arxiv.org/abs/2307.10169)\n- MiniLLM: Knowledge distillation of large language models \u003cbr\u003e ICLR 2023 [[Paper]](https://arxiv.org/abs/2306.08543) [[github]](https://github.com/Tebmer/Awesome-Knowledge-Distillation-of-LLMs)\n- Gptq: Accurate post-training quantization for generative pre-trained transformers \u003cbr\u003e ICLR 2023 [[Paper]](https://arxiv.org/abs/2210.17323) [[Github]](https://github.com/IST-DASLab/gptq)\n- Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale \u003cbr\u003e NeurIPS 2022 [[Paper]](https://arxiv.org/abs/2208.07339)\n\n### On-Device LLMs Training\n\n- OpenELM: An Efficient Language Model Family with Open Training and Inference Framework \u003cbr\u003e ICML 2024 [[Paper]](https://arxiv.org/abs/2404.14619) [[Github]](https://github.com/apple/corenet)\n\n### Limitations of Cloud-Based LLM Inference and Advantages of On-Device Inference\n\n- Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2404.07973)\n- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2404.14219)\n- Exploring post-training quantization in llms from comprehensive study to low rank compensation \u003cbr\u003e AAAI 2024 [[Paper]](https://arxiv.org/abs/2303.08302)\n- Matrix compression via randomized low rank and low precision factorization \u003cbr\u003e NeurIPS 2023 [[Paper]](https://arxiv.org/abs/2310.11028) [[Github]](https://github.com/pilancilab/matrix-compressor)\n\n### The Performance Indicator of On-Device LLMs\n\n- MNN: A lightweight deep neural network inference engine \u003cbr\u003e 2024 [[Github]](https://github.com/alibaba/MNN)\n- PowerInfer-2: Fast Large Language Model Inference on a Smartphone \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2406.06282) [[Github]](https://github.com/SJTU-IPADS/PowerInfer)\n- llama.cpp: Lightweight library for Approximate Nearest Neighbors and Maximum Inner Product Search \u003cbr\u003e 2023 [[Github]](https://github.com/ggerganov/llama.cpp)\n- Powerinfer: Fast large language model serving with a consumer-grade gpu \u003cbr\u003e arXiv 2023 [[Paper]](https://arxiv.org/abs/2312.12456) [[Github]](https://github.com/SJTU-IPADS/PowerInfer)\n\n## Efficient Architectures for On-Device LLMs\n\n| Model                           | Performance                                         | Computational Efficiency                                                    | Memory Requirements                                               |\n|---------------------------------|-----------------------------------------------------|----------------------------------------------------------------------------|-------------------------------------------------------------------|\n| **[MobileLLM](https://arxiv.org/abs/2402.14905)** | High accuracy, optimized for sub-billion parameter models | Embedding sharing, grouped-query attention                                  | Reduced model size due to deep and thin structures                 |\n| **[EdgeShard](https://arxiv.org/abs/2405.14371)** | Up to 50% latency reduction, 2× throughput improvement | Collaborative edge-cloud computing, optimal shard placement                  | Distributed model components reduce individual device load         |\n| **[LLMCad](https://arxiv.org/abs/2309.04255)**      | Up to 9.3× speedup in token generation              | Generate-then-verify, token tree generation                                 | Smaller LLM for token generation, larger LLM for verification      |\n| **[Any-Precision LLM](https://arxiv.org/abs/2402.10517)** | Supports multiple precisions efficiently            | Post-training quantization, memory-efficient design                         | Substantial memory savings with versatile model precisions         |\n| **[Breakthrough Memory](https://ieeexplore.ieee.org/abstract/document/10477465)** | Up to 4.5× performance improvement                  | PIM and PNM technologies enhance memory processing                          | Enhanced memory bandwidth and capacity                             |\n| **[MELTing Point](https://arxiv.org/abs/2403.12844)** | Provides systematic performance evaluation          | Analyzes impacts of quantization, efficient model evaluation                | Evaluates memory and computational efficiency trade-offs           |\n| **[LLMaaS on device](https://arxiv.org/abs/2403.11805)** | Reduces context switching latency significantly     | Stateful execution, fine-grained KV cache compression                       | Efficient memory management with tolerance-aware compression and swapping |\n| **[LocMoE](https://arxiv.org/abs/2401.13920)**     | Reduces training time per epoch by up to 22.24%     | Orthogonal gating weights, locality-based expert regularization             | Minimizes communication overhead with group-wise All-to-All and recompute pipeline |\n| **[EdgeMoE](https://arxiv.org/abs/2308.14352)**     | Significant performance improvements on edge devices | Expert-wise bitwidth adaptation, preloading experts                         | Efficient memory management through expert-by-expert computation reordering |\n|**[JetMoE](https://arxiv.org/abs/2404.07413)**| Outperforms Llama27B and 13B-Chat with fewer parameters | Reduces inference computation by 70% using sparse activation | 8B total parameters, only 2B activated per input token |\n|**[Pangu-$`\\pi`$ Pro](https://arxiv.org/abs/2402.02791)**| Neural architecture, parameter initialization, and optimization strategy for billion-level parameter models | Embedding sharing, tokenizer compression | Reduced model size via architecture tweaking |\n|**[Zamba2](https://www.zyphra.com/post/zamba2-small)**| 2x faster time-to-first-token, a 27% reduction in memory overhead, and a 1.29x lower generation latency compared to Phi3-3.8B. | Hybrid Mamba2/Attention architecture and shared transformer block | 2.7B parameters, fewer KV-states due to reduced attention |\n\n\n### Model Compression and Parameter Sharing\n\n- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2306.00978) [[Github]](https://github.com/mit-han-lab/llm-awq)\n- MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2402.14905) [[Github]](https://github.com/facebookresearch/MobileLLM)\n\n### Collaborative and Hierarchical Model Approaches\n\n- EdgeShard: Efficient LLM Inference via Collaborative Edge Computing \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2405.14371)\n- Llmcad: Fast and scalable on-device large language model inference \u003cbr\u003e arXiv 2023 [[Paper]](https://arxiv.org/abs/2309.04255)\n\n### Memory and Computational Efficiency\n\n- The Breakthrough Memory Solutions for Improved Performance on LLM Inference \u003cbr\u003e IEEE Micro 2024 [[Paper]](https://ieeexplore.ieee.org/document/10477465)\n- MELTing point: Mobile Evaluation of Language Transformers \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2403.12844) [[Github]](https://github.com/brave-experiments/MELT-public)\n\n### Mixture-of-Experts (MoE) Architectures\n\n- LLM as a system service on mobile devices \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2403.11805)\n- Locmoe: A low-overhead moe for large language model training \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2401.13920)\n- Edgemoe: Fast on-device inference of moe-based large language models \u003cbr\u003e arXiv 2023 [[Paper]](https://arxiv.org/abs/2308.14352)\n\n### Hybrid Architectures\n\n- Zamba2: Hybrid Mamba2 and attention models for on-device \u003cbr\u003e 2024 [[Zamba2-2.7B]](https://www.zyphra.com/post/zamba2-small) [[Zamba2-1.2B]](https://www.zyphra.com/post/zamba2-mini)\n\n### General Efficiency and Performance Improvements\n\n- Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs \u003cbr\u003e arXiv 2024 [[Paper]](https://www.arxiv.org/pdf/2402.10517) [[Github]](https://github.com/SNU-ARC/any-precision-llm)\n- On the viability of using llms for sw/hw co-design: An example in designing cim dnn accelerators \u003cbr\u003eIEEE SOCC 2023 [[Paper]](https://arxiv.org/abs/2306.06923)\n\n## Model Compression and Optimization Techniques for On-Device LLMs\n\n### Quantization\n\n- The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2402.17764)\n- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration \u003cbr\u003e arXiv 2024 [[Paper]](https://arxiv.org/abs/2306.00978) [[Github]](https://github.com/mit-han-lab/llm-awq)\n- Gptq: Accurate post-training quantization for generative pre-trained transformers \u003cbr\u003e ICLR 2023 [[Paper]](https://arxiv.org/abs/2210.17323) [[Github]](https://github.com/IST-DASLab/gptq)\n- Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale \u003cbr\u003e NeurIPS 2022 [[Paper]](https://arxiv.org/abs/2208.07339)\n\n### Pruning\n\n- Challenges and applications of large language models \u003cbr\u003e arXiv 2023 [[Paper]](https://arxiv.org/abs/2307.10169)\n\n### Knowledge Distillation\n\n- MiniLLM: Knowledge distillation of large language models \u003cbr\u003e ICLR 2024 [[Paper]](https://arxiv.org/abs/2306.08543)\n\n### Low-Rank Factorization\n\n- Exploring post-training quantization in llms from comprehensive study to low rank compensation \u003cbr\u003e AAAI 2024 [[Paper]](https://arxiv.org/abs/2303.08302)\n- Matrix compression via randomized low rank and low precision factorization \u003cbr\u003e NeurIPS 2023 [[Paper]](https://arxiv.org/abs/2310.11028) [[Github]](https://github.com/pilancilab/matrix-compressor)\n\n## Hardware Acceleration and Deployment Strategies\n\n### Popular On-Device LLMs Framework\n\n- llama.cpp: A lightweight library for efficient LLM inference on various hardware with minimal setup. [[Github]](https://github.com/ggerganov/llama.cpp)\n- MNN: A blazing fast, lightweight deep learning framework. [[Github]](https://github.com/alibaba/MNN)\n- PowerInfer: A CPU/GPU LLM inference engine leveraging activation locality for device. [[Github]](https://github.com/SJTU-IPADS/PowerInfer)\n- ExecuTorch: A platform for On-device AI across mobile, embedded and edge for PyTorch. [[Github]](https://github.com/pytorch/executorch)\n- MediaPipe: A suite of tools and libraries, enables quick application of AI and ML techniques. [[Github]](https://github.com/google-ai-edge/mediapipe)\n- MLC-LLM: A machine learning compiler and high-performance deployment engine for large language models. [[Github]](https://github.com/mlc-ai/mlc-llm)\n- VLLM: A fast and easy-to-use library for LLM inference and serving. [[Github]](https://github.com/vllm-project/vllm)\n- OpenLLM: An open platform for operating large language models (LLMs) in production. [[Github]](https://python.langchain.com/v0.2/docs/integrations/llms/openllm/)\n- mllm: Fast and lightweight multimodal LLM inference engine for mobile and edge devices. [[Github]](https://github.com/UbiquitousLearning/mllm)\n\n\n### Hardware Acceleration\n\n- The Breakthrough Memory Solutions for Improved Performance on LLM Inference \u003cbr\u003e IEEE Micro 2024 [[Paper]](https://ieeexplore.ieee.org/document/10477465)\n- Aquabolt-XL: Samsung HBM2-PIM with in-memory processing for ML accelerators and beyond \u003cbr\u003e IEEE Hot Chips 2021 [[Paper]](https://ieeexplore.ieee.org/abstract/document/9567191)\n\n## Applications\n- Text Generating For Messaging: [Gboard smart reply](https://developer.android.com/ai/aicore#gboard-smart)\n- Translation: [LLMCad](https://arxiv.org/abs/2309.04255)\n- Meeting Summarizing\n- Healthcare application: [BioMistral-7B](https://arxiv.org/abs/2402.10373), [HuatuoGPT](https://arxiv.org/abs/2311.09774)\n- Research Support\n- Companion Robot\n- Disability Support: [Octopus v3](https://arxiv.org/abs/2404.11459), [Talkback with Gemini Nano](https://store.google.com/intl/en/ideas/articles/gemini-nano-google-pixel/) \n- Autonomous Vehicles: [DriveVLM](https://arxiv.org/abs/2402.12289)\n\n## Model Reference\n\n|         Model         |      Institute      | Paper                                                                                                                                                                                                                                                                                                                                                                                                                 |\n| :-------------------: | :-----------------: | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n|      Gemini Nano      |       Google        | [Gemini: A Family of Highly Capable Multimodal Models](https://arxiv.org/pdf/2312.11805.pdf)                                                                                                                                                                                                                                                                                                                          |\n| Octopus series model  |       Nexa AI       | [Octopus v2: On-device language model for super agent](https://arxiv.org/pdf/2404.01744.pdf)\u003cbr\u003e[Octopus v3: Technical Report for On-device Sub-billion Multimodal AI Agent](https://arxiv.org/pdf/2404.11459.pdf)\u003cbr\u003e[Octopus v4: Graph of language models](https://arxiv.org/pdf/2404.19296.pdf)\u003cbr\u003e[Octopus: On-device language model for function calling of software APIs](https://arxiv.org/pdf/2404.01549.pdf) |\n| OpenELM and Ferret-v2 |        Apple        | [OpenELM is a significant large language model integrated within iOS to enhance application functionalities.](https://arxiv.org/abs/2404.14619) \u003cbr\u003e[Ferret-v2 significantly improves upon its predecessor, introducing enhanced visual processing capabilities and an advanced training regimen.](https://arxiv.org/abs/2404.07973)                                                                                                                                                          |\n|      Phi series       |      Microsoft      | [Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone](https://arxiv.org/pdf/2404.14219.pdf)                                                                                                                                                                                                                                                                                                 |\n|        MiniCPM        | Tsinghua University | [A GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone](https://huggingface.co/openbmb/MiniCPM-V-2_6)                                                                                                                                                                                                                                                                                                                    |\n|       Gemma2-9B       |       Google        | [Gemma 2: Improving Open Language Models at a Practical Size](https://storage.googleapis.com/deepmind-media/gemma/gemma-2-report.pdf)                                                                                                                                                                                                                                                                                 |\n|      Qwen2-0.5B       |    Alibaba Group    | [Qwen Technical Report](https://arxiv.org/pdf/2309.16609.pdf)                                                                                                                                                                                                                                                                                                                                                         |\n|      GLM-Edge       |    THUDM    | [GLM-Edge Github Page](https://github.com/THUDM/GLM-Edge)                                                                                                                                                                                                                                                                                                                                                         |\n\n## Tutorials and Learning Resources\n\n- MIT: [TinyML and Efficient Deep Learning Computing](https://efficientml.ai)\n- Harvard: [Machine Learning Systems](https://mlsysbook.ai/)\n- Deep Learning AI : [Introduction to on-device AI](https://www.deeplearning.ai/short-courses/introduction-to-on-device-ai/)\n\n# 🤝 Join the On-Device LLM Revolution\n\nWe believe in the power of community! If you're passionate about on-device AI and want to contribute to this ever-growing knowledge hub, here's how you can get involved:\n1. Fork the repository\n2. Create a new branch for your brilliant additions\n3. Make your updates and push your changes\n4. Submit a pull request and become part of the on-device LLM movement\n\n# ⭐ Star History ⭐\n\n[![Star History Chart](https://api.star-history.com/svg?repos=NexaAI/Awesome-LLMs-on-device\u0026type=Timeline)](https://star-history.com/#NexaAI/Awesome-LLMs-on-device\u0026Timeline)\n   \n#  📖 Cite Our Work\nIf our hub fuels your research or powers your projects, we'd be thrilled if you could cite our paper [here](https://arxiv.org/abs/2409.00088):\n\n```bibtex\n@article{xu2024device,\n  title={On-Device Language Models: A Comprehensive Review},\n  author={Xu, Jiajun and Li, Zhiyuan and Chen, Wei and Wang, Qun and Gao, Xin and Cai, Qi and Ling, Ziyuan},\n  journal={arXiv preprint arXiv:2409.00088},\n  year={2024}\n}\n```\n\n# 📄 License\n\nThis project is open-source and available under the MIT License. See the [LICENSE](LICENSE) file for more details.\n\nDon't just read about the future of AI – be part of it. Star this repo, spread the word, and let's push the boundaries of on-device LLMs together! 🚀🌟\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FNexaAI%2FAwesome-LLMs-on-device","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FNexaAI%2FAwesome-LLMs-on-device","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FNexaAI%2FAwesome-LLMs-on-device/lists"}