{"id":13476676,"url":"https://github.com/vikhyat/moondream","last_synced_at":"2025-05-15T00:00:36.457Z","repository":{"id":214598866,"uuid":"736812439","full_name":"vikhyat/moondream","owner":"vikhyat","description":"tiny vision language model","archived":false,"fork":false,"pushed_at":"2025-04-14T03:41:31.000Z","size":9521,"stargazers_count":7906,"open_issues_count":144,"forks_count":616,"subscribers_count":64,"default_branch":"main","last_synced_at":"2025-05-07T22:45:12.184Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"https://moondream.ai","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/vikhyat.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2023-12-29T00:27:18.000Z","updated_at":"2025-05-07T21:43:41.000Z","dependencies_parsed_at":null,"dependency_job_id":"f9f5cf10-960f-4bc7-b396-b3ecbeb2128c","html_url":"https://github.com/vikhyat/moondream","commit_stats":null,"previous_names":["vikhyat/moondream"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/vikhyat%2Fmoondream","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/vikhyat%2Fmoondream/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/vikhyat%2Fmoondream/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/vikhyat%2Fmoondream/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/vikhyat","download_url":"https://codeload.github.com/vikhyat/moondream/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254249199,"owners_count":22039029,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-07-31T16:01:33.265Z","updated_at":"2025-05-15T00:00:36.209Z","avatar_url":"https://github.com/vikhyat.png","language":"Python","funding_links":[],"categories":["Python","Jupyter Notebook","others","Repos","Voice \u0026 Multimodal (local) (16)","Foundation VLMs"],"sub_categories":["General-purpose open models"],"readme":"# 🌔 moondream\n\na tiny vision language model that kicks ass and runs anywhere\n\n[Website](https://moondream.ai/) | [Demo](https://moondream.ai/playground)\n\n## Examples\n\n| Image                  | Example                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |\n| ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| ![](assets/demo-1.jpg) | **What is the girl doing?**\u003cbr\u003eThe girl is sitting at a table and eating a large hamburger.\u003cbr\u003e\u003cbr\u003e**What color is the girl's hair?**\u003cbr\u003eThe girl's hair is white.                                                                                                                                                                                                                                                                                                                                                                                                    |\n| ![](assets/demo-2.jpg) | **What is this?**\u003cbr\u003eThis is a computer server rack, which is a device used to store and manage multiple computer servers. The rack is filled with various computer servers, each with their own dedicated space and power supply. The servers are connected to the rack via multiple cables, indicating that they are part of a larger system. The rack is placed on a carpeted floor, and there is a couch nearby, suggesting that the setup is in a living or entertainment area.\u003cbr\u003e\u003cbr\u003e**What is behind the stand?**\u003cbr\u003eBehind the stand, there is a brick wall. |\n\n## About\n\nMoondream is a highly efficient open-source vision language model that combines powerful image understanding capabilities with a remarkably small footprint. It's designed to be versatile and accessible, capable of running on a wide range of devices and platforms.\n\nThe project offers two model variants:\n\n- **Moondream 2B**: The primary model with 2 billion parameters, offering robust performance for general-purpose image understanding tasks including captioning, visual question answering, and object detection.\n- **Moondream 0.5B**: A compact 500 million parameter model specifically optimized as a distillation target for edge devices, enabling efficient deployment on resource-constrained hardware while maintaining impressive capabilities.\n\n## Getting Started\n\n### Latest Model Checkpoints\n\nThese are the latest bleeding-edge versions of both models, with all new features and improvements:\n\n| Model          | Precision | Download Size | Memory Usage | Best For                   | Download Link                                                                                                                                   |\n| -------------- | --------- | ------------- | ------------ | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |\n| Moondream 2B   | int8      | 1,733 MiB     | 2,624 MiB    | General use, best quality  | [Download](https://huggingface.co/vikhyatk/moondream2/resolve/9dddae84d54db4ac56fe37817aeaeb502ed083e2/moondream-2b-int8.mf.gz?download=true)   |\n| Moondream 0.5B | int8      | 593 MiB       | 996 MiB      | Edge devices, faster speed | [Download](https://huggingface.co/vikhyatk/moondream2/resolve/9dddae84d54db4ac56fe37817aeaeb502ed083e2/moondream-0_5b-int8.mf.gz?download=true) |\n\n### Python Client Library\n\nFirst, install the client library:\n\n```bash\npip install moondream==0.0.5\n```\n\nThe recommended way to use the latest version of Moondream is through our Python client library:\n\n```python\nimport moondream as md\nfrom PIL import Image\n\n# Initialize with local model path. Can also read .mf.gz files, but we recommend decompressing\n# up-front to avoid decompression overhead every time the model is initialized.\nmodel = md.vl(model=\"path/to/moondream-2b-int8.mf\")\n\n# Load and process image\nimage = Image.open(\"path/to/image.jpg\")\nencoded_image = model.encode_image(image)\n\n# Generate caption\ncaption = model.caption(encoded_image)[\"caption\"]\nprint(\"Caption:\", caption)\n\n# Ask questions\nanswer = model.query(encoded_image, \"What's in this image?\")[\"answer\"]\nprint(\"Answer:\", answer)\n```\n\n⚠️ Note: The Python client currently only supports CPU inference. CUDA (GPU) and MPS (Apple Silicon) optimization is coming soon. For GPU support, use the Hugging Face transformers implementation below.\n\nFor complete documentation of the Python client, including cloud API usage and additional features, see the [Python Client README](clients/python/README.md).\n\n### Node.js Client Library\n\nFor JavaScript/TypeScript developers, we offer a full-featured Node.js client library. See the [Node.js Client README](https://github.com/rohan-kulkarni-25/moondream/blob/main/clients/node/README.MD) for installation and usage instructions.\n\n### Hugging Face Transformers Integration\n\nThe Hugging Face hub version tracks the last official release of the 2B model. While more stable, it doesn't include the latest features or support for the 0.5B model. Use this if you need GPU acceleration or prefer the transformers ecosystem:\n\nFirst, install the required packages:\n\n```bash\npip install transformers torch einops\n```\n\n```python\nfrom transformers import AutoModelForCausalLM\nfrom PIL import Image\n\nmodel = AutoModelForCausalLM.from_pretrained(\n    \"vikhyatk/moondream2\",\n    revision=\"2025-01-09\",\n    trust_remote_code=True,\n    # Uncomment to run on GPU.\n    # device_map={\"\": \"cuda\"}\n)\n\nimage = Image.open('\u003cIMAGE_PATH\u003e')\nenc_image = model.encode_image(image)\nprint(model.query(enc_image, \"Describe this image.\"))\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fvikhyat%2Fmoondream","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fvikhyat%2Fmoondream","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fvikhyat%2Fmoondream/lists"}