{"id":25238694,"url":"https://github.com/theoddysey/visual-answering-transformers-model","last_synced_at":"2026-05-10T03:55:21.644Z","repository":{"id":276995438,"uuid":"928828506","full_name":"TheODDYSEY/Visual-Answering-Transformers-Model","owner":"TheODDYSEY","description":"An Intelligent Image Analysis Application using BLIP and Streamlit","archived":false,"fork":false,"pushed_at":"2025-02-11T14:52:27.000Z","size":49265,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-02-11T15:37:55.383Z","etag":null,"topics":["blip","huggingface-transformers","image-processing","image-to-text","python3","streamlit"],"latest_commit_sha":null,"homepage":"https://theoddysey-visual-answering-transformers-m-streamlit-app-vimg1u.streamlit.app/","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/TheODDYSEY.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2025-02-07T10:07:52.000Z","updated_at":"2025-02-11T14:52:31.000Z","dependencies_parsed_at":"2025-02-11T15:48:03.676Z","dependency_job_id":null,"html_url":"https://github.com/TheODDYSEY/Visual-Answering-Transformers-Model","commit_stats":null,"previous_names":["theoddysey/visual-answering-transformers-model"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/TheODDYSEY%2FVisual-Answering-Transformers-Model","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/TheODDYSEY%2FVisual-Answering-Transformers-Model/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/TheODDYSEY%2FVisual-Answering-Transformers-Model/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/TheODDYSEY%2FVisual-Answering-Transformers-Model/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/TheODDYSEY","download_url":"https://codeload.github.com/TheODDYSEY/Visual-Answering-Transformers-Model/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":238343258,"owners_count":19456202,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["blip","huggingface-transformers","image-processing","image-to-text","python3","streamlit"],"created_at":"2025-02-11T17:53:46.201Z","updated_at":"2026-05-10T03:55:21.610Z","avatar_url":"https://github.com/TheODDYSEY.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"\n\u003cdiv align=\"center\"\u003e \n  \u003ch1 align=\"center\"\u003e📸 Visual-Answering-Transformers-Model \u003c/h1\u003e  \n  \u003ch3 align=\"center\"\u003eAn intelligent image analysis application using BLIP and Streamlit\u003c/h3\u003e  \n  \u003cp align=\"center\"\u003eUpload an image, ask questions, and get AI-powered answers instantly.\u003c/p\u003e  \n\n\u003cdiv\u003e  \n    \u003cimg src=\"https://img.shields.io/badge/-Python-blue?style=for-the-badge\u0026logo=python\u0026logoColor=white\u0026color=3776AB\" alt=\"Python\" /\u003e  \n    \u003cimg src=\"https://img.shields.io/badge/-Streamlit-red?style=for-the-badge\u0026logo=streamlit\u0026logoColor=white\u0026color=FF4B4B\" alt=\"Streamlit\" /\u003e  \n    \u003cimg src=\"https://img.shields.io/badge/-Hugging_Face-yellow?style=for-the-badge\u0026logo=huggingface\u0026logoColor=black\u0026color=FFD700\" alt=\"Hugging Face\" /\u003e  \n  \u003c/div\u003e  \n\n  \u003ca href=\"\" target=\"_blank\"\u003e  \n    \u003cimg src=\"./project.png\" alt=\"Project Banner\" /\u003e  \n  \u003c/a\u003e  \n  \u003cbr /\u003e  \n\n\n\u003c/div\u003e  \n\n---\n\n## 📋 \u003ca name=\"table\"\u003eTable of Contents\u003c/a\u003e\n\n1. 🤖 [Introduction](#introduction)  \n2. ⚙️ [Tech Stack](#tech-stack)  \n3. 🔋 [Features](#features)  \n4. 🚀 [Quick Start](#quick-start)  \n5. 🕸️ [Code Snippets](#snippets)  \n6. 🔗 [Links](#links)  \n7. 📌 [More](#more)  \n\n---\n\n## \u003ca name=\"introduction\"\u003e🤖 Introduction\u003c/a\u003e\n\nThe **AI Image Question Answering** project is a **deep learning-powered web application** designed to process images and provide meaningful answers to user-posed questions. Built with **Streamlit**, this app leverages **Salesforce’s BLIP (Bootstrapped Language-Image Pretraining) model** to extract relevant insights from images.  \n\nKey functionalities include:  \n\n- **Uploading images (JPEG, PNG) for analysis.**  \n- **Asking both predefined and custom questions about the image.**  \n- **Generating AI-based answers using the BLIP model.**  \n- **Providing real-time inference with GPU acceleration (if available).**  \n- **Delivering a seamless and interactive user experience with Streamlit.**  \n\nWhether you're a researcher, developer, or AI enthusiast, this project serves as an excellent **introduction to multimodal AI** and **visual question answering (VQA) applications**.  \n\n---\n\n## \u003ca name=\"tech-stack\"\u003e⚙️ Tech Stack\u003c/a\u003e\n\n- **Python**  \n- **Streamlit**  \n- **Hugging Face Transformers**  \n- **PyTorch**  \n- **PIL (Pillow)**  \n\n---\n\n## \u003ca name=\"features\"\u003e🔋 Features\u003c/a\u003e\n\n👉 **AI-Powered Image Question Answering**: Upload an image and ask any question related to its content.  \n\n👉 **Predefined Questions for Instant Insights**: Select from a list of commonly asked questions for a quick analysis.  \n\n👉 **Custom Question Input**: Type your own question to get AI-generated responses tailored to your query.  \n\n👉 **Real-Time Processing with AI Feedback**: Experience **instantaneous results** with **dynamic loading indicators** while the AI processes your request.  \n\n👉 **Seamless Streamlit UI**: Intuitive **drag-and-drop** image upload and interactive **question submission** for smooth user experience.  \n\n👉 **Optimized for GPU Acceleration**: The model runs efficiently on CUDA-enabled devices for **faster inference**.  \n\n---\n\n## \u003ca name=\"quick-start\"\u003e🚀 Quick Start\u003c/a\u003e\n\nFollow these steps to set up and run the project on your local machine.  \n\n### **Prerequisites**  \n\nEnsure you have the following installed:  \n\n- [Python 3.8+](https://www.python.org/downloads/)  \n- [pip](https://pip.pypa.io/en/stable/installation/)  \n\n### **Cloning the Repository**  \n\n```bash\ngit clone https://github.com/TheODDYSEY/Visual-Answering-Transformers-Model.git\ncd ai-image-question-answering\n```\n\n### **Installation**  \n\nInstall all required dependencies using:  \n\n```bash\npip install -r requirements.txt\n```\n\n### **Running the Project**  \n\n```bash\nstreamlit run streamlit_app.py\n```\n\nOpen **[http://localhost:8501](http://localhost:8501)** in your browser to interact with the application.  \n\n---\n\n## \u003ca name=\"snippets\"\u003e🕸️ Code Snippets\u003c/a\u003e\n\n### **1️⃣ Loading the BLIP Model**  \n\n```python\nfrom transformers import BlipProcessor, BlipForQuestionAnswering\nimport torch\n\nDEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"\n\nprocessor = BlipProcessor.from_pretrained(\"Salesforce/blip-vqa-base\")\nmodel = BlipForQuestionAnswering.from_pretrained(\"Salesforce/blip-vqa-base\").to(DEVICE)\n```\n\n### **2️⃣ Processing Images and Questions**  \n\n```python\ndef get_answer(image, question):\n    \"\"\"Processes an image and question to return an AI-generated response.\"\"\"\n    inputs = processor(image, question, return_tensors=\"pt\").to(DEVICE)\n    output = model.generate(**inputs)\n    return processor.decode(output[0], skip_special_tokens=True)\n```\n\n### **3️⃣ Implementing Streamlit UI**  \n\n```python\nimport streamlit as st\nfrom PIL import Image\n\nst.title(\"📸 AI Image Question Answering\")\n\nuploaded_image = st.file_uploader(\"Upload an image\", type=[\"jpg\", \"png\"])\n\nif uploaded_image is not None:\n    image = Image.open(uploaded_image)\n    st.image(image, caption=\"Uploaded Image\", use_column_width=True)\n\n    question = st.text_input(\"Ask a question about the image:\")\n    if st.button(\"Get Answer\"):\n        answer = get_answer(image, question)\n        st.success(f\"**Q:** {question}\")\n        st.write(f\"**A:** {answer}\")\n```\n\n---\n\n## \u003ca name=\"links\"\u003e🔗 Links\u003c/a\u003e\n\n- 🔗 **Live Demo**: [Try it here](https://theoddysey-visual-answering-transformers-m-streamlit-app-vimg1u.streamlit.app/)  \n- 📜 **Project Repository**: [GitHub](https://github.com/TheODDYSEY/Visual-Answering-Transformers-Model.git)  \n- 📚 **BLIP Model**: [Hugging Face](https://huggingface.co/Salesforce/blip-vqa-base)  \n\n---\n\n## \u003ca name=\"more\"\u003e📌 More\u003c/a\u003e\n\n🔹 **Future Enhancements**  \n- ✅ Improve model inference speed for real-time responses.  \n- ✅ Extend support for **OCR-based text recognition**.  \n- ✅ Enhance UI with **interactive visualization** options.  \n\n## **License**  \nThis project is **open-source** under the **MIT License**.  \n\n---\n ","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftheoddysey%2Fvisual-answering-transformers-model","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Ftheoddysey%2Fvisual-answering-transformers-model","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftheoddysey%2Fvisual-answering-transformers-model/lists"}