{"id":29792278,"url":"https://github.com/armanjscript/hybrid-rag-chatbot","last_synced_at":"2026-04-07T09:31:03.634Z","repository":{"id":305862478,"uuid":"1024181589","full_name":"armanjscript/Hybrid-RAG-chatbot","owner":"armanjscript","description":"A powerful web-based application designed to answer questions based on the content of uploaded PDF documents. This project leverages a Hybrid Retrieval-Augmented Generation (RAG) approach, combining the strengths of vector-based semantic search and keyword-based search to deliver accurate and relevant responses","archived":false,"fork":false,"pushed_at":"2025-07-26T21:00:12.000Z","size":13,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2025-09-10T09:13:14.224Z","etag":null,"topics":["bm25","chroma","chromadb","ensemble-retriever","hybrid-rag","langchain","langchain-ollama","ollama","ollama-embeddings","pypdf","qwen2-5","rag","rag-chatbot","streamlit"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/armanjscript.png","metadata":{"files":{"readme":"README.markdown","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-07-22T09:55:47.000Z","updated_at":"2025-07-26T21:00:15.000Z","dependencies_parsed_at":"2025-07-22T12:26:10.218Z","dependency_job_id":null,"html_url":"https://github.com/armanjscript/Hybrid-RAG-chatbot","commit_stats":null,"previous_names":["armanjscript/hybrid-rag-chatbot"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/armanjscript/Hybrid-RAG-chatbot","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/armanjscript%2FHybrid-RAG-chatbot","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/armanjscript%2FHybrid-RAG-chatbot/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/armanjscript%2FHybrid-RAG-chatbot/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/armanjscript%2FHybrid-RAG-chatbot/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/armanjscript","download_url":"https://codeload.github.com/armanjscript/Hybrid-RAG-chatbot/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/armanjscript%2FHybrid-RAG-chatbot/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":31507904,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-04-07T03:10:19.677Z","status":"ssl_error","status_checked_at":"2026-04-07T03:10:13.982Z","response_time":105,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.5:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["bm25","chroma","chromadb","ensemble-retriever","hybrid-rag","langchain","langchain-ollama","ollama","ollama-embeddings","pypdf","qwen2-5","rag","rag-chatbot","streamlit"],"created_at":"2025-07-28T01:06:12.844Z","updated_at":"2026-04-07T09:31:03.613Z","avatar_url":"https://github.com/armanjscript.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Hybrid RAG Chatbot with Streamlit, LangChain, Chroma, and Ollama\n\n[![GitHub Stars](https://img.shields.io/github/stars/yourusername/hybrid-rag-chatbot?style=social)](https://github.com/armanjscript/Hybrid-RAG-chatbot)\n[![License](https://img.shields.io/badge/License-MIT-green.svg)](https://opensource.org/licenses/MIT)\n\n## Description\n\nWelcome to the **Hybrid RAG Chatbot**, a powerful web-based application designed to answer questions based on the content of uploaded PDF documents. This project leverages a **Hybrid Retrieval-Augmented Generation (RAG)** approach, combining the strengths of vector-based semantic search and keyword-based search to deliver accurate and relevant responses. Built with cutting-edge technologies like **Streamlit**, **LangChain**, **Chroma**, and **Ollama**, this chatbot is ideal for researchers, students, or professionals who need to extract insights from documents efficiently.\n\nThe application features a user-friendly interface where you can upload PDFs, adjust settings, and interact with the chatbot in a conversational manner. Responses are generated in real-time, with citations to the source documents for transparency and verifiability.\n\n## Features\n\n| Feature | Description |\n|---------|-------------|\n| **PDF Upload \u0026 Processing** | Upload multiple PDF files, which are automatically split into chunks and indexed for querying. |\n| **Conversational Interface** | Ask questions in a chat-like interface and receive detailed answers based on document content. |\n| **Hybrid RAG Technique** | Combines vector search (Chroma) and keyword search (BM25) for enhanced retrieval accuracy. |\n| **Adjustable Parameters** | Tune the **Temperature** (response creativity) and **Hybrid Search Ratio** (balance between vector and keyword search). |\n| **Real-Time Responses** | Answers are streamed in real-time for a seamless user experience. |\n| **Error Handling** | Robust error handling and fallback mechanisms ensure reliable operation. |\n| **Source Citations** | Responses include references to the source documents, enhancing trust and verifiability. |\n\n## How It Works\n\nThe Hybrid RAG Chatbot uses a sophisticated pipeline to process documents and generate answers. Here’s a detailed look at the process:\n\n### Document Processing\n- **PDF Loading**: Uploaded PDFs are loaded using `PyPDFLoader` from LangChain.\n- **Text Splitting**: Documents are split into manageable chunks (1000 characters, 200-character overlap) using `RecursiveCharacterTextSplitter`.\n- **Metadata Tagging**: Each chunk is tagged with metadata, such as the source file name, for traceability.\n\n### Retrieval\nThe chatbot employs a **Hybrid RAG** approach, integrating two retrieval methods:\n- **Vector Search**: Documents are embedded into a high-dimensional vector space using `OllamaEmbeddings` (model: `nomic-embed-text:latest`) and stored in a Chroma vector store. This enables semantic similarity searches, retrieving the top 5 most relevant chunks (`k=5`).\n- **Keyword Search**: A BM25 retriever ranks documents based on keyword matching, also retrieving the top 5 chunks (`k=5`).\n- **Hybrid Search**: An `EnsembleRetriever` combines the results of both methods, with a configurable **Hybrid Search Ratio** (e.g., 0.5 for equal weighting, 0.7 for 70% vector search). If hybrid search fails, it falls back to vector search for robustness.\n\n### Augmentation\n- Retrieved document chunks are formatted into a context string, including the source file name and content, using a custom `format_docs` function.\n- The context is combined with the user’s query in a `ChatPromptTemplate` to provide the language model with all necessary information.\n\n### Generation\n- The formatted context and query are passed to an `OllamaLLM` (model: `qwen2.5:latest`) to generate a detailed response.\n- The response is parsed using `StrOutputParser` and streamed to the user interface in real-time using Streamlit’s `st.write_stream`.\n\n### Diagram of the RAG Pipeline\nThe following diagram illustrates the flow of the Hybrid RAG pipeline:\n\n```mermaid\nflowchart LR\n    A([\"User Query\"])\n    A --\u003e B[\"Hybrid Retriever\"]\n    B --\u003e C[\"Vector Search (Chroma)\"]\n    C --\u003e D[\"Embeddings (OllamaEmbeddings)\"]\n    B --\u003e E[\"Keyword Search (BM25)\"]\n    C --\u003e F[\"Retrieved Documents\"]\n    E --\u003e F\n    F --\u003e G[\"Formatter\"]\n    G --\u003e H[\"Prompt Template\"]\n    H --\u003e I[\"LLM\"]\n    I --\u003e J[\"Output Parser\"]\n    J --\u003e K[\"Response\"]\n```\n\nThis diagram can be rendered in GitHub to visualize the pipeline from query to response.\n\n## Environment Setup\n\nTo run the Hybrid RAG Chatbot, you’ll need to set up the following:\n\n- **Python 3.8 or later**: Ensure Python is installed on your system. Download from [python.org](https://www.python.org/downloads/).\n- **Ollama**: A tool for running large language models locally. Install it based on your operating system:\n  - **Windows**: Download the installer from [Ollama Download](https://ollama.com/download) and run it.\n  - **macOS**: Download the installer from [Ollama Download](https://ollama.com/download), unzip it, and drag the `Ollama.app` to your Applications folder.\n  - **Linux**: Run the installation script as per the [Ollama GitHub repository](https://github.com/ollama/ollama).\n- **Python Libraries**: Install the required dependencies listed in `requirements.txt`.\n\n## Installation\n\nFollow these steps to set up the project:\n\n1. **Clone the Repository**:\n   ```bash\n   git clone https://github.com/armanjscript/Hybrid-RAG-chatbot.git\n   ```\n2. **Navigate to the Project Directory**:\n   ```bash\n   cd Hybrid-RAG-chatbot\n   ```\n3. **Install Dependencies**:\n   ```bash\n   pip install -r requirements.txt\n   ```\n   The `requirements.txt` file includes dependencies like `streamlit`, `langchain`, `chromadb`, `langchain-ollama`, and others.\n\n## Usage\n\n1. **Start Ollama**: Ensure the Ollama service is running on your system. Follow the instructions from the [Ollama GitHub repository](https://github.com/ollama/ollama) to start the service.\n2. **Run the Streamlit App**:\n   ```bash\n   streamlit run hybrid_rag.py\n   ```\n   This will launch the app in your default web browser.\n3. **Upload PDFs**:\n   - In the sidebar, use the file uploader to select one or more PDF files.\n   - The files are saved locally in the `uploaded_pdfs` directory and indexed in a Chroma database (`chroma_db`).\n4. **Adjust Settings**:\n   - Use the sliders in the sidebar to set the **Temperature** (0.0 to 1.0) and **Hybrid Search Ratio** (0.0 to 1.0).\n   - Optionally, clear all documents to reset the application.\n5. **Ask Questions**:\n   - Enter your query in the chat input field in the main interface.\n   - The chatbot will retrieve relevant document chunks, generate a response, and stream it in real-time.\n   - Responses include citations to the source documents for reference.\n\n## Configuration\n\nThe application allows you to fine-tune its behavior through two key parameters:\n\n| Parameter | Description | Range |\n|-----------|-------------|-------|\n| **Temperature** | Controls the randomness of the LLM’s responses. Lower values (e.g., 0.0) produce more deterministic answers, while higher values (e.g., 1.0) increase creativity. | 0.0 to 1.0 |\n| **Hybrid Search Ratio** | Balances the contribution of vector search (semantic) and keyword search (BM25). A value of 0.5 gives equal weight, while 0.7 prioritizes vector search. | 0.0 to 1.0 |\n\n## Contributing\n\nWe welcome contributions to enhance the Hybrid RAG Chatbot! To contribute:\n- Fork the repository.\n- Make your changes in a new branch.\n- Submit a pull request with a clear description of your changes.\n\nFor detailed guidelines, refer to the [CONTRIBUTING.md](CONTRIBUTING.md) file.\n\n## License\n\nThis project is licensed under the MIT License. See the [LICENSE](LICENSE) file for details.\n\n## Contact\n\nFor questions, feedback, or collaboration opportunities, please reach out:\n- **Email**: [armannew73@gmail.com]\n- **GitHub Issues**: Open an issue on this repository for bug reports or feature requests.\n\n## Acknowledgments\n\nThis project builds on the following open-source technologies:\n- [Streamlit](https://streamlit.io/) for the web interface\n- [LangChain](https://www.langchain.com/) for document processing and RAG pipeline\n- [Chroma](https://www.trychroma.com/) for vector storage\n- [Ollama](https://ollama.com/) for local language models and embeddings\n\nThank you for exploring the Hybrid RAG Chatbot! We hope it simplifies your document analysis tasks.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Farmanjscript%2Fhybrid-rag-chatbot","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Farmanjscript%2Fhybrid-rag-chatbot","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Farmanjscript%2Fhybrid-rag-chatbot/lists"}