{"id":19714780,"url":"https://github.com/baranzinilab/versa_kg_rag_documentation","last_synced_at":"2026-06-03T23:31:10.344Z","repository":{"id":261392381,"uuid":"883620591","full_name":"BaranziniLab/versa_kg_rag_documentation","owner":"BaranziniLab","description":null,"archived":false,"fork":false,"pushed_at":"2024-11-13T09:54:16.000Z","size":1993,"stargazers_count":1,"open_issues_count":0,"forks_count":0,"subscribers_count":2,"default_branch":"main","last_synced_at":"2025-02-27T21:49:15.180Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/BaranziniLab.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-11-05T09:38:03.000Z","updated_at":"2025-01-06T08:57:50.000Z","dependencies_parsed_at":"2024-11-06T10:56:50.417Z","dependency_job_id":"e0b287ad-712f-4cf9-bd42-f63d969b41cc","html_url":"https://github.com/BaranziniLab/versa_kg_rag_documentation","commit_stats":null,"previous_names":["baranzinilab/versa_kg_rag_documentation"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/BaranziniLab/versa_kg_rag_documentation","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/BaranziniLab%2Fversa_kg_rag_documentation","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/BaranziniLab%2Fversa_kg_rag_documentation/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/BaranziniLab%2Fversa_kg_rag_documentation/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/BaranziniLab%2Fversa_kg_rag_documentation/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/BaranziniLab","download_url":"https://codeload.github.com/BaranziniLab/versa_kg_rag_documentation/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/BaranziniLab%2Fversa_kg_rag_documentation/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":33884733,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-03T02:00:06.370Z","response_time":59,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-11-11T22:35:51.014Z","updated_at":"2026-06-03T23:31:10.327Z","avatar_url":"https://github.com/BaranziniLab.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"images/kg-rag-logo.png\" alt=\"KG-RAG Logo\" width=\"400\"/\u003e\n\n  # KG-RAG: Knowledge Graph-Enhanced Biomedical Assistant\n\n  A specialized assistant in Versa that enhances Large Language Models with SPOKE biomedical knowledge graph\n  \n  [Quick Start](#quick-start) • [KG-RAG](#what-is-kg-rag) • [Examples](docs/EXAMPLES.md) • [API Documentation](docs/API.md)\n\u003c/div\u003e\n\n## Table of Contents\n- [Quick Start](#quick-start)\n- [What is SPOKE?](#what-is-spoke)\n- [What is KG-RAG?](#what-is-kg-rag)\n- [Using KG-RAG in UCSF Versa](#using-kg-rag-in-ucsf-versa)\n- [KG-RAG API Access](#kg-rag-api-access)\n- [Additional Resources](#additional-resources)\n\n## Quick Start\n\n1. Access Versa (requires UCSF authentication)\n2. Select \"SPOKE - Knowledge Graph\" from Assistants dropdown\n3. Choose your preferred language model (e.g., GPT-4o)\n4. Ask your disease-related biomedical question\n5. Review the evidence-based response with sources\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"images/versa-ui.png\" alt=\"Versa Interface\" width=\"800\"/\u003e\n  \u003cp\u003e\u003ci\u003eVersa interface showing SPOKE Knowledge Graph selection (which uses KG-RAG in the backend)\u003c/i\u003e\u003c/p\u003e\n\u003c/div\u003e\n\n\n## What is SPOKE?\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"images/spoke-logo.png\" alt=\"SPOKE Logo\" width=\"200\"/\u003e\n  \u003cp\u003e\u003ci\u003eSPOKE: Scalable Precision Medicine Open Knowledge Engine\u003c/i\u003e\u003c/p\u003e\n\u003c/div\u003e\n\nSPOKE (Scalable Precision Medicine Open Knowledge Engine) is UCSF's comprehensive biomedical knowledge graph that integrates and connects information from over 40 specialized databases. It serves as a unified platform for biomedical knowledge, making complex relationships between different biological entities discoverable and accessible.\n\n### Key Statistics\n- **Nodes**: 27+ million nodes of 21 different types\n- **Edges**: 53+ million edges of 55 different types\n- **Sources**: Integrates 41+ specialized biomedical databases\n- **Updates**: Refreshed weekly to ensure current information\n\n### Data Quality\n- Prioritizes experimentally validated information\n- Maintains clear provenance for all relationships\n- Provides statistical evidence when available (p-values, z-score, confidence scores)\n- Focuses on curated databases rather than text mining\n\n### Example SPOKE Relationships\n- Disease-Disease Ontology associations\n- Disease-Gene associations\n- Disease-Symptom relationships\n- Disease-Compound associations\n- Compound-Protein interactions\n- Protein-Protein interactions\n- Anatomical hierarchies\nand many more\n\n\u003e 📚 For more detailed information about SPOKE, visit the [SPOKE Explorer](https://spoke.rbvi.ucsf.edu) or read the [SPOKE Paper](https://academic.oup.com/bioinformatics/article/39/2/btad080/7033465).\n\n\n\n## What is KG-RAG?\n\nKG-RAG (Knowledge Graph Retrieval Augmented Generation) is a specialized framework that enhances Large Language Models (LLMs) with SPOKE's biomedical knowledge. By combining the reasoning capabilities of LLMs with verified biomedical information from SPOKE, KG-RAG provides reliable, evidence-based responses to biomedical questions.\n\n### Core Components\n\n1. **Large Language Models (LLMs)**\n   - Advanced AI models like GPT-4\n   - Natural language understanding and generation\n   - Reasoning capabilities\n\n2. **SPOKE Knowledge Graph**\n   - Verified biomedical knowledge\n   - Structured relationships\n   - Statistical evidence and provenance\n\n3. **Sentence Transformers**\n   - Creates embeddings of biomedical context from SPOKE\n   - Creates embeddings of user queries\n   - Enables context pruning through semantic similarity, thereby optimizes the knowledge retrieval\n\n### How It Works\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"images/kg-rag-workflow.png\" alt=\"KG-RAG Workflow\" width=\"800\"/\u003e\n  \u003cp\u003e\u003ci\u003eKG-RAG's workflow for processing biomedical queries\u003c/i\u003e\u003c/p\u003e\n\u003c/div\u003e\n\n1. **Disease Recognition**\n   - Uses sentence transformers to embed user questions\n   - Matches with disease entities in SPOKE\n   - Ensures accurate disease identification\n\n2. **Context Retrieval \u0026 Pruning**\n   - Extracts relevant context from SPOKE\n   - Uses sentence transformers to embed biomedical context\n   - Prunes context based on semantic similarity to query\n\n3. **Context Enhancement**\n   - Combines pruned knowledge with LLM capabilities\n   - Pruning optimizes the token utilization for the LLM\n   - Preserves evidence and provenance information\n\n4. **Response Generation**\n   - Generates comprehensive answers grounded on factual biomedical knowledge from SPOKE\n   - Includes evidence-based support\n   - Maintains scientific accuracy\n   - Provided provenance in the generated text \n\n### Core Features \u0026 Benefits\n\n#### 🎯 Knowledge-Grounded Responses\n- **Verified Information**\n  - Responses backed by SPOKE's curated biomedical knowledge\n  - Clear provenance for all information\n\n- **Scientific Accuracy**\n  - Statistical evidence when available (Note: Versa maynot support this, but KG-RAG API does)\n  - Multiple source validation\n\n#### 🔍 Advanced Query Processing\n- **Intelligent Disease Recognition**\n  - Robust entity recognition using embeddings\n  - Handles variations in disease names\n  - Maps to standardized disease concepts\n\n- **Smart Context Retrieval**\n  - Semantic matching for relevant information\n  - Efficient pruning of knowledge graph data\n  - Optimal context selection\n\n#### ⚡ Enhanced Performance\n- **Token Efficiency**\n  - Optimized context selection\n  - Reduced token usage compared to traditional RAG\n  - Cost-effective implementation\n\n- **Consistent Results**\n  - Stable responses across LLM updates\n  - Evidence-based conclusions\n\n### Example KG-RAG Capabilities\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"images/example-queries.png\" alt=\"KG-RAG Example Queries\" width=\"800\"/\u003e\n  \u003cp\u003e\u003ci\u003eExamples of KG-RAG's comprehensive responses with statistical evidence. Note that these queries were run in March 2024 using GPT-4. When running the same queries today, GPT-4-only responses (blue box) may differ from those shown in the figure due to the possible model updates from OpenAI.\u003c/i\u003e\u003c/p\u003e\n\u003c/div\u003e\n\n#### Example 1: Drug-Disease Relationships\n**Query**: \"Are there any latest drugs used for weight management in patients with Bardet-Biedl Syndrome?\"\n\nKG-RAG provides:\n- Retrieves drug treatment information from SPOKE\n- Shows clinical trial phase (Phase 3)\n- Includes multiple source databases (ChEMBL, DrugCentral)\n- Provides clear provenance for information\n\n#### Example 2: Gene-Disease Associations\n**Query**: \"Is it PNPLA3 or HLA-B that has a significant association with the disease liver benign neoplasm?\"\n\nKG-RAG provides:\n- Comparative statistical analysis\n- Precise p-values (PNPLA3: 4e-14, HLA-B: 2e-08)\n- Source attribution (GWAS Catalog)\n- Evidence-based conclusion\n\n\u003e 💡 Note: These examples demonstrate KG-RAG's full capabilities with statistical evidence. The current Versa implementation may have different features, which we'll discuss in the [Using KG-RAG in Versa](#using-kg-rag-in-versa) section.\n\n\n\u003e 📚 For more technical details about KG-RAG's architecture and performance, read our [research paper](https://academic.oup.com/bioinformatics/article/40/9/btae560/7759620) published in Bioinformatics.\n\n\n\n## Using KG-RAG in UCSF Versa\n\nKG-RAG is available in UCSF Versa as a specialized assistant that enables biomedical question-answering using SPOKE knowledge. Here's how to effectively use KG-RAG in Versa:\n\n\n### Getting Started\n\n1. **Connect to UCSF VPN**\nTo access the Versa application, users must first connect to the UCSF VPN.\n\n2. **Select the Assistant**\n   - Choose \"SPOKE - Knowledge Graph\" from the Assistants dropdown menu of Versa\n   - Select your preferred language model (e.g., GPT-4o)\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"images/versa-ui.png\" alt=\"Versa Interface\" width=\"800\"/\u003e\n  \u003cp\u003e\u003ci\u003eVersa interface showing SPOKE Knowledge Graph selection (which uses KG-RAG in the backend)\u003c/i\u003e\u003c/p\u003e\n\u003c/div\u003e\n\n3. **Frame Your Question**\n   - Currently, Versa accepts only disease-related queries (i.e. queries that have disease names mentioned in it. e.g. \u003ci\u003ewhat are the genes associated with multiple sclerosis?\u003c/i\u003e)\n   - Be specific and clear in your questions\n\n\u003e 💡 **Disease Coverage**: SPOKE contains 11,697 disease concepts, providing comprehensive coverage across various medical domains. This means users can inquire about a wide spectrum of diseases, from common conditions to rare disorders, all backed by verified biomedical knowledge.\n\n\n### Current Implementation Notes\n- For now, Versa's KG-RAG implementation focuses on disease-centric questions\n- Responses will include information sourced from SPOKE\n- While statistical evidence is available in SPOKE, the current Versa implementation doesn't use that (you can get that information using [KG-RAG API](#api-access))\n\n### Response Structure in Versa\nKG-RAG in Versa provides structured responses in three sections:\n\n1. **SPOKE-Prioritized Response**\n   - Presents findings directly from SPOKE knowledge base\n   - Lists entities with their relationships\n   - Includes provenance information (data sources)\n   - Complemented with relevant LLM knowledge\n\n2. **Analysis Without SPOKE**\n   - Provides context from LLM's training\n   - Offers additional insights\n   - Helps validate and complement SPOKE information\n\n3. **Summary**\n   - Combines insights from both sources\n   - Highlights key findings from SPOKE\n   - Includes additional context from LLM\n   - Provides comprehensive conclusions\n\n\u003e 💡 **Example Response Format:**\n```text\nSECTION 1 - SPOKE-PRIORITIZED RESPONSE:\nBased on SPOKE knowledge base:\n* Information directly from SPOKE with provenance\n* Additional context from LLM\n\nSECTION 2 - ANALYSIS WITHOUT SPOKE:\nBased on trained biomedical knowledge:\n* LLM's knowledge about the topic without using SPOKE\n\nSECTION 3 - SUMMARY:\nFrom Section 1 (with SPOKE):\n* Key findings from SPOKE\n\nFrom Section 2 (without SPOKE):\n* Key findings from LLM\n\nFinal comprehensive conclusion combining insights from Section 1 and Section 2.\n```\n\n### Best Practices - Recommended Query Types\n\n#### ✅ Direct Queries\nBased on SPOKE's knowledge graph structure, you can ask questions about:\n\n🔎 **Gene-Disease Associations**\n*Example: \"What genes are associated with Acute Monocytic Leukemia?\"*\n\n🔎 **Disease-Disease Similarity**\n*Example: \"Which diseases are similar to Parkinson's disease?\"*\n\n🔎 **Disease-Disease Ontology**\n*Example: \"What is the disease ontology of Alzheimer's disease?\"*\n\n🔎 **Disease-Drug Treatments**\n*Example: \"What drugs are used to treat multiple sclerosis?\"*\n\n🔎 **Disease-Symptom Relationships**\n*Example: \"What are the symptoms of Bardet-Biedl Syndrome?\"*\n\n🔎 **Disease-Organism Associations**\n*Example: \"Which organisms can cause pneumonia?\"*\n\n🔎 **Disease-Anatomy Localization**\n*Example: \"Which anatomical structures are affected by diabetes?\"*\n\n#### ✅ Intersection Queries\nKG-RAG also supports queries that combine multiple relationship types or explore intersections between two or more diseases. Here are some examples:\n\n🔎 **Disease-Gene-Disease Connections**\n*Example: \"What genes are common between Parkinson's disease and Alzheimer's disease?\"*\n\n🔎 **Disease-Symptom-Disease Patterns**\n*Example: \"What symptoms are shared between multiple sclerosis and lupus?\"*\n\n🔎 **Disease-Drug-Disease Relationships**\n*Example: \"What drugs are used to treat both rheumatoid arthritis and psoriatic arthritis?\"*\n\n🔎 **Disease-Anatomy-Disease Associations**\n*Example: \"Which anatomical structures are affected by both diabetes and hypertension?\"*\n\n\u003e 💡 **Tip**: When forming intersection queries, clearly specify both diseases and the relationship type you're interested in exploring between them.\n\n### Limitations\n\n#### 🕒 Response Time\nCurrent average response time:\n - GPT-4o: 24.5 ± 17.7 seconds\n - GPT-4: 30.3 ± 9.6 seconds\n\nThese latencies are due to multiple API calls in the backend pipeline:\n1. User query → GPT API (disease entity extraction)\n2. Azure API (semantic search)\n3. KG-RAG API (context extraction from SPOKE)\n4. GPT API (response generation and summarization)\n\n#### 🎯 Query Scope\n- Currently limited to disease-centric questions\n- Queries must explicitly mention disease names\n- Other biomedical queries (e.g., drug-protein interactions without disease context) are not supported in the current Versa implementation\n- For broader biomedical queries, consider using the [KG-RAG API](#api-access)\n\n#### 🔍 Graph Search Depth\n- Versa's implementation uses single-hop graph search for optimal performance\n- While deeper graph searches are possible through the KG-RAG API, they result in:\n - Exponential increase in response time\n - Larger context volume\n - Higher API costs\n\n### Want to See More?\nFor a comprehensive collection of example queries and their responses, visit our [Examples Guide](docs/EXAMPLES.md).\n\n## KG-RAG API Access\n\nWhile KG-RAG is integrated into Versa for disease-centric queries, you can also access it directly through our REST APIs for broader biomedical questions. We provide two specialized endpoints:\n\n### Available Endpoints\n\n1. **Disease-Centric Endpoint** (`v1/kg_rag_context`)\n   - Optimized for disease-related queries\n   - Used by Versa integration\n   - Requires specific disease nodes\n\n2. **Extended Endpoint** (`v1/kg_rag_context_extended`)\n   - Supports broader biomedical queries\n   - No disease node requirement\n   - More flexible querying capabilities\n\n### Key Features\n- Access to complete SPOKE knowledge\n- Statistical evidence inclusion option\n- Configurable search depth\n- Detailed provenance information\n\nFor detailed documentation, including:\n- Complete API reference\n- Code examples\n- Response formats\n- Implementation guidelines\n\n📚 Please refer to our [API Documentation](docs/API.md)\n\n\n## Additional Resources\n\n### 📚 Publications\n- [KG-RAG Paper](https://academic.oup.com/bioinformatics/article/40/9/btae560/7759620) - Technical details about KG-RAG framework and its performance\n- [SPOKE Paper](https://academic.oup.com/bioinformatics/article/39/2/btad080/7033465) - Comprehensive overview of SPOKE knowledge graph\n\n### 🛠️ Development Resources\n- [KG-RAG GitHub Repository](https://github.com/BaranziniLab/KG_RAG) - Open-source code and setup instructions\n- [SPOKE API Documentation](https://spoke.rbvi.ucsf.edu/swagger/) - Complete SPOKE API reference\n\n### 🔍 Tools\n- [SPOKE Explorer](https://spoke.rbvi.ucsf.edu/) - Interactive interface to explore SPOKE knowledge graph\n\n\u003e 💡 **Want to run KG-RAG on your machine?** Follow the instructions in the [KG-RAG GitHub repository](https://github.com/BaranziniLab/KG_RAG) to set up and run KG-RAG locally.","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fbaranzinilab%2Fversa_kg_rag_documentation","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fbaranzinilab%2Fversa_kg_rag_documentation","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fbaranzinilab%2Fversa_kg_rag_documentation/lists"}