{"id":41959839,"url":"https://github.com/mitulgarg/env-doctor","last_synced_at":"2026-04-02T21:51:26.749Z","repository":{"id":325696087,"uuid":"1102054677","full_name":"mitulgarg/env-doctor","owner":"mitulgarg","description":"Debug your GPU, CUDA, and AI stacks across local, Docker, and CI/CD (CLI and MCP server)","archived":false,"fork":false,"pushed_at":"2026-02-17T19:47:50.000Z","size":2213,"stargazers_count":104,"open_issues_count":7,"forks_count":4,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-02-18T00:55:08.584Z","etag":null,"topics":["compatibility-tool","cuda","cuda-library","cuda-support","cuda-toolkit","cudnn","gpu-acceleration","mcp-server","nvidia-driver","nvidia-gpu","nvidia-smi","pytorch","wsl2"],"latest_commit_sha":null,"homepage":"https://mitulgarg.github.io/env-doctor/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/mitulgarg.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-11-22T18:25:58.000Z","updated_at":"2026-02-17T19:47:42.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/mitulgarg/env-doctor","commit_stats":null,"previous_names":["mitulgarg/env-doctor"],"tags_count":11,"template":false,"template_full_name":null,"purl":"pkg:github/mitulgarg/env-doctor","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mitulgarg%2Fenv-doctor","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mitulgarg%2Fenv-doctor/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mitulgarg%2Fenv-doctor/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mitulgarg%2Fenv-doctor/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/mitulgarg","download_url":"https://codeload.github.com/mitulgarg/env-doctor/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mitulgarg%2Fenv-doctor/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":29874556,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-02-26T21:05:00.265Z","status":"ssl_error","status_checked_at":"2026-02-26T20:57:13.669Z","response_time":89,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.5:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["compatibility-tool","cuda","cuda-library","cuda-support","cuda-toolkit","cudnn","gpu-acceleration","mcp-server","nvidia-driver","nvidia-gpu","nvidia-smi","pytorch","wsl2"],"created_at":"2026-01-25T22:54:34.936Z","updated_at":"2026-04-02T21:51:26.742Z","avatar_url":"https://github.com/mitulgarg.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://raw.githubusercontent.com/mitulgarg/env-doctor/main/docs/assets/logo.svg\" alt=\"Env-Doctor Logo\" width=\"80\" height=\"80\"\u003e\n\u003c/p\u003e\n\n\u003ch1 align=\"center\"\u003eEnv-Doctor\u003c/h1\u003e\n\n\n\u003cp align=\"center\"\u003e\n  \u003cstrong\u003eThe missing link between your GPU and Python AI libraries\u003c/strong\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://mitulgarg.github.io/env-doctor/\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/docs-github.io-blueviolet?style=flat-square\" alt=\"Documentation\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://pypi.org/project/env-doctor/\"\u003e\n    \u003cimg src=\"https://img.shields.io/pypi/v/env-doctor?style=flat-square\u0026color=blue\u0026label=PyPI\" alt=\"PyPI\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://pypi.org/project/env-doctor/\"\u003e\n    \u003cimg src=\"https://img.shields.io/pypi/dm/env-doctor?style=flat-square\u0026color=success\u0026label=Downloads\" alt=\"Downloads\"\u003e\n  \u003c/a\u003e\n  \u003cimg src=\"https://img.shields.io/badge/python-3.10%2B-blue?style=flat-square\" alt=\"Python\"\u003e\n  \u003ca href=\"https://github.com/mitulgarg/env-doctor/blob/main/LICENSE\"\u003e\n    \u003cimg src=\"https://img.shields.io/github/license/mitulgarg/env-doctor?style=flat-square\u0026color=green\" alt=\"License\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://github.com/mitulgarg/env-doctor/stargazers\"\u003e\n    \u003cimg src=\"https://img.shields.io/github/stars/mitulgarg/env-doctor?style=flat-square\u0026color=yellow\" alt=\"GitHub Stars\"\u003e\n  \u003c/a\u003e\n\u003c/p\u003e\n\n---\n\n\u003e **\"Why does my PyTorch crash with CUDA errors when I just installed it?\"**\n\u003e\n\u003e Because your driver supports CUDA 11.8, but `pip install torch` gave you CUDA 12.4 wheels.\n\n**Env-Doctor diagnoses and fixes the #1 frustration in GPU computing:** mismatched CUDA versions between your NVIDIA driver, system toolkit, cuDNN, and Python libraries.\n\nIt takes **5 seconds** to find out if your environment is broken - and exactly how to fix it.\n\n\n## Doctor \"Check\" (Diagnosis)\n\n![Env-Doctor Demo](https://raw.githubusercontent.com/mitulgarg/env-doctor/main/docs/assets/envdoctordemo.gif)\n\n\n## Features\n\n| Feature | What It Does |\n|---------|--------------|\n| **One-Command Diagnosis** | Check compatibility: GPU Driver → CUDA Toolkit → cuDNN → PyTorch/TensorFlow/JAX |\n| **Compute Capability Check** | Detect GPU architecture mismatches — catches why `torch.cuda.is_available()` returns `False` on new GPUs (e.g. Blackwell) even when driver and CUDA are healthy |\n| **Python Version Compatibility** | Detect Python version conflicts with AI libraries and dependency cascade impacts |\n| **CUDA Auto-Installer** | Execute CUDA Toolkit installation directly with `--run`; CI-friendly with `--yes`; preview with `--dry-run` |\n| **Safe Install Commands** | Get the exact `pip install` command that works with YOUR driver |\n| **Extension Library Support** | Install compilation packages (flash-attn, SageAttention, auto-gptq, apex, xformers) with CUDA version matching |\n| **AI Model Compatibility** | Check if LLMs, Diffusion, or Audio models fit on your GPU before downloading |\n| **WSL2 GPU Support** | Validate GPU forwarding, detect driver conflicts within WSL2 env for Windows users |\n| **Deep CUDA Analysis** | Find multiple installations, PATH issues, environment misconfigurations |\n| **Container Validation** | Catch GPU config errors in Dockerfiles before you build |\n| **MCP Server** | Expose diagnostics to AI assistants (Claude Desktop, Zed) via Model Context Protocol |\n| **CI/CD Ready** | JSON output, proper exit codes, and CI-aware env-var persistence (GitHub Actions, GitLab CI, CircleCI, Azure Pipelines, Jenkins) |\n\n## Installation\n\n```bash\npip install env-doctor\n```\n\n**Or with [uv](https://docs.astral.sh/uv/)** (a faster Python package manager):\n\n```bash\n# Install as an isolated tool (won't touch your project env)\nuv tool install env-doctor\n\n# Or run once without installing\nuvx env-doctor check\n```\n\nBoth methods install the same package from PyPI — pick whichever you prefer.\n\n## MCP Server (AI Assistant Integration)\n\nEnv-Doctor includes a built-in [Model Context Protocol (MCP)](https://modelcontextprotocol.io) server that exposes diagnostic tools to AI assistants like Claude Code and Claude Desktop.\n\n### Quick Setup for Claude Desktop\n\n1. **Install env-doctor:**\n   ```bash\n   pip install env-doctor\n   ```\n\n2. **Add to Claude Desktop config** (`~/Library/Application Support/Claude/claude_desktop_config.json`):\n   ```json\n   {\n     \"mcpServers\": {\n       \"env-doctor\": {\n         \"command\": \"env-doctor-mcp\"\n       }\n     }\n   }\n   ```\n\n3. **Restart Claude Desktop** - the tools will be available automatically.\n\n### Available Tools (11 Total)\n\n- `env_check` - Full GPU/CUDA environment diagnostics\n- `env_check_component` - Check specific component (driver, CUDA, cuDNN, etc.)\n- `python_compat_check` - Check Python version compatibility with installed AI libraries\n- `cuda_info` - Detailed CUDA toolkit information\n- `cudnn_info` - Detailed cuDNN library information\n- `cuda_install` - Step-by-step CUDA installation instructions\n- `install_command` - Get safe pip install commands for AI libraries\n- `model_check` - Analyze if AI models fit on your GPU\n- `model_list` - List all available models in database\n- `dockerfile_validate` - Validate Dockerfiles for GPU issues\n- `docker_compose_validate` - Validate docker-compose.yml for GPU configuration\n\n### Demo — Claude Code using env-doctor MCP tools\n\n\u003cvideo src=\"https://github.com/user-attachments/assets/7e761c28-1f44-44a0-8dfd-cf06cb9939a2\" autoplay loop muted playsinline width=\"100%\"\u003e\u003c/video\u003e\n\n### Example Usage\n\nAsk your AI assistant:\n- \"Check my GPU environment\"\n- \"Is my Python version compatible with my installed AI libraries?\"\n- \"How do I install CUDA Toolkit on Ubuntu?\"\n- \"Get me the pip install command for PyTorch\"\n- \"Can I run Llama 3 70B on my GPU?\"\n- \"Validate this Dockerfile for GPU issues\"\n- \"What CUDA version does my PyTorch require?\"\n- \"Show me detailed CUDA toolkit information\"\n\n**Learn more:** [MCP Integration Guide](docs/guides/mcp-integration.md)\n\n---\n\n## Usage\n\n### Diagnose Your Environment\n\n```bash\nenv-doctor check\n```\n\n**Example output:**\n```\n🩺 ENV-DOCTOR DIAGNOSIS\n============================================================\n\n🖥️  Environment: Native Linux\n\n🎮 GPU Driver\n   ✅ NVIDIA Driver: 535.146.02\n   └─ Max CUDA: 12.2\n\n🔧 CUDA Toolkit\n   ✅ System CUDA: 12.1.1\n\n📦 Python Libraries\n   ✅ torch 2.1.0+cu121\n\n✅ All checks passed!\n```\n\n**On new-generation GPUs** (e.g. RTX 5070 / Blackwell), env-doctor catches architecture mismatches and distinguishes between two failure modes:\n\n**Hard failure** — `torch.cuda.is_available()` returns `False`:\n```\n🎯  COMPUTE CAPABILITY CHECK\n    GPU: NVIDIA GeForce RTX 5070 (Compute 12.0, Blackwell, sm_120)\n    PyTorch compiled for: sm_50, sm_60, sm_70, sm_80, sm_90, compute_90\n    ❌ ARCHITECTURE MISMATCH: Your GPU needs sm_120 but PyTorch 2.5.1 doesn't include it.\n\n    This is likely why torch.cuda.is_available() returns False even though\n    your driver and CUDA toolkit are working correctly.\n\n    FIX: Install PyTorch nightly with sm_120 support:\n       pip install --pre torch torchvision torchaudio --index-url https://download.pytorch.org/whl/nightly/cu126\n```\n\n**Soft failure** — `torch.cuda.is_available()` returns `True` via NVIDIA's PTX JIT, but complex ops may silently degrade:\n```\n🎯  COMPUTE CAPABILITY CHECK\n    GPU: NVIDIA GeForce RTX 5070 (Compute 12.0, Blackwell, sm_120)\n    PyTorch compiled for: sm_50, sm_60, sm_70, sm_80, sm_90, compute_90\n    ⚠️  ARCHITECTURE MISMATCH (Soft): Your GPU needs sm_120 but PyTorch 2.5.1 doesn't include it.\n\n    torch.cuda.is_available() returned True via NVIDIA's driver-level PTX JIT,\n    but you may experience degraded performance or failures with complex CUDA ops.\n\n    FIX: Install a newer PyTorch with native sm_120 support for full compatibility:\n       pip install --pre torch torchvision torchaudio --index-url https://download.pytorch.org/whl/nightly/cu126\n```\n\n### Check Python Version Compatibility\n\n```bash\nenv-doctor python-compat\n```\n\n```\n🐍  PYTHON VERSION COMPATIBILITY CHECK\n============================================================\nPython Version: 3.13 (3.13.0)\nLibraries Checked: 2\n\n❌  2 compatibility issue(s) found:\n\n    tensorflow:\n      tensorflow supports Python \u003c=3.12, but you have Python 3.13\n      Note: TensorFlow 2.15+ requires Python 3.9-3.12. Python 3.13 not yet supported.\n\n    torch:\n      torch supports Python \u003c=3.12, but you have Python 3.13\n      Note: PyTorch 2.x supports Python 3.9-3.12. Python 3.13 support experimental.\n\n⚠️   Dependency Cascades:\n    tensorflow [high]: TensorFlow's Python ceiling propagates to keras and tensorboard\n      Affected: keras, tensorboard, tensorflow-estimator\n    torch [high]: PyTorch's Python version constraint affects all torch ecosystem packages\n      Affected: torchvision, torchaudio, triton\n\n💡  Consider using Python 3.12 or lower for full compatibility\n\n💡  Cascade: tensorflow constraint also affects: keras, tensorboard, tensorflow-estimator\n\n💡  Cascade: torch constraint also affects: torchvision, torchaudio, triton\n\n============================================================\n```\n\n### Get Safe Install Command\n\n```bash\nenv-doctor install torch\n```\n\n```\n⬇️ Run this command to install the SAFE version:\n---------------------------------------------------\npip install torch torchvision --index-url https://download.pytorch.org/whl/cu118\n---------------------------------------------------\n```\n\n### Install CUDA Toolkit\n\nDisplay instructions or execute the installation directly:\n\n```bash\n# Show platform-specific steps (default)\nenv-doctor cuda-install\n\n# Preview what would run — no changes made\nenv-doctor cuda-install --dry-run\n\n# Execute interactively (asks [y/N] before running)\nenv-doctor cuda-install --run\n\n# Execute headlessly — great for CI/scripts\nenv-doctor cuda-install --run --yes\n\n# Install a specific version, headless\nenv-doctor cuda-install 12.6 --run --yes\n```\n\n**Example dry-run output (Windows):**\n```\n[DRY RUN] [1/1] winget install Nvidia.CUDA --version 12.2\n\n[DRY RUN] [1/1] nvcc --version\n\nCUDA 12.2 installation completed successfully.\nVerification: PASSED\n\nFull log: C:\\Users\\you\\.env-doctor\\install.log\n```\n\nEvery run writes a timestamped log to `~/.env-doctor/install.log` for debugging.\n\n**Supported Platforms:**\n- Ubuntu 20.04, 22.04, 24.04\n- Debian 11, 12\n- RHEL 8, 9 / Rocky Linux / AlmaLinux\n- Fedora 39+\n- WSL2 (Ubuntu)\n- Windows 10/11 (via `winget`)\n- Conda (all platforms)\n\n**Exit codes for CI pipelines:**\n\n| Code | Meaning |\n|------|---------|\n| `0` | Installation succeeded and verified |\n| `1` | An installation step failed |\n| `2` | Installed but `nvcc --version` failed |\n\n### Install Compilation Packages (Extension Libraries)\n\nFor extension libraries like **flash-attn**, **SageAttention**, **auto-gptq**, **apex**, and **xformers** that require compilation from source, `env-doctor` provides special guidance to handle CUDA version mismatches:\n\n```bash\nenv-doctor install flash-attn\n```\n\n**Example output (with CUDA mismatch):**\n```\n🩺  PRESCRIPTION FOR: flash-attn\n\n⚠️   CUDA VERSION MISMATCH DETECTED\n     System nvcc: 12.1.1\n     PyTorch CUDA: 12.4.1\n\n🔧  flash-attn requires EXACT CUDA version match for compilation.\n    You have TWO options to fix this:\n\n============================================================\n📦  OPTION 1: Install PyTorch matching your nvcc (12.1)\n============================================================\n\nTrade-offs:\n  ✅ No system changes needed\n  ✅ Faster to implement\n  ❌ Older PyTorch version (may lack new features)\n\nCommands:\n  # Uninstall current PyTorch\n  pip uninstall torch torchvision torchaudio -y\n\n  # Install PyTorch for CUDA 12.1\n  pip install torch --index-url https://download.pytorch.org/whl/cu121\n\n  # Install flash-attn\n  pip install flash-attn --no-build-isolation\n\n============================================================\n⚙️   OPTION 2: Upgrade nvcc to match PyTorch (12.4)\n============================================================\n\nTrade-offs:\n  ✅ Keep latest PyTorch\n  ✅ Better long-term solution\n  ❌ Requires system-level changes\n  ❌ Verify driver supports CUDA 12.4\n\nSteps:\n  1. Check driver compatibility:\n     env-doctor check\n\n  2. Download CUDA Toolkit 12.4:\n     https://developer.nvidia.com/cuda-12-4-0-download-archive\n\n  3. Install CUDA Toolkit (follow NVIDIA's platform-specific guide)\n\n  4. Verify installation:\n     nvcc --version\n\n  5. Install flash-attn:\n     pip install flash-attn --no-build-isolation\n\n============================================================\n```\n\n### Check Model Compatibility\n\n```bash\nenv-doctor model llama-3-8b\n```\n\n```\n🤖  Checking: LLAMA-3-8B (8.0B params)\n\n🖥️   Your Hardware: RTX 3090 (24GB)\n\n💾  VRAM Requirements:\n  ✅  FP16: 19.2GB - fits with 4.8GB free\n  ✅  INT4:  4.8GB - fits with 19.2GB free\n\n✅  This model WILL FIT on your GPU!\n```\n\nList all models: `env-doctor model --list`\n\n**Cloud GPU Recommendations:**\n\n```bash\n# Get cloud GPU recommendations for a model that doesn't fit\nenv-doctor model llama-3-70b --recommend\n\n# Direct VRAM lookup (no model name needed)\nenv-doctor model --vram 80000 --recommend\n```\n\n```\n☁️   Cloud GPU Recommendations\n\n  FP16 (~140.0 GB):\n    $27.20 /hr  azure  ND96asr_v4              8x A100 (40GB each)          180.0GB free\n    $29.39 /hr  gcp    a2-highgpu-8g            8x A100 (40GB each)          180.0GB free\n    ...\n```\n\nAutomatic HuggingFace Support (New ✨)\nIf a model isn't found locally, env-doctor automatically checks the HuggingFace Hub, fetches its parameter metadata, and caches it locally for future runs — no manual setup required.\n\n```bash\n# Fetches from HuggingFace on first run, cached afterward\nenv-doctor model bert-base-uncased\nenv-doctor model sentence-transformers/all-MiniLM-L6-v2\n```\n\n**Output:**\n\n```\n🤖  Checking: BERT-BASE-UNCASED\n    (Fetched from HuggingFace API - cached for future use)\n    Parameters: 0.11B\n    HuggingFace: bert-base-uncased\n\n🖥️   Your Hardware:\n    RTX 3090 (24GB VRAM)\n\n💾  VRAM Requirements \u0026 Compatibility\n  ✅  FP16:  264 MB - Fits easily!\n\n💡  Recommendations:\n1. Use fp16 for best quality on your GPU\n```\n\n\n\n### Validate Dockerfiles\n\n```bash\nenv-doctor dockerfile\n```\n\n```\n🐳  DOCKERFILE VALIDATION\n\n❌  Line 1: CPU-only base image: python:3.10\n    Fix: FROM nvidia/cuda:12.1.0-runtime-ubuntu22.04\n\n❌  Line 8: PyTorch missing --index-url\n    Fix: pip install torch --index-url https://download.pytorch.org/whl/cu121\n```\n\n### More Commands\n\n| Command | Purpose |\n|---------|---------|\n| `env-doctor check` | Full environment diagnosis |\n| `env-doctor python-compat` | Check Python version compatibility with AI libraries |\n| `env-doctor cuda-install` | Step-by-step CUDA Toolkit installation guide |\n| `env-doctor install \u003clib\u003e` | Safe install command for PyTorch/TensorFlow/JAX, extension libraries (flash-attn, auto-gptq, apex, xformers, SageAttention, etc.) |\n| `env-doctor model \u003cname\u003e` | Check model VRAM requirements |\n| `env-doctor cuda-info` | Detailed CUDA toolkit analysis |\n| `env-doctor cudnn-info` | cuDNN library analysis |\n| `env-doctor dockerfile` | Validate Dockerfile |\n| `env-doctor docker-compose` | Validate docker-compose.yml |\n| `env-doctor init --github-actions` | Generate GitHub Actions workflow |\n| `env-doctor scan` | Scan for deprecated imports |\n| `env-doctor debug` | Verbose detector output |\n\n### CI/CD Integration\n\nGenerate a GitHub Actions workflow with one command:\n\n```bash\nenv-doctor init --github-actions\n```\n\nThis creates `.github/workflows/env-doctor.yml` — review, commit, and push. Your CI will validate the ML environment on every push and PR.\n\nOr add manually:\n\n```bash\n# JSON output for scripting\nenv-doctor check --json\n\n# CI mode with exit codes (0=pass, 1=warn, 2=error)\nenv-doctor check --ci\n```\n\n**GitHub Actions example:**\n```yaml\n- run: pip install env-doctor\n- run: env-doctor check --ci\n```\n\n## Documentation\n\n**Full documentation:** https://mitulgarg.github.io/env-doctor/\n\n- [Getting Started](docs/getting-started.md)\n- [Command Reference](docs/commands/check.md)\n- [MCP Integration Guide](docs/guides/mcp-integration.md)\n- [WSL2 GPU Guide](docs/guides/wsl2.md)\n- [CI/CD Integration](docs/guides/ci-cd.md)\n- [Architecture](docs/architecture.md)\n\n**Video Tutorial:** [Watch Demo on YouTube](https://youtu.be/mGAwxGuLpxk?si=Buf9yzNTSJmoirMU)\n\n## Contributing\n\nContributions welcome! See [CONTRIBUTING.md](CONTRIBUTING.md) for details.\n\n## License\n\nMIT License - see [LICENSE](LICENSE)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmitulgarg%2Fenv-doctor","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fmitulgarg%2Fenv-doctor","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmitulgarg%2Fenv-doctor/lists"}