{"id":49921167,"url":"https://github.com/saganaki22/pixal3d-comfyui","last_synced_at":"2026-05-21T00:01:34.089Z","repository":{"id":357986222,"uuid":"1239244570","full_name":"Saganaki22/Pixal3D-ComfyUI","owner":"Saganaki22","description":"Pixal3D image-to-3D nodes for ComfyUI - local TencentARC Pixal3D generation with textured GLB export + Windows support","archived":false,"fork":false,"pushed_at":"2026-05-17T15:56:17.000Z","size":1935,"stargazers_count":24,"open_issues_count":1,"forks_count":4,"subscribers_count":1,"default_branch":"main","last_synced_at":"2026-05-17T21:25:33.296Z","etag":null,"topics":["3d","3d-model","comfyui","comfyui-nodes","image-to-3d","imageto3d","trellis2"],"latest_commit_sha":null,"homepage":"https://ldyang694.github.io/projects/pixal3d/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Saganaki22.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-05-14T22:52:00.000Z","updated_at":"2026-05-17T20:23:16.000Z","dependencies_parsed_at":"2026-05-17T21:01:20.055Z","dependency_job_id":null,"html_url":"https://github.com/Saganaki22/Pixal3D-ComfyUI","commit_stats":null,"previous_names":["saganaki22/pixal3d-comfyui"],"tags_count":10,"template":false,"template_full_name":null,"purl":"pkg:github/Saganaki22/Pixal3D-ComfyUI","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Saganaki22%2FPixal3D-ComfyUI","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Saganaki22%2FPixal3D-ComfyUI/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Saganaki22%2FPixal3D-ComfyUI/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Saganaki22%2FPixal3D-ComfyUI/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Saganaki22","download_url":"https://codeload.github.com/Saganaki22/Pixal3D-ComfyUI/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Saganaki22%2FPixal3D-ComfyUI/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":33281294,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-20T15:12:43.734Z","status":"ssl_error","status_checked_at":"2026-05-20T15:12:42.300Z","response_time":356,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.6:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["3d","3d-model","comfyui","comfyui-nodes","image-to-3d","imageto3d","trellis2"],"created_at":"2026-05-16T20:03:32.929Z","updated_at":"2026-05-21T00:01:34.026Z","avatar_url":"https://github.com/Saganaki22.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n\n# Pixal3D: Pixel-Aligned 3D Generation from Images\n\n\u003ch3\u003eSIGGRAPH 2026\u003c/h3\u003e\n\n[Dong-Yang Li](https://ldyang694.github.io/)¹ · [Wang Zhao](https://thuzhaowang.github.io/)²* · [Yuxin Chen](https://orcid.org/0000-0002-7854-1072)² · [Wenbo Hu](https://wbhu.github.io/)² · [Meng-Hao Guo](https://menghaoguo.github.io/)¹ · [Fang-Lue Zhang](https://fanglue.github.io/)³ · [Ying Shan](https://www.linkedin.com/in/YingShanProfile)² · [Shi-Min Hu](https://cg.cs.tsinghua.edu.cn/shimin.htm)¹✉\n\n¹Tsinghua University (BNRist) \u0026nbsp;\u0026nbsp; ²Tencent ARC Lab \u0026nbsp;\u0026nbsp; ³Victoria University of Wellington\n\n*Project lead \u0026nbsp;\u0026nbsp; ✉Corresponding author\n\n\u003c/div\u003e\n\n\u003cdiv align=\"center\"\u003e\n  \u003ca href=\"https://ldyang694.github.io/projects/pixal3d/\"\u003e\u003cimg src=https://img.shields.io/badge/Project%20Page-333399.svg?logo=googlehome height=22px\u003e\u003c/a\u003e\n  \u003ca href=\"https://huggingface.co/spaces/TencentARC/Pixal3D\"\u003e\u003cimg src=https://img.shields.io/badge/%F0%9F%A4%97%20Demo-276cb4.svg height=22px\u003e\u003c/a\u003e\n  \u003ca href=\"https://huggingface.co/TencentARC/Pixal3D\"\u003e\u003cimg src=https://img.shields.io/badge/%F0%9F%A4%97%20Models-d96902.svg height=22px\u003e\u003c/a\u003e\n  \u003ca href=\"https://arxiv.org/abs/2605.10922\"\u003e\u003cimg src=https://img.shields.io/badge/Arxiv-b5212f.svg?logo=arxiv height=22px\u003e\u003c/a\u003e\n\u003c/div\u003e\n\n\u003cdiv align=\"center\"\u003e\n   \u003cimg width=\"3840\" height=\"2160\" alt=\"teaser-jpeg\" src=\"https://github.com/user-attachments/assets/80c31413-e51c-437f-9c5f-1c7fd7ee77f3\" /\u003e\n\n\u003c/div\u003e\n\n**Pixal3D** generates high-fidelity 3D assets from a single image. Unlike previous methods that loosely inject image features via attention, Pixal3D explicitly lifts pixel features into 3D through back-projection, establishing direct pixel-to-3D correspondences. This enables near-reconstruction-level fidelity with detailed geometry and PBR textures.\n\n---\n\n# Pixal3D-ComfyUI\n\n**Pixal3D image-to-3D nodes for ComfyUI** - local TencentARC Pixal3D generation with textured GLB export, FlashAttention 2/3 backend selection, and ComfyUI DynamicVRAM/Aimdo support.\n\n[![ComfyUI](https://img.shields.io/badge/ComfyUI-custom%20node-2f80ed)](https://github.com/comfyanonymous/ComfyUI)\n[![Windows CUDA](https://img.shields.io/badge/Windows-CUDA%20required-76b900)](docs/windows_wheels.md)\n[![Python](https://img.shields.io/badge/Python-3.10--3.13-3776ab)](docs/compatibility_matrix.md)\n[![PyTorch](https://img.shields.io/badge/PyTorch-2.8%2B-ee4c2c)](docs/compatibility_matrix.md)\n[![Pixal3D Model](https://img.shields.io/badge/HuggingFace-TencentARC%2FPixal3D-blue)](https://huggingface.co/TencentARC/Pixal3D)\n[![MoGe Weights](https://img.shields.io/badge/HuggingFace-Comfy--Org%2FMoGe-blue)](https://huggingface.co/Comfy-Org/MoGe)\n[![RMBG-2.0](https://img.shields.io/badge/RMBG--2.0-gated-orange)](https://huggingface.co/briaai/RMBG-2.0)\n[![License](https://img.shields.io/badge/License-see%20LICENSE-lightgrey)](LICENSE)\n\n[中文说明](README_ZH.md) | [Compatibility](docs/compatibility_matrix.md) | [Portable Install](docs/portable_standalone_install.md) | [Linux/WSL CUDA](docs/linux_wsl_cuda.md) | [Windows Wheels](docs/windows_wheels.md) | [Troubleshooting](docs/troubleshooting.md) | [Related Repos](docs/related_repos.md)\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://github.com/user-attachments/assets/45d596b4-9070-44d2-8e4f-1019169d3daa\" width=\"1200\"\u003e\u003cbr\u003e\u003cbr\u003e\n\n  \u003cimg src=\"https://github.com/user-attachments/assets/a2ef8b6e-ff68-4a81-a595-1e84eab2062c\" width=\"800\"\u003e\n\u003c/p\u003e\n\n\n\n## Features\n\n- Pixal3D image-to-3D generation directly inside ComfyUI\n- Textured `.glb` export from Pixal3D voxel attributes\n- FlashAttention 2 and FlashAttention 3 runtime selection\n- `auto` attention mode picks FlashAttention 3 when installed, otherwise FlashAttention 2\n- ComfyUI model management, unload, DynamicVRAM, and Aimdo/MemoryVisualization visibility\n- Native low-VRAM Pixal3D mode for staged CPU/GPU movement\n- **Pixal3D Camera Control** node for manual FOV, distance, and mesh-scale setup with Scene/POV preview\n\n\n\u003ctable\u003e\n  \u003ctr\u003e\n    \u003ctd align=\"center\"\u003e\n      \u003cimg src=\"https://github.com/user-attachments/assets/8b53ffe6-115c-4ab2-9170-dc7e8f0e69aa\" width=\"350\"\u003e\u003cbr\u003e\n      \u003cb\u003escene\u003c/b\u003e\n    \u003c/td\u003e\n    \u003ctd align=\"center\"\u003e\n      \u003cimg src=\"https://github.com/user-attachments/assets/ba6e7a92-89c8-449c-a81a-79a6d67c6011\" width=\"350\"\u003e\u003cbr\u003e\n      \u003cb\u003ePOV\u003c/b\u003e\n    \u003c/td\u003e\n  \u003c/tr\u003e\n\u003c/table\u003e\n\n\n- GLB path output connects directly to ComfyUI's native **Preview 3D \u0026 Animation**\n\n## Installation\n\n### Method 1: ComfyUI Manager\n\nSearch for `Pixal3D` or `Pixal3D-ComfyUI` in ComfyUI Manager and install it.\n\n### Method 2: Manual Install\n\n```bash\ncd ComfyUI/custom_nodes\ngit clone https://github.com/Saganaki22/Pixal3D-ComfyUI.git\ncd Pixal3D-ComfyUI\npython -m pip install -r requirements.txt\npython install.py --check\n```\n\n### Method 3: uv\n\n```bash\ncd ComfyUI/custom_nodes/Pixal3D-ComfyUI\nuv pip install -r requirements.txt\n```\n\nPackage-style installs also use the same dependency set:\n\n```bash\npython -m pip install .\nuv pip install .\n```\n\nRestart ComfyUI after installing or updating.\n\nPixal3D-ComfyUI includes a guarded `install.py` for ComfyUI Manager, portable ComfyUI, standalone venv installs, and Linux venv installs. By default it installs only the safe runtime requirements and prints an environment report. It does **not** change PyTorch and does **not** install CUDA wheels unless you explicitly enable an exact-match wheel path.\n\n`requirements.txt` is the safe runtime list. It intentionally does not include `torch`, `torchvision`, `flash-attn`, `triton`, `flex_gemm`, `cumesh`, `o_voxel`, `drtk`, or `nvdiffrast`. Those are binary/CUDA stack packages and must be installed from matching wheels or built for the active environment. See [requirements-cuda-manual.txt](requirements-cuda-manual.txt), [portable install guide](docs/portable_standalone_install.md), [Linux/WSL CUDA guide](docs/linux_wsl_cuda.md), and [Windows wheel guide](docs/windows_wheels.md).\n\nPlain `natten==0.21.6` is included as a baseline dependency because similar Pixal3D wrappers import it. Do not mistake that for strict NAF support. Pixal3D's strict NAF path requires `natten.HAS_LIBNATTEN == True`; a generic `natten-0.21.6-py3-none-any.whl` imports but does not provide CUDA libnatten.\n\nFlashAttention 2 or 3 is a prerequisite. Install a matching FlashAttention wheel for your Python, PyTorch, CUDA, and OS before loading Pixal3D.\n\nIf **Pixal3D Environment Check** reports missing `flex_gemm`, `cumesh`, `o_voxel`, or `drtk`, the normal install did not fail. Those are Pixal3D compiled CUDA wheels and are opt-in. On a known Windows stack, run:\n\n```bash\npython install.py --install-known-cuda\n```\n\nIf your stack is not in the bundled wheel map, install matching wheels or source builds manually from [Linux/WSL CUDA guide](docs/linux_wsl_cuda.md) or [Windows wheel guide](docs/windows_wheels.md).\n\n### Required Wheels And Model Files\n\n`requirements.txt` is not the full Pixal3D install. A working generation environment needs these too:\n\n| Type | Required item | What file/module should exist | Where to get it |\n|---|---|---|---|\n| Attention wheel | FlashAttention 2 or 3 | `flash_attn` or `flash_attn_interface` imports | [Windows wheels](docs/windows_wheels.md#attention-wheels) or [Linux/WSL guide](docs/linux_wsl_cuda.md#flashattention) |\n| CUDA wheel | Sparse GEMM | `flex_gemm_ap` or `flex_gemm` imports | [Windows wheels](docs/windows_wheels.md#required-pixal3d-cuda-wheels) or [Linux/WSL guide](docs/linux_wsl_cuda.md#required-pixal3d-cuda-extensions) |\n| CUDA wheel | Mesh ops | `cumesh_vb` or `cumesh` imports | Same wheel guide for your OS |\n| CUDA wheel | Voxel/remesh ops | `o_voxel_vb_ap` or `o_voxel` imports | Same wheel guide for your OS |\n| CUDA wheel | DRTK | `drtk` imports | Same wheel guide for your OS |\n| CUDA/runtime wheel | Triton | `triton` imports if your attention/sparse stack expects it | Windows usually uses `triton-windows`; Linux usually uses matching `triton` |\n| Optional renderer wheel | NVIDIA raster/render helpers | `nvdiffrast` / `nvdiffrec_render` imports | Optional; install only if your renderer path needs them |\n| Optional strict NAF wheel | CUDA NATTEN/libnatten | `natten.HAS_LIBNATTEN == True` | Linux/WSL usually has official wheels; Windows often needs fallback or a source build |\n| Main model files | TencentARC Pixal3D | `ComfyUI/models/Pixal3D/TencentARC_Pixal3D/pipeline.json` and `ckpts/*.safetensors` | [TencentARC/Pixal3D](https://huggingface.co/TencentARC/Pixal3D) or `download_if_missing=true` |\n| DINOv3 helper files | Pixal3D image encoder | `ComfyUI/models/Pixal3D/camenduru_dinov3-vitl16-pretrain-lvd1689m/model.safetensors` | Downloaded with helpers when enabled, or place a complete snapshot there |\n| MoGe camera files | Auto camera mode | `ComfyUI/models/moge/moge_2_vitl_normal_fp16.safetensors` | [Comfy-Org/MoGe](https://huggingface.co/Comfy-Org/MoGe), or skip with `camera_mode=manual` |\n| RMBG files | Built-in background removal | `ComfyUI/models/Pixal3D/briaai_RMBG-2.0/` complete snapshot | [briaai/RMBG-2.0](https://huggingface.co/briaai/RMBG-2.0) is gated; request access first, or use transparent PNG/WebP |\n\nFor the easiest low-VRAM/manual setup, you can skip MoGe and RMBG: set `load_moge=false`, `load_rembg=false`, use a transparent PNG/WebP with `background_mode=keep_alpha`, and connect **Pixal3D Camera Control** to `manual_fov`.\n\n### What Am I Missing?\n\nRun **Pixal3D Environment Check** first. Match the first missing line to this table:\n\n| Environment Check says | What it means | What to do |\n|---|---|---|\n| `flash_attn: MISSING` and `flash_attn_interface: MISSING` | Attention backend is missing | Install a matching FlashAttention 2 or 3 wheel for your Python/PyTorch/CUDA/OS |\n| `flex_gemm_ap: MISSING` and `flex_gemm: MISSING` | Pixal3D sparse GEMM extension is missing | Install/build matching `flex_gemm_ap` or `flex_gemm` |\n| `cumesh_vb: MISSING` and `cumesh: MISSING` | Pixal3D mesh extension is missing | Install/build matching `cumesh_vb` or `cumesh` |\n| `o_voxel_vb_ap: MISSING` and `o_voxel: MISSING` | Pixal3D voxel/remesh extension is missing | Install/build matching `o_voxel_vb_ap` or `o_voxel` |\n| `drtk: MISSING` | DRTK renderer dependency is missing | Install/build matching `drtk` |\n| `triton: MISSING` | Triton runtime is missing | Install matching `triton-windows` on Windows or matching `triton` on Linux if your wheel stack needs it |\n| `nvdiffrast: MISSING` or `nvdiffrec_render: MISSING` | Optional renderer packages are missing | Install only if your workflow needs those renderer paths; basic GLB export can work without them |\n| `natten: MISSING` | Baseline NATTEN import is missing | Reinstall `requirements.txt` or run `pip install natten==0.21.6` in ComfyUI's Python |\n| `natten.HAS_LIBNATTEN: False` | NATTEN imports, but strict CUDA NAF is not available | Use `naf_mode=fallback_if_missing`, or install/build CUDA NATTEN/libnatten for your stack |\n| `RMBG-2.0` missing or gated | Background remover weights are unavailable | Request access to [briaai/RMBG-2.0](https://huggingface.co/briaai/RMBG-2.0), download it, or use transparent PNG/WebP with `background_mode=keep_alpha` |\n| MoGe missing | Automatic camera estimator weights are unavailable | Download [Comfy-Org/MoGe](https://huggingface.co/Comfy-Org/MoGe), or use `camera_mode=manual` with **Pixal3D Camera Control** |\n\nWindows users: start with [Windows wheel guide](docs/windows_wheels.md). Linux/WSL users: start with [Linux/WSL CUDA guide](docs/linux_wsl_cuda.md). Do not let pip replace your working Torch install while fixing missing CUDA packages; use exact wheels or `--no-deps` where the guide says to.\n\n### Platform Reality Check\n\nFor the smoothest full upstream Pixal3D experience, Linux or WSL is recommended because upstream NATTEN publishes prebuilt NATTEN/libnatten wheels for recent official PyTorch CUDA stacks there.\n\nNative Windows is supported and can generate/export GLBs, but it may need fallback settings unless exact Windows CUDA wheels exist for your stack. In particular, for Python 3.12 + PyTorch 2.10 + CUDA 13.0, there is currently no known official `win_amd64` NATTEN/libnatten wheel for `natten==0.21.6+torch2100cu130`. Plain `natten==0.21.6` is installed for baseline imports, but if `natten.HAS_LIBNATTEN` is `False`, use `naf_mode=fallback_if_missing` instead of `strict`.\n\nRecommended default:\n\n| Environment | Recommendation |\n|---|---|\n| Linux/WSL NVIDIA | Best path for full upstream NAF if official NATTEN/libnatten wheels match your Torch/CUDA |\n| Native Windows NVIDIA | Works, but use exact CUDA extension wheels and `naf_mode=fallback_if_missing` unless `natten.HAS_LIBNATTEN` is `True` |\n\n\u003cdetails\u003e\n\u003csummary\u003eGuarded Installer Policy\u003c/summary\u003e\n\nSome 3D ComfyUI nodes use `comfy-env` with `install.py`, `prestartup_script.py`, and `comfy-env.toml` to build isolated CUDA environments automatically. That can be convenient, but it depends on the wheel map matching the user's exact stack.\n\nPixal3D-ComfyUI keeps the risky parts explicit:\n\n- No `prestartup_script.py`\n- No automatic Torch changes\n- No automatic CUDA wheel install unless an exact known wheel map is explicitly enabled\n- No automatic model download unless `download_if_missing` is enabled\n- Normal `pip` and `uv` installs for runtime requirements\n\nThis makes it less automatic, but safer for custom ComfyUI installs, portable ComfyUI, newer PyTorch/CUDA stacks, and users who already have working FlashAttention/Triton wheels.\n\n\u003c/details\u003e\n\n## Compatibility\n\nThis nodepack is written to import cleanly on normal ComfyUI, portable ComfyUI, venv installs, and uv-managed installs. Actual generation/export requires CUDA because Pixal3D depends on sparse attention and mesh/voxel extension wheels.\n\nThe node is **not pinned to one tiny stack**. It should work on any Python/PyTorch/CUDA combo where the required extension modules import successfully inside the same ComfyUI Python environment.\n\n| Component | Supported range | Notes |\n|-----------|-----------------|-------|\n| VRAM | **20–32 GB recommended** | `1536_cascade` needs ~32 GB; `native_low_vram` can run on much lower VRAM in some workflows |\n| System RAM | **20–40 GB recommended for native low-VRAM** | Pixal3D stages large CPU-side tensors before GPU transfer |\n| OS | Windows and Linux CUDA supported; macOS import-only | macOS should not break ComfyUI import, but CUDA generation/export is not supported unless compatible deps exist |\n| Python | `3.10`-`3.13` expected if wheels exist, `3.12.1` target-friendly | Wheels must match the Python ABI, for example `cp312` for Python 3.12.x |\n| PyTorch | `2.8+` expected if wheels exist, including `2.10` | The extension wheels must match the installed Torch ABI/build |\n| CUDA | `12.8` tested, `13.x` allowed if wheels exist | Use wheels matching `torch.version.cuda`, not just the system CUDA toolkit |\n| FlashAttention 2 | `flash-attn 2.8.3`, Torch `2.8.x`, CUDA `12.8`, Python `3.12` tested | Provides the `flash_attn` module |\n| FlashAttention 2 newer stacks | Torch `2.10` / CUDA `13.x` should work if your wheel imports | Select `flash_attn_2` or use `auto` |\n| FlashAttention 3 | Torch `\u003e=2.9` and matching wheel expected | Must provide the `flash_attn_interface` module |\n| Triton | Windows: matching `triton-windows`; Linux: matching `triton` | Required by several modern CUDA wheel stacks |\n| Required Pixal3D CUDA wheels | `flex_gemm_ap`/`flex_gemm`, `o_voxel_vb_ap`/`o_voxel`, `cumesh_vb`/`cumesh`, `drtk` | These must match Python, Torch, CUDA, and OS |\n| Optional CUDA wheels | `nvdiffrast`, `nvdiffrec_render` | Useful for renderer paths, not the basic GLB export path |\n| ComfyUI | Current ComfyUI with `CoreModelPatcher` and `load_models_gpu` | Needed for DynamicVRAM/Aimdo/MemoryVisualization visibility |\n| GPU | CUDA-capable NVIDIA GPU | Newer GPU architectures need extension wheels built for that Torch/CUDA stack |\n| CPU only | Not supported | Pixal3D sparse/mesh ops need CUDA |\n\nExample verified stack, not a required install path:\n\n```text\nWindows\nPython 3.12\nPyTorch 2.8.0+cu128\nCUDA wheel target cu128\nflash-attn 2.8.3+cu128torch2.8.0\ntriton-windows 3.5+\n```\n\nAnother valid stack shape:\n\n```text\nWindows\nPython 3.12.1\nPyTorch 2.10.x\nCUDA 13.x\nFlashAttention 2 or 3 matching Torch/CUDA/Python\nTriton matching Torch/CUDA/Python\nPixal3D CUDA extension wheels matching Torch/CUDA/Python\n```\n\nUse **Pixal3D Environment Check** inside ComfyUI before loading the model. It checks `torch`, CUDA, FlashAttention 2/3, Triton, `flex_gemm_ap`, `cumesh_vb`, `o_voxel_vb_ap`, and `drtk` without downloading the model.\n\nMore setup detail:\n\n- [Compatibility matrix](docs/compatibility_matrix.md)\n- [Portable and standalone install](docs/portable_standalone_install.md)\n- [Windows wheel guide](docs/windows_wheels.md)\n- [Related repo findings](docs/related_repos.md)\n- [Troubleshooting](docs/troubleshooting.md)\n\n\u003cdetails\u003e\n\u003csummary\u003eProduction Readiness\u003c/summary\u003e\n\nThis nodepack is close to production for Windows CUDA users who already have matching extension wheels installed, but it is not a one-click package for every environment. Release readiness depends on these checks:\n\n| Check | Status |\n|-------|--------|\n| ComfyUI import without model download | Ready |\n| Native ComfyUI MoGe folder support | Ready |\n| RMBG-2.0 gated-model note | Ready |\n| GLB export with Windows-viewer-friendly PNG textures | Ready |\n| DynamicVRAM/Aimdo model wrapper | Ready |\n| Automatic CUDA wheel installation | Opt-in exact-match installer only |\n| Exact upstream NAF on Windows | Requires a real CUDA NATTEN/libnatten wheel; fallback mode works without it |\n| Fresh-machine smoke test | Recommended before tagging a production release |\n\nRun **Pixal3D Environment Check** first. A production-capable Windows install must show the required CUDA modules importing in the same Python environment that launches ComfyUI.\n\n\u003c/details\u003e\n\n## Model Setup\n\nThe loader looks for the Pixal3D model in:\n\n```text\nComfyUI/models/Pixal3D/TencentARC_Pixal3D/\n```\n\nSet `download_if_missing` to `true` on **Pixal3D Model Loader** to download `TencentARC/Pixal3D` there. The default is `false`, so the node will not download anything unless you ask it to.\n\nHelper models also live under `ComfyUI/models/Pixal3D/`, except native ComfyUI MoGe:\n\n```text\nComfyUI/models/Pixal3D/briaai_RMBG-2.0/\nComfyUI/models/Pixal3D/camenduru_dinov3-vitl16-pretrain-lvd1689m/\n```\n\nPreferred clean folder names are `owner_repo`, because Windows folders cannot use the Hugging Face slash. The Pixal3D/RMBG/DINO helpers also check common manual-download names such as `RMBG-2.0`, `Pixal3D`, and Hugging Face cache-style folders like `models--owner--repo/snapshots/\u003chash\u003e/`.\n\nThe model folders may be normal directories, Windows junctions, or symlinks. Broken links will be treated as missing models. Linked folders should still expose normal files such as `pipeline.json`, `.json`, `.safetensors`, and `ckpts/*.safetensors`; blob-only Hugging Face cache folders are not enough.\n\nNative ComfyUI MoGe uses [Comfy-Org/MoGe](https://huggingface.co/Comfy-Org/MoGe). Place the files directly in `ComfyUI/models/moge/`:\n\n```text\nComfyUI/\n└── models/\n    └── moge/\n        ├── moge_1_vitl_fp16.safetensors\n        └── moge_2_vitl_normal_fp16.safetensors\n```\n\nPixal3D-ComfyUI uses `moge_2_vitl_normal_fp16.safetensors` from that native ComfyUI folder for `camera_mode=moge`. It does not use a `Ruicheng/moge-2-vitl` snapshot folder.\n\nIf `download_if_missing` is `true`, Pixal3D-ComfyUI downloads missing Comfy-Org/MoGe files into `ComfyUI/models/moge/` and missing Pixal3D helper snapshots into `ComfyUI/models/Pixal3D/`. If `download_if_missing` is `false`, Pixal3D-ComfyUI will not download these helper models. `hf_endpoint` defaults to `https://huggingface.co`; set it to a mirror such as `https://hf-mirror.com` when downloading from regions where the normal Hugging Face domain is blocked.\n\n`briaai/RMBG-2.0` is gated on Hugging Face. To use `background_mode=auto_remove`, accept the model terms and either log in with Hugging Face before launching ComfyUI, set `HF_TOKEN`, or place the downloaded snapshot in `ComfyUI/models/Pixal3D/briaai_RMBG-2.0/`.\n\nNode downloads remove Hugging Face `.cache` metadata and `.git` folders after use. MoGe files are placed directly in `ComfyUI/models/moge/`; other helper snapshots are normalized under `ComfyUI/models/Pixal3D/`. The runtime expects normal files such as `.safetensors`, `pipeline.json`, and `ckpts/*.safetensors`, not blob-only cache directories.\n\nTorch Hub helper code for Pixal3D's NAF upsampler is redirected to:\n\n```text\nComfyUI/models/Pixal3D/torch_hub/\n```\n\nOfficial TencentARC Pixal3D uses NAF for the shape and texture stages, and those released weights expect 2048-channel projected features. On Windows stacks without a matching CUDA NATTEN build, leave `naf_mode=fallback_if_missing`: the node keeps the expected tensor shape by duplicating DINO projection features and will not download NAF when `download_if_missing=false`. For exact upstream NAF behavior, install a CUDA-enabled NATTEN build for your Python/PyTorch/CUDA stack and set `naf_mode=strict`.\n\n`naf_target_size` controls real NAF only. Leave it on `upstream` for normal behavior. Lower values such as `512`, `256`, or `128` reduce NAF VRAM if strict NAF works, and are ignored by fallback mode.\n\n## Windows CUDA Wheels\n\nThis node uses ComfyUI's Python environment only. Do not install these packages into system Python.\n\nFor Windows, FlashAttention 2 is enough if the wheel matches your exact Python, PyTorch, and CUDA build:\n\n```bash\ncd C:\\path\\to\\ComfyUI\n.\\venv\\Scripts\\python.exe -m pip show flash-attn\n```\n\nPortable ComfyUI example:\n\n```bash\ncd ComfyUI_windows_portable\n.\\python_embeded\\python.exe -m pip install -r .\\ComfyUI\\custom_nodes\\Pixal3D-ComfyUI\\requirements.txt\n```\n\nuv example:\n\n```bash\nuv pip install --python C:\\path\\to\\ComfyUI\\venv\\Scripts\\python.exe -r C:\\path\\to\\ComfyUI\\custom_nodes\\Pixal3D-ComfyUI\\requirements.txt\n```\n\nIf you use package-style install instead of requirements files, `pip install .` and `uv pip install .` install the same baseline runtime dependencies from `pyproject.toml`, including plain `natten==0.21.6`. CUDA wheels still remain manual or opt-in because they must match the exact stack.\n\nFlashAttention 3 also works if you install a wheel that provides `flash_attn_interface` for your exact Python, CUDA, and PyTorch build. Use the loader's `attention_backend` dropdown:\n\n| Option | What it uses |\n|--------|--------------|\n| `auto` | FlashAttention 3 if present, otherwise FlashAttention 2 |\n| `flash_attn_2` | `flash_attn` |\n| `flash_attn_3` | `flash_attn_interface` |\n\nPixal3D sparse attention needs FlashAttention 2 or 3; plain PyTorch SDPA is not enough for the sparse stages.\n\n## Nodes\n\n### Pixal3D Model Loader\n\nLoads the Pixal3D pipeline and returns a Comfy-managed model handle.\n\n| Parameter | Default | Description |\n|-----------|---------|-------------|\n| `model_repo` | `TencentARC/Pixal3D` | Hugging Face repo or local Pixal3D model folder |\n| `hf_endpoint` | `https://huggingface.co` | Hugging Face endpoint or mirror used when `download_if_missing` is enabled |\n| `attention_backend` | `auto` | `auto`, `flash_attn_2`, or `flash_attn_3` |\n| `vram_mode` | `dynamic_vram` | Best-effort Comfy/Aimdo staging with Comfy-aware torch ops |\n| `download_if_missing` | `false` | Downloads Pixal3D/helper models into `ComfyUI/models/Pixal3D/` and native MoGe into `ComfyUI/models/moge/` only when enabled |\n| `load_moge` | `true` | Load MoGe for automatic camera estimation |\n| `load_rembg` | `true` | Load gated RMBG-2.0 for built-in background removal |\n| `naf_mode` | `fallback_if_missing` | Use duplicated DINO features if CUDA NATTEN/NAF is unavailable; `strict` requires real NAF |\n| `naf_target_size` | `upstream` | Real NAF upsample size; lower values reduce VRAM and are ignored by fallback mode |\n| `preload_naf` | `false` | Preload the NAF upsampler only when `naf_mode=strict` and CUDA NATTEN/libnatten is available |\n| `force_reload` | `false` | Rebuild the cached model handle |\n\nAdvanced source override: set `PIXAL3D_REPO_PATH` before launching ComfyUI if you want to use a different Pixal3D source checkout instead of the vendored source.\n\n### Pixal3D Camera Control\n\nOutputs one bundled native ComfyUI value for manual camera mode:\n\n| Output | Connect to |\n|--------|------------|\n| `manual_fov` | Optional single cable to `Pixal3D Image To 3D.manual_fov` |\n\nThe optional `image` input is preview-only for the camera widget. Connect the same `Load Image` node directly to `Pixal3D Image To 3D.image`.\n\nThis node only affects Pixal3D when `Pixal3D Image To 3D.camera_mode=manual`. Connect `manual_fov` to `Pixal3D Image To 3D.manual_fov`; when it is connected in manual mode, the scalar `manual_camera_angle_x`, `manual_distance`, and `mesh_scale` inputs on `Pixal3D Image To 3D` are ignored and the Camera Control values are used instead. If `camera_mode=moge`, the connected `manual_fov` is ignored, so MoGe still owns the camera estimate.\n\nThe widget has a Scene view for the camera rig and a POV view for the framed camera result. POV uses the same horizontal FOV, distance, and mesh scale values that the node sends to Pixal3D. Horizontal FOV is converted to radians for `manual_camera_angle_x`, and distance/scale are passed through unchanged. It does not expose fake yaw/elevation controls because the Pixal3D manual path does not consume those values.\n\n### Pixal3D Image To 3D\n\nRuns Pixal3D from a ComfyUI `IMAGE`.\n\nBackground handling is built in:\n\n| Mode | Description |\n|------|-------------|\n| `auto_remove` | Default. Uses Pixal3D/rembg to remove the background when the image has no alpha |\n| `keep_alpha` | Uses the input alpha mask when present; falls back to auto remove if there is no alpha |\n| `none` | Sends the RGB image through without background removal |\n\nOutputs:\n\n| Output | Description |\n|--------|-------------|\n| `pixal3d_result` | Pixal3D mesh, voxel attributes, latents, and camera metadata |\n\n### Pixal3D Export GLB\n\nExports a textured GLB from `pixal3d_result`.\n\nOutput files are written to `ComfyUI/output/`.\n\nExports use embedded PNG textures for broad Windows/glTF viewer compatibility. Older Pixal3D-ComfyUI builds wrote WebP textures with `EXT_texture_webp`; those GLBs could open in some tools but fail in Windows' built-in 3D viewer.\n\n`remesh` defaults to `true` to match the Pixal3D demo path and is passed through exactly as set in the node. If cleanup creates fragmented meshes, try `remesh=false` and keep the upstream-style export defaults: `decimation_target=1000000` and `texture_size=4096`.\n\nConnect:\n\n```text\nPixal3D Export GLB glb_path\n  -\u003e Preview 3D \u0026 Animation model_file\n```\n\n### Pixal3D Unload Model\n\nFully removes the loaded Pixal3D model handle from ComfyUI model management and Pixal3D-ComfyUI's Python cache. Use this when you want CPU RAM released, not only VRAM offloaded.\n\n## Recommended Workflow\n\n```text\nLoad Image\n  -\u003e Pixal3D Image To 3D image\n\nLoad Image\n  -\u003e Pixal3D Camera Control image  (optional preview only)\n\nPixal3D Model Loader\n  -\u003e Pixal3D Image To 3D model\n\nPixal3D Camera Control manual_fov\n  -\u003e Pixal3D Image To 3D manual_fov\n\nPixal3D Image To 3D\n  -\u003e Pixal3D Export GLB\n```\n\nOptional preview:\n\n```text\nPixal3D Export GLB glb_path\n  -\u003e Preview 3D \u0026 Animation model_file\n```\n\n\u003cdetails\u003e\n\u003csummary\u003eVRAM Modes\u003c/summary\u003e\n\n`dynamic_vram` is the default loader mode. Pixal3D-ComfyUI builds Pixal3D with Comfy/Aimdo-aware `Linear`, `Conv`, `LayerNorm`, `GroupNorm`, and `Embedding` ops where possible, then wraps the pipeline in ComfyUI's model-management path.\n\nPixal3D-ComfyUI keeps one active pipeline cache. Changing Model Loader settings such as `vram_mode`, `attention_backend`, helper-model toggles, or NAF settings unloads and destroys the previous handle before loading the new one. The **Pixal3D Unload Model** node also removes the active handle from the Python cache. Pixal3D-ComfyUI also hooks ComfyUI's global unload button so native unloads clear the Pixal3D cache too.\n\nTask Manager may still show some RAM held after unload because Python, PyTorch, memory-mapped safetensors, Hugging Face/Transformers imports, and Windows allocators can keep reserved pages for reuse. That is different from the Pixal3D model object still being referenced. A full ComfyUI restart is the only guaranteed way to return every reserved page to the OS immediately.\n\nPixal3D is still not fully Comfy-native: it has custom sparse kernel modules and large temporary tensors that Aimdo cannot virtualize like a normal Comfy UNet. If Comfy still reports a large `Force pre-loaded` value or a 1536 run OOMs, use `native_low_vram` as the fallback. That mode can run in low VRAM ranges such as **4–8 GB VRAM** for smaller workflows, but it needs a lot of host memory: plan for **20–40 GB system RAM** and slower runs. It bypasses Comfy's bulk model load and lets Pixal3D move stages to GPU only when needed. In that mode the background remover is moved to GPU only for preprocessing and then returned to CPU, MoGe is moved to GPU only for camera estimation and then returned to CPU, and the upstream Pixal3D pipeline stages its flow/decoder modules one at a time.\n\nRecommended low-VRAM setup:\n\n| Node | Setting |\n|------|---------|\n| Pixal3D Model Loader | `vram_mode=native_low_vram` |\n| Pixal3D Model Loader | `load_moge=false` |\n| Pixal3D Model Loader | `load_rembg=false` |\n| Pixal3D Image To 3D | `camera_mode=manual` |\n| Pixal3D Image To 3D | `background_mode=keep_alpha` for transparent PNG/WebP inputs |\n| Pixal3D Camera Control | Connect `manual_fov` to `Pixal3D Image To 3D.manual_fov` |\n\nFor this path, use a transparent-background PNG or WebP so Pixal3D does not need RMBG, and use **Pixal3D Camera Control** instead of MoGe for camera setup.\n\nUse `full_gpu` only when you want the whole model resident on the GPU and your card has enough free VRAM.\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003eSafety Notes\u003c/summary\u003e\n\n- No packages are installed into system Python.\n- `install.py` uses the Python that launches it; use ComfyUI's Python or portable `python_embeded`.\n- There is no `prestartup_script.py`.\n- CUDA wheels are not installed automatically unless an exact known wheel path is explicitly enabled.\n- The model is not downloaded during ComfyUI startup.\n- `download_if_missing` is off by default.\n- Pixal3D source imports are lazy; the heavy model code loads only when the loader node runs.\n- CUDA module aliases for `cumesh`, `flex_gemm`, and `o_voxel` are created only when Pixal3D loads, and point at the installed wheel modules.\n\n\u003c/details\u003e\n\n## Windows CUDA Wheel Resources\n\n- [PozzettiAndrea/cuda-wheels](https://github.com/PozzettiAndrea/cuda-wheels/releases) — direct Windows wheels for `flex_gemm_ap`, `cumesh_vb`, `o_voxel_vb_ap`, `drtk`, and some `flash_attn` builds.\n- [visualbruno/ComfyUI-Trellis2 wheels](https://github.com/visualbruno/ComfyUI-Trellis2/tree/main/wheels) — alternate Windows wheels for `flex_gemm`, `cumesh`, `o_voxel`, `nvdiffrast`, `nvdiffrec_render`, and some NATTEN builds.\n- [Wildminder/AI-windows-whl](https://huggingface.co/Wildminder/AI-windows-whl/tree/main) — Windows AI wheel index, especially useful for FlashAttention and related AI packages.\n- [lldacing/NATTEN-windows](https://huggingface.co/lldacing/NATTEN-windows/tree/main) — Windows NATTEN wheels where available; strict NAF still requires `natten.HAS_LIBNATTEN == True`.\n\n## 🤗 Acknowledgements\n\nThis project is heavily built upon [Trellis.2](https://github.com/microsoft/TRELLIS.2) and [Direct3D-S2](https://github.com/DreamTechAI/Direct3D-S2). We sincerely thank the authors for their outstanding work on scalable 3D generation, which serves as the foundation of our codebase and model architecture.\n\nWe also thank the following repos for their great contributions:\n\n- [Direct3D-S2](https://github.com/DreamTechAI/Direct3D-S2)\n- [Trellis](https://github.com/microsoft/TRELLIS)\n- [Trellis.2](https://github.com/microsoft/TRELLIS.2)\n\n## 📄 Citation\n\nIf you find this work useful, please consider citing:\n\n```bibtex\n@article{li2026pixal3d,\n    title={Pixal3D: Pixel-Aligned 3D Generation from Images},\n    author={Li, Dong-Yang and Zhao, Wang and Chen, Yuxin and Hu, Wenbo and Guo, Meng-Hao and Zhang, Fang-Lue and Shan, Ying and Hu, Shi-Min},\n    journal={arXiv preprint arXiv:2605.10922},\n    year={2026}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsaganaki22%2Fpixal3d-comfyui","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsaganaki22%2Fpixal3d-comfyui","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsaganaki22%2Fpixal3d-comfyui/lists"}