{"id":31505691,"url":"https://github.com/sashakolpakov/dire-rapids","last_synced_at":"2026-03-05T23:42:14.273Z","repository":{"id":312956900,"uuid":"1049255942","full_name":"sashakolpakov/dire-rapids","owner":"sashakolpakov","description":"DiRe accelerated by PyTorch, PyKeOps and cuVS ","archived":false,"fork":false,"pushed_at":"2026-02-24T01:30:21.000Z","size":35086,"stargazers_count":8,"open_issues_count":3,"forks_count":1,"subscribers_count":1,"default_branch":"main","last_synced_at":"2026-02-24T01:37:56.998Z","etag":null,"topics":["cuda","cuda-kernels","dimensionality-reduction","pykeops","pytorch","rapidsai","t-sne","umap"],"latest_commit_sha":null,"homepage":"https://arxiv.org/abs/2503.03156","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/sashakolpakov.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":"CITATION.cff","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-09-02T17:56:28.000Z","updated_at":"2026-02-24T01:23:08.000Z","dependencies_parsed_at":null,"dependency_job_id":"fe8f3ed1-313d-44eb-a625-d1cdb7853e3c","html_url":"https://github.com/sashakolpakov/dire-rapids","commit_stats":null,"previous_names":["sashakolpakov/dire-rapids"],"tags_count":3,"template":false,"template_full_name":null,"purl":"pkg:github/sashakolpakov/dire-rapids","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sashakolpakov%2Fdire-rapids","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sashakolpakov%2Fdire-rapids/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sashakolpakov%2Fdire-rapids/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sashakolpakov%2Fdire-rapids/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/sashakolpakov","download_url":"https://codeload.github.com/sashakolpakov/dire-rapids/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sashakolpakov%2Fdire-rapids/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":30156180,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-03-05T22:39:40.138Z","status":"ssl_error","status_checked_at":"2026-03-05T22:39:24.771Z","response_time":93,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.6:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["cuda","cuda-kernels","dimensionality-reduction","pykeops","pytorch","rapidsai","t-sne","umap"],"created_at":"2025-10-02T20:13:40.834Z","updated_at":"2026-03-05T23:42:14.234Z","avatar_url":"https://github.com/sashakolpakov.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003c!-- Logo + Project title --\u003e\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"images/dire_rapids_logo.png\" alt=\"DiRe-RAPIDS logo\" width=\"280\" style=\"margin-bottom:10px;\"\u003e\n\u003c/p\u003e\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://opensource.org/licenses/Apache-2.0\"\u003e\n    \u003cimg alt=\"License\" src=\"https://img.shields.io/badge/License-Apache%202.0-blue.svg\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://www.python.org/downloads/\"\u003e\n    \u003cimg alt=\"Python 3.8+\" src=\"https://img.shields.io/badge/python-3.8+-blue.svg\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://pypi.org/project/dire-rapids/\"\u003e\n    \u003cimg alt=\"PyPI\" src=\"https://img.shields.io/pypi/v/dire-rapids.svg\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://pepy.tech/projects/dire-rapids\"\u003e\n    \u003cimg src=\"https://static.pepy.tech/personalized-badge/dire-rapids?period=total\u0026units=ABBREVIATION\u0026left_color=GREY\u0026right_color=BLUE\u0026left_text=downloads\" alt=\"PyPI Downloads\"\u003e\n  \u003c/a\u003e\n\u003c/p\u003e\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://github.com/sashakolpakov/dire-rapids/actions/workflows/pylint.yml\"\u003e\n    \u003cimg alt=\"CI\" src=\"https://img.shields.io/github/actions/workflow/status/sashakolpakov/dire-rapids/pylint.yml?branch=main\u0026label=CI\u0026logo=github\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://github.com/sashakolpakov/dire-rapids/actions/workflows/deploy_docs.yml\"\u003e\n    \u003cimg alt=\"Docs\" src=\"https://img.shields.io/github/actions/workflow/status/sashakolpakov/dire-rapids/deploy_docs.yml?branch=main\u0026label=Docs\u0026logo=github\"\u003e\n  \u003c/a\u003e\n  \u003ca href=\"https://sashakolpakov.github.io/dire-rapids/\"\u003e\n    \u003cimg alt=\"Docs Live\" src=\"https://img.shields.io/website-up-down-green-red/https/sashakolpakov.github.io/dire-rapids?label=API%20Documentation\"\u003e\n  \u003c/a\u003e\n\u003c/p\u003e\n\n# DiRe Rapids\n\nGPU-accelerated implementation of [DiRe](https://github.com/sashakolpakov/dire-jax) using PyTorch and optionally NVIDIA RAPIDS for massive-scale datasets.\n\n## Installation\n\n### From Repository (development)\n\n```bash\n# Clone the repository\ngit clone https://github.com/sashakolpakov/dire-rapids.git\ncd dire-rapids\n\n# Basic installation (CPU + PyTorch)\npip install -e .\n\n# With CUDA support\npip install -e .[cuda]\n\n# For development (includes testing and dev tools)\npip install -e .[dev]\n```\n\n#### With RAPIDS Support (Optional, GPU only)\n\nFirst, install RAPIDS following provided [installation instructions](https://docs.rapids.ai/install/). \n```bash\n# Then install dire-rapids with RAPIDS support\npip install -e .[rapids]\n```\n\n### From PyPI (stable)\n\n```bash\n# Use the above installation options\npip install dire-rapids[options]\n```\n\n## Quick Start [![Open in Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/sashakolpakov/dire-rapids/blob/main/benchmarking/dire_rapids_benchmarks.ipynb)\n\nYou can import the standard or memory-efficient backend for DiRe. Also, some datasets is needed: we shall use higher-dimensional Blobs as a simple visual test. \n\n```python\nfrom dire_rapids import DiRePyTorch, DiRePyTorchMemoryEfficient\nfrom sklearn.datasets import make_blobs\n```\nThe standard backend will work for the example below, but not necessarily for a larger (100x) dataset. \n\n```python\n# Generate sample data\nX, _ = make_blobs(n_samples=1_000, centers=12, n_features=10, random_state=42)\n\n# Standard PyTorch implementation\nreducer = DiRePyTorch(n_components=2, n_neighbors=16, verbose=True)\nX_embedded = reducer.fit_transform(X)\n```\nThe memory-efficient version gets you there (how soon, depends on the hardware). \n\n```python\nreducer = DiRePyTorchMemoryEfficient(n_components=2, n_neighbors=16, verbose=True)\nX_embedded = reducer.fit_transform(X)\n```\n\n### Custom Distance Metrics\n\nDiRe Rapids now supports custom distance metrics for k-nearest neighbor computation, while keeping the layout forces Euclidean for optimal embedding quality:\n\n```python\n# Using L1 (Manhattan) distance for k-NN\nreducer = DiRePyTorch(\n    metric='(x - y).abs().sum(-1)',\n    n_neighbors=32,\n    verbose=True\n)\nX_embedded = reducer.fit_transform(X)\n\n# Using cosine distance for k-NN\ndef cosine_distance(x, y):\n    return 1 - (x * y).sum(-1) / (x.norm(dim=-1, keepdim=True) * y.norm(dim=-1, keepdim=True) + 1e-8)\n\nreducer = DiRePyTorch(\n    metric=cosine_distance,\n    n_neighbors=32\n)\nX_embedded = reducer.fit_transform(X)\n\n# Custom metrics work with all backends\nreducer = DiRePyTorchMemoryEfficient(\n    metric='(x - y).abs().sum(-1)',  # L1 distance\n    use_fp16=True,\n    n_neighbors=32\n)\nX_embedded = reducer.fit_transform(X)\n```\n\n**Supported Metric Types:**\n- **None** or `'euclidean'`/`'l2'`: Fast built-in Euclidean distance (default)\n- **String expressions**: Evaluated tensor expressions (e.g., `'(x - y).abs().sum(-1)'` for L1)\n- **Callable functions**: Custom Python functions taking (x, y) tensors\n\nAfter starting the above example, you should see a verbose output similar to the below:\n\n```python\n[KeOps] Compiling cuda jit compiler engine ... OK\n[pyKeOps] Compiling nvrtc binder for python ... OK\n2025-09-04 16:03:54.409 | INFO     | dire_rapids.dire_cuvs:\u003cmodule\u003e:25 - cuVS available - GPU-accelerated k-NN enabled\n2025-09-04 16:03:59.060 | INFO     | dire_rapids.dire_cuvs:\u003cmodule\u003e:36 - cuML available - GPU-accelerated PCA enabled\n2025-09-04 16:03:59.581 | INFO     | dire_rapids.dire_pytorch:__init__:105 - Using CUDA device: Tesla T4\n2025-09-04 16:03:59.581 | INFO     | dire_rapids.dire_pytorch_memory_efficient:__init__:89 - Memory-efficient mode enabled\n2025-09-04 16:03:59.582 | INFO     | dire_rapids.dire_pytorch_memory_efficient:__init__:91 - FP16 enabled for k-NN computation\n2025-09-04 16:03:59.583 | INFO     | dire_rapids.dire_pytorch_memory_efficient:__init__:93 - PyKeOps repulsion enabled (threshold: 50000 points)\n2025-09-04 16:03:59.598 | INFO     | dire_rapids.dire_pytorch_memory_efficient:fit_transform:302 - Memory-efficient processing: 100000 samples, 100 features\n2025-09-04 16:03:59.599 | INFO     | dire_rapids.dire_pytorch_memory_efficient:fit_transform:306 - Large dataset (100000 \u003e 50000): using random sampling for repulsion\n2025-09-04 16:03:59.614 | INFO     | dire_rapids.dire_pytorch:fit_transform:476 - Processing 100000 samples with 100 features\n2025-09-04 16:03:59.619 | INFO     | dire_rapids.dire_pytorch:_find_ab_params:123 - Found kernel params: a=1.8956, b=0.8006\n2025-09-04 16:03:59.619 | INFO     | dire_rapids.dire_pytorch_memory_efficient:_compute_knn:109 - Forcing FP16 for large dataset (100000 samples, 100D)\n2025-09-04 16:03:59.834 | INFO     | dire_rapids.dire_pytorch_memory_efficient:_compute_knn:123 - Memory-efficient k-NN: chunk_size=11790, FP16=True\n2025-09-04 16:03:59.834 | INFO     | dire_rapids.dire_pytorch:_compute_knn:138 - Computing 16-NN graph for 100000 points in 100D...\n2025-09-04 16:03:59.835 | INFO     | dire_rapids.dire_pytorch:_compute_knn:150 - Using FP16 for k-NN (2x memory, faster on H100/A100)\n2025-09-04 16:03:59.893 | INFO     | dire_rapids.dire_pytorch:_compute_knn:166 - Using PyTorch for k-NN\n2025-09-04 16:03:59.893 | INFO     | dire_rapids.dire_pytorch:_compute_knn:186 - Using chunk size: 23580 (GPU memory: 14.6GB, dtype: torch.float16)\n2025-09-04 16:03:59.894 | INFO     | dire_rapids.dire_pytorch:_compute_knn:197 - Processing chunk 1/5\n2025-09-04 16:04:00.665 | INFO     | dire_rapids.dire_pytorch:_compute_knn:197 - Processing chunk 2/5\n2025-09-04 16:04:00.962 | INFO     | dire_rapids.dire_pytorch:_compute_knn:197 - Processing chunk 3/5\n2025-09-04 16:04:01.259 | INFO     | dire_rapids.dire_pytorch:_compute_knn:197 - Processing chunk 4/5\n2025-09-04 16:04:01.556 | INFO     | dire_rapids.dire_pytorch:_compute_knn:197 - Processing chunk 5/5\n2025-09-04 16:04:01.636 | INFO     | dire_rapids.dire_pytorch:_compute_knn:237 - k-NN graph computed: shape (100000, 16)\n2025-09-04 16:04:01.833 | INFO     | dire_rapids.dire_pytorch:_initialize_embedding:243 - Initializing with PCA\n2025-09-04 16:04:01.908 | INFO     | dire_rapids.dire_pytorch_memory_efficient:_optimize_layout:253 - Memory-efficient optimization for 100000 points...\n2025-09-04 16:04:01.921 | INFO     | dire_rapids.dire_pytorch_memory_efficient:_optimize_layout:259 - Initial GPU memory: 0.01/15.8 GB\n2025-09-04 16:04:02.097 | DEBUG    | dire_rapids.dire_pytorch_memory_efficient:_compute_forces:207 - Using random sampling for repulsion\n2025-09-04 16:04:02.272 | INFO     | dire_rapids.dire_pytorch_memory_efficient:_optimize_layout:272 - Iteration 0/128, avg force: 14.770476\n2025-09-04 16:04:02.288 | DEBUG    | dire_rapids.dire_pytorch_memory_efficient:_optimize_layout:281 - GPU memory: 0.01 GB\n2025-09-04 16:04:02.295 | DEBUG    | dire_rapids.dire_pytorch_memory_efficient:_compute_forces:207 - Using random sampling for repulsion\n2025-09-04 16:04:02.313 | DEBUG    | dire_rapids.dire_pytorch_memory_efficient:_compute_forces:207 - Using random sampling for repulsion\n2025-09-04 16:04:02.330 | DEBUG    | dire_rapids.dire_pytorch_memory_efficient:_compute_forces:207 - Using random sampling for repulsion\n2025-09-04 16:04:02.347 | DEBUG    | dire_rapids.dire_pytorch_memory_efficient:_compute_forces:207 - Using random sampling for repulsion\n```\n\nThe final result is the expected image of 2D blobs\n\n![12 blobs with 100k points embedded in dimension 2](images/blobs_layout.png)\n\n### Available Backends\n\n- **DiRePyTorch**: Standard PyTorch implementation with adaptive chunking\n- **DiRePyTorchMemoryEfficient**: Memory-optimized version with:\n  - FP16 support for 2x memory savings\n  - Point-by-point force computation\n  - More aggressive memory management\n  - PyKeOps LazyTensors for efficient repulsion (when available)\n- **DiReCuVS**: RAPIDS cuVS backend for massive-scale datasets\n\n### Auto Backend Selection\n\nUse the `create_dire()` function for automatic backend selection based on available hardware:\n\n```python\nfrom dire_rapids import create_dire\n\n# Auto-select optimal backend\nreducer = create_dire(\n    n_neighbors=32,\n    metric='(x - y).abs().sum(-1)',  # Custom L1 metric\n    verbose=True\n)\nX_embedded = reducer.fit_transform(X)\n\n# Force memory-efficient backend\nreducer = create_dire(\n    memory_efficient=True,\n    use_fp16=True,\n    metric=cosine_distance  # Custom callable metric\n)\nX_embedded = reducer.fit_transform(X)\n```\n\n## Testing\n\n```bash\n# Run basic CPU tests\npytest tests/test_cpu_basic.py -v\n\n# Run all tests\npytest tests/ -v\n```\n\n### Citation\n\nIf you use this work, please cite it as:\n\n**BibTeX:**\n```bibtex\n@misc{kolpakov-rivin-2025dimensionality,\n  title={Dimensionality reduction for homological stability and global structure preservation},\n  author={Kolpakov, Alexander and Rivin, Igor},\n  year={2025},\n  eprint={2503.03156},\n  archivePrefix={arXiv},\n  primaryClass={cs.LG},\n  url={https://arxiv.org/abs/2503.03156}\n}\n```\n\n**APA Style:**\n```\nKolpakov, A., \u0026 Rivin, I. (2025). Dimensionality reduction for homological stability and global structure preservation. arXiv preprint arXiv:2503.03156. https://arxiv.org/abs/2503.03156\n```\n\n## Requirements\n\n- Python 3.8-3.12\n- PyTorch 2.0+\n- PyKeOps 2.1+\n- NumPy, SciPy, scikit-learn\n- (Optional) CUDA 12.x+ for GPU acceleration\n- (Optional) RAPIDS 23.08+ for cuVS backend\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://submitaitools.org/github-com-sashakolpakov-dire-rapids/\"\u003e\n    \u003cimg src=\"https://submitaitools.org/static_submitaitools/images/submitaitools.png\"\n         alt=\"DiRe-RAPIDS: Fast Dimensionality Reduction on the GPU\" height=\"60\" /\u003e\n  \u003c/a\u003e\n\u003c/p\u003e\n\n                                       \n            \n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsashakolpakov%2Fdire-rapids","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsashakolpakov%2Fdire-rapids","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsashakolpakov%2Fdire-rapids/lists"}