{"id":22731070,"url":"https://github.com/genentech/islander","last_synced_at":"2025-04-14T00:31:18.197Z","repository":{"id":231364304,"uuid":"727744084","full_name":"Genentech/Islander","owner":"Genentech","description":null,"archived":false,"fork":false,"pushed_at":"2024-07-11T04:19:07.000Z","size":7356,"stargazers_count":5,"open_issues_count":0,"forks_count":0,"subscribers_count":4,"default_branch":"main","last_synced_at":"2024-07-11T05:30:09.043Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Genentech.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-12-05T13:42:16.000Z","updated_at":"2024-07-11T05:30:14.530Z","dependencies_parsed_at":"2024-07-11T05:40:17.258Z","dependency_job_id":null,"html_url":"https://github.com/Genentech/Islander","commit_stats":null,"previous_names":["genentech/islander"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Genentech%2FIslander","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Genentech%2FIslander/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Genentech%2FIslander/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Genentech%2FIslander/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Genentech","download_url":"https://codeload.github.com/Genentech/Islander/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":229117070,"owners_count":18022819,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-12-10T19:19:20.391Z","updated_at":"2025-04-14T00:31:18.191Z","avatar_url":"https://github.com/Genentech.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"## Islander\nThis repository is the official implementation for the paper **Metric Mirages in Cell Embeddings**. \n\nPlease contact wang.hanchen@gene.com or hanchenw@cs.stanford.edu if you have any questions.\n\n\n\n![teaser](teaser.png)\n\n\n\n### Citation\n\n```bibtex\n@article {Islander,\n\tauthor = {Hanchen Wang and Jure Leskovec and Aviv Regev},\n\ttitle = {Metric Mirages in Cell Embeddings},\n\tdoi = {10.1101/2024.04.02.587824},\n\tpublisher = {Cold Spring Harbor Laboratory},\n\tURL = {https://www.biorxiv.org/content/early/2024/04/02/2024.04.02.587824}\n\tjournal = {bioRxiv},\n\tyear = {2024},\n}\n```\n\n\n\n---\n\n\n\n### Usage\n\nWe include scripts and logs to reproduce the results in the \u003ca href=\"scripts/\"\u003escripts\u003c/a\u003e folder. You can also follow the step-by-step instructions below:\n\n**Step 0**: Set up the environment.\n\n```bash\nconda env create -f env.yml\n```\n\n**NOTE**: The default setup uses GPU-compiled packages (for PyTorch, JAXlib, *etc*.). Please adjust them according to your local CUDA version *or* switch to the CPU version as needed. The calculation of scGraph scores does not require GPU access.\n\n**Step 1**: Preprocessing. Data can be downloaded from:\n\n| Brain                                                        | Breast                                                       | COVID                                                        | Eye                                                          | FetalGut                                                     | FetalLung                                                    | Heart                                                        | Lung                                                         | Pancreas                                                     | Skin                                                         |\n| ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ |\n| [Paper](https://www.science.org/doi/10.1126/science.add7046) | [Paper](https://www.nature.com/articles/s41586-023-06252-9)  | [Paper](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7402042/) | [Paper](https://www.sciencedirect.com/science/article/pii/S2666979X22001069?via%3Dihub) | [Paper](https://www.sciencedirect.com/science/article/pii/S1534580720308868?via%3Dihub) | [Paper](https://linkinghub.elsevier.com/retrieve/pii/S0092867422014155) | [Paper](https://www.nature.com/articles/s44161-022-00183-w)  | [Paper](https://www.nature.com/articles/s41591-023-02327-2)  | [Paper](https://www.nature.com/articles/s41592-021-01336-8)  | [Paper](https://www.nature.com/articles/s42003-020-0922-4)   |\n| [Data](https://cellxgene.cziscience.com/collections/283d65eb-dd53-496d-adb7-7570c7caa443) | [Data](https://cellxgene.cziscience.com/collections/4195ab4c-20bd-4cd3-8b3d-65601277e731) | [Data](https://atlas.fredhutch.org/fredhutch/covid/)         | [Data](https://cellxgene.cziscience.com/collections/348da6dc-5bf6-435d-adc5-37747b9ae38a) | [Data](https://cellxgene.cziscience.com/collections/17481d16-ee44-49e5-bcf0-28c0780d8c4a) | [Data](https://cellxgene.cziscience.com/collections/2d2e2acd-dade-489f-a2da-6c11aa654028) | [Data](https://cellxgene.cziscience.com/collections/43b45a20-a969-49ac-a8e8-8c84b211bd01) | [Data](https://cellxgene.cziscience.com/collections/6f6d381a-7701-4781-935c-db10d30de293) | [Data](https://figshare.com/articles/dataset/Benchmarking_atlas-level_data_integration_in_single-cell_genomics_-_integration_task_datasets_Immune_and_pancreas_/12420968?file=24539828) | [Data](https://cellxgene.cziscience.com/collections/c353707f-09a4-4f12-92a0-cb741e57e5f0) |\n\n\n\nWe applied quality control to each dataset by filtering out cell profiles with fewer than 1,000 reads or fewer than 500 detected genes. Genes present in fewer than five cells were also excluded. Normalization was performed using Scanpy, where each cell’s read counts were scaled to a total of 10,000, followed by a log1p transformation:\n\n```python\n# download via \"wget -O data/breast/local.h5ad https://datasets.cellxgene.cziscience.com/b8b5be07-061b-4390-af0a-f9ced877a068.h5ad\"\nadata = sc.read_h5ad(dh.DATA_RAW_[\"breast\"])\nadata.X = adata.raw.X\nadata.layers[\"raw_counts\"] = adata.raw.X\ndel adata.raw\nuh.preprocess(adata)\n\n[Output]\nfiltered out 9954 cells that have less than 1000 counts\nfiltered out 865 cells that have less than 500 genes expressed\nfiltered out 3803 genes that are detected in less than 5 cells\n=============================================================================\n29431 genes x 703512 cells after quality control.\n=============================================================================\nnormalizing by total count per cell\n    finished (0:00:06): normalized adata.X and added    'n_counts', counts per cell before normalization (adata.obs)\n```\n\n\n\nThe top 1000 highly variable genes are selected through:\n\n```python\nsc.pp.highly_variable_genes(adata, subset=True, flavor=\"seurat_v3\", n_top_genes=1000)\n```\n\nThen metadata is saved as JSON files. See the minimal example: [Process_Breast.ipynb](jupyter_nb/Process_Breast.ipynb).\n\n\n\n**Step 2**: Run Islander and benchmark with scIB\n\n```bash\ncd ${HOME}/Islander/src\n\nexport LR=0.001\nexport EPOCH=10\nexport MODE=\"mixup\"\nexport LEAKAGE=16\nexport MLPSIZE=\"128 128\"\nexport DATASET_List=(\"lung\" \"lung_fetal_donor\" \"lung_fetal_organoid\" \\\n    \"brain\" \"breast\" \"heart\" \"eye\" \"gut_fetal\" \"skin\" \"COVID\" \"pancreas\")\n\nfor DATASET in \"${DATASET_List[@]}\"; do\nexport PROJECT=\"_${DATASET}_\"\nexport SavePrefix=\"${HOME}/Islander/models/${PROJECT}\"\nexport RUNNAME=\"MODE-${MODE}-ONLY_LEAK-${LEAKAGE}_MLP-${MLPSIZE}\"\necho \"DATASET-${DATASET}_${RUNNAME}\"\nmkdir -p $SavePrefix\n\n# === Training ===\npython scTrain.py \\\n    --gpu 3 \\\n    --lr ${LR} \\\n    --mode ${MODE} \\\n    --epoch ${EPOCH} \\\n    --dataset ${DATASET} \\\n    --leakage ${LEAKAGE} \\\n    --project ${PROJECT} \\\n    --mlp_size ${MLPSIZE} \\\n    --runname \"${RUNNAME}\" \\\n    --savename \"${SavePrefix}/${RUNNAME}\";\n\n# === Benchmarking ===\npython scBenchmarker.py \\\n    --islander \\\n    --saveadata \\\n    --dataset \"${DATASET}\" \\\n    --save_path \"${SavePrefix}/${RUNNAME}\";\ndone\n```\n\n\n\nWe have also provided variants of Islander, which make use of different forms of semi-supervised learning loss (triplet and supervised contrastive loss). See [scripts/_Islander_SCL.sh](scripts/_Islander_SCL.sh) and [scripts/_Islander_Triplet.sh](scripts/_Islander_Triplet.sh) for details.\n\n\n\n**Step 3**: Run integration methods and benchmark with scIB\n\n```bash\nexport DATASET_List=\"lung_fetal_donor\"\necho -e \"\\n\\n\"\n\necho \"DATASET-${DATASET}_HVG\"\nexport CUDA_VISIBLE_DEVICES=2 \u0026 python scBenchmarker.py \\\n    --all \\\n    --highvar \\\n    --saveadata \\\n    --dataset \"${DATASET}\" \\\n    --savecsv \"${DATASET}_FULL\" \\\n    --save_path \"${HOME}/Islander_dev/models/_${DATASET}_/MODE-mixup-ONLY_LEAK-16_MLP-128 128\";\n   \n# === highly variable genes ===\necho \"DATASET-${DATASET}_HVG\"\nexport CUDA_VISIBLE_DEVICES=2 \u0026 python scBenchmarker.py \\\n\t--all \\\n\t--highvar \\\n\t--saveadata \\\n\t--dataset \"${DATASET}\" \\\n\t--savecsv \"${DATASET}_HVG\" \\\n\t--save_path \"${HOME}/Islander_dev/models/_${DATASET}_/MODE-mixup-ONLY_LEAK-16_MLP-128 128\";\n```\n\n\n\n**Step 4**: Run and benchmark foundation models\n\nPlease refer to the authors' original tutorials ([scGPT](https://github.com/bowang-lab/scGPT/tree/main/tutorials/zero-shot), [Geneformer](https://huggingface.co/ctheodoris/Geneformer/tree/main/examples), [scFoundation](https://github.com/biomap-research/scFoundation/tree/main/model), [UCE](https://github.com/snap-stanford/UCE)) for extracting zero-shot and fine-tuned cell embeddings. We provide a minimal example notebook [_nb/Geneformer_Skin.ipynb](jupyter_nb/Geneformer_Skin.ipynb) to extract zero-shot cell embeddings for the skin dataset using pre-trained Geneformer. To evaluate such embedding with scIB:\n\n```bash\ncd ${HOME}/Islander/src\n\nexport DATASET=\"brain\"\necho -e \"\\n\\n\"\necho \"DATASET-${DATASET}_Geneformer\"\npython scBenchmarker.py \\\n    --obsm_keys Geneformer \\\n    --dataset \"${DATASET}\" \\\n    --savecsv \"${DATASET}_Geneformer\" \\\n    --save_path \"${HOME}/Islander/models/_${skin}_/MODE-mixup-ONLY_LEAK-16_MLP-128 128\";\n\n```\n\n\n\n**Step 5**: Benchmark with scGraph (can be replaced on customized AnnData file)\n\n```bash\ncd ${HOME}/Islander/src\n\npython scGraph.py \\\n    --adata_path ${HOME}/Islander/data/lung/emb.h5ad \\\n    --batch_key sample \\\n    --label_key cell_type \\\n    --savename ${HOME}/lung_scGraph;\n```\n\nThe output file is in the format **Corr-Weights**, reported as scGraph scores in the paper. It is based on weighted rank correlation, where the weights are inversely proportional to the inter-cluster centroid distances. **Corr-PCA** represents the rank correlation using equal weights. **Rank-PCA** represents rank differences.\n\n|             | Rank-PCA | Corr-PCA | Corr-Weights |\n| :---------- | -------: | -------: | -----------: |\n| Geneformer  |    0.610 |    0.799 |        0.498 |\n| Harmony     |    0.670 |    0.924 |        0.678 |\n| Harmony_hvg |    0.709 |    0.941 |        0.724 |\n| Islander    |    0.292 |    0.847 |        0.160 |\n\n\n\n**Parameter Settings in scGraph**:\n\nscGraph uses PCA for each cell type within each batch to represent cluster-cluster relationships. Batches with fewer than 100 cells or cell types with fewer than 10 cells are excluded. PCA is calculated on the 1,000 highly variable genes, after removing 10% of the cells (5% from each extreme).\n\nAll the numerical values mentioned above are adjustable. For further details, please refer to `scGraph.py`.\n\n---\n\n\n\n**Case study**:  scGraph vs scIB on fibroblast cells from the human fetal lung\n\nPlease see [Fibroblast_Case.ipynb](jupyter_nb/Fibroblast_Case.ipynb) to reproduce the results reported in the paper.\n\n\n\n### File Organization\n\n```\n├── LICENSE.txt\n├── README.md\n├── data\n├── env.yml  # for GPU environments\n├── jupyter_nb\n│   ├── Fibroblast_Case.ipynb\n│   ├── Geneformer_Skin.ipynb\n│   └── Process_Breast.ipynb\n├── meta\n│   ├── COVID\n│   │   ├── batch2cat.json\n│   │   └── cell2cat.json\n│   ├── ...\n├── res\n│   └── scGraph\n│   └── scIB\n├── scripts\n│   ├── Lung_Mixup.log \n│   ├── Lung_scGraph.log\n│   ├── _Islander_MixUp.sh\n│   ├── _Islander_SCL.sh\n│   ├── _Islander_Triplet.sh\n│   ├── _download_data.sh\n│   ├── _scBenchmark.sh\n│   ├── _scGraph.sh\n│   ├── _scIB_Geneformer.sh\n│   └── _scIB_Islander.sh  # scib benchmark on cell islands embeddings\n├── src\n│   ├── ArgParser.py\n│   ├── Data_Handler.py\n│   ├── Utils_Handler.py\n│   ├── Vis_Handler.py\n│   ├── __init__.py\n│   ├── scBenchmarker.py\n│   ├── scDataset.py\n│   ├── scFinetuner.py\n│   ├── scGraph.py\n│   ├── scLoss.py\n│   ├── scModel.py\n│   └── scTrain.py\n└── teaser.png\n```\n\n\n\n### Stand-alone scGraph package for evaluations\n\nWe also provided a standalone [scgraph](https://pypi.org/project/scgraph-eval/) python package, that can be installed via:\n\n```\n# conda create -n scgraph python=3.10 # to create another conda environment if necessary\npip install scgraph-eval\n```\n\n#### Python\n\n```python3\nfrom scgraph import scGraph\n\n# Initialize the graph analyzer\nscgraph = scGraph(\n    adata_path=\"path/to/your/data.h5ad\",   # Path to AnnData object\n    batch_key=\"batch\",                     # Column name for batch information\n    label_key=\"cell_type\",                 # Column name for cell type labels\n    trim_rate=0.05,                        # Trim rate for robust mean calculation\n    thres_batch=100,                       # Minimum number of cells per batch\n    thres_celltype=10,                     # Minimum number of cells per cell type\n    only_umap=True,                        # Only evaluate 2D embeddings (mostly umaps)\n)\n\n# Run the analysis, return a pandas dataframe\nresults = scgraph.main()\n\n# Save the results\nresults.to_csv(\"embedding_evaluation_results.csv\")\n```\n\n#### Command line\n\n```bash\nscgraph --adata_path path/to/data.h5ad --batch_key batch --label_key cell_type --savename results\n```\n\n#### Notebook\n\nWe provide a notebook on how to install the scgraph on a labtop, and reproduce the results on using scgraph to evaluate cell embeddings of fibroblast family in human fetal lung.\n\n- [Data](https://drive.google.com/file/d/1a2UF4V_INGMKayCoMErZG-_kq_KuZmjA/view?usp=drive_link) \n\n- [Notebook](jupyter_nb/scGraph-minimal.ipynb)\n\n#### Contributing\n\nContributions are welcome! Please feel free to submit a Pull Request.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fgenentech%2Fislander","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fgenentech%2Fislander","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fgenentech%2Fislander/lists"}