awesome-single-cell
Community-curated list of software packages and data resources for single-cell, including RNA-seq, ATAC-seq, etc.
https://github.com/seandavi/awesome-single-cell
Last synced: 5 days ago
JSON representation
-
Citation
-
Image-based profiling
-
Spatial transcriptomics
- Pycytominer - [Python] Pycytominer is a suite of common functions used to process high dimensional readouts from high-throughput cell experiments. Manuscript: [Reproducible image-based profiling with Pycytominer](https://www.nature.com/articles/s41592-025-02611-8)
-
-
Journal articles of general interest
-
Big data approach overview
-
Experimental design
- Sensei - type abundance change estimated from scRNA-seq experiment](https://www.biorxiv.org/content/10.1101/2020.05.31.126565v1).
- How to design a single-cell RNA-sequencing experiment: pitfalls, challenges and perspectives
- How to design a single-cell RNA-sequencing experiment: pitfalls, challenges and perspectives
- Design and computational analysis of single-cell RNA-sequencing experiments
-
Methods comparisons
- Bias, Robustness And Scalability In Differential Expression Analysis Of Single-Cell RNA-Seq Data - comparison of 36 statistical methods to detect differentially expressed genes between two annotated populations from the [conquer](http://imlspenticton.uzh.ch:3838/conquer/) database of consistently processed scRNA-seq datasets.
- Single-Cell RNA-Sequencing: Assessment of Differential Expression Analysis Methods - an assessment of main bulk and single-cell differential analysis methods used to analyze scRNA-seq data.
- A comparison of single-cell trajectory inference methods - Unsure which of the more than 70 **trajectory inference** methods to use for your single-cell dataset? We evaluated 45 methods based on four criteria: the accuracy of the trajectory, how scalable the method is, how stable its outputs are, and the usability of the tool. These are summarised in a *"funky heatmap"* (Figures 2 & 3). Check out [dynverse.org](https://dynverse.org) for more information.
- Evaluation of methods to assign cell type labels to cell clusters from single-cell RNA-sequencing data - In this study, we benchmarked five methods (CIBERSORT, GSEA, GSVA, ORA and METANEIGHBOR) for the task of assigning cell type labels to cell clusters from scRNA-seq data. We used five scRNA-seq datasets: human liver, 11 Tabula Muris mouse tissues, two human peripheral blood mononuclear cell datasets, and mouse retinal neurons, for which reference cell type signatures were available. Our results show that, in general, all five methods perform well in the task as evaluated by receiver operating characteristic curve analysis (average area under the curve (AUC) = 0.91, sd = 0.06), whereas precision-recall analyses show a wide variation depending on the method and dataset (average AUC = 0.53, sd = 0.24). GSVA was the overall top performer and was more robust in cell type signature subsampling simulations, although different methods performed well using different datasets. METANEIGHBOR and GSVA were the fastest methods.
- Evaluation of single-cell classifiers for single-cell RNA sequencing data sets - In this article, nine tools have been systematically compared. The article provides a guideline for researchers to select and apply suitable single cell and cluster classification tools in their analysis workflows and sheds some lights on potential direction of future improvement on classification tools.
- Benchmarking algorithms for gene regulatory network inference from single-cell transcriptomic data - a comparison of gene regulatory network inference methods using simulated and real single-cell RNA-seq datasets
- Comparison of computational methods for imputing single-cell RNA-sequencing data - We compared eight imputation methods, evaluated their power in recovering original real data, and performed broad analyses to explore their effects on clustering cell types, detecting differentially expressed genes, and reconstructing lineage trajectories in the context of both simulated and real data. Simulated datasets and case studies highlight that there are no one method performs the best in all the situations.
- Comparison of methods to detect differentially expressed genes between single-cell populations - comparison of five statistical methods to detect differentially expressed genes between two distinct single-cell populations.
- Evaluation of single-cell classifiers for single-cell RNA sequencing data sets - In this article, nine tools have been systematically compared. The article provides a guideline for researchers to select and apply suitable single cell and cluster classification tools in their analysis workflows and sheds some lights on potential direction of future improvement on classification tools.
- Comparative analysis of single-cell RNA sequencing methods - a comparison of wet lab protocols for scRNA sequencing.
- Single-Cell RNA-Sequencing: Assessment of Differential Expression Analysis Methods - an assessment of main bulk and single-cell differential analysis methods used to analyze scRNA-seq data.
-
Paper collections
- Mendeley Single Cell Sequencing Analysis
- Single-Cell Genomics in the Journal Science - Special issue on Single-Cell Genomics
- The emerging field of single-cell analysis - Special issue on single cell analysis
-
-
People
-
Female
- Rhonda Bacher (University of Wisconsin-Madison, USA)
- Barbara Di Camillo (Information Engineering Department, University of Padova, Italy
- Lana X. Garmire, (University of Hawaii Cancer Center, USA)
- Christina Kendziorski (University of Wisconsin–Madison, USA)
- Ning Leng (Morgridge Institute for Research, USA)
- Alicia Oshlack (Murdoch Children's Research Institute, Australia)
- Dana Pe'er (Memorial Sloan Kettering Cancer Center, USA)
- Emma Pierson (Stanford University, USA)
- Charlotte Soneson (Institute of Molecular Life Sciences, University of Zurich)
- Sarah Teichmann (Wellcome Trust Sanger Institute, UK)
- Barbara Treutlein (Max Planck Institute for Evolutionary Anthropology, Germany)
- Catalina Vallejos (The Alan Turing Institute & UCL, UK)
- Aviv Regev (Broad Institute, USA)
- Jinmiao Chen (Singapore Immunology Network, A\*STAR, Singapore)
- Samantha Morris (Depts of Dev. Bio. and Genetics, Washington University, St. Louis)
- Jean Fan (Johns Hopkins University, USA)
- Brooke Fridley (Children's Mercy Hospital, USA)
- Lana X. Garmire, (University of Hawaii Cancer Center, USA)
- Keegan Korthauer (Dana Farber Cancer Institute, USA)
- Sarah Snelling (University of Oxford, UK)
- Sarah Teichmann (Wellcome Trust Sanger Institute, UK)
- Barbara Treutlein (ETH Zurich, CH)
- Sanja Vickovic (New York Genome Center & Columbia University, USA)
- Elisabetta Mereu (Centre for Genomic Regulation, Barcelona)
- Sandrine Dudoit (UC Berkeley, USA)
- Rhonda Bacher (University of Wisconsin-Madison, USA)
- Brooke Fridley (Children's Mercy Hospital, USA)
- Laleh Haghverdi (EMBL, Germany)
- Stephanie Hicks (Johns Hopkins Bloomberg School of Public Health, USA)
- Smita Krishnaswamy (Yale University)
- Emma Pierson (Stanford University, USA)
-
Male
- Bart DePlancke (EPFL, School of Life sciences, Institute of Bioengineering, Switzerland)
- Raphael Gottardo (Fred Hutchinson Cancer Research Center, USA)
- Holger Heyn (Centre for Genomic Regulation, Barcelona)
- Peter Kharchenko (Department of Biomedical Informatics, Harvard Medical School, USA)
- Sten Linnarson (Karolinska Institutet, Sweden)
- Davis McCarthy (EBI, UK)
- John Reid (MRC Biostatistics Unit, Cambridge University, UK)
- Peter Sims (Columbia University, Department of Systems Biology)
- Fabian Theis (Institute of Computational Biology, Helmholtz Zentrum München)
- Cole Trapnell (University of Washington, Department of Genome Sciences)
- Itai Yanai (New York University, School of Medicine, Institute for Computational Medicine, USA)
- Ahmet Coskun (Georgia Tech, USA)
- Yanxiang Deng (University of Pennsylvania, USA)
- John Hickey (Duke University, USA)
- Peter Kharchenko (Department of Biomedical Informatics, Harvard Medical School, USA)
- Sten Linnarson (Karolinska Institutet, Sweden)
- Davis McCarthy (EBI, UK)
- John Reid (MRC Biostatistics Unit, Cambridge University, UK)
- Rickard Sandberg (Karolinska Institutet, SE)
- Neville Sanjana (New York Genome Center & NYU)
- Peter Sims (Columbia University, Department of Systems Biology)
- Cole Trapnell (University of Washington, Department of Genome Sciences)
- Davis McCarthy (EBI, UK)
- Neville Sanjana (New York Genome Center & NYU)
- Davis McCarthy (EBI, UK)
- Davis McCarthy (EBI, UK)
- Davis McCarthy (EBI, UK)
- Davis McCarthy (EBI, UK)
- Davis McCarthy (EBI, UK)
- Davis McCarthy (EBI, UK)
- John Marioni (EBI, UK)
- Davis McCarthy (EBI, UK)
- Davis McCarthy (EBI, UK)
- Davis McCarthy (EBI, UK)
- Davis McCarthy (EBI, UK)
- Davis McCarthy (EBI, UK)
- Stein Aerts (KU Leuven Center for Human Genetics, Belgium)
- Bart DePlancke (EPFL, School of Life sciences, Institute of Bioengineering, Switzerland)
- Raphael Gottardo (Fred Hutchinson Cancer Research Center, USA)
- Chung Chau Hon (RIKEN Centre for Integrative Medical Sciences, Yokohama)
- Martin Hemberg (Sanger Institute, UK)
- John Hickey (Duke University, USA)
- Aaron Lun (Cancer Research UK, UK)
- Davis McCarthy (EBI, UK)
- Mark Robinson (Institute of Molecular Life Sciences, University of Zurich)
- Yvan Saeys (Vlaams Instituut voor Biotechnologie, Ghent, Belgium)
- Rahul Satija (New York Genome Center)
- Oliver Stegle (EBI, UK)
- Fabian Theis (Institute of Computational Biology, Helmholtz Zentrum München)
-
Methods comparisons
-
-
Similar lists and collections
-
Methods comparisons
- CrazyHotTommy's RNA-seq analysis list - Very broad list that includes some single cell RNA-seq packages and papers.
- Museum of Spatial Transcriptomics - A comprehensive catalog of spatial transcriptomics data sets and methods.
- agitter's Pseudotime estimation list - An overview of algorithms for estimating pseudotime in single-cell RNA-seq data.
- scRNA-tools.org - Database of scRNA-seq analysis tools and their functions. Managed through this [Github repository](https://github.com/Oshlack/scRNA-tools).
-
-
Software packages
-
Archetypal analysis
- scAAnet - [Python] - scAAnet performs non-linear archetypal analysis through autoencoder networks to identify shared gene expression programs (GEPs) among heterogenous cell populations and infer relative activity of each GEP across cells.
- scAAnet - [Python] - scAAnet performs non-linear archetypal analysis through autoencoder networks to identify shared gene expression programs (GEPs) among heterogenous cell populations and infer relative activity of each GEP across cells.
-
Batch-effect removal
- BatchEffectRemoval - [Python] - [Removal of Batch Effects using Distribution-Matching Residual Networks](https://doi.org/10.1093/bioinformatics/btx196)
- ResPAN - [Python] - ResPAN is a light structured **Res**idual autoencoder and mutual nearest neighbor **P**aring guided **A**dversarial **N**etwork for scRNA-seq batch correction.
- scPLS - [C++, R] - A normalization method to remove unwanted variation using both control and target genes. It takes advantage of the fact that genes in a scRNAseq study often can be naturally classified into two sets: a control set of genes that are free of effects of the predictor variables and a target set of genes that are of primary interest. By modeling the two sets of genes jointly using the partial least squares regression, scPLS is capable of making full use of the data to improve the inference of confounding effects. https://www.nature.com/articles/s41598-017-13665-w
- TASC - [C++, python] - To account for cell-to-cell technical differences, we propose a statistical framework, TASC (Toolkit for Analysis of Single Cell RNA-seq), an empirical Bayes approach to reliably model the cell-specific dropout rates and amplification bias by use of external RNA spike-ins. TASC incorporates the technical parameters, which reflect cell-to-cell batch effects, into a hierarchical mixture model to estimate the biological variance of a gene and detect differentially expressed genes. More importantly, TASC is able to adjust for covariates to further eliminate confounding that may originate from cell size and cell cycle differences.
- UNCURL - [Python] - Unsupervised and semi-supervised sampling effect removal for single-cell RNA-seq data.
-
Cell clustering
- BackSPIN - [Python] - Biclustering algorithm developed taking into account intrinsic features of single-cell RNA-seq experiments.
- dropClust - [R/Python] - Efficient clustering of ultra-large scRNA-seq data.
- SC3 - [R] - SC3 is a tool for the unsupervised clustering of cells from single cell RNA-Seq experiments.
- TooManyCells - [Haskell, CLI program] - [Suite of graph-based tools for efficient, global, and unbiased identification and visualization of cell clades.](https://www.biorxiv.org/content/10.1101/519660v1).
-
Cell projection and unimodal integration
-
Cell subsampling
- geosketch - [Python] - Method to subsample massive scRNA-seq datasets while preserving rare cell states. Resulting “sketch” accelerates clustering, visualization, and integration analyses. [Paper](https://doi.org/10.1016/j.cels.2019.05.003)
-
Cell type identification and classification
- CIPR - [R] - (Cluster Identity PRedictor-pronounced cy-per). A Shiny web applet (and R-package) that helps annotating the cluster identities in single-cell RNA-sequencing (scRNA-seq) experiments. The algorithm compares gene expression signature of experimental clusters with known reference datasets. In addition to 7 reference datasets implemented in CIPR (2 from mouse and 5 from human), users can upload custom high-throughput reference data for specialized studies. The CIPR pipeline can be further tailored to different analytical contexts by excluding irrelevant reference subsets and low-variance reference genes from the analysis. The manuscript describing CIPR and comparing its performance against other similar software was published in [BMC Bioinformatics](https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-020-3538-2). CIPR's fast and computationally efficient calculations and graphical outputs will facilitate scRNA-seq analysis where the user wants to try different clustering parameters iteratively and examine the cluster identities. Source code for the [Shiny](https://github.com/atakanekiz/CIPR-Shiny) and [R-package](https://github.com/atakanekiz/CIPR-Package) implementations are available on GitHub.
- SingleR - [R] - SingleR leverages reference transcriptomic datasets of pure cell types to infer the cell of origin of each of the single cells independently. [Reference-based analysis of lung single-cell sequencing reveals a transitional profibrotic macrophage. Nature Immunology (2019)](https://www.nature.com/articles/s41590-018-0276-y)
- Celltypist - [Python] - Celltypist is an automated cell type annotation tool for scRNA-seq datasets on the basis of logistic regression classifiers optimized by the stochastic gradient descent algorithm. Celltypist provides several different models for predictions, with a current focus on immune sub-populations, in order to assist in the accurate classification of different cell types and subtypes.
- scExtract - [Python] - scExtract is a tool for automating annotation and integration of published single-cell RNA-seq datas. This tool uses LLMs agents to extract relevant information from scientific articles, process the data, and use annotations to guide multi-datasets integration.
- CyteType - [Python] - CyteType is a Python package for deep chracterization of cell clusters from single-cell RNA-seq data. This package interfaces with Anndata objects to call CyteType API.
- ceLLama - [R/Python] - ceLLama is a streamlined automation pipeline for cell type annotations using local large-language models (LLMs).
- CHETAH - [R] - CHETAH: a selective, hierarchical cell type identification method for single-cell RNA sequencing. CHETAH (CHaracterization of cEll Types Aided by Hierarchical clustering) is an accurate cell type identification algorithm that is rapid and selective, including the possibility of intermediate or unassigned categories. Evidence for assignment is based on a classification tree of previously available scRNA-seq reference data and includes a confidence score based on the variance in gene expression per cell type. For cell types represented in the reference data, CHETAH's accuracy is as good as existing methods. Its specificity is superior when cells of an unknown type are encountered, such as malignant cells in tumor samples which it pinpoints as intermediate or unassigned. [bioRxiv](https://doi.org/10.1101/558908)
- easybio - [R] - easybio is an R pacakge for cell type annotation using the CellMarker2.0 database. [bioRxiv](https://doi.org/10.1101/2024.09.14.609619)
- Garnett - [R] - Garnett is a software package that facilitates automated cell type classification from single-cell expression data. Garnett works by taking single-cell data, along with a cell type definition (marker) file, and training a regression-based classifier. Once a classifier is trained for a tissue/sample type, it can be applied to classify future datasets from similar tissues. In addition to describing training and classifying functions, this website aims to be a repository of previously trained classifiers. [Supervised Classification Enables Rapid Annotation of Cell Atlases](https://www.nature.com/articles/s41592-019-0535-3)
- SignacX - [R] - Signac classifies the cellular phenotype for each individual cell in scRNA-seq data using neural networks trained with sorted bulk gene expression data from the Human Primary Cell Atlas. Signac can: map cells from one data set to another, classify non-human single cell data, identify novel cell types, and classify single cell data across many tissues, diseases and technologies. [Cell type classification and discovery across diseases, technologies and tissues reveals conserved gene signatures and enables standardized single-cell readouts](https://www.biorxiv.org/content/10.1101/2021.02.01.429207v3.full)
- scCATCH - [R] - A single cell cluster-based annotation package from cluster marker genes identification to cluster annotation based on evidence-based score by matching the identified potential marker genes with known cell markers in tissue-specific cell taxonomy reference database (CellMatch) [Automatic Annotation on Cell Types of Clusters from Single-Cell RNA Sequencing Data. iScience (2020)](https://www.sciencedirect.com/science/article/pii/S2589004220300663)
- ImmClassifier - [R,python,Docker] - A cell type annotation algorithm that employs a knowledge-based approach to annotating cells based on their underlying ontology and multitudes of previously-published data. By encoding immune cell hierarchy in a neural network, ImmClassifier is able to identify fine-grained cell types with high accuracy. By running in Docker the tool is platform-agnostic. [bioRxiv](https://www.biorxiv.org/content/10.1101/2020.03.23.002758v1)
- mLLMCelltype - [R/Python] - A multi-model framework for single-cell RNA-seq cell type annotation using large language models (LLMs). It implements an interactive consensus mechanism where multiple LLMs collaborate to reach agreement on cell type annotations, with uncertainty quantification through consensus proportion and entropy metrics. Supports OpenAI, Anthropic, Google, and Alibaba models.
- CASSIA - [R/Python/Web] - CASSIA is a multi-agent large language model (LLM) framework for automated, reference-free, and interpretable cell type annotation of single-cell RNA-seq data. It includes dedicated agents for annotation, validation, formatting, quality scoring, and reporting, along with optional modules for subclustering, uncertainty quantification, retrieval-augmented generation (RAG), and annotation refinement via the Annotation Boost agent. CASSIA has been applied to correct errors in gold-standard annotations, detect mixed cell types, and accurately annotate rare cell populations across diverse species. [CASSIA: a multi-agent large language model for reference free, interpretable, and automated cell annotation of single-cell RNA-sequencing data](https://www.biorxiv.org/content/10.1101/2024.12.04.626476v2)
- ScType - [Web/R/Python] - ScType is an automated ultra-fast, marker-based cell type annotation tool for single-cell and spatial transcriptomics data. [Fully-automated and ultra-fast cell-type identification using specific marker combinations from single-cell transcriptomic data](https://www.nature.com/articles/s41467-022-28803-w)
- Compocyte - [python] - Compocyte is a composite classifier for modular hierarchical cell type annotation of single cell data. Using Compocyte you can build different hierarchical classifier architectures (local classifier per parent node, local classifer per node and local classifier per level) using all relevant models from pytorch, TensorFlow and keras. Local classifiers can be individually modified to account for alterations in classification taxonomies or selectively improve specific annotations in human-in-the-loop approaches.
- mtSC - [Python] - A multitask deep metric learning framework that integrates multiple references for single-cell assignment, including cross-species settings. [Integrating multiple references for single-cell assignment](https://doi.org/10.1093/nar/gkab380).
- clustifyr - [R] - Classifies cells and clusters in single-cell RNA-seq experiments using external reference data, gene signatures, or marker gene lists.
- cellassign - [R] - Automated, probabilistic assignment of scRNA-seq to known types. `cellassign` automatically assigns single-cell RNA-seq data to known cell types across thousands of cells accounting for patient and batch specific effects. Information about a priori known markers for cell types is provided as input to the model. cellassign then probabilistically assigns each cell to a cell type, removing subjective biases from typical unsupervised clustering workflows. [bioRxiv](https://www.biorxiv.org/content/early/2019/01/16/521914)
- singleCellNet - [R] - A near-universal step in the analysis of single cell RNA-Seq data is to hypothesize the identity of each cell. Often, this is achieved by finding cells that express combinations of marker genes that had previously been implicated as being cell-type specific, an approach that is not quantitative and does not explicitly take advantage of other single cell RNA-Seq studies. SingleCellNet, which addresses these issues and enables the classification of query single cell RNA-Seq data in comparison to reference single cell RNA-Seq data. [bioRxiv](https://www.biorxiv.org/content/early/2018/12/31/508085)
- DeepSort - [python] - A reference-free cell-type annotation tool for single-cell RNA-seq data using deep learning with a weighted graph neural network, which is learned based on the most comprehensive single-cell transcriptomics atlases involving 764,741 cells across 88 tissues of human and mouse. [bioRxiv](https://www.biorxiv.org/content/10.1101/2020.05.13.094953v1)
- Celltypist - [Python] - Celltypist is an automated cell type annotation tool for scRNA-seq datasets on the basis of logistic regression classifiers optimized by the stochastic gradient descent algorithm. Celltypist provides several different models for predictions, with a current focus on immune sub-populations, in order to assist in the accurate classification of different cell types and subtypes.
-
Cellular interactions/communication
- CellPhoneDB - [python] - Publicly available repository of curated receptors, ligands and their interactions in humam (subunit architecture is included for both ligands and receptors, representing heteromeric complexes accurately). [Paper](https://www.nature.com/articles/s41596-020-0292-x)
- Celcomen - [python] - Causal generative model that disentangles intra- and inter-cellular gene regulation programs in spatial transcriptomics and single-cell data through a generative graph neural network. Can generate post-perturbation counterfactual spatial transcriptomics. [Paper](https://arxiv.org/abs/2409.05804)
- NicheNet - [R] - To study intercellular communication from a computational perspective. It uses human or mouse gene expression data of interacting cells as input and combines this with a prior model that integrates existing knowledge on ligand-to-target signaling paths. This allows to predict ligand-receptor interactions that might drive gene expression changes in cells of interest. [Paper](https://www.nature.com/articles/s41592-019-0667-5)
- NICHES - [R] - Computational toolset that analyzes cell-cell signaling by creating unique one-to-one cell pairs rather than using prior knowledge networks. Unlike NicheNet, NICHES enables low-dimensional embedding of cellular interactions in signal-space and specifically supports spatial datasets through nearest-neighbor constraints. Integrates with standard single-cell analysis packages for visualization and trajectory analysis. [Paper](https://doi.org/10.1093/bioinformatics/btac775)
- COMUNET - [python] - It streamlines the interpretation of the results from cell-cell communication analyses by using multiplex networks to represent and cluster all potential communication pathways between cell types. [Paper](https://academic.oup.com/bioinformatics/article/36/15/4296/5836497)
- CellChat - [R] - It predicts major signaling inputs and outputs for cells and how those cells and signals coordinate for functions using network analysis and pattern recognition approaches. Through manifold learning and quantitative contrasts, CellChat classifies signaling pathways and delineates conserved and context-specific pathways across different datasets. [Paper](https://www.nature.com/articles/s41467-021-21246-9)
- CellNEST - [Python] - Cell Neural Networks on Spatial Transcriptomics (CellNEST) deciphers patterns of cell-cell communication by introducing relay-network detection that identifies ligand-receptor-ligand-receptor communication chains. Uses attention mechanisms to analyze spatial transcriptomics data, detect T cell homing signals, and predict new communication patterns in various cancer types. [Paper](https://www.nature.com/articles/s41592-025-02721-3)
- Connectome - [R] - Software package that facilitates calculation and visualization of cell-cell signaling network topologies in single-cell RNA-seq data. Supports analysis of ligand-receptor interactions, differential connectomics between tissue systems, and interactive exploration of cellular communication patterns. [Paper](https://www.nature.com/articles/s41598-022-07959-x)
- GEARS - [Python] - Graph-enhanced gene activation and repression simulator that predicts transcriptional responses to both single and multigene perturbations. Integrates deep learning with knowledge graphs of gene-gene relationships to predict outcomes of novel gene perturbations not seen experimentally. Shows high precision in predicting genetic interaction subtypes. [Paper](https://www.nature.com/articles/s41587-023-01905-6)
- LIANA - [R, python] - LIANA enables the use of any combination of ligand-receptor methods and resources, and their consensus. [Paper](https://www.nature.com/articles/s41467-022-30755-0)
-
Copy number analysis
- aneufinder - [R] - Bioconductor module for copy-number detection in single-cell whole genome sequencing (scWGS) and strand-seq data using a Hidden Markov Model or binary bisection method.
- CopyKAT - [R] - Inference of genomic copy number and subclonal structure from scRNA-seq data. Outperforms *inferCNV*. [Paper](https://doi.org/10.1038/s41587-020-00795-2)
- Ginkgo - [R, C] - Ginkgo is a web application for single-cell copy-number variation analysis.
- inferCNV - [R] - Part of the TrinityCTAT (Trinity Cancer Transcriptome Analysis Toolkit). Provides tools for copy-number inference from single-cell RNA-seq data.
- inferCNVpy - [Python] - A Python/Scanpy re-implementation of `inferCNV`. Significantly faster than the R version.
- MEDALT - [R, Python] - This package performs lineage tracing using copy number profile from single cell sequencing technology. It will infer: 1. An rooted directed minimal spanning tree (RDMST) to represent aneuploidy evolution of tumor cells. 2. The focal and broad copy number alterations associated with lineage expansion.
- Numbat - [R] - Numbat is a haplotype-aware CNV caller from single-cell and spatial transcriptomics data. It integrates signals from gene expression, allelic ratio, and population-derived haplotype information to accurately infer allele-specific CNVs in single cells and reconstruct their lineage relationship. [Paper](https://www.nature.com/articles/s41587-022-01468-y)
- SCEVAN - [R] - Easy-to-use package that starting from the raw count matrix of scRNA data automatically classifies the cells present in the biopsy by segregating non-malignant cells of tumor microenviroment from the malignant cells, outperforms *copyKAT*. It also infers the copy number profile of malignant cells, identifies subclonal structures and automatically analyses the specific and shared alterations of each subpopulation. [Preprint](https://www.biorxiv.org/content/10.1101/2021.11.20.469390v1)
- SCICoNE - [C++, Python] - Single-cell copy number calling and event history reconstruction. SCICoNE reconstructs the history of copy number events in the tumour and uses these evolutionary relationships to identify the copy number profiles of the individual cells.
- HoneyBADGER - [R] - HoneyBADGER identifies and infers the presence of CNV and LOH events in single cells and reconstructs subclonal architecture using allele and expression information from single-cell RNA-sequencing data.
-
Count modelling and normalization
- BEARscc - [R] - BEARscc makes use of ERCC spike-in measurements to model technical variance as a function of gene expression and technical dropout effects on lowly expressed genes.
- BASiCS - [R] - Bayesian Analysis of single-cell RNA-seq data. Estimates cell-specific normalization constants. Technical variability is quantified based on spike-in genes. The total variability of the expression counts is decomposed into technical and biological components. BASiCS can also identify genes with differential expression/over-dispersion between two or more groups of cells.
- BPSC - [R] - Beta-Poisson model for single-cell RNA-seq data analyses
- dsb - [R or Python] - a method for normalizing and denoising protein data from antibody derived tags (ADT). Compatible with CITE-seq, ASAP-seq, TEA-seq, ICICLE-seq, MissionBio etc. Removes ambient and cell to cell technical noise from ADTs see vignettes on [CRAN](https://CRAN.R-project.org/package=dsb). Manuscript open access: [Normalizing and denoising protein expression data from droplet-sed single cell profiling. *Nature Communications* (2022)](https://www.nature.com/articles/s41467-022-29356-8)
- MAST - [R] - Model-based Analysis of Single-cell Transcriptomics (MAST) fits a two-part, generalized linear models that are specially adapted for bimodal and/or zero-inflated single cell gene expression data
- SCnorm - [R] - A quantile regression based approach for robust normalization of single cell RNA-seq data.
- zinbwaveZinger - [R] - We introduce a weighting strategy, based on a zero-inflated negative binomial model, that identifies excess zero counts and generates gene- and cell-specific weights to unlock bulk RNA-seq DE pipelines for zero-inflated data, boosting performance for scRNA-seq. https://doi.org/10.1186/s13059-018-1406-4
- Dino - [R] - normalizes single-cell RNA-seq data by constructing a flexible negative-binomial mixture model of gene expression and sampling from the posterior distribution of expected expression conditional on observed sequencing depth. [Normalization by distributional resampling of high throughput single-cell RNA-sequencing data. *Bioinformatics* (2021)](https://doi.org/10.1093/bioinformatics/btab450)
- Sanity - [C] - (SAmpling-Noise-corrected Inference of Transcription ActivitY) is a Bayesian procedure that infers the log expression levels (log transcription quotients) of genes by filtering out Poisson noise from UMI count matrices. It estimates expression values and error bars directly without tunable parameters. [Bayesian inference of gene expression states from single-cell RNA-seq data. *Nature Biotechnology* (2021](https://doi.org/10.1038/s41587-021-00875-x)
-
Dimension reduction
- scvis - [python] - [Interpretable dimensionality reduction of single cell transcriptome data with deep generative models](https://doi.org/10.1101/178624)
- torchdr - [python] - Dimensionality reduction toolbox using PyTorch, featuring various algorithms such as TSNE, UMAP, and more. Supports GPU acceleration to maximize computational efficiency.
- PHATE - Potential of Heat-diffusion for Affinity-based Transition Embedding - [Python, R, matlab] - PHATE is a tool for visualizing high dimensional single-cell data with natural progressions or trajectories. PHATE uses a novel conceptual framework for learning and visualizing the manifold inherent to biological systems in which smooth transitions mark the progressions of cells from one state to another.
- SWNE - [R] - [Visualizing single-cell RNA-seq datasets with Similarity Weighted Nonnegative Embedding (SWNE)](https://www.biorxiv.org/content/early/2018/03/05/276261)
- ZIFA - [Python] - Zero-inflated dimensionality reduction algorithm for single-cell data.
- scDEED - [R] optimizing hyperparameters of UMAP/t-SNE, assigning each embedding a “reliability score” by permutation , manuscript open access: [Statistical method scDEED for detecting dubious 2D single-cell embeddings and optimizing t-SNE and UMAP hyperparameters](https://www.nature.com/articles/s41467-024-45891-y)
- p-SNE - [Python] - Poisson Stochastic Neighbor Embedding, a nonlinear dimensionality reduction method for sparse count data using Poisson KL divergence and Hellinger distance. [Paper](https://arxiv.org/abs/2604.16932).
- destiny - [R] - Diffusion maps are spectral method for non-linear dimension reduction introduced by Coifman et al.(2005). Diffusion maps are based on a distance metric (diffusion distance) which is conceptually relevant to how differentiating cells follow noisy diffusion-like dynamics, moving from a pluripotent state towards more differentiated states.
-
Doublet Identification
- AMULET - [shell, Python, R] - A count based method for detecting multiplets from single nucleus ATAC-seq (snATAC-seq) data. [Genome Biology](https://doi.org/10.1186/s13059-021-02469-x)
- demuxlet - [shell] - [Multiplexed droplet single-cell RNA-sequencing using natural genetic variation](https://www.nature.com/articles/nbt.4042)
- DoubletFinder - [R] - Doublet detection in single-cell RNA sequencing data using artificial nearest neighbors. [BioRxiv](https://www.biorxiv.org/content/early/2018/06/20/352484)
- DoubletDecon - [R] - Cell-State Aware Removal of Single-Cell RNA-Seq Doublets. [BioRxiv](DoubletDecon: Cell-State Aware Removal of Single-Cell RNA-Seq Doublets)
- DoubletDetection - [R, Python] - A Python3 package to detect doublets (technical errors) in single-cell RNA-seq count matrices. An [R implementation](https://github.com/TomKellyGenetics/DoubletDetection) is in development.
- Scrublet - [Python] - Computational identification of cell doublets in single-cell transcriptomic data. [BioRxiv](https://www.biorxiv.org/content/early/2018/07/09/357368)
- solo - [Python] - Doublet detection via semi-supervised deep learning.
-
Epigenomics
- ArchR - [R] - ArchR is a full-featured R package for processing and analyzing single-cell ATAC-seq data. ArchR provides the most extensive suite of scATAC-seq analysis tools of any software available. [ArchR: An integrative and scalable software package for single-cell chromatin accessibility analysis](https://www.biorxiv.org/content/10.1101/2020.04.28.066498v1).
- ChromVAR - [R] - Determine variations in chromatin accessibility across sets of annotations or peaks. Designed primarily for single-cell or sparse chromatin accessibility data, e.g. from scATAC-seq or sparse bulk ATAC or DNAse-seq experiments. [BioRxiv](https://www.biorxiv.org/content/early/2017/02/21/110346)
- EpiScanpy - [python] - EpiScanpy is the epigenomic extension of scRNA-seq analysis tool Scanpy. It analyses single-cell open chromatin (scATAC-seq) and single-cell DNA methylation (for example scBS-seq) data. [EpiScanpy: integrated single-cell epigenomic analysis](https://www.nature.com/articles/s41467-021-25131-3)
- Signac - [R] - Signac is an extension of Seurat for the analysis, interpretation, and exploration of single-cell chromatin datasets.
- ATACdemultiplex - [Go] - Suites of low-level multi-threaded utilities to efficiently manipulate large single-cell ATAC-Seq data (BAM, BED/fragments, Fastq files). Very efficient for creating sparse matrices, subset fragment/BED files, annotate peaks, estimate FDR corrected fisher features, create bigwig files and compute TSS enrichments (global and at the single-cell level).
- AtacWorks - [python] - AtacWorks is a deep learning tool to denoise and identify accessible chromatin regions from low-coverage, low cell count, or low-quality ATAC-seq data. AtacWorks can denoise signal and identify peaks from rare cellular subtypes in a mixed population. [Biorxiv](https://www.biorxiv.org/content/10.1101/829481v2)
-
Programming Languages
Categories
Sub Categories
RNA-seq
104
Spatial transcriptomics
90
Male
49
Rare cell detection
46
Pseudotime and trajectory inference
31
Female
31
Interactive visualization and analysis
29
Web portals and databases
29
Cell type identification and classification
22
Other applications
21
Epigenomics
21
Multi-assay data integration
17
Methods comparisons
16
Marker and differential gene expression identification
11
Copy number analysis
10
Cellular interactions/communication
10
Count modelling and normalization
9
Variant calling
8
Dimension reduction
8
Gene regulatory network identification
8
Immune receptor profiling
8
Quality control
7
Doublet Identification
7
Feature (Gene) imputation
6
Batch-effect removal
5
Experimental design
4
Cell clustering
4
Single cell large model
4
Simulation
4
Paper collections
3
Archetypal analysis
2
Cell projection and unimodal integration
2
Malignant cell identification
1
Cell subsampling
1
Big data approach overview
1
Keywords
single-cell
39
single-cell-rna-seq
22
rna-seq
15
bioinformatics
15
scrna-seq
13
r
10
single-cell-genomics
8
single-cell-analysis
8
python
7
gene-expression
6
clustering
6
visualization
6
spatial-transcriptomics
6
scrnaseq
5
transcriptomics
5
data-visualization
5
scrna-seq-analysis
4
human-cell-atlas
4
seurat
4
dimensionality-reduction
4
scanpy
4
computational-biology
4
genomics
4
deep-learning
4
machine-learning
3
gene-regulatory-network
3
10x
3
network-analysis
3
network-inference
3
kallisto
3
visium
3
cell-cell-communication
3
nonnegative-matrix-factorization
2
single-cell-multiomics
2
variational-autoencoder
2
differential-expression
2
bayesian-inference
2
statistical-methods
2
r-package
2
simulation
2
imputation
2
bioinformatics-pipeline
2
spatial-analysis
2
data-integration
2
bustools
2
scverse
2
anndata
2
cell-cell-interaction
2
factor-analysis
2
enrichment-analysis
2