{"id":13490563,"url":"https://github.com/Bin-Chen-Lab/Awesome_BigData_AI_DrugDiscovery","last_synced_at":"2025-03-28T06:31:36.632Z","repository":{"id":50545711,"uuid":"91730303","full_name":"Bin-Chen-Lab/Awesome_BigData_AI_DrugDiscovery","owner":"Bin-Chen-Lab","description":"A collection of resources useful for leveraging big data and AI for drug discovery. It mainly serves as an orientation for new lab folks. It may be biased towards my lab interest.","archived":false,"fork":false,"pushed_at":"2023-09-07T20:22:59.000Z","size":1153,"stargazers_count":163,"open_issues_count":0,"forks_count":59,"subscribers_count":12,"default_branch":"master","last_synced_at":"2024-05-23T03:01:38.967Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Bin-Chen-Lab.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null}},"created_at":"2017-05-18T19:28:45.000Z","updated_at":"2024-05-09T06:50:42.000Z","dependencies_parsed_at":"2024-01-14T18:03:42.367Z","dependency_job_id":"e16f1a8f-5f19-44f8-9766-52d850e6378d","html_url":"https://github.com/Bin-Chen-Lab/Awesome_BigData_AI_DrugDiscovery","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Bin-Chen-Lab%2FAwesome_BigData_AI_DrugDiscovery","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Bin-Chen-Lab%2FAwesome_BigData_AI_DrugDiscovery/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Bin-Chen-Lab%2FAwesome_BigData_AI_DrugDiscovery/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Bin-Chen-Lab%2FAwesome_BigData_AI_DrugDiscovery/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Bin-Chen-Lab","download_url":"https://codeload.github.com/Bin-Chen-Lab/Awesome_BigData_AI_DrugDiscovery/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":245984569,"owners_count":20704792,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-07-31T19:00:48.643Z","updated_at":"2025-03-28T06:31:35.463Z","avatar_url":"https://github.com/Bin-Chen-Lab.png","language":null,"funding_links":[],"categories":["Sources:","Medicine","Other Lists","General Directories"],"sub_categories":["TeX Lists"],"readme":"# Introduction to Bioinformatics/Cheminformatics\n(***) [An Introduction to Statistical Learning](http://www-bcf.usc.edu/~gareth/ISL/) by Robert Tibshirani.  \nIf you have a little statistical background, read this book first.\n\n(***) [Machine Learning course at Coursera](https://www.coursera.org/learn/machine-learning) by Andrew Ng.  \nA required course for any people who are interested in machine learning.\n\n(***) [R \u0026 Bioconductor Manual](http://manuals.bioinformatics.ucr.edu/home/R_BioCondManual/).  \nEnsure you have run the code before taking any bioinformatics project.\n\n(***) [STATQUEST] (https://statquest.org/video-index/).  \nVideo collection for statistics, machine learning and bioinformatics.  \n\n(***) [HT Sequence Analysis with R and Bioconductor](http://manuals.bioinformatics.ucr.edu/home/ht-seq).  \nEnsure you have run the code before taking any NGS project.\n\n(***) [ChemmineR: Cheminformatics Toolkit for R](http://www.bioconductor.org/packages/devel/bioc/vignettes/ChemmineR/inst/doc/ChemmineR.html).  \nSuggest to run the code before taking any cheminformatics project.\n\n(**) [Step by Step to practice deep learning](http://pytorch.org/tutorials/beginner/deep_learning_60min_blitz.html).   \nPyTorch tutorial for deep learning.\n\n(***) [Introduction to Bioinformatics and Computational Biology](https://liulab-dfci.github.io/bioinfo-combio/).   \nVery comprenhensive video tutorials in computational biology by Shirley Liu.\n\n(***) [HarvardX Biomedical Data Science Open Online Training](http://rafalab.github.io/pages/harvardx.html).   \ncomprehensive tutorial on data science with code from rafalab.\n\n(***) [Statistics for biologists from StatQuest](https://statquest.org/video-index/).   \noutstanding video tutorials by Josh Starmer.\n\n(**) [Data Science Cheat Sheet](https://github.com/Bin-Chen-Lab/BigData_AI_DrugDiscovery/blob/master/data_science_cheatsheet.pdf).\nA quick check list of basics in data science (credit to Maverick Lin)\n\n(***) Single cell RNA-Seq analysis [osca](https://osca.bioconductor.org/introduction.html), [seurat](https://satijalab.org/seurat/vignettes.html). \noutstanding framework for scRNA-Seq analysis.\n\n# Fundamental papers\n**These papers I read at least 10 times, including supplementary materials! All of them are three stars!**\n\n## Field review\n(***) [Hallmarks of Cancer: The Next Generation](http://www.cell.com/abstract/S0092-8674%2811%2900127-9), by Robert A. Weinberg.  \nFundamental to understand cancer. \n\n(***) [Tumor Metastasis: Molecular Insights and Evolving Paradigms](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3261217/), by Robert A. Weinberg.  \nFundamental to understand cancer metastasis.\n\n(***) [Cancer genome landscapes](http://science.sciencemag.org/content/339/6127/1546.long), by Bert Vogelstein.   \nFundamental to understand cancer genomics.\n\n(***) [Cancer transcriptome profiling at the juncture of clinical translation](https://www.nature.com/articles/nrg.2017.96), by Arul M. Chinnaiyan.   \nreview on cancer transcriptomics.\n\n(***) [Ewing sarcoma: historical perspectives, current state-of-the-art, and opportunities for targeted therapy in the future.](https://www.ncbi.nlm.nih.gov/pubmed/18525337).  \nA typical review on the therapeutic discovery of one cancer.\n\n(***) [Opportunities and challenges in phenotypic drug discovery: an industry perspective](https://www.nature.com/nrd/journal/v16/n8/abs/nrd.2017.111.html).  \nOur drug discovery approach is one type of phenotypic screening.\n\n(***) [Ten Years of Pathway Analysis: Current Approaches and Outstanding Challenges](http://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1002375), by Purvesh Khatri.  \nA very nice summary of method development in pathway analysis.\n\n(***) [Deep learning](https://www.nature.com/nature/journal/v521/n7553/full/nature14539.html) by Yann LeCun,\tYoshua Bengio \u0026 Geoffrey Hinton.  \nDeep learning review.\n\n(***) [High-performance medicine: the convergence of human and artificial intelligence](https://www.nature.com/articles/s41591-018-0300-7) by Eric Topol.\nCurrent progress and challenges in applying DL into biomedical research.\n\n## Statistical method development\n(***) [Significance analysis of microarrays applied to the ionizing radiation response](http://www.pnas.org/content/98/9/5116.full), by Robert Tibshirani.  \nDevelopment of SAM, a popular method to perform differential expression analysis using microarray data.  \n\n(***) [limma: Linear Models for Microarray Data](https://link.springer.com/chapter/10.1007/0-387-29362-0_23).  \nDevelopment of LIMMA, another popular method to perform differential expression analysis using microarray data.\n\n(***) [Differential expression analysis for sequence count data](https://genomebiology.biomedcentral.com/articles/10.1186/gb-2010-11-10-r106), by Simon Anders.  \nDevelopment of DEseq, a popular method to perform differential expression analysis using RNA-SEQ data.  \n\n(***) [Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles](http://www.pnas.org/content/102/43/15545.long), by Jill P. Mesirov.  \nDevelopment of GSEA, the most popular gene set enrichment analysis method and the fundamental to understand our drug discovery method.\n\n(***) [Adjusting batch effects in microarray expression data using Empirical Bayes methods](https://academic.oup.com/biostatistics/article/8/1/118/252073/Adjusting-batch-effects-in-microarray-expression).  \nDevelopment of Combat, a method to correct batch effects.\n\n(***) [Emergence of Scaling in Random Networks](http://science.sciencemag.org/content/286/5439/509.full) by Albert-László Barabási.  \nDiscovery of scale-free networks.\n\n(***) [Pathsim: Meta path-based top-k similarity search in heterogeneous information networks](http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.220.2455) by Jiawei Han.  \nA typical machine learning approach to mining heterogeneous networks.\n\n(***) [MuSiC: Identifying mutational significance in cancer genomes](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3409272/), by Li Ding.  \nDevelopment of MuSic, a popular method to identify mutations.\n\n## Informatics method development and application\n(***) [The Connectivity Map: using gene-expression signatures to connect small molecules, genes, and disease.](http://science.sciencemag.org/content/313/5795/1929.long) by Justin Lamb.  \nThe first paper to describe our drug discovery approach.\n\n(***) [Discovery and Preclinical Validation of Drug Indications Using Compendia of Public Gene Expression Data](http://stm.sciencemag.org/content/3/96/96ra77) from Atul's lab.  \nThe basic of our drug discovery work, and a great demonstration of writing a computational paper (from method development to experimental validation).\n\n(***) [Relating protein pharmacology by ligand chemistry](http://www.nature.com/nbt/journal/v25/n2/full/nbt1284.html) by Michael J Keiser and Brian K Shoichet.  \nThe development of SEA, a method to predict drug-target interactions, and another great demonstration of writing a computational paper.\n\n(***) [Characterization of drug-induced transcriptional modules: towards drug repositioning and functional understanding](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3658274/) by Peer Bork.  \nstart with data analysis and end with a few biological experiments.\n\n(***) [Cross-Species Regulatory Network Analysis Identifies a Synergistic Interaction between FOXM1 and CENPF that Drives Prostate Cancer Malignancy](http://www.cell.com/cancer-cell/fulltext/S1535-6108(14)00125-1) by Andrea Califano.  \nstart with data analysis and end with a few biological experiments.\n\n(***) [Elucidating compound mechanism of action by network perturbation analysis](http://www.sciencedirect.com/science/article/pii/S0092867415006996) by Andrea Califano.  \nstart with data analysis and end with a few biological experiments.\n\n(***) [Discovery of drug mode of action and drug repositioning from transcriptional responses](http://www.pnas.org/content/107/33/14621.long).  \nstart with data analysis and end with a few biological experiments.\n\n(***) [Imagenet classification with deep convolutional neural networks](http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf).  \nDevelopment of convolutional neural networks (CNN), the popular deep learning method.\n\n## Computational analysis \n(***) [Drug-target network](https://www.nature.com/nbt/journal/v25/n10/full/nbt1338.html) by Barabási.  \nNetwork analysis of drug-target interactions.\n\n(***) [Comprehensive molecular portraits of human breast tumours](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3465532/) from TCGA.  \nA typical genomic analysis paper from TCGA.\n\n(***) [Mutational landscape and significance across 12 major cancer types.](http://dx.doi.org/10.1038/nature12634), by Li Ding.  \nA phenomenal paper on pan-cancer genomic analysis.\n\n(***) [Comprehensive Characterization of Molecular Differences in Cancer between Male and Female Patients](http://www.cell.com/cancer-cell/fulltext/S1535-6108(16)30111-8), by Han Liang.  \nA phenomenal paper on pan-cancer genomic analysis.\n\n(***) [Genetics of rheumatoid arthritis contributes to biology and drug discovery.](https://www.nature.com/nature/journal/v506/n7488/full/nature12873.html) by Robert M. Plenge.  \ngreat work using genetics for drug discovery.\n\n(***) [The Cancer Cell Line Encyclopedia enables predictive modelling of anticancer drug sensitivity](https://www.nature.com/nature/journal/v483/n7391/full/nature11003.html).  \nphenomenal work using cell line data to discover biomarkers.\n\n(***) [A comprehensive time-course–based multicohort analysis of sepsis and sterile inflammation reveals a robust diagnostic gene set](http://stm.sciencemag.org/content/7/287/287ra71.short) by Purvesh Khatri.  \nPhenomenal work using public microarray data to discover biomarkers.\n\n(***) [Prediction of biological targets for compounds using multiple-category Bayesian models trained on chemogenomics databases](http://pubs.acs.org/doi/10.1021/ci060003g) by Jeremy Jenkins.  \nA typical machine learning paper in cheminformatics.\n\n(***) [Do structurally similar molecules have similar biological activity](https://dx.doi.org/10.1021/jm020155c).  \nA typical data analysis paper in cheminformatics.\n\n## Deep-learning based drug discovery\n(***) [Predicting Drug Response and Synergy Using a Deep Learning Model of Human Cancer Cells](https://pubmed.ncbi.nlm.nih.gov/33096023/) by Ideker.  \nDevelop a model to predict drug activity based on a huge pharmacogenomics dataset, propose novel ways to model cells based on Gene Ontology,  and experimentally validate some hits.\n\n(***) [A Deep Learning Approach to Antibiotic Discovery](https://pubmed.ncbi.nlm.nih.gov/32084340/) by Barzilay and Collins.  \nDevelop a model to predict antibiotic activity based on chemical structure, screen millions of compounds and extensively validate one drug candidate.\n\n(***) [Deep reinforcement learning for de novo drug design](http://advances.sciencemag.org/content/4/7/eaap7885) by Tropsha.  \nDevelop a model to generate targeted chemical libraries of novel compounds optimized for either a single desired property or multiple properties.\n\n(***) [Convolutional Networks on Graphs for Learning Molecular Fingerprints](https://arxiv.org/abs/1509.09292) by Ryan Adams.  \nDesigned DL based fingerprints inspired by ECFP\n\n(***) [Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules](https://pubs.acs.org/doi/full/10.1021/acscentsci.7b00572). \nUsed Variational autoencoder to encode SMILES and optimize compounds from the latent space.\n\n## Shape our future\n(***) [Single-cell RNA-seq highlights intratumoral heterogeneity in primary glioblastoma](http://science.sciencemag.org/content/344/6190/1396).  \nApplication of single cell in a cancer study.\n\n(***) [Single-cell transcriptomics uncovers distinct molecular signatures of stem cells in chronic myeloid leukemia](https://www.nature.com/nm/journal/v23/n6/full/nm.4336.html).  \nApplication of single cell analysis toward personalized cancer therapy.\n\n(***) [Brown Adipogenic Reprogramming Induced by a Small Molecule](http://www.sciencedirect.com/science/article/pii/S2211124716317697) by Sheng Ding.  \nUsing small molecules to control cell development.\n\n(***) [Correlating chemical sensitivity and basal gene expression reveals mechanism of action.](https://www.nature.com/nchembio/journal/v12/n2/full/nchembio.1986.html) from Stuart Schreiber.  \nUsage of pharmacogenomics data to understand drug mechanisms.\n\n(***) [A Next Generation Connectivity Map: L1000 Platform And The First 1,000,000 Profiles](https://www.biorxiv.org/content/early/2017/05/10/136168) from Todd R. Golub.  \nLINCS, the dataset we primarily used for drug discovery.\n\n(***) [Integrative clinical genomics of metastatic cancer](http://www.nature.com/nature/journal/v548/n7667/full/nature23306.html).  \nWe have lots of experience working on primary cancer, now it's time to place our interest to metastatic cancer, which the majority of patients die from.\n\n(***) [Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks](https://arxiv.org/pdf/1703.10593.pdf).  \nUsing deep learning GAN to realize domain knowledge translation.\n\n# Outstanding tools and datasets for translational drug discovery\n\n**use liver cancer as an example, can be applied to other cancers, only list outstanding tools/datasets for liver cancer drug discovery**\ntwo review articles from the lab:  \n[Harnessing big ‘omics’ data and AI for drug discovery in hepatocellular carcinoma](https://www.nature.com/articles/s41575-019-0240-9).    \n[Leveraging big data to transform target selection and drug discovery](http://www.ncbi.nlm.nih.gov/pubmed/26659699).    \n\n## Diseases/Patients\n(***) [ClinicalTrials.gov](https://clinicaltrials.gov).  \nSearch drugs used in liver cancer clinical trials.\n\n(**) [Cancer Today (Globocan): Data visualization tools that present current national estimates of cancer incidence, mortality, and prevalence](http://gco.iarc.fr/today/home).  \nSearch liver cancer incidence.\n\n(*) [UK Biobank](http://www.ukbiobank.ac.uk/). [UK Biobank Engine](https://biobankengine.stanford.edu/).  \nSearch public clinical/molecular data for liver cancer and search genetic variants for liver cancer via Stanford Biobank Engine.\n\n(**) [COSMIC](http://cancer.sanger.ac.uk/cosmic).  \nSearch somatic mutation for liver cancer.\n\n## Target Discovery\n(***) [cBioPortal](http://www.cbioportal.org/).  \nSearch molecular alterations for liver cancer from public datasets including TCGA.\n\n(***) [GTEx](http://www.gtexportal.org).  \nSearch gene expression in normal liver tissues.\n\n(***) [The Human Protein Atlas](https://www.proteinatlas.org/).  \nSearch protein expression and pathology for liver cancer.\n\n(***) [Cancer Cell Line Encyclopedia](https://portals.broadinstitute.org/ccle).  \nSearch gene expression in liver cancer cell lines.\n\n(*) [Project Achilles](https://portals.broadinstitute.org/achilles).  \nSearch essential genes in liver cancer cells.\n\n(***) [DepMap](https://depmap.org/portal/).\ncreate a comprehensive preclinical reference map connecting tumor features with tumor dependencies to accelerate the development of precision treatments.\n\n(***) [GEO](https://www.ncbi.nlm.nih.gov/geo/).  \nSearch functional genomics data for liver cancer, requiring additional computational analysis to create a liver cancer signature.\n\n(***) [Enrichr](http://amp.pharm.mssm.edu/Enrichr/).  \nSearch enriched TS/pathways/biological processes/cell types given a list of genes.\n\n(***) [STRING DB](https://string-db.org/).  \nVisualize protein-protein interactions.\n\n## Drug Discovery\n(***) [PubChem](https://pubchem.ncbi.nlm.nih.gov/).  \nEverything needed to know about a compound/drug.\n\n(**) [DrugBank](https://www.drugbank.ca/).  \nSearch drug-target-indication.\n\n(**) [SEA](http://sea.bkslab.org/).  \nPredict targets of a given compound.\n\n(***) [LINCS](https://clue.io/).  \nPredict drugs given a liver cancer signature.\n\n(**) [ChemMine](http://chemmine.ucr.edu/).  \nvery useful for chemical structure enrichment analysis.\n\n# NGS analysis\n(***) [RNASEQ blog](http://www.rna-seqblog.com/).  \nA great collection of RNA-SEQ analysis methods/applications.\n\n(*) [RPKM, FPKM and TPM, clearly explained](http://www.rna-seqblog.com/rpkm-fpkm-and-tpm-clearly-explained/)\n\n(*) [RNA-seq workflow: gene-level exploratory analysis and differential expression](http://www.bioconductor.org/help/workflows/rnaseqGene/)\n\n# Python packages\n(***) [anaconda](https://anaconda.org/)\nSuggest using anaconda to manage python packages.\n\n(***) [scikit: a popular python machine learning packages](http://scikit-learn.org/stable/).  \n\n(**) [rdkit](http://www.rdkit.org/docs/index.html).  \nfree python library to process chemical structures.\n\n(***) [PyTorch](http://pytorch.org/).  \nDeep learning framework.\n\n# R/Bioconductor packages\n(**) [ggplot cheatsheet](http://zevross.com/blog/2014/08/04/beautiful-plotting-in-r-a-ggplot2-cheatsheet-3/). A must read to visualize data using R.  \n\n(**) [ChemmineR: Cheminformatics Toolkit for R](http://www.bioconductor.org/packages/devel/bioc/vignettes/ChemmineR/inst/doc/ChemmineR.html)\n\n(**) [biomaRt](http://bioconductor.org/packages/release/bioc/html/biomaRt.html).  \nA great package to map IDs.\n\n(***) [GEOquery](http://bioconductor.org/packages/release/bioc/html/GEOquery.html).  \nSearch and download data from GEO.\n\n(**) [cgdsr](https://cran.r-project.org/web/packages/cgdsr/index.html). \nAPI to access cBioportal data.\n\n(**) [pheatmap](https://cran.r-project.org/web/packages/pheatmap/index.html).  \nVisualize heatmap.\n\n(*) [Easy Way to Mix Multiple Graphs on The Same Page](http://www.sthda.com/english/articles/24-ggpubr-publication-ready-plots/81-ggplot2-easy-way-to-mix-multiple-graphs-on-the-same-page/).  \n\n\n\n\n\n\n\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FBin-Chen-Lab%2FAwesome_BigData_AI_DrugDiscovery","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FBin-Chen-Lab%2FAwesome_BigData_AI_DrugDiscovery","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FBin-Chen-Lab%2FAwesome_BigData_AI_DrugDiscovery/lists"}