{"id":13703973,"url":"https://github.com/deweylab/RSEM","last_synced_at":"2025-05-05T09:32:37.585Z","repository":{"id":1324923,"uuid":"1345163","full_name":"deweylab/RSEM","owner":"deweylab","description":"RSEM: accurate quantification of gene and isoform expression from RNA-Seq data","archived":false,"fork":false,"pushed_at":"2024-03-13T22:24:21.000Z","size":53120,"stargazers_count":405,"open_issues_count":134,"forks_count":117,"subscribers_count":22,"default_branch":"master","last_synced_at":"2024-08-03T21:04:27.157Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"http://deweylab.biostat.wisc.edu/rsem/","language":"C++","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"gpl-3.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/deweylab.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"COPYING","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2011-02-09T05:30:43.000Z","updated_at":"2024-07-27T07:52:18.000Z","dependencies_parsed_at":"2023-01-11T16:03:31.863Z","dependency_job_id":"914d4dc8-439e-4639-843d-bc4c181158de","html_url":"https://github.com/deweylab/RSEM","commit_stats":{"total_commits":469,"total_committers":12,"mean_commits":"39.083333333333336","dds":0.6247334754797441,"last_synced_commit":"8bc1e2115493c0cdf3c6bee80ef7a21a91b2acce"},"previous_names":[],"tags_count":59,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/deweylab%2FRSEM","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/deweylab%2FRSEM/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/deweylab%2FRSEM/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/deweylab%2FRSEM/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/deweylab","download_url":"https://codeload.github.com/deweylab/RSEM/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":224439640,"owners_count":17311491,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-02T21:01:02.431Z","updated_at":"2024-11-13T11:30:34.430Z","avatar_url":"https://github.com/deweylab.png","language":"C++","funding_links":[],"categories":["Next Generation Sequencing","Counting tools"],"sub_categories":["Quantification","Correction tools"],"readme":"README for RSEM\n===============\n\n[Bo Li](https://lilab-bcb.github.io/) \\(bli28 at mgh dot harvard dot edu\\)\n\n* * *\n\nTable of Contents\n-----------------\n\n* [Introduction](#introduction)\n* [Compilation \u0026 Installation](#compilation)\n* [Usage](#usage)\n    * [Build RSEM references using RefSeq, Ensembl, or GENCODE annotations](#built)\n    * [Build RSEM references for untypical organisms](#untypical)\n* [Example](#example-main)\n* [Simulation](#simulation)\n* [Generate Transcript-to-Gene-Map from Trinity Output](#gen_trinity)\n* [Differential Expression Analysis](#de)\n* [Prior-Enhanced RSEM (pRSEM)](#pRSEM)\n* [Authors](#authors)\n* [Acknowledgements](#acknowledgements)\n* [License](#license)\n\n* * *\n\n## \u003ca name=\"introduction\"\u003e\u003c/a\u003e Introduction\n\nRSEM is a software package for estimating gene and isoform expression\nlevels from RNA-Seq data. The RSEM package provides an user-friendly\ninterface, supports threads for parallel computation of the EM\nalgorithm, single-end and paired-end read data, quality scores,\nvariable-length reads and RSPD estimation. In addition, it provides\nposterior mean and 95% credibility interval estimates for expression\nlevels. For visualization, It can generate BAM and Wiggle files in\nboth transcript-coordinate and genomic-coordinate. Genomic-coordinate\nfiles can be visualized by both UCSC Genome browser and Broad\nInstitute's Integrative Genomics Viewer (IGV). Transcript-coordinate\nfiles can be visualized by IGV. RSEM also has its own scripts to\ngenerate transcript read depth plots in pdf format. The unique feature\nof RSEM is, the read depth plots can be stacked, with read depth\ncontributed to unique reads shown in black and contributed to\nmulti-reads shown in red. In addition, models learned from data can\nalso be visualized. Last but not least, RSEM contains a simulator.\n\n## \u003ca name=\"compilation\"\u003e\u003c/a\u003e Compilation \u0026 Installation\n\nTo compile RSEM, simply run\n   \n    make\n\nFor Cygwin users, run\n\n    make cygwin=true\n\nTo compile EBSeq, which is included in the RSEM package, run\n\n    make ebseq\n\nTo install RSEM, simply put the RSEM directory in your environment's PATH\nvariable. Alternatively, run\n\n    make install\n\nBy default, RSEM executables are installed to `/usr/local/bin`. You\ncan change the installation location by setting `DESTDIR` and/or\n`prefix` variables. The RSEM executables will be installed to\n`${DESTDIR}${prefix}/bin`. The default values of `DESTDIR` and\n`prefix` are `DESTDIR=` and `prefix=/usr/local`. For example,\n\n    make install DESTDIR=/home/my_name prefix=/software\n\nwill install RSEM executables to `/home/my_name/software/bin`.\n\n**Note** that `make install` does not install `EBSeq` related scripts,\nsuch as `rsem-generate-ngvector`, `rsem-run-ebseq`, and\n`rsem-control-fdr`. But `rsem-generate-data-matrix`, which generates\ncount matrix for differential expression analysis, is installed.\n\n### Prerequisites\n\nC++, Perl and R are required to be installed. \n\nTo use the `--gff3` option of `rsem-prepare-reference`, Python is also\nrequired to be installed.\n\nTo take advantage of RSEM's built-in support for the Bowtie/Bowtie\n2/STAR/HISAT2 alignment program, you must have\n[Bowtie](http://bowtie-bio.sourceforge.net)/[Bowtie2](http://bowtie-bio.sourceforge.net/bowtie2)/[STAR](https://github.com/alexdobin/STAR)/[HISAT2](http://daehwankimlab.github.io/hisat2/)\ninstalled.\n\n## \u003ca name=\"usage\"\u003e\u003c/a\u003e Usage\n\n### I. Preparing Reference Sequences\n\nRSEM can extract reference transcripts from a genome if you provide it\nwith gene annotations in a GTF/GFF3 file.  Alternatively, you can provide\nRSEM with transcript sequences directly.\n\nPlease note that GTF files generated from the UCSC Table Browser do not\ncontain isoform-gene relationship information.  However, if you use the\nUCSC Genes annotation track, this information can be recovered by\ndownloading the knownIsoforms.txt file for the appropriate genome.\n \nTo prepare the reference sequences, you should run the\n`rsem-prepare-reference` program.  Run \n\n    rsem-prepare-reference --help\n\nto get usage information or visit the [rsem-prepare-reference\ndocumentation page](http://deweylab.github.io/RSEM/rsem-prepare-reference.html).\n\n#### \u003ca name=\"built\"\u003e\u003c/a\u003e Build RSEM references using RefSeq, Ensembl, or GENCODE annotations\n\nRefSeq and Ensembl are two frequently used annotations. For human and\nmouse, GENCODE annotaions are also available. In this section, we show\nhow to build RSEM references using these annotations. Note that it is\nimportant to pair the genome with the annotation file for each\nannotation source. In addition, we recommend users to use the primary\nassemblies of genomes. Without loss of generality, we use human genome as\nan example and in addition build Bowtie indices. \n\nFor **RefSeq**, the genome and annotation file in GFF3 format can be found\nat RefSeq genomes FTP:\n\n```\nftp://ftp.ncbi.nlm.nih.gov/genomes/refseq/\n```\n\nFor example, the human genome and GFF3 file locate at the subdirectory\n`vertebrate_mammalian/Homo_sapiens/all_assembly_versions/GCF_000001405.31_GRCh38.p5`. `GCF_000001405.31_GRCh38.p5`\nis the latest annotation version when this section was written.\n\nDownload and decompress the genome and annotation files to your working directory:\n\n```\nftp://ftp.ncbi.nlm.nih.gov/genomes/refseq/vertebrate_mammalian/Homo_sapiens/all_assembly_versions/GCF_000001405.31_GRCh38.p5/GCF_000001405.31_GRCh38.p5_genomic.fna.gz\nftp://ftp.ncbi.nlm.nih.gov/genomes/refseq/vertebrate_mammalian/Homo_sapiens/all_assembly_versions/GCF_000001405.31_GRCh38.p5/GCF_000001405.31_GRCh38.p5_genomic.gff.gz\n```\n\n`GCF_000001405.31_GRCh38.p5_genomic.fna` contains all top level\nsequences, including patches and haplotypes. To obtain the primary\nassembly, run the following RSEM python script:\n\n```\nrsem-refseq-extract-primary-assembly GCF_000001405.31_GRCh38.p5_genomic.fna GCF_000001405.31_GRCh38.p5_genomic.primary_assembly.fna\n```\n\nThen type the following command to build RSEM references:\n\n```\nrsem-prepare-reference --gff3 GCF_000001405.31_GRCh38.p5_genomic.gff \\\n\t\t       --trusted-sources BestRefSeq,Curated\\ Genomic \\\n\t\t       --bowtie \\\n\t\t       GCF_000001405.31_GRCh38.p5_genomic.primary_assembly.fna \\\n\t\t       ref/human_refseq\n```\n\nIn the above command, `--trusted-sources` tells RSEM to only extract\ntranscripts from RefSeq sources like `BestRefSeq` or `Curated Genomic`. By\ndefault, RSEM trust all sources. There is also an\n`--gff3-RNA-patterns` option and its default is `mRNA`. Setting\n`--gff3-RNA-patterns mRNA,rRNA` will allow RSEM to extract all mRNAs\nand rRNAs from the genome. Visit [here](rsem-prepare-reference.html)\nfor more details.\n\nBecause the gene and transcript IDs (e.g. gene1000, rna28655)\nextracted from RefSeq GFF3 files are hard to understand, it is\nrecommended to turn on the `--append-names` option in\n`rsem-calculate-expression` for better interpretation of\nquantification results.\n\nFor **Ensembl**, the genome and annotation files can be found at\n[Ensembl FTP](http://uswest.ensembl.org/info/data/ftp/index.html).\n\nDownload and decompress the human genome and GTF files:\n\n```\nftp://ftp.ensembl.org/pub/release-83/fasta/homo_sapiens/dna/Homo_sapiens.GRCh38.dna.primary_assembly.fa.gz\nftp://ftp.ensembl.org/pub/release-83/gtf/homo_sapiens/Homo_sapiens.GRCh38.83.gtf.gz\n```\n\nThen use the following command to build RSEM references:\n\n```\nrsem-prepare-reference --gtf Homo_sapiens.GRCh38.83.gtf \\\n\t\t       --bowtie \\\n\t\t       Homo_sapiens.GRCh38.dna.primary_assembly.fa \\\n\t\t       ref/human_ensembl\n```\n\nIf you want to use GFF3 file instead, which is unnecessary and not\nrecommended, you should add option `--gff3-RNA-patterns transcript`\nbecause `mRNA` is replaced by `transcript` in Ensembl GFF3 files.\n\n**GENCODE** only provides human and mouse annotations. The genome and\n  annotation files can be found from [GENCODE\n  website](http://www.gencodegenes.org/).\n\nDownload and decompress the human genome and GTF files:\n\n```\nftp://ftp.sanger.ac.uk/pub/gencode/Gencode_human/release_24/GRCh38.primary_assembly.genome.fa.gz\nftp://ftp.sanger.ac.uk/pub/gencode/Gencode_human/release_24/gencode.v24.annotation.gtf.gz\n```\n\nThen type the following command:\n\n```\nrsem-prepare-reference --gtf gencode.v24.annotation.gtf \\\n\t\t       --bowtie \\\n\t\t       GRCh38.primary_assembly.genome.fa \\\n\t\t       ref/human_gencode\n```\n\nSimilar to Ensembl annotation, if you want to use GFF3 files (not\nrecommended), add option `--gff3-RNA-patterns transcript`.\n\n#### \u003ca name=\"untypical\"\u003e\u003c/a\u003e Build RSEM references for untypical organisms\n\nFor untypical organisms, such as viruses, you may only have a GFF3 file that containing only genes but not any transcripts. You need to turn on `--gff3-genes-as-transcripts` so that RSEM will make each gene as a unique transcript.\n\nHere is an example command:\n\n```\nrsem-prepare-reference --gff3 virus.gff \\\n               --gff3-genes-as-transcripts \\\n               --bowtie \\\n               virus.genome.fa \\\n               ref/virus\n```\n\n### II. Calculating Expression Values\n\nTo calculate expression values, you should run the\n`rsem-calculate-expression` program.  Run \n\n    rsem-calculate-expression --help\n\nto get usage information or visit the [rsem-calculate-expression\ndocumentation page](http://deweylab.github.io/RSEM/rsem-calculate-expression.html).\n\n#### Calculating expression values from single-end data\n\nFor single-end models, users have the option of providing a fragment\nlength distribution via the `--fragment-length-mean` and\n`--fragment-length-sd` options.  The specification of an accurate fragment\nlength distribution is important for the accuracy of expression level\nestimates from single-end data.  If the fragment length mean and sd are\nnot provided, RSEM will not take a fragment length distribution into\nconsideration.\n\n#### Using an alternative aligner\n\nBy default, RSEM automates the alignment of reads to reference\ntranscripts using the Bowtie aligner. Turn on `--bowtie2` for\n`rsem-prepare-reference` and `rsem-calculate-expression` will allow\nRSEM to use the Bowtie 2 alignment program instead. Please note that\nindel alignments, local alignments and discordant alignments are\ndisallowed when RSEM uses Bowtie 2 since RSEM currently cannot handle\nthem. See the description of `--bowtie2` option in\n`rsem-calculate-expression` for more details. Similarly, turn on\n`--star` will allow RSEM to use the STAR aligner. Turn on `--hisat2-hca`\nwill allow RSEM to use the HISAT2 aligner according to Human Cell\nAtals SMART-Seq2 pipeline. To use an alternative alignment program,\nalign the input reads against the file\n`reference_name.idx.fa` generated by `rsem-prepare-reference`, and\nformat the alignment output in SAM/BAM/CRAM format.  Then, instead of\nproviding reads to `rsem-calculate-expression`, specify the\n`--alignments` option and provide the SAM/BAM/CRAM file as an\nargument.\n\nRSEM requires the alignments of a read to be adjacent. For paired-end\nreads, RSEM also requires the two mates of any alignment be\nadjacent. To check if your SAM/BAM/CRAM file satisfy the requirements,\nrun\n\n    rsem-sam-validator \u003cinput.sam/input.bam/input.cram\u003e\n\nIf your file does not satisfy the requirements, you can use\n`convert-sam-for-rsem` to convert it into a BAM file which RSEM can\nprocess. Run\n \n    convert-sam-for-rsem --help\n\nto get usage information or visit the [convert-sam-for-rsem\ndocumentation\npage](http://deweylab.github.io/RSEM/convert-sam-for-rsem.html).\n\nNote that RSEM does ** not ** support gapped alignments. So make sure\nthat your aligner does not produce alignments with\nintersions/deletions. In addition, you should make sure that you use\n`reference_name.idx.fa`, which is generated by RSEM, to build your\naligner's indices.\n\n### III. Visualization\n\nRSEM includes a copy of SAMtools. When `--no-bam-output` is not\nspecified and `--sort-bam-by-coordinate` is specified, RSEM will\nproduce these three files:`sample_name.transcript.bam`, the unsorted\nBAM file, `sample_name.transcript.sorted.bam` and\n`sample_name.transcript.sorted.bam.bai` the sorted BAM file and\nindices generated by the SAMtools included. All three files are in\ntranscript coordinates. When users in addition specify the\n`--output-genome-bam` option, RSEM will produce three more files:\n`sample_name.genome.bam`, the unsorted BAM file,\n`sample_name.genome.sorted.bam` and\n`sample_name.genome.sorted.bam.bai` the sorted BAM file and\nindices. All these files are in genomic coordinates.\n\n#### a) Converting transcript BAM file into genome BAM file\n\nNormally, RSEM will do this for you via `--output-genome-bam` option\nof `rsem-calculate-expression`. However, if you have run\n`rsem-prepare-reference` and use `reference_name.idx.fa` to build\nindices for your aligner, you can use `rsem-tbam2gbam` to convert your\ntranscript coordinate BAM alignments file into a genomic coordinate\nBAM alignments file without the need to run the whole RSEM\npipeline.\n\nUsage:\n\n    rsem-tbam2gbam reference_name unsorted_transcript_bam_input genome_bam_output\n\nreference_name\t   \t\t  : The name of reference built by `rsem-prepare-reference`\t\t\t\t\nunsorted_transcript_bam_input\t  : This file should satisfy: 1) the alignments of a same read are grouped together, 2) for any paired-end alignment, the two mates should be adjacent to each other, 3) this file should not be sorted by samtools \ngenome_bam_output\t\t  : The output genomic coordinate BAM file's name\n\n#### b) Generating a Wiggle file\n\nA wiggle plot representing the expected number of reads overlapping\neach position in the genome/transcript set can be generated from the\nsorted genome/transcript BAM file output.  To generate the wiggle\nplot, run the `rsem-bam2wig` program on the\n`sample_name.genome.sorted.bam`/`sample_name.transcript.sorted.bam` file.\n\nUsage:    \n\n    rsem-bam2wig sorted_bam_input wig_output wiggle_name [--no-fractional-weight]\n\nsorted_bam_input        : Input BAM format file, must be sorted  \nwig_output              : Output wiggle file's name, e.g. output.wig  \nwiggle_name             : The name of this wiggle plot  \n--no-fractional-weight  : If this is set, RSEM will not look for \"ZW\" tag and each alignment appeared in the BAM file has weight 1. Set this if your BAM file is not generated by RSEM. Please note that this option must be at the end of the command line\n\n#### c) Loading a BAM and/or Wiggle file into the UCSC Genome Browser or Integrative Genomics Viewer(IGV)\n\nFor UCSC genome browser, please refer to the [UCSC custom track help page](http://genome.ucsc.edu/goldenPath/help/customTrack.html).\n\nFor integrative genomics viewer, please refer to the [IGV home page](http://www.broadinstitute.org/software/igv/home). Note: Although IGV can generate read depth plot from the BAM file given, it cannot recognize \"ZW\" tag RSEM puts. Therefore IGV counts each alignment as weight 1 instead of the expected weight for the plot it generates. So we recommend to use the wiggle file generated by RSEM for read depth visualization.\n\nHere are some guidance for visualizing transcript coordinate files using IGV:\n\n1) Import the transcript sequences as a genome \n\nSelect File -\u003e Import Genome, then fill in ID, Name and Fasta file. Fasta file should be `reference_name.idx.fa`. After that, click Save button. Suppose ID is filled as `reference_name`, a file called `reference_name.genome` will be generated. Next time, we can use: File -\u003e Load Genome, then select `reference_name.genome`.\n\n2) Load visualization files\n\nSelect File -\u003e Load from File, then choose one transcript coordinate visualization file generated by RSEM. IGV might require you to convert wiggle file to tdf file. You should use igvtools to perform this task. One way to perform the conversion is to use the following command:\n\n    igvtools tile reference_name.transcript.wig reference_name.transcript.tdf reference_name.genome   \n \n#### d) Generating Transcript Wiggle Plots\n\nTo generate transcript wiggle plots, you should run the\n`rsem-plot-transcript-wiggles` program.  Run \n\n    rsem-plot-transcript-wiggles --help\n\nto get usage information or visit the [rsem-plot-transcript-wiggles\ndocumentation page](http://deweylab.github.io/RSEM/rsem-plot-transcript-wiggles.html).\n\n#### e) Visualize the model learned by RSEM\n\nRSEM provides an R script, `rsem-plot-model`, for visulazing the model learned.\n\nUsage:\n    \n    rsem-plot-model sample_name output_plot_file\n\nsample_name: the name of the sample analyzed    \noutput_plot_file: the file name for plots generated from the model. It is a pdf file    \n\nThe plots generated depends on read type and user configuration. It\nmay include fragment length distribution, mate length distribution,\nread start position distribution (RSPD), quality score vs observed\nquality given a reference base, position vs percentage of sequencing\nerror given a reference base and alignment statistics.\n\nfragment length distribution and mate length distribution: x-axis is fragment/mate length, y axis is the probability of generating a fragment/mate with the associated length\n\nRSPD: Read Start Position Distribution. x-axis is bin number, y-axis is the probability of each bin. RSPD can be used as an indicator of 3' bias\n\nQuality score vs. observed quality given a reference base: x-axis is Phred quality scores associated with data, y-axis is the \"observed quality\", Phred quality scores learned by RSEM from the data. Q = -10log_10(P), where Q is Phred quality score and P is the probability of sequencing error for a particular base\n\nPosition vs. percentage sequencing error given a reference base: x-axis is position and y-axis is percentage sequencing error\n\nAlignment statistics: It includes a histogram and a pie chart. For the histogram, x-axis shows the number of **isoform-level** alignments a read has and y-axis provides the number of reads with that many alignments. The inf in x-axis means number of reads filtered due to too many alignments. For the pie chart, four categories of reads --- unalignable, unique, **isoform-level**multi-mapping, filtered -- are plotted and their percentages are noted. In both the histogram and the piechart, numbers belong to unalignable, unique, multi-mapping, and filtered are colored as green, blue, gray and red. \n \n## \u003ca name=\"example-main\"\u003e\u003c/a\u003e Example\n\nSuppose we download the mouse genome from UCSC Genome Browser.  We do\nnot add poly(A) tails and use `/ref/mouse_0` as the reference name.\nWe have a FASTQ-formatted file, `mmliver.fq`, containing single-end\nreads from one sample, which we call `mmliver_single_quals`.  We want\nto estimate expression values by using the single-end model with a\nfragment length distribution. We know that the fragment length\ndistribution is approximated by a normal distribution with a mean of\n150 and a standard deviation of 35. We wish to generate 95%\ncredibility intervals in addition to maximum likelihood estimates.\nRSEM will be allowed 1G of memory for the credibility interval\ncalculation.  We will visualize the probabilistic read mappings\ngenerated by RSEM on UCSC genome browser. We will generate a list of\ntranscript wiggle plots (`output.pdf`) for the genes provided in `gene_ids.txt`.\nWe will visualize the models learned in\n`mmliver_single_quals.models.pdf`\n\nThe commands for this scenario are as follows:\n\n    rsem-prepare-reference --gtf mm9.gtf --transcript-to-gene-map knownIsoforms.txt --bowtie --bowtie-path /sw/bowtie /data/mm9 /ref/mouse_0\n    rsem-calculate-expression --bowtie-path /sw/bowtie --phred64-quals --fragment-length-mean 150.0 --fragment-length-sd 35.0 -p 8 --output-genome-bam --calc-ci --ci-memory 1024 /data/mmliver.fq /ref/mouse_0 mmliver_single_quals\n    rsem-bam2wig mmliver_single_quals.sorted.bam mmliver_single_quals.sorted.wig mmliver_single_quals\n    rsem-plot-transcript-wiggles --gene-list --show-unique mmliver_single_quals gene_ids.txt output.pdf \n    rsem-plot-model mmliver_single_quals mmliver_single_quals.models.pdf\n\n## \u003ca name=\"simulation\"\u003e\u003c/a\u003e Simulation\n\nRSEM provides users the `rsem-simulate-reads` program to simulate RNA-Seq data based on parameters learned from real data sets. Run\n\n    rsem-simulate-reads\n\nto get usage information or read the following subsections.\n \n### Usage: \n\n    rsem-simulate-reads reference_name estimated_model_file estimated_isoform_results theta0 N output_name [-q]\n\n__reference_name:__ The name of RSEM references, which should be already generated by `rsem-prepare-reference`   \t     \n\n__estimated_model_file:__ This file describes how the RNA-Seq reads will be sequenced given the expression levels. It determines what kind of reads will be simulated (single-end/paired-end, w/o quality score) and includes parameters for fragment length distribution, read start position distribution, sequencing error models, etc. Normally, this file should be learned from real data using `rsem-calculate-expression`. The file can be found under the `sample_name.stat` folder with the name of `sample_name.model`. `model_file_description.txt` provides the format and meanings of this file.    \n\n__estimated_isoform_results:__ This file contains expression levels for all isoforms recorded in the reference. It can be learned using `rsem-calculate-expression` from real data. The corresponding file users want to use is `sample_name.isoforms.results`. If simulating from user-designed expression profile is desired, start from a learned `sample_name.isoforms.results` file and only modify the `TPM` column. The simulator only reads the TPM column. But keeping the file format the same is required. If the RSEM references built are aware of allele-specific transcripts, `sample_name.alleles.results` should be used instead.   \n\n__theta0:__ This parameter determines the fraction of reads that are coming from background \"noise\" (instead of from a transcript). It can also be estimated using `rsem-calculate-expression` from real data. Users can find it as the first value of the third line of the file `sample_name.stat/sample_name.theta`.   \n\n__N:__ The total number of reads to be simulated. If `rsem-calculate-expression` is executed on a real data set, the total number of reads can be found as the 4th number of the first line of the file `sample_name.stat/sample_name.cnt`.   \n\n__output_name:__ Prefix for all output files.   \n\n__--seed seed:__ Set seed for the random number generator used in simulation. The seed should be a 32-bit unsigned integer.\n\n__-q:__ Set it will stop outputting intermediate information.   \n\n### Outputs:\n\noutput_name.sim.isoforms.results, output_name.sim.genes.results: Expression levels estimated by counting where each simulated read comes from.\noutput_name.sim.alleles.results: Allele-specific expression levels estimated by counting where each simulated read comes from.\n\noutput_name.fa if single-end without quality score;   \noutput_name.fq if single-end with quality score;   \noutput_name_1.fa \u0026 output_name_2.fa if paired-end without quality\nscore;   \noutput_name_1.fq \u0026 output_name_2.fq if paired-end with quality score.   \n\n**Format of the header line**: Each simulated read's header line encodes where it comes from. The header line has the format:\n\n    {\u003e/@}_rid_dir_sid_pos[_insertL]\n\n__{\u003e/@}:__ Either '\u003e' or '@' must appear. '\u003e' appears if FASTA files are generated and '@' appears if FASTQ files are generated\n\n__rid:__ Simulated read's index, numbered from 0   \n\n__dir:__ The direction of the simulated read. 0 refers to forward strand ('+') and 1 refers to reverse strand ('-')   \n\n__sid:__ Represent which transcript this read is simulated from. It ranges between 0 and M, where M is the total number of transcripts. If sid=0, the read is simulated from the background noise. Otherwise, the read is simulated from a transcript with index sid. Transcript sid's transcript name can be found in the `transcript_id` column of the `sample_name.isoforms.results` file (at line sid + 1, line 1 is for column names)   \n\n__pos:__ The start position of the simulated read in strand dir of transcript sid. It is numbered from 0   \n\n__insertL:__ Only appear for paired-end reads. It gives the insert length of the simulated read.   \n\n### Example:\n\nSuppose we want to simulate 50 millon single-end reads with quality scores and use the parameters learned from [Example](#example-main). In addition, we set theta0 as 0.2 and output_name as `simulated_reads`. The command is:\n\n    rsem-simulate-reads /ref/mouse_0 mmliver_single_quals.stat/mmliver_single_quals.model mmliver_single_quals.isoforms.results 0.2 50000000 simulated_reads\n\n## \u003ca name=\"gen_trinity\"\u003e\u003c/a\u003e Generate Transcript-to-Gene-Map from Trinity Output\n\nFor Trinity users, RSEM provides a perl script to generate transcript-to-gene-map file from the fasta file produced by Trinity.\n\n### Usage:\n\n    extract-transcript-to-gene-map-from-trinity trinity_fasta_file map_file\n\ntrinity_fasta_file: the fasta file produced by trinity, which contains all transcripts assembled.    \nmap_file: transcript-to-gene-map file's name.    \n\n## \u003ca name=\"de\"\u003e\u003c/a\u003e Differential Expression Analysis\n\nPopular differential expression (DE) analysis tools such as edgeR and\nDESeq do not take variance due to read mapping uncertainty into\nconsideration. Because read mapping ambiguity is prevalent among\nisoforms and de novo assembled transcripts, these tools are not ideal\nfor DE detection in such conditions.\n\nEBSeq, an empirical Bayesian DE analysis tool developed in UW-Madison,\ncan take variance due to read mapping ambiguity into consideration by\ngrouping isoforms with parent gene's number of isoforms. In addition,\nit is more robust to outliers. For more information about EBSeq\n(including the paper describing their method), please visit [EBSeq's\nwebsite](http://www.biostat.wisc.edu/~ningleng/EBSeq_Package).\n\n\nRSEM includes EBSeq in its folder named `EBSeq`. To use it, first type\n\n    make ebseq\n\nto compile the EBSeq related codes. \n\nEBSeq requires gene-isoform relationship for its isoform DE\ndetection. However, for de novo assembled transcriptome, it is hard to\nobtain an accurate gene-isoform relationship. Instead, RSEM provides a\nscript `rsem-generate-ngvector`, which clusters transcripts based on\nmeasures directly relating to read mappaing ambiguity. First, it\ncalculates the 'unmappability' of each transcript. The 'unmappability'\nof a transcript is the ratio between the number of k mers with at\nleast one perfect match to other transcripts and the total number of k\nmers of this transcript, where k is a parameter. Then, Ng vector is\ngenerated by applying Kmeans algorithm to the 'unmappability' values\nwith number of clusters set as 3. This program will make sure the mean\n'unmappability' scores for clusters are in ascending order. All\ntranscripts whose lengths are less than k are assigned to cluster\n3. Run\n\n    rsem-generate-ngvector --help\n\nto get usage information or visit the [rsem-generate-ngvector\ndocumentation\npage](http://deweylab.github.io/RSEM/rsem-generate-ngvector.html).\n\nIf your reference is a de novo assembled transcript set, you should\nrun `rsem-generate-ngvector` first. Then load the resulting\n`output_name.ngvec` into R. For example, you can use \n\n    NgVec \u003c- scan(file=\"output_name.ngvec\", what=0, sep=\"\\n\")\n\n. After that, set \"NgVector = NgVec\" for your differential expression\ntest (either `EBTest` or `EBMultiTest`).\n\n\nFor users' convenience, RSEM also provides a script\n`rsem-generate-data-matrix` to extract input matrix from expression\nresults:\n\n    rsem-generate-data-matrix sampleA.[genes/isoforms].results sampleB.[genes/isoforms].results ... \u003e output_name.counts.matrix\n\nThe results files are required to be either all gene level results or\nall isoform level results. You can load the matrix into R by\n\n    IsoMat \u003c- data.matrix(read.table(file=\"output_name.counts.matrix\"))\n\nbefore running either `EBTest` or `EBMultiTest`.\n\nLastly, RSEM provides two scripts, `rsem-run-ebseq` and\n`rsem-control-fdr`, to help users find differential expressed\ngenes/transcripts. First, `rsem-run-ebseq` calls EBSeq to calculate related statistics\nfor all genes/transcripts. Run \n\n    rsem-run-ebseq --help\n\nto get usage information or visit the [rsem-run-ebseq documentation\npage](http://deweylab.github.io/RSEM/rsem-run-ebseq.html). Second,\n`rsem-control-fdr` takes `rsem-run-ebseq` 's result and reports called\ndifferentially expressed genes/transcripts by controlling the false\ndiscovery rate. Run\n\n    rsem-control-fdr --help\n\nto get usage information or visit the [rsem-control-fdr documentation\npage](http://deweylab.github.io/RSEM/rsem-control-fdr.html). These\ntwo scripts can perform DE analysis on either 2 conditions or multiple\nconditions.\n\nPlease note that `rsem-run-ebseq` and `rsem-control-fdr` use EBSeq's\ndefault parameters. For advanced use of EBSeq or information about how\nEBSeq works, please refer to [EBSeq's\nmanual](http://www.bioconductor.org/packages/devel/bioc/vignettes/EBSeq/inst/doc/EBSeq_Vignette.pdf).\n\nQuestions related to EBSeq should\nbe sent to \u003ca href=\"mailto:nleng@wisc.edu\"\u003eNing Leng\u003c/a\u003e.\n\n## \u003ca name=\"pRSEM\"\u003e\u003c/a\u003e Prior-Enhanced RSEM (pRSEM)\n\n### I. Overview\n\n[Prior-enhanced RSEM (pRSEM)](https://deweylab.github.io/pRSEM/) uses complementary information (e.g. ChIP-seq data) to allocate RNA-seq multi-mapping fragments. We included pRSEM code in the subfolder `pRSEM/` as well as in RSEM's scripts `rsem-prepare-reference` and `rsem-calculate-expression`. \n\n### II. Demo\n\nTo get a quick idea on how to use pRSEM, you can try [this demo](https://github.com/pliu55/pRSEM_demo). It provides a single script, named `run_pRSEM_demo.sh`, which allows you to run all pRSEM's functions. It also contains detailed descriptions of pRSEM's workflow, input and output files.\n\n### III. Installation\n\nTo compile pRSEM, type\n\n    make pRSEM\n\nNote that you need to first compile RSEM before compiling pRSEM. Currently, pRSEM has only been tested on Linux.\n\n\n### IV. Example\n\nTo run pRSEM on the [RSEM example above](#example-main), you need to provide:\n- __ChIP-seq sequencing file(s) in FASTQ format__ or __a ChIP-seq peak file in BED format__. They will be used by pRSEM to obtain complementatry information for allocating RNA-seq multi-mapping fragments.\n- __a genome mappability file in bigWig format__ to let pRSEM build a training\n  set of isoforms to learn prior. Mappability can be obtained from UCSC's \n  ENCODE composite track for [human hg19](http://hgdownload.cse.ucsc.edu/goldenPath/hg19/encodeDCC/wgEncodeMapability/wgEncodeCrgMapabilityAlign36mer.bigWig) \n  and [mouse mm9](http://hgdownload.cse.ucsc.edu/goldenPath/mm9/encodeDCC/wgEncodeMapability/wgEncodeCrgMapabilityAlign36mer.bigWig). For other genomes, you \n  can generate the mappability file by following [this tutorial] (http://wiki.bits.vib.be/index.php/Create_a_mappability_track#Install_and_run_the_GEM_library_tools).\n\nAssuming you would like to use RNA Pol II's ChIP-seq sequencing files `/data/mmliver_PolIIRep1.fq.gz` and `/data/mmliver_PolIIRep2.fq.gz`, with ChIP-seq control `/data/mmliver_ChIPseqCtrl.fq.gz`. Also, assuming the mappability file for mouse genome is `/data/mm9.bigWig` and you prefer to use STAR located at `/sw/STAR` to align RNA-seq fragments and use Bowtie to align ChIP-seq reads. Then, you can use the following commands to run pRSEM:\n\n    rsem-prepare-reference --gtf mm9.gtf \\\n                           --star \\\n                           --star-path /sw/STAR \\\n                           -p 8 \\\n                           --prep-pRSEM \\\n                           --bowtie-path /sw/bowtie \\\n                           --mappability-bigwig-file /data/mm9.bigWig \\\n                           /data/mm9 \\\n                           /ref/mouse_0\n  \n    rsem-calculate-expression --star \\\n                              --star-path /sw/STAR \\\n                              --calc-pme \\\n                              --run-pRSEM \\\n                              --chipseq-target-read-files /data/mmliver_PolIIRep1.fq.gz,/data/mmliver_PolIIRep2.fq.gz \\\n                              --chipseq-control-read-files /data/mmliver_ChIPseqCtrl.fq.gz \\\n                              --bowtie-path /sw/bowtie \\\n                              -p 8 \\\n                              /data/mmliver.fq \\\n                              /ref/mouse_0 \\\n                              mmliver_single_quals\n\n\nTo find out more about pRSEM options and examples, you can use the commands:\n\n    rsem-prepare-reference --help\n\nand \n\n    rsem-calculate-expression --help\n\n\n### V. System Requirements\n- Linux\n- Perl version \u003e= 5.8.8\n- Python version \u003e= 2.7.3\n- R version \u003e= 3.3.1\n- Bioconductor 3.3\n\n\n### VI. Required External Packages\nAll the following packages will be automatically installed when compiling pRSEM.\n- [data.table 1.9.6](https://cran.r-project.org/web/packages/data.table/index.html): an extension of R's data.frame, heavily used by pRSEM.\n- [GenomicRanges 1.24.3](https://bioconductor.org/packages/release/bioc/html/GenomicRanges.html): efficient representing and manipulating genomic intervals, heavily used by pRSEM.\n- [ShortRead 1.30.0](https://bioconductor.org/packages/release/bioc/html/ShortRead.html): guessing the encoding of ChIP-seq FASTQ file's quality score.\n- [caTools 1.17.1](https://cran.r-project.org/web/packages/caTools/index.html): used for SPP Peak Caller.\n- [SPP Peak Caller](https://code.google.com/archive/p/phantompeakqualtools/):\n  ChIP-seq peak caller. Source code was slightly modified in terms of included headers in order to be compiled under R v3.3.1.\n- [IDR](https://sites.google.com/site/anshulkundaje/projects/idr/idrCode.tar.gz?attredirects=0):\n  calculating Irreproducible Discovery Rate to call peaks from multiple ChIP-seq replicates.\n\n\n## \u003ca name=\"authors\"\u003e\u003c/a\u003e Authors\n\n[Bo Li](http://bli25ucb.github.io/) and [Colin Dewey](https://www.biostat.wisc.edu/~cdewey/) designed the RSEM algorithm. [Bo Li](http://bli25ucb.github.io/) implemented the RSEM software. [Peng Liu](https://www.biostat.wisc.edu/~cdewey/group.html) contributed the STAR aligner options and prior-enhanced RSEM (pRSEM).\n\n## \u003ca name=\"acknowledgements\"\u003e\u003c/a\u003e Acknowledgements\n\nRSEM uses the [Boost C++](http://www.boost.org/) and\n[SAMtools](http://www.htslib.org/) libraries. RSEM includes\n[EBSeq](http://www.biostat.wisc.edu/~ningleng/EBSeq_Package/) for\ndifferential expression analysis.\n\nWe thank earonesty, Dr. Samuel Arvidsson, John Marshall, and Michael\nR. Crusoe for contributing patches.\n\nWe thank Han Lin, j.miller, Jo\u0026euml;l Fillon, Dr. Samuel G. Younkin,\nMalcolm Cook, Christina Wells, Uro\u0026#353; \u0026#352;ipeti\u0026#263;,\noutpaddling, rekado, and Josh Richer for suggesting possible fixes.\n\n**Note** that `bam_sort.c` of SAMtools is slightly modified so that\n  `samtools sort -n` will not move the two mates of paired-end\n  alignments apart. In addition, we turn on the `--without-curses`\n  option when configuring SAMtools and thus SAMtools' curses-based\n  `tview` subcommand is not built.\n\n## \u003ca name=\"license\"\u003e\u003c/a\u003e License\n\nRSEM is licensed under the [GNU General Public License\nv3](http://www.gnu.org/licenses/gpl-3.0.html).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdeweylab%2FRSEM","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fdeweylab%2FRSEM","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdeweylab%2FRSEM/lists"}