RNA sequencing (RNA-seq) is a technology that reads out the RNA molecules in a biological sample, giving researchers a snapshot of which genes are active and how active they are at a given moment. Unlike DNA, which is essentially the same in every cell of your body, RNA changes constantly depending on cell type, health status, and environmental signals, making it a far more dynamic source of biological information. The technology has become a workhorse of modern biology and medicine, but its rapid evolution over the past fifteen years means the term “RNA-seq” now covers a surprisingly wide family of methods, from bulk tissue profiling to reading individual cells one at a time.
What RNA Sequencing Actually Measures
Your DNA is a static blueprint. RNA is the working copy that cells produce when they need to use a gene’s instructions. By capturing and reading those RNA molecules, RNA-seq tells you which genes are turned on, how strongly they are expressed, and in some cases which specific versions of a gene product a cell is making. That last point matters because a single gene can produce multiple RNA variants through a process called alternative splicing, and those variants can have very different biological effects.
Before RNA-seq existed, the standard tool for measuring gene expression was the DNA microarray, which works by attaching known DNA sequences to a chip and seeing which RNAs stick to them. RNA-seq replaced this approach for most applications because it does not require you to know in advance what you are looking for, and it captures a much wider range of expression levels. A head-to-head comparison in activated T cells found that RNA-seq’s dynamic range spanned roughly five orders of magnitude, compared to about three and a half orders for microarrays on the same set of genes.1PubMed Central. Comparison of RNA-Seq and Microarray in Transcriptome Profiling of Activated T Cells In practical terms, that means RNA-seq can detect very lowly expressed genes that a microarray would miss, while still accurately measuring highly expressed ones.
How the Process Works, From Sample to Data
A typical RNA-seq experiment has three broad stages: preparing the RNA for sequencing, running it through a sequencing machine, and computationally making sense of the resulting data. Each stage introduces choices that affect the final results.
The first decision is how to select the RNA you care about. Most of the RNA in a cell is ribosomal RNA, which is essential for protein production but usually not what researchers want to study. Two common strategies deal with this: you can fish out the messenger RNA (mRNA) that carries protein-coding instructions by grabbing its poly(A) tail, or you can remove the ribosomal RNA and sequence everything else. These two approaches produce meaningfully different results. Poly(A) selection concentrates your sequencing effort on protein-coding transcripts, giving better coverage of those genes per dollar spent. Ribosomal depletion captures a broader set of RNA types, including non-coding RNAs, but requires substantially more sequencing to reach the same depth on coding genes. One study found that for blood-derived RNA, you would need roughly three times as many sequencing reads with ribosomal depletion to match the coding-gene coverage of poly(A) selection.2PubMed Central. Evaluation of two main RNA-seq approaches for gene quantification in clinical RNA sequencing: polyA+ selection versus rRNA depletion The trade-off is straightforward: if you mainly care about protein-coding genes, poly(A) selection is more economical; if you need to see non-coding RNA or are working with degraded samples where poly(A) tails may be damaged, ribosomal depletion is the way to go.
Once the RNA is selected, it gets converted into a form the sequencing machine can read. This typically involves making a DNA copy of each RNA molecule (since most sequencers read DNA), adding short adapter sequences to the ends, and amplifying the material. Some newer protocols streamline this process. The SHERRY method, for example, works by directly tagging RNA/DNA hybrid molecules, which simplifies library preparation and can work with as little as 200 nanograms of starting RNA.3PubMed Central. Protocol for RNA-seq library preparation from low-volume total RNA by RNA/cDNA hybrid tagmentation
Making Sense of Millions of Short Reads
A single RNA-seq run produces millions of short sequence fragments, each typically 75 to 150 bases long. Turning those fragments into useful biological information requires computational tools that map each fragment back to the genome or transcriptome and then count how many fragments came from each gene. The choice of mapping software used to feel high-stakes, but benchmarking studies have found that most widely used tools produce remarkably similar results. A comparison of seven different alignment tools on plant data found similarity scores above 0.98 across all pairwise comparisons, with some tools like Salmon and kallisto being nearly identical.4PubMed Central. Evaluation of Seven Different RNA-Seq Alignment Tools Based on Experimental Data from the Model Plant Arabidopsis thaliana Newer “alignment-free” methods that skip the step of placing each read onto the genome are both fast and accurate,5PubMed Central. Evaluation and comparison of computational tools for RNA-seq isoform quantification and direct comparisons have shown that genome alignment and transcriptome pseudoalignment are largely interchangeable for most quantification purposes.6bioRxiv. A direct comparison of genome alignment and transcriptome pseudoalignment
After counting, the most common analysis is differential gene expression: figuring out which genes are turned up or down between two conditions, say healthy tissue versus a tumor. The statistical challenge here is that RNA-seq count data do not follow the simple distributions you might expect. Counts for lowly expressed genes are noisy, and the variance often exceeds what standard models predict. Specialized statistical approaches based on the negative binomial distribution were developed specifically to handle this overdispersion.7Statistical Applications in Genetics and Molecular Biology. The NBP Negative Binomial Model for Assessing Differential Gene Expression from RNA-Seq These methods, implemented in widely used software packages, have become the standard way to call genes differentially expressed.
RNA Quality Matters More Than You Might Think
The quality of your starting RNA has a surprisingly large impact on results. RNA degrades quickly after a cell dies or a tissue is removed, and degraded RNA produces biased measurements. The RNA Integrity Number (RIN), a standardized score from 1 to 10, was developed because older measures of RNA quality were inconsistent.8PubMed Central. The RIN: an RNA integrity number for assigning integrity values to RNA measurements Studies have documented widespread effects of RNA degradation on measured gene expression levels, along with a loss of library complexity in more degraded samples.9PubMed Central. RNA-seq: impact of RNA degradation on transcript quantification In practical terms, this means that if you are comparing two groups of samples and one group had poorer tissue handling, the resulting gene expression differences might reflect degradation rather than biology.
The US FDA coordinated a large-scale benchmarking effort, the Sequencing Quality Control (SEQC) project, which generated over 100 billion reads across multiple sequencing platforms and laboratory sites to systematically assess RNA-seq accuracy and reproducibility.10PubMed Central. A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the Sequencing Quality Control Consortium That kind of standardization effort is essential for clinical use, where regulatory agencies need confidence that results from different laboratories are comparable.
Single-Cell RNA Sequencing Changed the Game
Traditional “bulk” RNA-seq grinds up a tissue sample and measures the average gene expression across millions of cells. Single-cell RNA-seq (scRNA-seq) isolates individual cells and profiles each one separately, revealing the diversity hidden within what bulk methods treat as a uniform population. This distinction turned out to be transformative. A tumor that looks like one thing in bulk might contain dozens of distinct cell populations with different gene expression programs, and understanding that heterogeneity matters for treatment.
The throughput of single-cell methods has increased dramatically thanks to droplet-based microfluidic approaches. The inDrop method, for instance, can index over 15,000 cells in an hour by encapsulating individual cells into tiny droplets along with barcoded beads that tag each cell’s RNA with a unique identifier.11Nature Protocols. Single-cell barcoding and sequencing using droplet microfluidics Newer platforms like spinDrop add fluorescence-based sorting to enrich for droplets containing single viable cells, further improving sensitivity while reducing cost.12Nature Communications. spinDrop: a droplet microfluidic platform to maximise single-cell sequencing information content
A major technical concern in single-cell work is amplification noise. Because each cell contains only tiny amounts of RNA, the material must be amplified extensively before sequencing, and that amplification can distort the true abundances. Unique molecular identifiers (UMIs), short random sequences attached to each original RNA molecule before amplification, solve this by letting you count original molecules rather than amplified copies. UMIs can nearly eliminate amplification noise and, combined with optimized reagents, have produced roughly a fivefold improvement in RNA capture efficiency.13PubMed. Quantitative single-cell RNA-seq with unique molecular identifiers Computational tools for identifying duplicated molecules from UMI sequences have become a standard part of single-cell analysis pipelines.14PubMed Central. BUTTERFLY: addressing the pooled amplification paradox with unique molecular identifiers in single-cell RNA-seq
When You Cannot Dissociate Cells
Not every tissue can be easily broken into individual cells. Adipose tissue, brain, and many frozen clinical specimens are difficult to dissociate without damaging or destroying the cells. Single-nucleus RNA-seq (snRNA-seq) addresses this by isolating nuclei rather than whole cells and profiling their RNA content. This approach achieves comparable gene detection to whole-cell methods while avoiding dissociation-induced stress responses that can alter gene expression during sample preparation.15PubMed Central. Advantages of Single-Nucleus over Single-Cell RNA Sequencing of Adult Kidney: Rare Cell Types and Novel Cell States Revealed in Fibrosis For cancer research, where biobanks store large numbers of frozen tumor specimens, snRNA-seq has been particularly valuable because it is compatible with frozen samples that cannot be used for conventional single-cell approaches.16Nature Medicine. A single-cell and single-nucleus RNA-Seq toolbox for fresh and frozen human tumors
Keeping Cells in Their Spatial Context
Both single-cell and single-nucleus methods destroy the tissue architecture in the process of isolating cells. Spatial transcriptomics technologies preserve it, measuring gene expression while retaining information about where each measurement came from within the tissue. These approaches have been applied across neuroscience, developmental biology, plant biology, and cancer research, providing insights that neither bulk nor dissociation-based single-cell methods can offer.17PubMed Central. Exploring tissue architecture using spatial transcriptomics Knowing that a gene is expressed is useful; knowing that it is expressed specifically at the interface between tumor and immune cells, for example, is a different kind of information entirely.
Long-Read Sequencing Reveals Full-Length Transcripts
Standard RNA-seq platforms produce short reads that must be computationally stitched back together, which makes it hard to determine exactly which combination of exons a given RNA molecule contains. Long-read platforms from Oxford Nanopore and Pacific Biosciences can sequence full-length transcripts in a single read, which is a fundamental advantage for studying alternative splicing and isoform diversity.18PubMed Central. LIQA: long-read isoform quantification and analysis
The results can be eye-opening. A study using Oxford Nanopore native RNA sequencing on human monocytes identified over 24,000 expressed isoforms, and about 62% of those were absent from the standard reference annotation, meaning they were previously unrecognized transcript variants.19PubMed Central. Native long-read RNA sequencing of human monocytes reveals activation-induced alternative splicing toward functional isoforms When the monocytes were activated by an immune stimulus, widespread isoform switching occurred: activated cells preferentially expressed longer, coding-competent isoforms with more complex protein domain structures. That level of isoform-specific resolution is extremely difficult to achieve with short reads alone.
Long-read platforms also open the door to detecting RNA modifications directly. Oxford Nanopore sequencing works by threading an RNA or DNA strand through a tiny pore and measuring electrical current changes. Chemical modifications to RNA bases, such as N6-methyladenosine (m6A), alter the current signal in detectable ways, enabling researchers to map these modifications without additional chemical treatment steps.20PubMed Central. Detecting m6A RNA modification from nanopore sequencing using a semi-supervised learning framework This field, sometimes called epitranscriptomics, is revealing a layer of gene regulation that was largely invisible to earlier methods.
Predicting Where a Cell Is Headed
One of the more creative uses of single-cell RNA-seq data exploits a biological quirk: newly made RNA transcripts contain introns that are spliced out as the molecule matures. By measuring the ratio of unspliced to spliced RNA for each gene, researchers can estimate “RNA velocity,” essentially a prediction of each cell’s future gene expression state on a timescale of hours.21PubMed Central. RNA velocity of single cells If a cell has lots of unspliced RNA for a particular gene, that gene’s expression is likely increasing; if the unspliced pool is depleted, expression is winding down. Dynamical modeling extensions have refined these estimates by accounting for transient cell states and varying transcription rates over time.22PubMed Central. Generalizing RNA velocity to transient cell states through dynamical modeling The result is a tool that can infer developmental trajectories and cell-fate decisions from a single snapshot experiment, which is remarkably useful for studying processes like how stem cells differentiate or how cancer cells evolve within a tumor.
Combining RNA With Other Measurements in the Same Cell
Knowing a cell’s gene expression is valuable, but cells are defined by more than their RNA. CITE-seq (cellular indexing of transcriptomes and epitopes by sequencing) measures both RNA and surface protein levels in the same individual cell by using antibodies labeled with short DNA barcodes.23Nature Methods. Simultaneous epitope and transcriptome measurement in single cells This matters because RNA levels and protein levels do not always agree, and for many biological questions, especially in immunology, the proteins on a cell’s surface are what determine its function. CITE-seq has been widely applied in immune-related research, including studies of influenza and COVID-19.24PubMed Central. Dynamic thresholding and tissue dissociation optimization for CITE-seq identifies differential surface protein abundance in metastatic melanoma Most early applications focused on blood-derived samples because liquid biopsies are easier to work with, but methods are being developed to extend multi-omics profiling to solid tissues as well.
Reading the RNA of Entire Microbial Communities
RNA-seq does not have to be applied to a single organism’s cells. Metatranscriptomics applies sequencing to complex microbial communities, like the bacteria in your gut or the microbes in a soil sample, to determine which organisms are active and what they are doing. Standard approaches to studying microbiomes, such as 16S ribosomal gene sequencing or shotgun DNA sequencing, tell you which organisms are present but not whether they are metabolically active. Metatranscriptomics fills that gap by capturing the RNA being actively produced, revealing how a community responds to changing conditions over time.25PubMed Central. Advances and Challenges in Metatranscriptomic Analysis This distinction between presence and activity turns out to be crucial for understanding how microbiomes influence human health and disease.26Trends in Molecular Medicine. Metatranscriptomics in human health and disease
Clinical Uses and Fusion Gene Detection
RNA-seq is moving steadily from research laboratories into clinical practice. One of its most established clinical applications is the detection of fusion genes in cancer, where two genes that are normally separate get joined together, producing abnormal proteins that can drive tumor growth. Targeted RNA-seq panels can detect these fusions with high sensitivity while simultaneously profiling additional genes with prognostic value, including transcription factors, cell-type markers, and immune genes that might inform treatment decisions.27Nature Communications. Diagnosis of fusion genes using targeted RNA sequencing Long-read RNA-seq is proving particularly useful here because sequencing full-length transcripts makes it easier to identify fusion boundaries and distinguish true fusions from artifacts.28PubMed. Gene Fusion Detection in Long-Read Transcriptome Datasets from Multiple Cancer Cell Lines
The Batch Effect Problem
When RNA-seq data from different experiments, laboratories, or even different days in the same lab are combined, technical variation between batches can swamp the biological signal. Correcting these batch effects is one of the trickiest computational challenges in the field, and a recent study found that many widely used correction methods are poorly calibrated, meaning the correction process itself creates measurable artifacts in the data.29PubMed Central. Batch correction methods used in single-cell RNA sequencing analyses are often poorly calibrated Benchmarking efforts have compared over a dozen correction methods on their ability to remove batch effects while preserving genuine biological differences between cell types,30PubMed Central. A benchmark of batch-effect correction methods for single-cell RNA sequencing data and the “best” method depends on the specific dataset characteristics, such as whether batch sizes are balanced.31PubMed Central. Batch-effect correction in single-cell RNA sequencing data using JIVE There is no universally safe default choice, which means analysts need to evaluate correction quality on their own data rather than blindly trusting a popular tool.
Privacy Risks in RNA-Seq Data
An underappreciated issue is that RNA-seq data carry genetic information that can identify individuals. Even though the experiments are designed to study gene expression rather than genotypes, the raw sequencing reads contain the genetic variants of the person whose tissue was sequenced. These expressed variants can be extracted and used to re-identify individuals or link supposedly anonymous datasets back to a specific person.32PubMed Central. Data sanitization to reduce private information leakage from functional genomics As the balance in genomic data acquisition shifts toward large-scale functional genomics and away from pure DNA sequencing, this privacy concern becomes more pressing. Data sanitization approaches have been developed to strip identifying variants from public datasets, but the tension between open data sharing and individual privacy remains an active area of work. For researchers who contribute tissue samples or blood for RNA-seq studies, this is worth knowing: even “gene expression data” can contain identifying genetic fingerprints.
Small RNA Sequencing
Not all RNA-seq focuses on messenger RNA. Small RNA-seq targets the short regulatory molecules, typically 18 to 30 nucleotides long, that play outsized roles in gene regulation. MicroRNAs are the best-known members of this class, but the technique also captures piRNAs, snoRNAs, and other small species across diverse sample types including cells, tissues, and cell-free fluids like blood plasma.33Nature Biotechnology. Comprehensive multi-center assessment of small RNA-seq methods for quantitative miRNA profiling Small RNA-seq enables genome-wide profiling of both known and novel microRNA variants,34PubMed Central. Small RNA-Sequencing: Approaches and Considerations for miRNA Analysis which has practical implications for biomarker discovery. MicroRNAs circulating in blood have attracted interest as potential non-invasive markers for cancer, cardiovascular disease, and other conditions, and small RNA-seq is the tool that makes those discovery efforts possible at scale.

