ATAC-seq (Assay for Transposase-Accessible Chromatin using sequencing) is a laboratory technique that maps which stretches of DNA in a cell are physically open and available for use. It works by deploying a bacterial enzyme called Tn5 transposase, which cuts DNA and inserts sequencing tags only where the chromatin is loose and accessible, essentially flagging the genome’s active regulatory zones in one step.1PubMed Central. ATAC-seq: A Method for Assaying Chromatin Accessibility Genome-Wide Since its introduction in 2013, ATAC-seq has become one of the most widely used tools in genomics, reshaping how researchers study gene regulation, disease, and development.
What “Chromatin Accessibility” Actually Means
Your DNA is not a loose string floating around inside a cell. It is wound tightly around protein spools called histones, and the resulting fiber, chromatin, is further packed and folded to fit roughly two meters of DNA into a microscopic nucleus. But not all of it is packed equally. Genes that a cell is actively using, or getting ready to use, tend to sit in looser, more open stretches of chromatin. Regulatory elements like promoters and enhancers, the switches that turn genes on or off, also live in these open zones. The pattern of which regions are open and which are closed varies dramatically between cell types. A liver cell and a neuron carry the same DNA but keep different parts of it accessible, which is a large part of why they look and function so differently.
Mapping these open regions gives researchers a snapshot of a cell’s regulatory state. Before ATAC-seq, techniques like DNase-seq and MNase-seq could do this, but they required millions of cells, involved laborious multi-day protocols, and demanded significant expertise. ATAC-seq changed the equation by drastically reducing the amount of starting material needed and compressing the workflow.
How the Technique Works
The core of ATAC-seq is surprisingly elegant. The Tn5 transposase is a hyperactive enzyme originally derived from a bacterial transposon. Researchers pre-load it with short DNA sequences called sequencing adapters, then let it loose on a sample of cells. The enzyme can only access and cut DNA where chromatin is open. Wherever it cuts, it simultaneously pastes in those adapter sequences, a process called tagmentation. After tagmentation, researchers amplify the tagged fragments and feed them into a sequencing machine. The result is a genome-wide map showing every spot where Tn5 was able to get in, which corresponds to the regions of open chromatin.2PubMed Central. ATAC-seq: A Method for Assaying Chromatin Accessibility Genome-Wide
Because the enzyme does the cutting and tagging in a single reaction, the wet-lab portion is fast. A researcher with standard molecular biology skills can prepare libraries for about a dozen samples in a single working day.3PubMed Central. Chromatin accessibility profiling by ATAC-seq Simplified versions of the protocol have pushed the bar even lower, enabling work with organisms as small as individual water fleas or parasitic worms, using roughly 10,000 cells and minimal hands-on expertise.4PubMed Central. A simple ATAC-seq protocol for population epigenetics
The Mitochondrial DNA Problem
One of the biggest practical headaches with ATAC-seq has nothing to do with chromatin. Mitochondria, the energy-producing organelles inside cells, carry their own small circular genomes. These genomes lack the protective histone packaging of nuclear DNA, making them wide open to Tn5. In early ATAC-seq experiments, mitochondrial reads routinely consumed around half of all sequencing output, meaning researchers were paying to sequence junk data that told them nothing about gene regulation.5PubMed Central. ATAC-seq Assay with Low Mitochondrial DNA Contamination from Primary Human CD4+ T Lymphocytes
Several strategies have emerged to deal with this. Optimized cell lysis buffers can strip away mitochondria before Tn5 is added, bringing contamination down from roughly 50% to as low as 3% of reads.6PubMed Central. ATAC-seq Assay with Low Mitochondrial DNA Contamination from Primary Human CD4+ T Lymphocytes Another approach uses CRISPR to selectively chop up mitochondrial DNA fragments after tagmentation. Testing both methods in the same cell line, one group found that CRISPR treatment gave the cleanest data, reducing mitochondrial reads while preserving the number and quality of chromatin peaks. Removing detergent from the lysis buffer, by contrast, cut mitochondrial reads but also increased background noise and lost real signal.7Scientific Reports. Reducing mitochondrial reads in ATAC-seq using CRISPR/Cas9 A separate chemistry-based approach achieved a reduction from 52% down to about 6%.8PubMed Central. Significant reduction in mitochondrial DNA contamination improves sequencing power for ATAC-seq libraries The choice of strategy often depends on the cell type and downstream goals, but all of them substantially improve how much useful data you get per sequencing dollar.
Footprinting and Transcription Factor Binding
ATAC-seq can do more than just identify open regions. Within those open stretches, wherever a protein is physically sitting on the DNA, Tn5 is blocked from cutting. This creates tiny “footprints,” dips in the cutting signal at the exact spots where transcription factors are bound. Computational tools can detect these footprints and infer which regulatory proteins are active in a given cell type, without needing an antibody or any advance knowledge of which factors to look for.
Getting footprinting right, however, requires careful handling of Tn5’s own biases. The enzyme does not cut DNA with perfect randomness. It has sequence preferences and shows strand-specific cleavage patterns that are shaped by nearby nucleosomes. Early footprinting methods designed for DNase-seq did not account for these quirks. A tool called HINT-ATAC was the first to build a model of Tn5’s specific insertion preferences, learning its biases so they could be subtracted out. This significantly improved the accuracy of predicted transcription factor binding sites compared to older approaches.9PubMed Central. Identification of transcription factor binding sites using ATAC-seq A later framework called TOBIAS expanded on this, enabling researchers to track how transcription factor binding changes over time across hundreds of factors simultaneously. When validated against independent protein-binding data, TOBIAS outperformed existing methods for bias correction and footprint detection.10Nature Communications. ATAC-seq footprinting unravels kinetics of transcription factor binding during zygotic genome activation
A more recent refinement involves adding unique molecular identifiers, short random barcodes attached to each DNA fragment before amplification. Because PCR amplification introduces duplicate reads that can distort the signal, tagging each original molecule lets researchers distinguish true biological reads from duplicates. In rice, this UMI-based ATAC-seq approach rescued only about 6% more total reads, but those rescued reads contributed to the identification of over 50% more footprints, a disproportionate gain that reflects how sensitive footprinting is to data depth and accuracy.11Communications Biology. ATAC-seq with unique molecular identifiers improves quantification and footprinting
Going Single Cell
Bulk ATAC-seq averages the chromatin landscape across thousands or millions of cells. That is fine when you are studying a relatively uniform cell population, but tissues like the brain or a tumor contain dozens of distinct cell types mixed together. Single-cell and single-nucleus ATAC-seq (scATAC-seq and snATAC-seq) solve this by profiling chromatin accessibility one cell at a time, revealing which regulatory elements are active in each cell type even before the cells look or behave differently from one another.12PubMed. Single-Nucleus ATAC-seq for Mapping Chromatin Accessibility in Individual Cells of Murine Hearts
The data generated by these experiments is extremely sparse, since each individual cell yields only a fraction of the genome’s open regions, and the datasets are massive. Specialized software is needed to handle both challenges. SnapATAC, for example, uses a mathematical shortcut called the Nyström method to process datasets from up to a million cells. Applied to roughly 55,000 single-nucleus profiles from the mouse brain, it resolved 31 distinct cell populations and identified around 370,000 candidate regulatory elements.13Nature Communications. Comprehensive analysis of single cell ATAC-seq data with SnapATAC
In the pancreas, single-cell ATAC-seq has been used to separate the chromatin profiles of different hormone-producing cell types and uncover regulatory signatures specific to type 2 diabetes, work that would be impossible in a bulk experiment where the signals from rare cell types are drowned out.14PubMed Central. Single-cell ATAC-Seq in human pancreatic islets and deep learning upscaling of rare cells reveals cell-specific type 2 diabetes regulatory signatures
Cancer and the Regulatory Landscape of Tumors
One of the largest ATAC-seq efforts in cancer biology profiled 410 tumor samples across 23 cancer types as part of The Cancer Genome Atlas project. That work identified over 560,000 accessible DNA elements, vastly expanding the known catalog of regulatory regions in human cancers. By combining chromatin accessibility with mutation, gene expression, and methylation data from the same tumors, the study revealed how different molecular subtypes of cancer use distinct enhancers, identified the transcription factors driving those subtypes through protein-DNA footprints, and linked noncoding mutations to enhancer activation with potential effects on patient survival.15PubMed Central. The chromatin accessibility landscape of primary human cancers
ATAC-seq has also illuminated why the immune system sometimes fails to destroy tumors. In mice, profiling of CD8+ T cells that had infiltrated tumors revealed a distinct pattern of chromatin accessibility tied to T-cell exhaustion, the state where immune cells are present but functionally impaired. The exhaustion-specific open regions were enriched for binding sites of particular transcription factor families, providing molecular targets that could potentially be manipulated to reinvigorate the immune response.16PubMed Central. Exhaustion-associated regulatory regions in CD8(+) tumor-infiltrating T cells
Mapping the Brain
A brain atlas project used single-nucleus ATAC-seq to profile chromatin accessibility across 1.1 million cells from 42 regions of the adult human brain. The analysis resolved 107 distinct cell types and mapped over 540,000 candidate regulatory elements. Perhaps more striking, the study found strong connections between specific brain cell types and neuropsychiatric disorders including schizophrenia, bipolar disorder, Alzheimer’s disease, and major depression, with deep learning models predicting how noncoding genetic risk variants might disrupt regulation in those cell types.17PubMed Central. A comparative atlas of single-cell chromatin accessibility in the human brain This kind of work moves the field beyond knowing that a genetic variant is associated with a disease toward understanding in which cell type and through which regulatory mechanism it acts.
Connecting Noncoding Variants to Disease
The vast majority of genetic variants linked to complex diseases by genome-wide association studies sit outside of genes, in the noncoding regions of the genome. For years, interpreting these hits was a bottleneck: researchers could see that a stretch of DNA was statistically linked to, say, heart disease, but had no idea what that DNA actually did. ATAC-seq has become a central tool for cracking these cases open.
Because Tn5 cuts with base-pair resolution, ATAC-seq footprints can narrow down a disease-associated region to the exact spot where a transcription factor binds. Footprint quantitative trait loci, which measure how a genetic variant alters the depth of a Tn5 footprint, can pinpoint the causal variant within a broad region where many variants are correlated due to linkage disequilibrium.18PubMed Central. Characterization of non-coding variants associated with transcription factor binding through ATAC-seq-defined footprint QTLs in liver A complementary approach looks at allelic imbalance within ATAC-seq peaks, asking whether one copy of a chromosome is more open than the other at a particular spot. When that imbalance colocalizes with a disease-associated variant, it provides strong evidence that the variant is directly affecting chromatin accessibility and, by extension, gene regulation.19Trends in Genetics. Leveraging allelic imbalance in ATAC-seq to identify functional noncoding variants and dissect gene regulation in complex diseases
In a concrete example from cardiology, researchers integrated ATAC-seq with three-dimensional chromatin contact maps and gene expression data in heart fibroblast cells. They identified noncoding variants sitting in open chromatin that made physical contact, via long-range chromatin loops, with the promoters of expressed genes. This kind of multi-layered analysis can reconstruct entire regulatory circuits connecting a disease-associated variant to the gene it controls and the transcription factor that mediates the effect.20Nature Communications. Dissecting regulatory non-coding GWAS loci reveals fibroblast causal genes with pathophysiological relevance to heart failure
Combining ATAC-seq with Gene Expression Data
Single-cell ATAC-seq tells you where the genome is open; single-cell RNA-seq tells you which genes are active. Combining the two in the same analysis creates a much richer picture, linking regulatory elements to the genes they control. Several computational frameworks have been developed to do this. FigR, for example, pairs single-cell ATAC-seq and RNA-seq data computationally, connects distal regulatory elements to their target genes, and infers gene-regulatory networks that identify the transcription factors orchestrating a cell’s behavior.21Cell Genomics. High-throughput chromatin accessibility and gene-regulatory analyses of human blood cell stimulation
Other tools take different approaches to the same problem. ScReNI uses a machine-learning algorithm to capture nonlinear relationships between chromatin accessibility and gene expression at the single-cell level, building regulatory networks cell by cell rather than averaging across clusters.22PubMed Central. ScReNI: Single-cell Regulatory Network Inference Through Integrating scRNA-seq and scATAC-seq Data A method called scMI models the two data types as a network with different node types and uses attention-based learning to capture long-range interactions between genes and chromatin peaks, embedding everything into a shared space where cells are grouped by their combined regulatory and transcriptional profiles.23Briefings in Bioinformatics. Integrating scRNA-seq and scATAC-seq with inter-type attention heterogeneous graph neural networks The field is actively iterating on these integration tools because no single approach yet handles every dataset or biological question equally well.
Toward Liquid Biopsy and Clinical Use
When cells die, they release fragments of DNA into the bloodstream. These cell-free DNA fragments carry an imprint of the chromatin structure of the cell they came from, because nucleosome-bound DNA is protected from degradation while open regions are not. Researchers have shown that transcription factor accessibility patterns derived from ATAC-seq data in tumors can be reliably inferred from the fragmentation patterns in plasma DNA, raising the possibility of detecting and classifying cancers from a blood draw.24Nature Communications. Inference of transcription factor binding from cell-free DNA enables tumor subtype prediction and early detection
Building on this concept, machine learning approaches have been used to identify chromatin regions that are specifically open in particular cancer types but closed in blood cells. A scoring system that quantifies this difference can prioritize which genomic regions to include in targeted sequencing panels for liquid biopsy, focusing resources on the markers most likely to be detectable in a patient’s blood.25Scientific Reports. Integrating chromatin accessibility states in the design of targeted sequencing panels for liquid biopsy This work is still in the development phase, but it represents one of the more direct paths from ATAC-seq data to clinical decision-making.
Peak Calling and Computational Analysis
Raw ATAC-seq data is a pile of short sequencing reads. Turning those reads into a meaningful map of open chromatin requires a computational pipeline. After aligning reads to a reference genome and removing duplicates, the critical step is peak calling: identifying regions where reads pile up significantly above background, indicating true open chromatin. The most widely used peak caller, MACS2, was originally designed for a different type of experiment and defaults to treating reads as single-ended. Using that default on paired-end ATAC-seq data produces inaccurate coverage estimates and suboptimal peaks.26bioRxiv. Improved peak-calling with MACS2 Corrected versions and proper parameter settings have addressed this, and standardized pipelines now chain together alignment, duplicate removal, peak calling, and downstream motif analysis into reproducible workflows.27Briefings in Bioinformatics. ATAC-pipe: general analysis of genome-wide chromatin accessibility
Deep learning has begun to enter the picture as well. LanceOtron, a neural-network-based peak caller, outperformed MACS2 across all quality metrics in at least one head-to-head comparison on ATAC-seq data, suggesting that learned models may eventually replace the statistical assumptions baked into traditional tools.28Bioinformatics. LanceOtron: a deep learning peak caller for genome sequencing experiments
Adapting ATAC-seq for Plants
Plant cells pose a unique problem. In addition to mitochondria, they contain chloroplasts with their own genomes, and both organellar genomes soak up Tn5 activity. In Arabidopsis, more than half of ATAC-seq reads mapped to organellar DNA when the standard protocol was used. A workaround called FANS-ATAC-seq sorts intact nuclei away from organelle debris using fluorescence-activated sorting before adding Tn5. This brought organellar contamination down to about 30%, comparable to what was seen with older DNase-seq methods in the same species, and the protocol worked with as few as 500 sorted nuclei.29Oxford Academic (Nucleic Acids Research). Combining ATAC-seq with nuclei sorting for discovery of cis-regulatory regions in plant genomes
A hybrid method called TAC-C has pushed plant chromatin studies further by combining ATAC-seq chemistry with chromosome conformation capture, a technique that detects which distant parts of the genome physically touch each other in three-dimensional space. Applied to rice, sorghum, maize, and wheat, TAC-C captures interactions specifically among open chromatin regions, filtering out the noise from tightly packed, inactive parts of these large and repetitive genomes.30bioRxiv. TAC-C uncovers open chromatin interaction in crops and SPL-mediated photosynthesis regulation
Spatial ATAC-seq and Archived Tissue
Standard ATAC-seq, even in single-cell form, requires dissociating a tissue into individual cells or nuclei, which destroys spatial context. You learn what each cell type looks like epigenetically, but you lose where in the tissue each cell was sitting. Spatial-ATAC-seq aims to preserve both. One approach performs the Tn5 tagmentation reaction directly on a tissue section, then uses microfluidic barcoding to stamp each fragment with coordinates encoding its position, producing a genome-wide chromatin accessibility map that retains the tissue’s architecture.31PubMed Central. Spatial-ATAC-seq: spatially resolved chromatin accessibility profiling of tissues at genome scale and cellular level
An especially promising extension is spatial FFPE-ATAC-seq, designed for formalin-fixed, paraffin-embedded tissue, the standard way hospitals archive biopsy and surgical samples. Formalin fixation crosslinks proteins to DNA, which normally ruins chromatin accessibility assays. The new method overcomes these crosslinks, opening up decades’ worth of stored clinical material to chromatin analysis while preserving spatial resolution.32Nature Communications. Spatial profiling of chromatin accessibility in formalin-fixed paraffin-embedded tissues For cancer research in particular, where tumor heterogeneity within a single biopsy is clinically important, this could eventually allow pathologists to overlay regulatory information onto the same tissue sections they already examine under a microscope.
Tracing Regulatory Evolution Across Species
Because ATAC-seq maps regulatory elements genome-wide without requiring prior knowledge of what to look for, it is well suited for comparative work across species. In flatworms of the genus Schmidtea, ATAC-seq was used to compare chromatin accessibility across four species. Of roughly 55,000 accessible regions in one species, about 14% were highly conserved across the genus, 40% were partially conserved, and 46% showed no conservation, reflecting substantial regulatory divergence despite relatively recent common ancestry.33Nature Communications. A comparative analysis of planarian genomes reveals regulatory conservation in the face of rapid structural divergence
On a much deeper evolutionary timescale, ATAC-seq profiling of sea star and sea urchin embryos across multiple developmental stages revealed several thousand regulatory elements conserved for hundreds of millions of years. The pattern of regulatory element evolution in these echinoderms resembled what is seen in vertebrates, suggesting that the basic logic of how regulatory DNA changes over time may be an ancestral feature of animal genomes rather than a vertebrate innovation.34Nature Ecology & Evolution. Deep conservation of cis-regulatory elements and chromatin organization in echinoderms uncover ancestral regulatory features of animal genomes

