Topologically associating domains, usually called TADs, are stretches of DNA that fold together into distinct three-dimensional neighborhoods inside the nucleus. They were first identified in 2012 using a technique called Hi-C, which maps how often different parts of the genome physically touch each other in millions of cells at once. What stood out was a consistent pattern: sequences within a TAD contact each other far more frequently than they contact sequences in neighboring TADs, creating a kind of insulated compartment along the chromosome. Understanding these structures has reshaped how biologists think about gene regulation, because TADs help determine which switches in the genome can reach which genes.
How TADs Were Discovered
The genome does not float around the nucleus as a loose strand. It is tightly packaged and elaborately folded. But until the development of chromosome conformation capture methods, researchers had limited tools for seeing exactly how that folding worked across the whole genome. The original chromosome conformation capture technique, known as 3C, could measure how often two specific locations in the genome touched each other in three-dimensional space. Over time, the method was scaled up from measuring pairs of sites to mapping contacts across entire genomes. Hi-C, a genome-wide version, was the technology that first revealed TADs: megabase-scale blocks of self-interacting chromatin visible as prominent squares along the diagonal of a contact-frequency heatmap. These domains were discovered empirically from ensemble data, meaning the patterns emerged from averaging the interactions of millions of cells.
The Loop Extrusion Machine
TADs do not form randomly. Their boundaries are actively created by a molecular machine involving two key proteins: cohesin and CTCF. The dominant model for how this works is called loop extrusion. Cohesin is a ring-shaped protein complex that threads DNA through itself, progressively enlarging a loop of chromatin. As cohesin slides along the DNA, it pulls sequences that were far apart in the linear genome into close physical proximity. This continues until cohesin encounters a CTCF protein bound to the DNA, which acts as a stop sign.
What makes this system remarkably precise is that CTCF works in a directional way. The DNA-binding motif for CTCF has an orientation, and loop extrusion is only halted effectively when two CTCF sites face each other in a convergent arrangement. This explains several observations that puzzled researchers: why loops rarely overlap, and why CTCF motifs at loop anchors almost always point toward each other. A model in which CTCF simply acts as a generic roadblock, regardless of direction, could not account for those patterns.
Recent structural work has clarified the physical basis for this polarity. When cohesin approaches a CTCF molecule from the direction of CTCF’s N-terminal end, a short stretch of CTCF called the YDF motif engages with a cohesin subunit called STAG1 or STAG2 that sits at the leading edge of the translocating complex. That interaction jams the machine, arresting loop extrusion at that site. Approaching from the opposite direction, cohesin does not encounter the YDF motif and passes through without stopping.
What TAD Boundaries Look Like
TAD boundaries are not empty stretches of DNA. They are typically packed with active features. CTCF-binding sites concentrate heavily at boundaries, and these sites often sit near gene promoters. The boundaries also tend to be enriched for histone modifications associated with active gene expression, such as H3K4me3 and H3K27ac. This overlap between insulating elements and transcriptional activity is not coincidental: CTCF sites and active promoters frequently co-locate in the genome, and TAD boundaries often coincide with regions where transcription is happening.
Not all boundaries are equally strong. When researchers compare TAD boundaries across many different cell types, some boundaries show up consistently while others are unique to one cell type or another. Boundaries that are stable across cell types tend to sit in regions of the genome under stronger evolutionary constraint, with more conserved DNA sequence compared to boundaries found in only one or two cell types. Stable boundaries also overlap more with sequences flagged as evolutionarily constrained by comparative genomics tools, suggesting that mutations disrupting these boundaries are selected against.
What TADs Do for Gene Regulation
The genome is full of regulatory elements called enhancers that can boost gene expression, sometimes from great distances along the DNA. TADs act as organizing units that keep enhancers close to their appropriate target genes while insulating them from genes in adjacent domains. Within a TAD, an enhancer can reach its target promoter through the physical proximity created by chromatin folding. The TAD boundary, reinforced by CTCF and cohesin, limits how far that enhancer’s influence extends.
This insulation is not absolute, though. Research using the fruit fly gene twist as a test case has shown that enhancer-promoter interactions can form across TAD boundaries and even between different chromosomes. But these cross-boundary interactions produce weaker transcriptional activation than interactions within the same TAD. So TADs do not create impenetrable walls; they create a strong preference for local regulatory interactions, tuning the volume rather than flipping a binary switch.
When Boundaries Break and Genes Go Wrong
If TAD boundaries serve as insulation between regulatory neighborhoods, disrupting them should cause regulatory elements to reach the wrong genes. This is exactly what happens in several human diseases. Some of the clearest examples come from limb malformations. At a region of the genome containing the genes WNT6, IHH, EPHA4, and PAX3, different structural rearrangements produce different limb defects depending on exactly which TAD boundary is disrupted. Deletions, inversions, and duplications in this region shuffle the positions of a cluster of limb enhancers that normally regulate EPHA4. When a rearrangement moves those enhancers across a TAD boundary, they activate the wrong gene in developing limbs. Crucially, this “enhancer hijacking” only occurs when the rearrangement disrupts a CTCF-associated boundary. If the boundary remains intact, the enhancers stay in their proper regulatory domain and the limb develops normally.
These findings established a principle that researchers now call TADopathies: diseases caused not by mutations in genes themselves, but by structural variants that rearrange the three-dimensional regulatory landscape. The gene’s coding sequence is perfectly fine; the problem is that the wrong enhancer now has access to it.
TAD Disruption in Cancer
The same logic that explains limb malformations applies to cancer, but at a much broader scale. Tumor genomes are often riddled with structural rearrangements, and many of those rearrangements scramble TAD architecture. There are two general routes by which this drives oncogenesis. In one, a TAD boundary itself is deleted or mutated, fusing two adjacent domains into one and exposing genes on one side to regulatory elements from the other. In the other, a large-scale genomic rearrangement like a translocation or inversion breaks up existing TADs and creates new ones, placing oncogenes near powerful enhancers they would never normally encounter.
Boundary disruption does not always require a structural mutation. In gastrointestinal stromal tumors, a TAD boundary that normally insulates fibroblast growth factor oncogenes is silenced by DNA hypermethylation rather than deletion. The boundary’s CTCF sites become methylated, preventing CTCF from binding, and the insulation collapses. This epigenetic route to boundary loss has expanded the concept of TAD disruption beyond chromosomal rearrangements.
More recently, researchers have found that breakdown of the hierarchical organization within TADs can activate ancient viral sequences embedded in the genome. These sequences, called long terminal repeats or LTRs, are remnants of past retroviral infections and are normally kept silent. When the nested structure of sub-TADs within a larger domain is disrupted, some LTRs get co-opted as alternative promoters, driving inappropriate activation of oncogenes through a mechanism that does not require a boundary deletion at all.
TADs in Individual Cells Look Different Than Expected
The original discovery of TADs came from Hi-C, a technique that averages contact frequencies across millions of cells. A natural question followed: do individual cells actually have sharp TAD boundaries, or is the neat pattern an artifact of averaging many slightly different folding states?
Super-resolution imaging of chromatin in single cells has answered this definitively. TAD-like structures with globular conformations and sharp boundaries do exist in individual cells. But the exact positions of those boundaries vary from cell to cell. Boundaries occur with some probability at every position along the genome, but they are most likely to appear at CTCF- and cohesin-binding sites. So the Hi-C heatmap does not lie about where boundaries tend to form, but it overstates the crispness of the boundary. In any given cell at any given moment, the domain structure is real but somewhat fuzzy compared to the population average.
This cell-to-cell variability has implications for how we think about gene regulation. If TAD boundaries shift slightly from one cell to the next, then regulatory insulation is probabilistic rather than deterministic. An enhancer on the wrong side of a typical boundary might still reach its target in a fraction of cells where that boundary happens to be absent.
What Happens When You Remove CTCF
One of the most striking experiments in the field used an engineered degradation system to rapidly destroy all CTCF protein in mouse embryonic stem cells. The result was dramatic: within about 24 hours (roughly two cell divisions), Hi-C maps showed extensive new contacts across positions that had been TAD boundaries, visible as a widespread loss of insulation. Genome-wide quantification showed that more than 80% of boundaries lost their insulating capacity. The median TAD size in untreated cells was about 340 kilobases; after CTCF depletion, those organized domains largely dissolved.
Two details from these experiments are worth noting. First, the changes were fully reversible. When the drug causing CTCF degradation was washed away, TAD boundaries reformed. Second, removing CTCF disrupted local insulation between adjacent TADs but did not abolish the larger-scale separation of the genome into active and inactive compartments. This means TADs and compartments, while both features of three-dimensional genome organization, are maintained by at least partially independent mechanisms. CTCF drives the local loop-based insulation, while compartmentalization seems to arise from other forces, possibly related to the biochemical properties of active versus inactive chromatin.
TADs Through the Cell Cycle
When a cell divides, it has to pack its chromosomes into the compact mitotic form you might remember from textbook images of X-shaped chromosomes. This condensation was thought to completely erase all TAD structure, which would then need to be rebuilt from scratch after division. The reality is more nuanced. Tracking chromatin organization through the transition from mitosis back into the growth phase of the cell cycle, researchers identified over 8,000 contact domains that were progressively gained from pro-metaphase through mid-G1. Insulation at domain boundaries strengthened gradually over this period, suggesting TADs are rebuilt incrementally.
The surprise was that even in pro-metaphase, when chromosomes are at their most condensed, faint residual domain-like structures were still detectable. These were not distributed uniformly across the genome, arguing against the possibility that they were just contamination from interphase cells in the sample. So the mitotic chromosome does not completely forget its interphase folding. Some structural memory persists even through the most dramatic reorganization the genome undergoes, which may help cells re-establish proper gene regulation quickly after division.
Evolution of TADs Across Vertebrates
If TADs are truly important for keeping genes regulated correctly, you would expect evolution to preserve them. And it does. Comparisons across vertebrate genomes show that TAD structures are conserved well beyond what you would expect from neutral drift, with evidence pointing to stabilizing selection as a key driver. This conservation is especially strong around developmental genes, where misregulation would have severe consequences for body formation. But the selective pressure extends beyond development, suggesting TADs play important roles in genome function more broadly.
The evolutionary constraint on TAD boundaries themselves is also measurable at the DNA sequence level. Boundaries that remain stable across many cell types show stronger purifying selection, meaning mutations that disrupt them are weeded out more efficiently. Compared to cell-type-specific boundaries, the most stable boundaries overlap with about 500 additional base pairs of evolutionarily constrained sequence per 100-kilobase window. This is consistent with the disease data: if disrupting a stable boundary causes developmental malformations or cancer, natural selection will act to keep those boundaries intact.
TADs in Plants and Other Non-Mammalian Genomes
TADs are not a universal feature of all genomes, and the differences are informative. In mammals, flies, and bacteria, self-interacting chromatin domains are readily apparent in contact maps, though the proteins involved differ. Plants present a more complicated picture. In the model plant Arabidopsis thaliana, classic TADs are essentially absent from Hi-C maps. Instead, researchers find weaker “insulator-like” regions that are somewhat analogous to animal TAD boundaries but lack the sharp domain structure.
Other plants tell a different story. In rice, cotton, and Brassica species, TADs are clearly present and structurally prominent. The critical difference may relate to genome size and transposable element content, though the mechanism is not fully resolved. Notably, no plant genome studied to date contains a CTCF homologue. Whatever is maintaining TAD-like structures in rice and cotton, it is not the CTCF-cohesin loop extrusion system that dominates in mammals. This tells us that the three-dimensional folding pattern we call a TAD can arise through different molecular routes, and that CTCF-based loop extrusion is one solution rather than the only one.
High-resolution mapping technologies adapted for plants have begun to fill in the details. Micro-C-XL, which uses micrococcal nuclease instead of restriction enzymes to achieve nucleosome-level resolution, has revealed fine-scale chromatin organization and enhancer-promoter loops in Arabidopsis that were invisible to standard Hi-C. Even in a genome without obvious TADs, there is more three-dimensional structure than earlier low-resolution methods could detect.
Compartments Versus TADs
It is easy to conflate TADs with the larger-scale division of the genome into active (A) and inactive (B) compartments, since both are visible on Hi-C maps. But they are distinct phenomena maintained by different mechanisms. Compartments reflect the tendency of chromatin with similar biochemical properties to cluster together: active regions associate with other active regions, and inactive with inactive. This partitioning appears to be driven by phase-separation-like behavior and the intrinsic properties of chromatin rather than by loop extrusion.
The CTCF degradation experiments described earlier showed this separation cleanly. Removing CTCF destroyed local TAD insulation but left compartmentalization intact, and in some cases even strengthened it. Molecular dynamics simulations trained on high-resolution compartment data have shown that fine-scale compartmental patterns, alternating between A and B at the kilobase scale, can be accurately predicted by physical models based on chromatin type alone, without accounting for CTCF loops. The patterns generated by compartmentalization and loop extrusion overlap and interact, but they arise from fundamentally different forces.
Predicting TAD Boundaries from DNA Sequence
Because TAD boundaries are enriched for specific protein-binding motifs and other sequence features, researchers have asked whether boundaries can be predicted directly from the DNA sequence without any experimental contact data. Deep learning models trained on fruit fly genomes have shown that the answer is yes, at least to a useful degree. Models combining convolutional neural network layers with recurrent layers that capture long-range sequence dependencies can accurately distinguish TAD boundary sequences from internal TAD sequences.
The practical value of such models extends beyond academic curiosity. If you can predict where TAD boundaries should be from sequence alone, you can also predict which patient mutations are likely to disrupt a boundary and potentially cause disease. This is especially relevant for interpreting structural variants identified by clinical genome sequencing, where the pathogenic significance of a deletion or inversion is often unclear. A model that flags “this rearrangement removes a predicted TAD boundary near a developmental gene cluster” could help prioritize which variants to investigate further.
The Ongoing Debate About TAD Functionality
Despite the evidence described above, some researchers have pushed back on the idea that TADs are discrete functional units. The original observation was, after all, an empirical pattern extracted from noisy population-averaged data. The boundaries identified by different algorithms applied to the same Hi-C dataset do not always agree, and the choice of resolution and normalization method can shift boundary calls substantially. Some groups have argued that what we call TADs might be better described as a continuum of insulation strengths rather than a set of well-defined blocks with hard edges.
The single-cell imaging data sharpened this debate. TAD-like domains clearly exist in individual cells, but their boundaries are variable. If a “TAD boundary” is simply the position where insulation is most probable rather than a fixed structural feature, then the concept of a TAD as a defined unit becomes somewhat fuzzy. On the other hand, the disease evidence is hard to dismiss: disrupting the most conserved boundaries has clear and reproducible phenotypic consequences, from limb malformations to cancer. Whether you call the structures TADs, contact domains, or insulated neighborhoods, the functional relevance of their boundaries is well supported.
This tension between the messiness of the underlying biology and the usefulness of the TAD concept is unlikely to be resolved with a clean answer. TADs are a useful framework for understanding how genomes organize regulatory interactions, and the loop extrusion model provides a concrete mechanism. But as with most biological categories, the boundaries of the category itself are softer than the name implies.

