How Nucleic Acid Bases Shape DNA Structure and Function

Nucleic acid bases are the molecular letters that encode genetic information in every living thing. DNA uses four: adenine (A), guanine (G), cytosine (C), and thymine (T). RNA swaps thymine for a close relative, uracil (U). These five molecules are remarkably small compared to the enormous strands they help build, yet they do far more than simply store a genetic code. They participate in energy transfer, help regulate which genes get turned on or off, form unexpected three-dimensional structures, and even show up on asteroids. The deeper you look at nucleic acid bases, the more they reveal about chemistry, evolution, and modern medicine.

How Bases Pair and What Actually Holds DNA Together

The standard picture of DNA is two strands zipped together by base pairs: A pairs with T through two hydrogen bonds, and G pairs with C through three. This selective pairing is what makes DNA replication possible, because each strand serves as a template for the other. But the question of what physically stabilizes the double helix turns out to be more contentious than textbooks let on.

One line of research concluded that stacking interactions between neighboring base pairs, not the hydrogen bonds between paired bases, are the dominant force keeping DNA intact. That work found that A-T pairing is actually slightly destabilizing, and G-C pairing contributes little to overall stability, with stacking doing the heavy lifting across all temperature and salt conditions tested.1Nucleic Acids Research. Base-stacking and base-pairing contributions into thermal stability of the DNA double helix A later reanalysis challenged that conclusion, arguing that the original model underestimated pairing contributions. Using a different framework for how double-stranded DNA forms, the reanalysis concluded that base pairing actually drives the assembly of the double helix, while stacking is already present in the single strands before they come together.2PubMed. Base-Pairing and Base-Stacking Contributions to Double-Stranded DNA Formation The practical takeaway is that both forces matter, and disentangling them depends heavily on the assumptions built into the analysis. For the reader, the useful point is that DNA stability is not just about hydrogen bonds between A-T and G-C. The way bases stack on top of each other like coins in a roll contributes substantially, and the field is still sorting out the precise balance.

Hoogsteen Pairing and the Flexibility of Adenine-Thymine Bonds

Watson-Crick pairing is the default geometry in DNA, but it is not the only one. In 1959, crystallographer Karst Hoogsteen discovered that adenine and thymine could pair using a different face of the adenine ring. In this arrangement, thymine’s hydrogen bond donor connects to a nitrogen on adenine’s major groove side rather than the position used in the classic Watson-Crick pair.3The Journal of Chemical Physics. Precision neutron diffraction structure determination of protein and nucleic acid components. XII. A study of hydrogen bonding in the purine-pyrimidine base pair 9-methyladenine · 1-methylthymine

For decades, researchers expected Hoogsteen pairs to be rare curiosities. Instead, they kept turning up: in AT-rich sequences, in DNA bound to certain antibiotics and proteins, in damaged DNA, and in polymerases that replicate DNA using Hoogsteen geometry. NMR studies eventually showed that base pairs in normal duplex DNA exist in a dynamic equilibrium, flipping back and forth between Watson-Crick and Hoogsteen forms. Hoogsteen pairs now appear to exist in meaningful abundance in genomic DNA, expanding the structural and functional versatility of the double helix well beyond what Watson-Crick pairing alone could achieve.4PubMed Central. A historical account of Hoogsteen base-pairs in duplex DNA

G-Quadruplexes at the Ends of Chromosomes

Guanine has a special talent that the other bases lack: four guanines can arrange themselves into a flat quartet, held together by a network of Hoogsteen-type hydrogen bonds. Stack several of these quartets on top of one another and you get a G-quadruplex, a bulky four-stranded structure that forms in guanine-rich regions of the genome. Human telomeres, the protective caps at the ends of chromosomes, are one of the most studied locations for these structures.

Telomeric G-quadruplexes appear to serve a dual role. On one hand, they can block the replication machinery, creating a barrier for replication forks and potentially causing telomere instability. They can also inhibit telomerase, the enzyme that maintains telomere length, leading to telomere shortening. On the other hand, G-quadruplex formation can promote the activity of protective proteins at chromosome ends, blocking competitors and enhancing the shielding effects of the proteins that cap telomeres.5Chem. Biological Functions and Therapeutic Potential of G-Quadruplexes The first direct evidence for G-quadruplexes at telomeres in a living organism came from studies using antibodies that specifically recognize these structures, which showed that G-quadruplexes form and are then resolved during DNA replication.6Nucleic Acids Research. G-quadruplexes and their regulatory roles in biology

G-quadruplexes also form in the promoter regions of certain cancer-related genes, and their structures show a striking polymorphism. In potassium solution, which mimics the cell interior, telomeric G-quadruplexes adopt at least two hybrid forms along with a two-layer intermediate structure. This shape-shifting quality has made them a target for drug design: small molecules that stabilize G-quadruplexes in oncogene promoters could, in principle, silence those genes and slow tumor growth.7PubMed Central. DNA G-Quadruplex in Human Telomeres and Oncogene Promoters: Structures, Functions, and Small Molecule Targeting

Tautomers, Damage, and Spontaneous Mutations

Each nucleic acid base can exist in more than one form. A proton can shift from one position on the ring to another, producing what chemists call a tautomer. The dominant form of each base is the one that participates in standard Watson-Crick pairing, but rare tautomeric forms pop up transiently. When they do so at exactly the wrong moment, during DNA replication, the result can be a mispair that looks nearly identical to a correct pair. Structural studies have shown that a single proton shift on a mismatched base can create a hydrogen-bonding pattern that is virtually indistinguishable from a canonical Watson-Crick pair, providing direct evidence for a long-hypothesized mechanism of spontaneous mutation.8Proceedings of the National Academy of Sciences. Structural evidence for the rare tautomer hypothesis of spontaneous mutagenesis

How often does this actually cause a permanent mutation? Computational work on the G-C tautomer pair suggests the answer is: almost never. The fraction of tautomeric G-C pairs that would persist through replication and become a permanent G-to-A mutation works out to less than one base pair per human genome replication, which is effectively negligible.9PubMed Central. The influence of base pair tautomerism on single point mutations in aqueous DNA Tautomeric shifts are fleeting, and proofreading enzymes catch most of the errors they introduce. Still, these rare events are one ingredient in the background mutation rate that fuels evolution.

Bases also suffer outright chemical damage. The most common oxidative lesion in DNA is 8-oxoguanine, a modified form of guanine produced by reactive oxygen species. It is dangerous because it readily pairs with adenine instead of cytosine, which would turn a G-C pair into a T-A pair after two rounds of replication. Cells rely on a dedicated repair enzyme to find and remove 8-oxoguanine, initiating a repair pathway that restores the correct base.10PubMed Central. Lost in the Crowd: How Does Human 8-Oxoguanine DNA Glycosylase 1 (OGG1) Find 8-Oxoguanine in the Genome?

Why DNA Uses Thymine Instead of Uracil

RNA uses uracil where DNA uses thymine, and the two molecules are nearly identical: thymine is just uracil with an extra methyl group. So why bother making thymine at all? The answer comes down to a maintenance problem. Cytosine spontaneously loses an amino group and turns into uracil at a low but steady rate. If DNA already contained uracil as a normal base, the repair system would have no way to tell a legitimate uracil from one that was really a damaged cytosine. By reserving uracil for RNA and using thymine in DNA, cells can treat every uracil found in DNA as an error and remove it. Two routes produce unwanted uracil in DNA: direct misincorporation in place of thymine, and spontaneous deamination of cytosine. The deamination pathway is especially dangerous because it creates a U-G mismatch that, if left unrepaired, becomes a permanent C-to-T mutation. Repair enzymes therefore also remove uracil from U-A pairs, even though those are not mutagenic, to keep the system clean.11PubMed Central. Keeping uracil out of DNA: physiological role, structure and catalytic mechanism of dUTPases

Chemical Marks That Change What Bases Mean

The genetic code is written in the sequence of bases, but cells add chemical tags to those bases that change how the code is read without altering the sequence itself. The best-known example in DNA is methylation of cytosine at its fifth carbon position, producing 5-methylcytosine. This modification is one of the central mechanisms of epigenetics: it helps silence genes, stabilize the genome, inactivate one of the two X chromosomes in female mammals, and enforce genomic imprinting, the process by which certain genes are expressed from only one parental copy.12PubMed Central. 5-methylcytosine turnover: Mechanisms and therapeutic implications in cancer Methylation also serves as a genome defense against transposons, the parasitic DNA sequences that can jump around and disrupt genes.13Proceedings of the National Academy of Sciences. DEMETER and REPRESSOR OF SILENCING 1 encode 5-methylcytosine DNA glycosylases

Removing methyl marks is just as important as placing them. In mammals, removal involves oxidizing 5-methylcytosine through a series of intermediates, eventually leading to its replacement with a plain cytosine through repair pathways. Researchers have recently fused a plant enzyme that directly excises methylated cytosine to a CRISPR-guided system, achieving targeted demethylation that can reactivate silenced genes without relying on DNA replication to dilute out the marks.14PubMed. DNA Methylation Editing by CRISPR-guided Excision of 5-Methylcytosine

RNA has its own suite of chemical modifications. The most abundant internal modification on messenger RNA is N6-methyladenosine, a methyl group added to adenine. This tag influences nearly every stage of an mRNA molecule’s life: how it is spliced, exported from the nucleus, translated into protein, and eventually destroyed. A reader protein recognizes the tag and shuttles the marked mRNA to degradation sites, controlling how long the message lasts in the cell.15PubMed Central. N6-methyladenosine-dependent regulation of messenger RNA stability The modification is reversible, making it a dynamic switch rather than a permanent mark.16PubMed Central. Dynamic regulation and functions of mRNA m6A modification

Where the Bases Came From

One of the more remarkable facts about nucleic acid bases is how easily they form under conditions that plausibly existed on the early Earth. Formamide, a simple nitrogen-containing molecule, yields all five standard nucleic acid bases when heated in the presence of common mineral catalysts.17BMC Evolutionary Biology. Formamide as the main building block in the origin of nucleic acids Another route starts with hydrogen cyanide, which can be generated in a planetary atmosphere by high-velocity impacts. Through a series of reactions involving dimerization, polymerization, and interaction with ultraviolet light and small molecules like urea, hydrogen cyanide chemistry can produce both purines and pyrimidines, including guanine.18Astrobiology. One-Pot Hydrogen Cyanide-Based Prebiotic Synthesis of Canonical Nucleobases and Glycine Initiated by High-Velocity Impacts on Early Earth

These bases are not confined to Earth. Samples returned from the carbonaceous asteroid Ryugu by Japan’s Hayabusa2 mission contained uracil at concentrations of roughly 11 to 32 parts per billion, depending on the sample.19Nature Communications. Uracil in the carbonaceous asteroid (162173) Ryugu The concentrations were lower than those found in certain meteorites, but the detection was significant because the Ryugu samples were collected in a way that avoided Earthly contamination. Finding a nucleobase on an asteroid strengthens the idea that some of the building blocks of life were delivered to Earth from space, supplementing whatever was being produced locally.

Bases as Drug Targets

Since the 1950s, chemists have exploited the resemblance between natural nucleic acid bases and synthetic look-alikes to create a broad class of drugs. Nucleoside analogs mimic the natural building blocks of DNA and RNA closely enough to be incorporated by viral or cellular enzymes, but they carry modifications that disrupt the replication process. These drugs are used against both viruses and cancers. Antiviral nucleoside analogs, including those deployed against HIV, hepatitis B and C, and SARS-CoV-2, work by getting inserted into the growing viral genome and then stalling or terminating the viral polymerase.20PubMed Central. Nucleotide and nucleoside-based drugs: past, present, and future Anticancer versions operate on a similar principle, interfering with DNA synthesis in rapidly dividing tumor cells. Many of these drugs require activation inside the cell through a series of phosphorylation steps before they become active.21PubMed. Nucleoside-based anticancer drugs: Mechanism of action and drug resistance

A different approach involves not mimicking bases but editing them directly in the genome. CRISPR base editors fuse a guide RNA system to an enzyme that chemically converts one base into another without cutting both strands of DNA. Two major classes exist: cytosine base editors, which convert C-G pairs into T-A pairs, and adenine base editors, which convert A-T pairs into G-C pairs.22PubMed Central. CRISPR-Cas9 DNA Base-Editing and Prime-Editing These tools install precise single-letter changes without making double-stranded breaks, which reduces unwanted insertions and deletions.23PubMed Central. Base editing: precision chemistry on the genome and transcriptome of living cells The adenine editor was a particular technical challenge because no natural enzyme deaminates adenine in DNA. Researchers solved this by evolving a transfer RNA deaminase to work on DNA, and subsequent rounds of evolution produced a version that catalyzes the reaction more than a thousand times faster than the original.24Science. DNA capture by a CRISPR-Cas9–guided adenine base editor

Expanding the Alphabet Beyond Four Letters

Nature settled on a four-letter genetic alphabet, but researchers have asked whether that number is a chemical inevitability or an evolutionary accident. Several groups have designed artificial base pairs that can function alongside the natural four in replication, transcription, and even translation. These unnatural base pairs expand the genetic code, in principle allowing the creation of proteins containing amino acids that do not exist in nature.25PubMed. Creation of unnatural base pairs for genetic alphabet expansion toward synthetic xenobiology One early optimized pair, d5SICS paired with dMMO2, was shown to be efficiently and selectively copied by a DNA polymerase within the context of natural DNA.26PubMed Central. Discovery, characterization, and optimization of an unnatural base pair for expansion of the genetic alphabet

Nature itself has experimented with alternatives. A group of bacterial viruses completely replace adenine in their DNA with diaminopurine, sometimes called base Z. Unlike adenine, which forms two hydrogen bonds with thymine, diaminopurine forms three. This substitution helps the phage evade bacterial restriction enzymes, which evolved to recognize and cut DNA containing the standard bases.27Science. A widespread pathway for substitution of adenine by diaminopurine in phage genomes The existence of a fully functional genome built on a non-standard base in the wild underscores that the four-letter code is not the only workable solution.

When a Broken Recycling Pathway Causes Disease

Cells constantly build and break down nucleotides, and a key part of this economy is the salvage pathway, which recycles free bases back into usable nucleotides rather than synthesizing them from scratch. When this recycling system fails, the consequences can be severe. A deficiency in the enzyme that recycles the purine bases hypoxanthine and guanine causes overproduction of uric acid, leading to kidney stones, kidney failure, and gout. In the most severe form, patients also develop neurological symptoms including self-injurious behavior. Milder variants of the disease correlate with some residual enzyme function, while the full syndrome is associated with complete loss of activity.28PubMed Central. Genotypic and phenotypic spectrum in attenuated variants of Lesch-Nyhan disease

Bases Outside of Biology

The self-assembling tendencies of nucleic acid bases have attracted attention from materials scientists. The same hydrogen-bonding patterns that hold DNA together can be harnessed outside of living systems to build gels, nanostructures, and functional materials. Nucleobase-containing molecules, from simple mononucleosides up to G-quadruplex-forming sequences, can self-assemble into hydrogels with potential applications in drug delivery and tissue engineering.29PubMed Central. Dancing with Nucleobases: Unveiling the Self-Assembly Properties of DNA and RNA Base-Containing Molecules for Gel Formation Watson-Crick pairing has also been used to control the structure of materials designed for energy transfer, charge transport, and even self-healing adhesives.30ChemistryOpen. Functional Systems Derived from Nucleobase Self‐assembly In these applications, the selectivity of base pairing, A with T and G with C, provides a built-in recognition system that is hard to replicate with other small molecules. The information content that makes DNA useful in biology turns out to be equally useful for programming the assembly of synthetic materials.