What Is an Allele? How Gene Variants Shape Traits

An allele is simply one version of a gene. Every gene in your body occupies a specific spot on a chromosome, and the particular version sitting at that spot is an allele. You carry two copies of most genes, one inherited from each parent, and those two copies can be identical or different. When they differ, the interplay between those alleles shapes everything from your blood type to your disease risk. The concept sounds straightforward, but alleles are the source of almost all the genetic variation that makes individuals distinct from one another.

Why You Have Two, and Why They Can Differ

Humans are diploid, meaning we carry two sets of chromosomes, one maternal and one paternal. At any given gene, the allele your mother passed down might be the same version your father passed down, or it might not. When both copies match, you are homozygous at that gene. When they differ, you are heterozygous. That heterozygous state is where most of the interesting biology happens, because the two alleles can interact in ways that determine what you actually look like, how your cells function, and how susceptible you are to disease.

The underlying sequence differences between alleles are often tiny. Comparing any two human genomes, roughly 99.9% of the DNA is identical; the variation that distinguishes one person’s alleles from another’s lives in that remaining fraction.1Nature / Journal of Human Genetics. SNP alleles in human disease and evolution The most common type of difference is a single-nucleotide change, where one DNA letter is swapped for another. These single-letter variants are scattered across the genome and are stable enough to be passed faithfully from parent to child. A single swap can be trivial, or it can change a protein’s shape enough to cause disease.

Dominant, Recessive, and Everything in Between

The terms “dominant” and “recessive” describe what happens when two different alleles meet in the same person. A dominant allele produces its effect even when only one copy is present. A recessive allele only shows up in your traits when both copies are the recessive version, because the dominant partner masks it. These labels apply specifically to heterozygous individuals; if you have two copies of the same allele, the question of dominance does not arise.2PubMed. Dominant versus recessive: molecular mechanisms in metabolic disease

Textbook examples make this look clean. Brown-eye alleles are broadly dominant over blue-eye alleles. The allele for Huntington’s disease is dominant: one copy is enough to cause the condition. But biology is rarely that tidy. Dominance is not a fixed property stamped onto an allele; it emerges from how the gene’s product interacts with everything else in the cell. Researchers have found that phenotypic dominance can arise at many different levels of biological organization, from the molecule to the tissue to the whole organism.3PubMed Central. The integrative biology of genetic dominance A single allele might behave as dominant for one measurable trait and recessive for another, depending on where you look.

There are also patterns that fall outside the simple dominant-recessive binary. In codominance, both alleles contribute visibly to the trait: the ABO blood group system is a classic case, where someone with one A allele and one B allele expresses both on their red blood cells, giving them type AB blood. In incomplete dominance, the heterozygote lands somewhere between the two homozygous states, like a red-flowered plant crossed with a white-flowered plant producing pink offspring. These patterns are common enough that treating dominant-versus-recessive as the default can be misleading.

Where New Alleles Come From

Every allele that exists today started as a mutation in an older allele. Mutations are copying errors: when a cell divides and replicates its DNA, the molecular machinery occasionally drops in the wrong nucleotide, skips one, or inserts an extra. Most of these errors get caught and repaired. The ones that slip through become new alleles. In bacteria whose repair systems have been disabled experimentally, researchers can watch these errors accumulate and see that specific types of base substitutions arise in predictable patterns depending on whether the error happened on the leading or lagging strand of replication.4PubMed Central. Detection of DNA replication errors and 8-oxo-dGTP-mediated mutations in E. coli by Duplex DNA Sequencing

Some stretches of DNA are especially mutation-prone. Repetitive sequences called microsatellites, where a short motif repeats over and over, are hot spots for slippage during copying. DNA polymerases produce errors within microsatellites at rates ten to a hundred times higher than in ordinary coding regions, and the kinds of errors vary depending on the motif being repeated.5PubMed Central. Every microsatellite is different: Intrinsic DNA features dictate mutagenesis of common microsatellites present in the human genome This is one reason microsatellites are so variable between individuals and are useful in forensic DNA profiling.

Recombination also shuffles alleles into new arrangements. During the formation of eggs and sperm, matching chromosomes physically cross over and exchange segments. This does not create brand-new sequence variants the way point mutations do, but it recombines existing alleles into novel combinations, effectively creating new multi-site allelic haplotypes that natural selection has never tested before.

When a Gene Has Dozens or Hundreds of Versions

The simple picture of a gene with two alleles, one dominant and one recessive, is a teaching simplification. Many genes have dozens or even thousands of alleles circulating in the global population. You still carry only two at a time, but the pool of possible versions is enormous.

The most extreme example in humans is the HLA system, a set of genes on chromosome 6 that encode proteins your immune system uses to distinguish your own cells from invaders. HLA is the most polymorphic genetic system known in humans.6PubMed Central. The HLA system: genetics, immunology, clinical testing, and clinical implications Population databases catalog HLA data from over 1,200 human populations worldwide, and the diversity is staggering.7PubMed. A snapshot of human leukocyte antigen (HLA) diversity using data from the Allele Frequency Net Database This extreme variation is not an accident. It is maintained by a form of natural selection called balancing selection: populations benefit from having many different HLA alleles because a wider range of immune-recognition molecules means the population as a whole can fend off a wider range of pathogens.

Interestingly, most HLA variation exists within populations rather than between them. When researchers divided human groups by continent and asked how much of the total HLA diversity separates one continental group from another, the answer was: very little. The overwhelming majority of allelic diversity is found among individuals within the same population.8PubMed Central. How HLA diversity is apportioned: influence of selection and relevance to transplantation This is consistent with a broader pattern across the genome and has practical implications for organ transplantation: finding a tissue-type match depends more on the luck of individual allele combinations than on the donor’s geographic ancestry.

Plants have their own spectacularly diverse multi-allelic systems. Many flowering species use self-incompatibility genes to avoid inbreeding. The pistil recognizes pollen that carries the same allelic type and rejects it, forcing cross-pollination. This mechanism generates extremely high allelic diversity at the self-incompatibility locus because any rare allele has an automatic advantage: pollen carrying a rare type is rejected by fewer pistils and therefore succeeds more often.9PubMed. Plant self-incompatibility in natural populations: a critical assessment of recent theoretical and empirical advances In some plant families, recombination within the gene itself contributes to generating new allelic variants.10Plant Physiology. Evidence That Intragenic Recombination Contributes to Allelic Diversity of the S-RNase Gene at the Self-Incompatibility (S) Locus in Petunia inflata

How Alleles Rise and Fall in Populations

An allele’s frequency in a population is not static. It drifts up or down over generations, pushed by two main forces: selection and chance. In a long-studied population of wild Soay sheep, researchers tracked allele frequency changes across generations using detailed pedigree data and found that the shifts were primarily driven by differences in survival and reproductive success among individuals, with migration playing a smaller role.11PubMed Central. Allele frequency dynamics in a pedigreed natural population Alleles carried by animals that lived longer and had more offspring became more common. That is selection in action, operating on a timescale researchers could observe directly.

Genetic drift, the random component, matters most when populations are small. A population bottleneck, even a brief one, can drastically reduce the number of alleles present at a gene while barely denting the overall heterozygosity.12Zoo Biology. Genetic drift and the loss of alleles versus heterozygosity The rare alleles are the first casualties: if only a handful of individuals survive a crash, any allele that happened to be carried by just one or two members is likely gone forever. This is why small, isolated populations, whether endangered species in the wild or captive breeding programs in zoos, tend to lose allelic diversity faster than they lose average genetic variation per individual.

Teasing apart selection from drift at the genomic level is difficult, but experiments in controlled populations suggest that at least 17 to 37% of allele frequency change over short timescales is driven by selection rather than random drift.13PubMed Central. Estimating the genome-wide contribution of selection to temporal allele frequency change The rest is noise. For most individual alleles, drift is the dominant force; selection mainly shapes the trajectory of alleles that happen to affect fitness-relevant traits.

The Sickle Cell Example and Balancing Selection

Perhaps the most famous allele in human biology is the sickle cell variant of the hemoglobin gene. People who inherit two copies of the sickle allele develop sickle cell disease, a serious condition in which red blood cells deform and clog small blood vessels. On the surface, you might expect natural selection to push this allele toward extinction. But in regions where malaria is endemic, people who carry one sickle allele and one normal allele have a survival advantage: their infected red blood cells tend to sickle preferentially, which flags them for removal by immune cells, reducing the parasite load without causing the full-blown disease.14PubMed Central. Sickle cell anaemia and malaria

This is balancing selection: the allele is harmful when homozygous but protective when heterozygous, so it persists in the population at a stable intermediate frequency. It is a vivid example of why labeling an allele “good” or “bad” in isolation makes little sense. The same variant can be a lifesaver or a liability depending on what the other copy says and what environment you are living in.

Alleles and Complex Traits

For traits like height, blood pressure, or susceptibility to type 2 diabetes, the simple one-gene, two-allele model breaks down completely. These traits are polygenic, meaning they are influenced by alleles at hundreds or thousands of genes simultaneously, each contributing a small nudge. Genome-wide association studies work by scanning across the entire genome to find alleles that are slightly more common in people with a given trait than in those without it. The picture that has emerged is that for most complex traits, each individual allele contributes so little that it is effectively invisible on its own.15PubMed Central. Evolutionary perspectives on polygenic selection, missing heritability, and GWAS

Estimates from yeast genetics, where researchers can systematically delete genes and measure the effects, suggest that somewhere between 1,200 and 1,900 genes may contribute to the heritability of a single quantitative trait. For human height, the best-characterized complex trait, hundreds of genomic regions have been identified, yet they collectively explain only a fraction of the known heritability.16bioRxiv. Small effect-size mutations cumulatively affect yeast quantitative traits The rest remains scattered among alleles with effects too small to detect individually. This “missing heritability” problem has been one of the central puzzles of modern genetics.

On top of the sheer number of alleles involved, their effects depend on context. Modifier genes can amplify or suppress the phenotypic consequences of other alleles, meaning two people who carry the exact same disease-associated variant might have completely different outcomes because their genetic backgrounds differ.17PubMed Central. From Peas to Disease: Modifier Genes, Network Resilience, and the Genetics of Health Some individuals carry clearly harmful alleles yet remain healthy, apparently protected by favorable combinations of modifiers elsewhere in the genome. Identifying these modifier effects has proven methodologically challenging, but they clearly exist and clearly matter.18Cold Spring Harbor Perspectives in Medicine. Genetic Modifiers and Oligogenic Inheritance

Alleles and Recessive Disease

Many serious genetic diseases are recessive, meaning you need two copies of the disease-causing allele to become ill. Carriers, who have one disease allele and one normal allele, are typically healthy and often have no idea they carry the variant. Cystic fibrosis is the most common severe autosomal recessive disease in populations of Northern European descent, with a carrier frequency of about 1 in 25 and a disease prevalence of roughly 1 in 2,500 to 3,500 live births.19PubMed. Population-based carrier screening for cystic fibrosis: a systematic review of 23 years of research

Those numbers shift dramatically across populations. In South African Black populations, a specific mutation in the cystic fibrosis gene has a carrier frequency that, when corrected for the fraction of disease-causing alleles it represents, suggests carrier rates could range from about 1 in 14 to 1 in 59.20Journal of Medical Genetics. Cystic fibrosis carrier frequencies in populations of African origin The wide range reflects uncertainty about how many different mutations contribute to cystic fibrosis in that population, a problem known as allelic heterogeneity. When a disease can be caused by many different alleles at the same gene, screening programs that test for only the most common mutations in one population can miss carriers from another.

Not all disease-associated alleles work through simple loss of function. Some produce a protein that actively interferes with the normal copy, called a dominant-negative effect. Structural studies comparing different mutation types have found that loss-of-function mutations tend to be far more disruptive to protein structure than dominant-negative or gain-of-function mutations, which are milder in their structural impact but can still cause disease because the defective protein poisons the normal one.21Nature Communications. Loss-of-function, gain-of-function and dominant-negative mutations have profoundly different effects on protein structure This distinction matters for treatment: a loss-of-function allele might be addressed by supplying the missing protein, while a dominant-negative allele requires silencing the bad copy.

Same Allele, Different Behavior

Your two alleles at any gene are not always treated equally by the cell, even when neither is mutated. Through a process called allele-specific DNA methylation, chemical tags are placed on one copy of a gene but not the other, influencing how much protein each allele produces. This is the mechanism behind genomic imprinting, where certain genes are expressed only from the copy inherited from one parent. The methylation patterns that distinguish maternal from paternal alleles are established during egg and sperm development and are faithfully maintained throughout life in most tissues.22PubMed Central. Genomic landscape of human allele-specific DNA methylation

Regions of allele-specific methylation identified across multiple human samples cluster tightly around known imprinted genes and precisely mark the boundaries of regulatory regions that control imprinting. This means two people can carry identical alleles at an imprinted gene yet have different expression levels depending on which parent provided which copy.23Nucleic Acids Research. ASMdb: a comprehensive database for allele-specific DNA methylation in diverse organisms Imprinting disorders like Prader-Willi syndrome and Angelman syndrome arise not from having a “bad” allele per se, but from inheriting the wrong parent’s copy or from errors in the methylation marks that distinguish the two alleles.

Editing One Allele and Leaving the Other Alone

One of the more exciting recent developments in genetics is the ability to target a single allele with gene-editing tools like CRISPR while leaving its partner intact. This matters enormously for dominant-negative diseases, where the harmful allele actively disrupts the protein made by the healthy copy. If you can selectively knock out the bad allele, the good one can function normally.

Researchers recently demonstrated this approach in a form of congenital muscular dystrophy caused by a dominant-negative variant in the COL6A1 gene. Using guide RNAs designed to match only the mutant allele’s sequence, they introduced disabling edits in patient-derived cells. The challenge was specificity: because the mutant and normal alleles differ by a single nucleotide, the initial guide RNAs showed some activity against the healthy copy too. By engineering deliberate additional mismatches into the guide RNA, the team developed a version that showed minimal activity at the normal allele while still efficiently editing the mutant one, and cultured cells treated this way showed improved collagen production.24PubMed Central. Allele-specific CRISPR/Cas9 editing inactivates a single nucleotide variant associated with collagen VI muscular dystrophy This is still lab-bench work, not a therapy you can receive, but it illustrates the principle: the goal is allele-level precision, disabling the troublemaker while leaving the functional copy alone.

Alleles Shared With Archaic Humans

Modern humans share alleles not only with one another but with our extinct relatives. When the Neanderthal and Denisovan genomes were sequenced and compared to ours, it became clear that a substantial amount of allelic variation in living people traces back to interbreeding events tens of thousands of years ago. An analysis building an ancestral recombination graph across all three lineages estimated that only about 1.5 to 7% of the modern human genome is uniquely ours, meaning the rest is shared, to varying degrees, with one or both archaic groups.25Science Advances. An ancestral recombination graph of human, Neanderthal, and Denisovan genomes The study also identified bursts of adaptive change specific to modern humans over the past 600,000 years, concentrated in genes related to brain development.

Some of the archaic alleles that persist in modern populations appear to have been useful. Variants inherited from Denisovans help Tibetan populations tolerate high altitude. Neanderthal-derived alleles in immune genes may have bolstered defenses against local pathogens when modern humans first arrived in Europe and Asia. Other archaic alleles have been linked to increased disease risk. The point is that the allelic pool in any living population is a palimpsest, layers of variants accumulated across hundreds of thousands of years of mutation, drift, migration, and selection, some of it from lineages that no longer exist as distinct species.

Why “The Gene For X” Is Almost Always Wrong

Popular media frequently describes alleles as though each one controls a single trait: a gene “for” intelligence, a gene “for” obesity, a gene “for” depression. The research paints a very different picture. For complex diseases, predictions about whether a given allele will even be detectable depend heavily on how many different disease-causing variants exist at the same gene, a property called allelic heterogeneity. When many different alleles at the same locus can each independently raise disease risk, standard mapping approaches perform poorly because no single allele stands out against the noise.26Oxford Academic. The allelic architecture of human disease genes: common disease–common variant… or not?

This is the reality behind headlines announcing the discovery of a gene “for” some trait. What has usually been found is one allele, at one gene, associated with a modest statistical increase in risk or a tiny nudge in a measurable trait, in one study population, with the caveat that the effect depends on genetic background and environment. The allele is real. Its effect is real. But calling it “the gene for X” strips away so much context that the statement becomes functionally misleading. A person’s phenotype is the cumulative product of thousands of alleles interacting with one another and with the world, not a readout of any single variant.