What Is the Proteome? How It Differs From the Genome

A proteome is the entire set of proteins produced by a cell, tissue, or organism at a given moment. While a genome is relatively fixed, the proteome shifts constantly in response to signals from the environment, the cell cycle, disease states, and aging. Understanding the proteome has become one of the central challenges in biology because proteins do the vast majority of functional work in living systems, and the gap between knowing which genes an organism carries and knowing what its proteins are actually doing turns out to be enormous.

Why the Proteome Is Far More Complex Than the Genome

Humans have roughly 20,000 protein-coding genes, but the number of distinct protein forms circulating in the body is vastly larger. A single gene can give rise to multiple protein variants through several mechanisms: the DNA itself may carry different versions of a gene across a population, the RNA transcript can be spliced in alternative ways, and the finished protein can be chemically modified after it is built. These modifications, which include the addition of sugar chains, phosphate groups, and lipid tags, change how a protein behaves, where it goes inside the cell, and which other molecules it interacts with.1Nature Methods. Proteoform: a single term describing protein complexity Researchers coined the term “proteoform” to describe each of these molecularly distinct versions. Current estimates of the total number of human proteoforms range into the millions, which is why cataloging even a single snapshot of the proteome is a much harder task than sequencing the genome.

Cataloging the Human Proteome

The Human Proteome Project, coordinated by the Human Proteome Organization, has been working for over a decade to confirm the existence of at least one protein product from every human protein-coding gene. Progress has been steady but slow. In its early years, the project had confirmed about 13,664 proteins at the highest level of evidence; by 2013-2014, that number climbed to roughly 15,646, with an estimated 3,844 “missing proteins” still undetected.2PubMed Central. Metrics for the Human Proteome Project 2013-2014 and strategies for finding missing proteins A decade later, the 2024 report shows that protein expression has been detected for 18,138 of the 19,411 recognized protein-coding genes, a rate of about 93%. That leaves 1,273 proteins still unaccounted for.3PubMed Central. The 2024 Report on the Human Proteome from the HUPO Human Proteome Project

Those remaining proteins are “missing” not because they do not exist but because they are extremely rare, appear only in unusual tissues or developmental stages, or are so unstable that they degrade before instruments can capture them. Tracking down the last few percent of the proteome has become a painstaking hunt, each new confirmed protein harder to find than the last.

How Scientists Read the Proteome

Mass spectrometry is the workhorse technology for proteomics. In broad terms, two strategies dominate. In the “bottom-up” approach, proteins are first chopped into small fragments called peptides, which are then fed into a mass spectrometer that measures their masses and uses that information to identify the parent proteins. In the “top-down” approach, intact proteins are introduced directly into the instrument and fragmented inside it, preserving information about modifications and variant forms that bottom-up methods can miss.4PubMed. Proteomics by FTICR mass spectrometry: top down and bottom up

Each strategy has trade-offs. Bottom-up methods are mature and high-throughput; they can identify thousands of proteins in a single experiment. But because they work with fragments, they sometimes struggle to tell apart closely related proteoforms or to cover certain regions of a protein. Membrane proteins, for instance, sit partly inside the oily cell membrane, and those embedded stretches tend to be poorly represented in bottom-up data. Top-down methods give better coverage of those buried regions, and recent advances have begun making top-down proteomics practical at larger scale.5PubMed Central. Integral membrane proteins: bottom-up, top-down and structural proteomics

Beyond mass spectrometry, affinity-based platforms have emerged as a complementary approach, especially for blood samples. These use engineered molecules, often short DNA-like strands called aptamers, that each latch onto a specific protein. One such platform can measure hundreds of proteins simultaneously with very low detection limits and broad dynamic range.6PubMed Central. Aptamer-Based Multiplexed Proteomic Technology for Biomarker Discovery Newer versions of these technologies now profile thousands of proteins from a single drop of blood, making large-scale population studies feasible for the first time.7PubMed Central. Emerging Affinity-Based Proteomic Technologies for Large-Scale Plasma Profiling in Cardiovascular Disease

Single-Cell Proteomics

Until recently, proteomics required pooling thousands or millions of cells together, which meant the data reflected an average across the whole sample. If a tumor contains five different cell types, each with its own protein profile, that detail gets blurred. Technological advances in instrument sensitivity have now pushed proteomics down to the single-cell level, allowing researchers to measure proteins in individual cells and map how protein landscapes differ among neighbors in a tissue.8PubMed. Mass-spectrometry-based proteomics: from single cells to clinical applications This is particularly valuable in cancer research, immunology, and developmental biology, where cell-to-cell variation often determines how a disease progresses or how a patient responds to treatment.

Where Proteins Live Inside Cells

Knowing that a protein exists is only part of the picture. Where it sits inside a cell matters enormously for what it does. A protein in the nucleus might regulate gene activity; the same protein found at the cell surface could be receiving signals from other cells. The Human Protein Atlas project built a comprehensive image-based map, the Cell Atlas, by combining microscopy and mass spectrometry to pin down the locations of more than 12,000 human proteins across 30 distinct structures inside the cell. That effort defined the protein contents of 13 major organelles and found that roughly half of all proteins show up in more than one compartment, sometimes with different abundance from one individual cell to the next.9PubMed. A subcellular map of the human proteome

Newer methods are pushing spatial resolution even further. One technique, called cross-link assisted spatial proteomics, can distinguish proteins in different sub-compartments of the same organelle. Applied to mitochondria, for example, it separates the proteins sitting in the outer membrane from those in the inner membrane or the interior space, and it can determine which direction a membrane protein faces. The approach has already led to corrections of earlier assumptions about where certain mitochondrial proteins reside.10Nature Communications. Cross-link assisted spatial proteomics to map sub-organelle proteomes and membrane protein topologies

The Dark Proteome

A substantial fraction of known proteins have no experimentally determined three-dimensional structure, and many resist the standard techniques used to solve structures. Researchers call this the “dark proteome.” A large-scale analysis across nearly a thousand species found that dark proteomes tend to be enriched in intrinsically disordered proteins, stretches of amino acids that do not fold into a stable shape but instead remain flexible. These disordered regions are not broken or defective; they often play critical roles in signaling and regulation, precisely because their flexibility lets them interact with many different partners.11PubMed. Taxonomic Landscape of the Dark Proteomes: Whole-Proteome Scale Interplay Between Structural Darkness, Intrinsic Disorder, and Crystallization Propensity The challenge is that traditional structure-determination methods depend on proteins holding still, and disordered proteins do not cooperate.

AI and Protein Structure Prediction

The release of AlphaFold2 in 2021 dramatically expanded the structural coverage of the proteome. The deep-learning system predicted three-dimensional structures for about 98.5% of human proteins, including many that had resisted experimental approaches for decades.12Nature. Highly accurate protein structure prediction for the human proteome Structure prediction alone does not replace experiments, and confidence varies across predictions, particularly for disordered regions. But having a predicted shape for nearly every human protein has accelerated research across the board, from understanding disease mechanisms to designing new drugs.

A more recent frontier is using AI to predict not just single protein structures but which proteins physically interact with each other. A 2025 study overcame a long-standing barrier in human interactome prediction by building much deeper evolutionary comparisons drawn from 30 petabytes of unassembled genomic data and training a new network on interactions from 200 million predicted structures.13PubMed Central. Predicting protein-protein interactions in the human proteome Mapping the full network of human protein interactions remains one of the field’s major goals, because diseases often stem from disrupted connections between proteins rather than from any single protein acting alone.14PubMed Central. Fundamentals of protein interaction network mapping

Proteomics in Cancer

Genomics transformed cancer diagnosis and treatment by identifying driver mutations, but knowing that a tumor carries a particular genetic change does not always predict how aggressive it will be or whether a targeted drug will work. Proteogenomics, which combines genomic, transcriptomic, and proteomic measurements from the same tumor samples, adds a layer of information that genetics alone misses. Protein-level data can reveal which mutated genes are actually producing active protein, which signaling pathways are switched on, and which therapeutic targets are accessible on the surface of cancer cells.15PubMed Central. Mass Spectrometry-Based Proteogenomics: New Therapeutic Opportunities for Precision Medicine Large-scale proteogenomic studies of breast, colon, lung, and ovarian cancers have already identified new tumor subtypes that are invisible to genomic analysis alone, opening the door to more finely tailored treatments.

Protein Misfolding and Neurodegenerative Disease

Alzheimer’s, Parkinson’s, and amyotrophic lateral sclerosis all share a common proteome-level problem: specific proteins aggregate into clumps that the cell cannot clear. In Alzheimer’s, amyloid-beta plaques and tau tangles accumulate in the brain. In Parkinson’s, alpha-synuclein forms aggregates called Lewy bodies. In ALS, the protein TDP-43 is frequently found in abnormal deposits. Many patients show aggregates of multiple proteins at once, pointing to a broader breakdown in the cell’s quality-control machinery rather than a failure involving just one molecule. That breakdown is influenced by the lipid environment of cell membranes, metal ion concentrations, post-translational modifications, and genetic mutations.16PubMed Central. Protein Turnover in Aging and Longevity – Section: 8 Protein Turnover in Aging Mammals Proteomic studies of brain tissue from patients with these diseases are helping researchers identify which proteins begin aggregating earliest and which cellular systems fail first, information that could guide the development of interventions aimed at the root of the problem rather than the downstream symptoms.

Proteome Changes with Aging

Even in the absence of disease, the proteome changes as an organism ages. One of the hallmarks of aging is a decline in proteostasis, the set of processes that fold, maintain, and recycle proteins. Studies tracking how fast proteins are replaced in aging mice have found that while the bulk rate of protein turnover across the whole body does not shift dramatically, individual proteins show significant changes, and the brain stands out as uniquely vulnerable. Brain protein turnover fluctuates across the aging continuum in ways not seen in heart or liver tissue, and these fluctuations differ between males and females.17PubMed Central. Derailed protein turnover in the aging mammalian brain The insoluble fraction of the brain proteome, proteins that have aggregated or misfolded, shows particularly hampered turnover in older animals, consistent with the idea that aging brains lose the ability to clear damaged proteins efficiently.

Finding Drug Targets Through Chemoproteomics

One of the most practical applications of proteomics is in drug discovery. Chemoproteomics uses chemical probes and mass spectrometry together to figure out exactly which proteins a drug binds in a living cell, and where on those proteins it attaches. This matters because many drugs have effects beyond their intended target, and understanding off-target binding helps explain side effects and can reveal unexpected therapeutic opportunities.

A method called LiP-Quant, for example, uses limited protein digestion combined with machine learning to identify drug targets across species, including in human cells, and approximate the binding site on each protein.18Nature Communications. A machine learning-based chemoproteomic approach to identify drug targets and binding sites in complex proteomes A newer technique called SEE-CITE uses photoreactive chemical handles that can be cleanly removed after the experiment, enabling precise identification of where a small molecule binds and allowing head-to-head comparisons of how different drug candidates engage the same site. When applied to fragments of FDA-approved drugs, it confirmed known binding sites and also uncovered previously unknown sites on proteins involved in cell function.19PubMed Central. Small-molecule binding-site discovery using silyl ether-enabled chemoproteomics

How Viruses Rewrite the Host Proteome

When a virus infects a cell, the battle plays out largely at the protein level. Viruses hijack host proteins to replicate themselves and deploy their own proteins to neutralize the cell’s defenses. Proteomic methods have made it possible to track these interactions over the course of an infection, revealing which host proteins a virus recruits, which it degrades, and how the timing of these events determines whether the cell mounts a successful immune response or gets overwhelmed.20PubMed Central. Proteomic approaches to uncovering virus-host protein interactions during the progression of viral infection

Adenoviruses provide a well-studied example. They encode proteins from a genomic region called E4 that specifically target host antiviral proteins for destruction or mislocation, effectively disarming the cell’s defenses. Proteome-wide surveys of adenovirus-infected cells have cataloged the full scope of this remodeling, identifying dozens of host proteins whose levels change during infection.21PubMed Central. Adenovirus Remodeling of the Host Proteome and Host Factors Associated with Viral Genomes This kind of mapping has direct practical value: every protein a virus depends on for replication is a potential drug target.

The Gut Microbiome’s Proteome

Metaproteomics extends the concept of proteome analysis beyond a single organism to entire microbial communities. In the human gut, trillions of bacteria, archaea, fungi, and viruses coexist, and their collective protein output influences digestion, immune regulation, and metabolism. Unlike DNA-based microbiome studies, which tell you what organisms are present and what genes they carry, metaproteomics reveals what those organisms are actually doing at any given moment.22PubMed Central. Microbial metaproteomics for characterizing the range of metabolic functions and activities of human gut microbiota

An added advantage is that metaproteomic samples from the gut contain both microbial and human proteins. The human fraction, roughly 15% of the total protein in these samples, provides a window into how the gut lining is responding to its microbial neighbors, making it possible to study host-microbe interactions in a single experiment.23Molecular & Cellular Proteomics. The Landscape and Perspectives of the Human Gut Metaproteomics – Section: The Advantage of Metaproteomics Compared to Other Omics

Proteomics in Agriculture and Ecology

Proteomics is not only a medical tool. In agriculture, understanding how crop plants reshape their proteomes under heat, drought, or salinity stress is informing breeding strategies for climate resilience. Multi-omics approaches that include proteomics alongside genomics and metabolomics give breeders a more complete picture of which biological pathways a plant activates when conditions turn harsh, and which protein changes actually correlate with survival.24PubMed Central. Increase Crop Resilience to Heat Stress Using Omic Strategies

In ecology and evolution, proteomic comparisons across species and populations have revealed something unexpected. When fruit fly populations were experimentally bred for resistance to six different stressors, their proteomes converged on a shared set of changes, a “common stress response” that overlapped heavily regardless of whether the stressor was starvation, desiccation, or cold. Each stress also triggered some unique protein changes, but the degree of overlap suggested that organisms may have evolved a general-purpose proteomic toolkit for dealing with environmental challenge.25Evolution. Evolutionary adaptation to environmental stressors: a common response at the proteomic level Similar comparative proteomics work in cyanobacteria found that even closely related species employ different protein strategies to tolerate salt stress, highlighting how proteomic diversity can exist beneath genetic similarity.26PubMed. Comparative proteomics unveils cross species variations in Anabaena under salt stress

The field has grown rapidly enough that entire national scientific communities have organized around it. The broader trajectory is clear: as instruments become more sensitive and computational tools more powerful, the proteome is shifting from a research curiosity to a routine layer of biological analysis, one that adds functional depth to the genetic blueprint in ways that are reshaping medicine, agriculture, and our understanding of how living systems actually work.