What Is a Random Barcode in Single-Cell Genomics?

A random barcode, in molecular biology, is a short stretch of synthetic DNA made up of nucleotides assembled in a random order so that each molecule in a mixture receives its own unique tag. These sequences typically range from about 10 to 30 nucleotides long, and because each position can be one of four bases, even a short barcode generates an enormous diversity of possible sequences. Researchers attach these tags to individual cells, DNA fragments, or RNA molecules, then read them back later with sequencing technology to figure out which molecule came from where, which cell divided into which descendants, or how many distinct molecules were in a sample to begin with.

How Random Barcodes Are Made

Most random DNA barcodes start life on an oligonucleotide synthesizer, the same machine that builds custom DNA sequences for labs. At the barcode positions, the synthesizer is fed a mixture of all four nucleotide building blocks instead of just one, so each molecule coming off the machine gets a different random sequence. The rest of the oligonucleotide is fixed, providing a known handle for later amplification and sequencing. The result is a library of molecules that all share the same flanking sequences but carry a unique random tag in the middle.

This sounds clean in theory, but the chemistry introduces real biases. The position of the random bases along the strand matters: one study using eight degenerate bases found that bases synthesized near the end of the strand showed more uneven representation than those placed earlier, with the end position having a higher Gini coefficient (a measure of inequality) than the front position.1ACS Synthetic Biology. Highly Accurate Sequence- and Position-Independent Error Profiling of DNA Synthesis and Sequencing The GC content of barcodes also creates problems. Sequences rich in G and C bases tend to amplify differently during PCR than those heavy in A and T, and this bias can vary unpredictably from one experiment to the next. Researchers have found that constraining the GC content of all barcodes to a narrow range, using carefully chosen “anchor” sequences between the random positions, helps level the playing field.2PubMed Central. Best Practices in Designing, Sequencing, and Identifying Random DNA Barcodes

When barcodes are used in codon-level libraries for protein studies, some amino acids are encoded by more codons than others, which skews representation. Certain codons rich in thymine or guanine consistently show higher barcode diversity, while codons for amino acids like proline, threonine, and alanine tend to be underrepresented.3PLOS Biology. A cost-effective and scalable barcoded library construction method for deep mutational scanning studies These biases are manageable, but ignoring them leads to misleading results downstream.

The Collision Problem

The whole system depends on every molecule getting a unique tag. When two different molecules happen to receive the same barcode, that is called a collision, and it is the central headache of random barcoding. Two molecules wearing identical labels become indistinguishable, which means you either miss one entirely or mistakenly merge data from two different sources.

Whether collisions matter depends on the math: how many unique barcodes are in your library versus how many molecules you are trying to tag. A well-designed barcode library with diversity exceeding 20 million unique sequences can label 100,000 cells with less than 0.3% collision probability.4Scientific Reports. Genomic barcoding for clonal diversity monitoring and control in cell-based complex antibody production But push the ratio too far and collisions spike. Computer simulations of barcoding in mutation-detection experiments have shown that collision rates can range from 0% to 100% depending on the number of sample molecules, the number of available barcodes, and the frequency of the mutation being sought. At very low mutation frequencies, collisions cause researchers to miss rare variants entirely and underestimate how common mutations really are.5Cancer Informatics. The Impact of Collisions on the Ability to Detect Rare Mutant Alleles Using Barcode-Type Next-Generation Sequencing Techniques

The practical lesson is that barcode diversity needs to be at least an order of magnitude larger than the number of molecules being tagged. Researchers who cut corners on library diversity often end up with data that looks fine on the surface but contains invisible errors.

Tracking Cells and Their Descendants

One of the most powerful uses of random barcodes is following individual cells as they grow, divide, and change over time. The idea is straightforward: deliver a unique barcode into each cell’s genome, let the cells do their thing, and then read back the barcodes to see which cells descended from which ancestor. This is called clonal tracking, and it has reshaped how scientists study blood formation, immune responses, and tissue development.

The barcodes are usually delivered by lentiviral vectors, engineered viruses that insert a small piece of DNA into a host cell’s genome. Each virus particle carries a different random barcode, so each infected cell and all its progeny share the same tag. A detailed protocol for creating viral barcode libraries, delivering them, recovering the sequences, and analyzing the data has been established for labs to follow.6PubMed Central. Clonal tracking using embedded viral barcoding and high-throughput sequencing – Section: Abstract An alternative approach uses targeted genome editing (zinc-finger nucleases) to insert barcodes at a specific location, but this method has been found to surprisingly reduce clonal diversity compared to the viral approach.7PubMed Central. Lentiviral and targeted cellular barcoding reveals ongoing clonal dynamics of cell lines in vitro and in vivo

This technology made it possible to track individual blood-forming stem cells as they differentiated into mature blood cells inside a living animal, work that was among the first to combine viral barcoding with high-throughput sequencing at the single-cell level.8PubMed Central. Tracking single hematopoietic stem cells in vivo using high-throughput sequencing in conjunction with viral genetic barcoding More recently, researchers tracked hundreds of thousands of blood stem cell clones in primates for over 10 years after transplantation, documenting stable, multi-lineage output from a highly diverse pool of stem cells.9PubMed Central. Long-term tracking of haematopoietic clonal dynamics and mutations in non-human primate undergoing transplantation of lentivirally barcoded haematopoietic stem and progenitor cells That kind of long-duration clonal data simply did not exist before barcoding.

Revealing Drug Resistance in Cancer

Random barcodes have given oncologists a clearer picture of how tumors resist treatment. By tagging individual cancer cells with unique barcodes before exposing them to a drug, researchers can ask a pointed question: did the surviving cells acquire resistance during treatment, or were resistant cells already hiding in the population beforehand?

The answer, at least for targeted therapies, turns out to be the latter. In experiments where barcoded cancer cell populations were treated with targeted drugs, the same barcode-tagged subclones expanded massively and reproducibly across independent replicates. Three biological replicates seeded per drug showed nearly identical compositions of selected barcodes, meaning those resistant subclones existed before the drug ever arrived and were simply selected by the treatment.10Cancer Research. Epigenetic Heritability of Cell Plasticity Drives Cancer Drug Resistance through a One-to-Many Genotype-to-Phenotype Paradigm Without barcoding, this kind of finding would be extremely difficult to demonstrate, because you would have no way to match post-treatment survivors to their pre-treatment identities.

Counting Molecules Without Double-Counting

A conceptually different application uses very short random barcodes, often called unique molecular identifiers (UMIs), to solve a basic problem in sequencing: PCR amplification makes copies of molecules, and you cannot tell from the sequencing data alone whether you are looking at ten copies of one original molecule or ten genuinely different originals. UMIs fix this by tagging each original molecule before amplification. After sequencing, you collapse all reads sharing the same UMI into a single count, giving you an absolute molecule count rather than an inflated, amplification-distorted number.11PubMed Central. Counting absolute numbers of molecules using unique molecular identifiers

UMIs have become standard in RNA sequencing, where accurate quantification of transcript abundance is the whole point. They also find use in detecting rare mutations in clinical samples, where a single mutant molecule among millions of normal ones needs to be counted reliably.

Single-Cell Genomics at Scale

Random barcodes are the engine behind the explosion in single-cell RNA sequencing over the past decade. The challenge is this: you have a tissue containing millions of cells, and you want to know which genes each individual cell is expressing. Physically isolating each cell and sequencing it separately would be impossibly slow and expensive. Instead, barcoding strategies label each cell’s RNA with a unique identifier so that after all the RNA is pooled and sequenced together, a computer can sort every transcript back to its cell of origin.

Two major approaches dominate. Droplet-based methods encapsulate single cells into tiny water-in-oil droplets alongside hydrogel beads coated with barcoded primers. Inside each droplet, a cell’s RNA is captured and tagged with the bead’s barcode during reverse transcription.12Nature Protocols. Single-cell barcoding and sequencing using droplet microfluidics Split-pool methods take a different route, passing cells through multiple rounds of random splitting and barcoding so that each cell accumulates a unique combination of barcode segments. This method, called SPLiT-seq, was used to profile the developing mouse brain and spinal cord without requiring any specialized microfluidic hardware.13PubMed Central. Single-cell profiling of the developing mouse brain and spinal cord with split-pool barcoding The split-pool concept has since been extended to bacteria, where the approach was adapted to handle the tougher cell walls of microbes.14PubMed Central. Microbial single-cell RNA sequencing by split-pool barcoding

CRISPR-Based Barcodes That Evolve Over Time

Static barcodes, whether delivered by virus or attached to beads, record a cell’s identity at one moment. They tell you what a cell’s ancestor was but not the branching history of decisions made along the way. A newer generation of barcodes solves this by using CRISPR-Cas9 to continuously edit a barcode sequence over time, so that the pattern of accumulated mutations encodes a record of the cell’s lineage history.15PubMed Central. A New Generation of Lineage Tracing Dynamically Records Cell Fate Choices

Think of it like a family tree written in DNA: every time a cell divides, CRISPR introduces small changes to the barcode, and daughter cells inherit those changes plus accumulate new ones. By comparing the barcode sequences of cells at the end of an experiment, researchers can reconstruct which cells are more closely related and build a developmental tree. This approach has been combined with single-cell RNA sequencing in zebrafish, allowing scientists to simultaneously read a cell’s lineage history and its current gene expression profile across organs including the heart, liver, pancreas, and brain.16PubMed Central. Simultaneous lineage tracing and cell-type identification using CRISPR–Cas9-induced genetic scars

Mapping Gene Activity Across a Tissue

Standard single-cell sequencing destroys the spatial arrangement of cells: you dissociate a tissue into individual cells and lose all information about which cell sat next to which. Spatial transcriptomics restores that information by using barcoded arrays, where each tiny spot on a slide carries a unique barcode sequence that tags the RNA from whatever tissue section sits on top of it. One approach deposits barcoded beads into densely packed microwells on a slide, and RNA from a tissue section placed over the array is captured and tagged by the nearest bead’s barcode.17Nature Methods. High-definition spatial transcriptomics for in situ tissue profiling Because each bead’s position is known and its barcode is unique, the resulting sequencing data can be mapped back to the tissue’s original geography.

Barcoding Synaptic Connections in the Brain

An especially creative application uses barcoded rabies viruses to map which neurons are connected to which. Rabies virus naturally jumps backward across synapses, from a postsynaptic neuron to its presynaptic partners. By engineering a library of rabies viruses, each carrying a unique transcribed barcode, researchers can reconstruct thousands of synaptic circuits in parallel. The method, called SBARRO, combines this monosynaptic tracing with single-cell RNA sequencing so that each connected neuron’s identity, gene expression, and place in the circuit are captured simultaneously.18Nature Communications. Ascertaining cells’ synaptic connections and RNA expression simultaneously with barcoded rabies virus libraries Tracking the path of clonal viral infections, each originating from a single postsynaptic starter cell, allows thousands of independent networks to be reconstructed from one experiment.

Fitness Profiling Across Thousands of Mutants

Random barcodes also underpin large-scale genetic screens in microorganisms. In a typical setup, every gene-deletion mutant in a yeast collection carries its own unique barcode. Researchers pool all the mutants together, grow them under some condition of interest, and then sequence the barcodes to see which mutants grew well and which disappeared. This “bar-seq” approach enables parallel measurement of thousands of mutants’ fitness in a single flask.19PubMed Central. Design and analysis of Bar-seq experiments It has been applied across both budding and fission yeast to profile the fitness consequences of gene deletions genome-wide.20PubMed Central. Global fitness profiling of fission yeast deletion strains by barcode sequencing

Taking this further, an ultra-high-resolution lineage tracking system in yeast can monitor the relative frequencies of roughly 500,000 lineages simultaneously, making it possible to observe the dynamics of evolution in real time: how often beneficial mutations arise, how they compete, and how quickly they spread.21PubMed Central. Quantitative evolutionary dynamics using high-resolution lineage tracking

When Barcodes Go Wrong

Even a perfectly designed barcode library runs into trouble during the experiment. The biggest culprit is PCR, the amplification step that makes enough copies of the barcoded DNA for sequencing. PCR can introduce two kinds of errors that corrupt barcode data. First, base substitutions during copying can change a barcode’s sequence, making one real barcode look like two different ones. Second, and more insidiously, template switching during PCR can create hybrid molecules where one half of the sequence comes from one original molecule and the other half from a different one. This produces chimeric reads that look like novel barcodes but are pure artifacts.22Nucleic Acids Research. Sources of PCR-induced distortions in high-throughput sequencing data sets – Section: Template switching

Emulsion PCR, where each molecule is amplified in its own tiny droplet, dramatically reduces recombination compared to standard PCR. Testing showed that under normal PCR conditions, recombination occurs to a significant degree regardless of cycle number or extension time, but emulsion PCR makes these events far less frequent and much less sensitive to reaction parameters.23Scientific Reports. A novel process of viral vector barcoding and library preparation enables high-diversity library generation and recombination-free paired-end sequencing

Biological artifacts also creep in. Even well-maintained cell lines kept under optimal conditions show ongoing clonal dynamics, meaning the barcode composition of a population keeps shifting over time as some clones outcompete others. Starting from a single clone reduces but does not eliminate this phenomenon.24PubMed Central. Lentiviral and targeted cellular barcoding reveals ongoing clonal dynamics of cell lines in vitro and in vivo Researchers need to account for this background drift when interpreting barcode data from cell-based experiments.

Computational Error Correction

Because sequencing and PCR errors are inevitable, computational methods for cleaning up barcode data are essential. The simplest approach compares observed barcodes to known reference sequences using edit-distance measures that count how many single-base changes separate two sequences. Traditional Hamming codes only handle substitutions, but DNA sequencing also produces insertions and deletions. Adapted Levenshtein-based error-correcting codes have been developed specifically for DNA, redefining sequence length whenever an insertion or deletion is detected. In simulations, these adapted codes outperform both standard Levenshtein and Hamming-based approaches when multiple errors are present.25PubMed Central. Levenshtein error-correcting barcodes for multiplexed DNA sequencing A separate effort produced indel-correcting DNA barcodes with user-friendly software for creating barcode libraries and decoding sequenced barcodes.26PubMed Central. Indel-correcting DNA barcodes for high-throughput sequencing

When no reference set of “true” barcodes exists, as is the case with fully random libraries, error correction becomes a clustering problem: group together all sequences that are probably noisy versions of the same original. Shepherd, a clustering method that indexes barcode sequences using short subsequences and applies a statistical test based on expected substitution error rates, was developed to distinguish genuine barcodes from sequencing errors in this setting.27Bioinformatics. Shepherd: accurate clustering for correcting DNA barcode errors

Random Barcodes Beyond DNA Sequencing

The concept of embedding random, readable codes into materials has extended well past the sequencing lab. Hydrogel beads fabricated in microfluidic systems can be loaded with unique combinations of molecular components, creating physical particles that each carry a distinct identity. These multiplexed beads are cross-linked by photopolymerization, with biomolecules bound directly into the hydrogel network during formation, and they can serve as identifiers in biomedical assays.28ACS Applied Materials & Interfaces. Particle ID: A Multiplexed Hydrogel Bead Platform for Biomedical Applications

DNA-based anti-counterfeiting has also attracted interest. Because synthetic DNA sequences can encode vast amounts of information in a tiny physical space and are difficult to replicate without specialized equipment, they have potential as molecular tags for verifying the authenticity of products. The advantages are clear: DNA is relatively easy to synthesize, stores huge amounts of data, and offers encryption possibilities. But real-world deployment faces unsolved challenges, including protecting DNA from harsh environments, developing rapid and cheap reading methods, and physically attaching tags to products in a durable way.29PubMed Central. Application of DNA sequences in anti-counterfeiting: Current progress and challenges Meanwhile, the proliferation of synthetic DNA sequences in open environments and commercial supply chains raises its own concerns around attribution, intellectual property, and biosecurity, since recipients may lack knowledge of a sequence’s origin over time.30Cell Press (Trends in Biotechnology). Cryptography and DNA: securing the future of the bioeconomy