How t-SNE Works for High-Dimensional Data Visualization

t-SNE, short for t-distributed Stochastic Neighbor Embedding, is a dimensionality reduction technique that converts high-dimensional data into two-dimensional scatter plots, making it possible to see patterns that would otherwise be invisible. Proposed by Laurens van der Maaten and Geoffrey Hinton in 2008, it quickly became the go-to visualization tool in fields like single-cell genomics, deep learning, and forensic science, largely because it excels at pulling apart clusters of similar data points and arranging them in visually intuitive layouts.1PubMed Central. Clustering with t-SNE, provably Its popularity also brought widespread misuse, and understanding what t-SNE actually shows you, and what it doesn’t, is just as important as knowing how to run it.

What t-SNE Actually Does

Imagine you have a spreadsheet where each row is a cell from a tissue sample and each column records the activity level of a different gene. That spreadsheet might have 20,000 columns. You cannot visualize a 20,000-dimensional space, so you need to compress it down to two dimensions that you can plot on a screen. t-SNE does this by focusing on local neighborhoods: it asks, for every data point, “which other points are closest to me in the original high-dimensional space?” and then tries to arrange all the points on a flat surface so that those neighborhood relationships are preserved as faithfully as possible.

The technique works in two stages. First, it measures the similarity between every pair of data points in the original space using a bell-shaped (Gaussian) curve, so nearby points get high similarity scores and distant points get scores near zero. Second, it places points randomly on a 2D canvas and iteratively adjusts their positions, trying to make the similarities on the canvas match the similarities in the original data. The key innovation is that the canvas-side similarities use a different, heavier-tailed distribution called the Student-t distribution rather than a Gaussian.2arXiv. Stochastic Neighbor Embedding with Gaussian and Student-t Distributions: Tutorial and Survey That heavier tail gives distant points more room to spread apart, which is why t-SNE plots tend to show well-separated clusters with clear gaps between them rather than one amorphous blob.

Why Single-Cell Biology Embraced It

The field where t-SNE arguably had its biggest impact is single-cell RNA sequencing (scRNA-seq). In these experiments, researchers measure gene expression in thousands or even millions of individual cells. The goal is often to identify distinct cell types or states within a tissue, and t-SNE’s tendency to pull apart clusters made it a natural fit. When applied to well-clustered high-dimensional data, t-SNE tends to produce a visualization with distinctly isolated clusters that often agree well with the clusters found by dedicated clustering algorithms.3Nature Communications. The art of using t-SNE for single-cell transcriptomics That visual clarity, combined with the lack of serious competitors until recently, made t-SNE the de facto standard for visual exploration of scRNA-seq data.

Before t-SNE, principal component analysis (PCA) was the main approach for visualizing scRNA-seq data. PCA is fast and deterministic, but it is a linear method, meaning it can only capture relationships that lie along straight axes. Biological data rarely cooperates with that assumption. t-SNE’s nonlinear approach lets it capture the curved, branching structure of cell populations far more effectively.4PubMed. Visualization of Single Cell RNA-Seq Data Using t-SNE in R Standard workflows in popular toolkits like Seurat now include t-SNE as a built-in step, making it accessible to biologists without much computational background.

The Perplexity Parameter

If you have ever run t-SNE, you’ve encountered the perplexity setting. Perplexity roughly controls how many neighbors each point considers when computing its similarity scores in the original space. A low perplexity (say, 5) means each point cares only about its closest handful of neighbors, which tends to produce many small, tight clusters. A high perplexity (say, 100 or more) means each point takes a wider view of the data, which can reveal larger-scale structure but may blur fine distinctions between small subgroups.

Most implementations default to a perplexity of 30, and the standard recommended range is roughly 10 to 100.5PubMed Central. The art of using t-SNE for single-cell transcriptomics That default works reasonably well for many datasets, but it is not universally appropriate. For very large datasets with millions of points, perplexity values in the standard range may not capture the global geometry well. For very small datasets, a perplexity of 30 might be close to the total number of data points, which distorts the results. A useful rule of thumb is that perplexity should be substantially smaller than the total number of data points but large enough to cover the typical size of the clusters you expect.

Some researchers have developed methods to sidestep this choice entirely. One approach uses a multi-scale scheme that combines information from many different perplexity values simultaneously, removing the need to pick a single setting. This multi-scale parametric version has been shown to produce reliable embeddings competitive with the best manually tuned perplexity values, while also allowing the model to generalize to new, unseen data points.6arXiv. Perplexity-free Parametric t-SNE Another approach sets perplexity locally for each data point based on the connectivity structure of the dataset itself, rather than using one global number.7Elsevier. Automating t-SNE parameterization with prototype-based learning of manifold connectivity

What t-SNE Plots Do Not Tell You

This is where a lot of people get into trouble. t-SNE is optimized to preserve local neighborhoods, and it does that well. But it makes no guarantees about global structure. The distances between clusters on a t-SNE plot are essentially meaningless. Two clusters that appear far apart may not actually be very different in the original data, and two clusters that appear close together may not be particularly similar. The algorithm arranges clusters to fill available space, not to reflect their true separations.

Misuses along these lines have become increasingly common in visual analytics. Practitioners frequently use t-SNE projections to investigate inter-cluster relationships, drawing conclusions about which groups of data points are most similar or different based on how close they appear in the plot, even though the projections often do not faithfully reflect the original distances between clusters.8IEEE Transactions on Visualization and Computer Graphics. Stop Misusing t-SNE and UMAP for Visual Analytics This is roughly equivalent to measuring the distance between cities on a subway map: the map is designed to show you which stops are connected, not how far apart they are geographically.

Another pitfall involves cluster sizes and densities. Standard t-SNE neglects the local density of data points in the original space, which can result in misleading visualizations where densely populated subsets of cells are given more visual space than their actual transcriptional diversity warrants.9Nature Biotechnology. Assessing single-cell transcriptomic variability through density-preserving data visualization In other words, a cluster that looks enormous on the plot might just contain a lot of very similar cells, while a small-looking cluster might contain cells with wildly diverse gene expression. Density-preserving variants of t-SNE have been developed specifically to address this problem, adjusting the embedding so that areas with high density in the original space also appear dense in the visualization.

A few concrete rules for reading t-SNE plots honestly:

  • Cluster existence: If points form a visually distinct group, there is probably a real cluster in the original data. t-SNE is quite good at this.
  • Cluster distance: The gap between two clusters says almost nothing about how different they are. Do not interpret it.
  • Cluster size: A big cluster is not necessarily more diverse than a small one. The visual area is not proportional to internal variability.
  • Cluster shape: Elongated or branching structures within a cluster can be meaningful, but they can also be artifacts of the random initialization or the perplexity setting. Run the algorithm multiple times and see which features persist.

Scaling t-SNE to Large Datasets

The original t-SNE algorithm computes pairwise similarities between every pair of data points, which means its computational cost grows with the square of the number of points. For a dataset of 10,000 cells, that’s 100 million pairwise comparisons, manageable on a laptop. For a million cells, it’s a trillion comparisons, far beyond what most machines can handle.

The first major speedup came from Barnes-Hut t-SNE, which borrows a technique from astrophysics originally designed for N-body simulations. Instead of computing exact forces between every pair of points, it groups distant points together and treats each group as a single entity. This reduces the computational cost from growing with the square of the number of points to growing at a rate proportional to N times the logarithm of N, a dramatic improvement that made datasets of tens of thousands of points practical.10arXiv. Barnes-Hut-SNE

For truly massive datasets like those in modern single-cell genomics, even Barnes-Hut is not always fast enough. FIt-SNE (Fast interpolation-based t-SNE) pushed the boundary further by using interpolation on a grid to approximate the expensive gradient calculations, dramatically accelerating t-SNE and removing the need for downsampling. This made it possible to visualize rare cell populations that would have been lost if the data had been subsampled before running the algorithm.11Nature Methods. Fast interpolation-based t-SNE for improved visualization of single-cell RNA-seq data Additional GPU-accelerated implementations have pushed the technique’s reach to millions of data points, using strategies like CUDA parallelization to exploit the massive parallelism of modern graphics cards.12Elsevier. Global and local structure preserving GPU t-SNE methods for large-scale applications

For ultra-large datasets where even these fast implementations strain, a practical pipeline recommended by researchers involves downsampling the data to a manageable size, running t-SNE on the subsample with settings chosen to preserve global geometry, positioning the remaining points on the plot using nearest neighbors, and then using the result as an initialization to run t-SNE on the whole dataset.13PubMed Central. The art of using t-SNE for single-cell transcriptomics This approach combines the global structure awareness of large perplexity values with the practical speed limits of current hardware.

How t-SNE Compares to PCA and UMAP

PCA is fast, linear, and fully deterministic: run it twice on the same data and you get the same result. It’s a good first pass for dimensionality reduction and is often used as a preprocessing step before running t-SNE. But PCA captures only the directions of greatest variance, which may not correspond to the groupings a biologist or data scientist actually cares about. In direct comparison, t-SNE has been shown to achieve better metrics than PCA for reducing the dimensionality of hyperspectral data, with downstream regression models performing better on t-SNE-reduced features than PCA-reduced ones.14Artificial Intelligence in Agriculture. t-SNE: A study on reducing the dimensionality of hyperspectral data for the regression problem of estimating oenological parameters Similar superiority in clustering quality has been observed in forensic analysis of ink spectral data.15PubMed. Dimensionality reduction and visualisation of hyperspectral ink data using t-SNE

UMAP (Uniform Manifold Approximation and Projection) arrived as the main challenger to t-SNE’s dominance around 2018. The two methods produce broadly similar cluster arrangements, and in side-by-side comparisons on multiplex immunofluorescence data, both provided comparable visualization capabilities. UMAP’s primary advantage is speed: it runs considerably faster, which matters when your datasets have hundreds of thousands of points.16bioRxiv. Comparison Between UMAP and t-SNE for Multiplex-Immunofluorescence Derived Single-Cell Data from Tissue Sections UMAP also tends to preserve more of the global structure of the data, meaning the relative positions of clusters are somewhat more meaningful than in t-SNE, though still not fully trustworthy as distance measures. On the other hand, UMAP produces characteristic branching patterns that some researchers find harder to interpret. Neither method is strictly better; the choice often comes down to dataset size and what structural features matter most for a particular analysis.

One underappreciated difference: t-SNE is a purely visualization-oriented tool with no natural way to embed new data points that were not part of the original run. If you get new cells after generating a t-SNE plot, you cannot simply add them; you have to rerun the entire computation. Parametric variants of t-SNE address this by training a neural network to learn the mapping, so new points can be projected without rerunning the algorithm.17arXiv. Perplexity-free Parametric t-SNE UMAP handles out-of-sample projection more naturally in its standard implementations, which is another practical reason for its growing adoption.

The Early Exaggeration Phase

One detail that matters for practitioners but rarely gets explained clearly is early exaggeration. During the first few hundred iterations of a t-SNE run, the algorithm artificially multiplies the similarities in the high-dimensional space, usually by a factor of 4 or 12. This forces the clusters apart early, before the algorithm settles into fine-tuning local structure. Without early exaggeration, clusters that should be well separated can remain tangled together, because the gradient that pushes distant points apart is too weak relative to the gradient that pulls neighbors together.

This phase has been rigorously analyzed and shown to be the mechanism by which t-SNE provably recovers well-separated clusters.18PubMed Central. Clustering with t-SNE, provably The practical takeaway is that if your t-SNE plot looks like one giant undifferentiated mass, the early exaggeration factor or the number of iterations during the exaggeration phase may need adjustment. Most modern implementations handle this automatically, but it’s useful to know the lever exists if the default settings produce unsatisfying results.

Uses Beyond Biology

Although single-cell genomics is t-SNE’s most visible application, the technique is widely used wherever high-dimensional data needs to be visually explored. In deep learning, t-SNE is commonly used to visualize the internal representations learned by neural networks. Researchers have used it to depict how convolutional neural networks organize histomorphologic information when classifying pathology images, revealing whether the network has learned categories that correspond to meaningful biological distinctions or is relying on artifacts.19PubMed Central. Visualizing histopathologic deep learning classification and anomaly detection using nonlinear feature space dimensionality reduction

In forensic science, t-SNE has been applied to hyperspectral imaging data from inks to distinguish between different writing instruments, a task that matters for document authentication and fraud detection. The technique’s ability to pull apart subtle spectral differences between inks that look identical to the human eye gives it a practical edge over linear methods in this setting.20PubMed. Dimensionality reduction and visualisation of hyperspectral ink data using t-SNE In agriculture, it has been used to reduce hyperspectral data from grapevines to estimate wine-relevant chemical parameters, outperforming PCA-based approaches for downstream regression models.21Artificial Intelligence in Agriculture. t-SNE: A study on reducing the dimensionality of hyperspectral data for the regression problem of estimating oenological parameters

Natural language processing, cybersecurity, and social network analysis have all seen t-SNE used to give researchers a visual handle on data that would otherwise exist only as abstract matrices of numbers. In each case, the same interpretive cautions apply: clusters are real, distances between them are not, and the plot should be a starting point for analysis rather than the conclusion.

When Not to Use t-SNE

t-SNE is a visualization technique, not a general-purpose analysis method. Using its output as input features for a machine learning model, for instance, is risky because the coordinates are influenced by random initialization and parameter choices. Two runs of t-SNE on the same data will produce different-looking plots, and the coordinates carry no consistent meaning across runs. Quantitative analysis should be done in the original high-dimensional space or in a PCA-reduced space where the transformations are well-defined and reproducible.

t-SNE also struggles when the data lacks clear cluster structure. If the underlying distribution is a smooth continuum rather than a set of discrete groups, t-SNE will still create apparent clusters in the visualization, because that is what its objective function rewards. This can lead researchers to see categories where none exist, a particularly dangerous failure mode when the goal is exploratory. If you suspect your data is continuous rather than clustered, methods that preserve density information or global distances are likely more appropriate for visual inspection.

Datasets with very few points, say under a hundred, also tend to produce unreliable t-SNE results. The algorithm needs enough data to form meaningful neighborhood estimates, and with too few points, the output is dominated by noise. Similarly, if your high-dimensional features are mostly noise rather than signal, t-SNE will happily visualize the noise and produce a plot that looks just as authoritative as one built from clean data. Garbage in, attractive garbage out.

Running t-SNE With Multiple Random Seeds

Because t-SNE uses stochastic gradient descent with random initialization, every run produces a slightly different layout. This is a feature, not a bug, but it means you should never draw strong conclusions from a single t-SNE plot. Running the algorithm three to five times with different random seeds and looking for features that persist across runs is the minimum standard for responsible use. If a cluster appears in one run but dissolves into a neighboring group in another, it is not a robust finding.

Some implementations now offer deterministic initialization strategies, such as initializing the 2D positions using the first two PCA components. This makes results more reproducible and often produces better-separated layouts than random initialization, because the algorithm starts from a globally sensible arrangement rather than having to discover it from scratch. If your toolkit offers PCA initialization as an option, it is generally worth using.