What Is 3D Gaussian Splatting and How Does It Work?

3D Gaussian Splatting is a technique for reconstructing and rendering three-dimensional scenes from ordinary photographs, and it does so fast enough to display those scenes in real time. Introduced in a 2023 paper that achieved at least 30 frames per second at 1080p resolution while matching or exceeding the visual quality of prior methods, it quickly became one of the most talked-about developments in computer graphics and 3D reconstruction.1ACM Transactions on Graphics. 3D Gaussian Splatting for Real-Time Radiance Field Rendering The approach sidesteps the heavy neural networks that earlier radiance-field methods relied on, replacing them with something more direct and surprisingly intuitive.

What Problem Does It Solve

Before Gaussian splatting, the dominant approach for creating photorealistic 3D scenes from photos was Neural Radiance Fields, commonly called NeRFs. NeRFs work by training a neural network to predict what color and density exist at every point in a volume of space. The results can look stunning, but there is a catch: rendering a single frame means querying that neural network millions of times, which is slow. Even with various speedup tricks, getting a NeRF to run at real-time frame rates on a full HD display meant accepting visible quality trade-offs.

Gaussian splatting attacks this bottleneck by ditching the neural network at render time entirely. Instead of storing a scene as a learned function inside a network, it stores the scene as a collection of millions of small, semi-transparent 3D blobs, each one a Gaussian-shaped ellipsoid with its own position, size, orientation, color, and opacity. Rendering becomes a matter of projecting those blobs onto the screen and blending them together, which is the kind of parallel work that graphics cards are built to do efficiently.

How the Scene Gets Built

The pipeline starts the same way many 3D reconstruction methods do: you take a set of photographs of a scene from different angles, and a standard photogrammetry step called Structure from Motion figures out where the camera was for each shot. That process also produces a sparse point cloud, a rough scattering of 3D points that mark features the algorithm could track across images.

Those sparse points become the seeds for the Gaussians. Each point is turned into a small 3D Gaussian, and then an optimization loop begins. The system renders the scene from the known camera positions, compares the rendered images to the actual photographs, and adjusts the Gaussians’ properties to reduce the difference. Over thousands of iterations, the Gaussians shift position, change shape, grow or shrink, and adjust their color and transparency until the rendered views closely match the input photos.

A critical part of this process is density control: deciding when to add new Gaussians and when to remove unhelpful ones. The original method uses an adaptive approach that clones Gaussians in under-reconstructed areas and splits overly large ones into smaller pieces. Follow-up research has noted that these clone and split operations are not always efficient, sometimes slowing optimization and making it harder to recover fine details.2arXiv.org. Efficient Density Control for 3D Gaussian Splatting Improving density control remains an active area of work, since the number and placement of Gaussians directly affects both quality and speed.

Why It Renders So Fast

The speed advantage comes from how the scene is drawn on screen. Each Gaussian is projected from 3D into a 2D “splat” on the image plane, and these splats are sorted by depth and blended front to back. This is a tile-based rasterization process, meaning the screen is divided into small tiles, each tile figures out which Gaussians overlap it, and a GPU thread processes that tile. There is no ray marching, no repeated neural network queries, and no volumetric sampling along each pixel’s line of sight.

The original implementation already hit real-time rates, but researchers continue to squeeze out more performance. One recent approach reorganizes how Gaussians are sorted within each tile, replacing a single long globally sorted list with shorter depth-local ranges processed front to back. On a high-end desktop GPU, this delivered a roughly 1.44 times speedup in the core rasterization step while producing output that is essentially pixel-identical to the baseline.3arXiv. TileGS: Tile-Local Depth Binning for Gaussian Splatting Rasterization These kinds of engineering refinements matter because they expand the range of hardware that can run Gaussian splatting smoothly, from workstation GPUs down to laptop-class chips.

Common Visual Artifacts

Gaussian splatting produces impressive results, but it is not artifact-free. Three recurring problems show up across many reconstructed scenes. The first is floaters: stray Gaussians that hang in empty space, appearing as ghostly blobs when the camera moves to a novel viewpoint. These typically form in regions seen by only a few input photos, where the optimization does not have enough information to place Gaussians correctly. The second is blur artifacts, where Gaussians become excessively large and diffuse, smearing detail across an area instead of capturing it crisply. The third is needle-like artifacts, extremely elongated Gaussians that look like thin spikes or streaks, especially along edges or in regions with complex geometry.4arXiv. Vis4GS: A Visual Analytic Tool for 3D Gaussian Splatting Reconstruction

These artifacts are not random glitches. They reflect fundamental limitations in how Gaussians approximate real surfaces. A Gaussian is a smooth blob, not a flat polygon, so it inherently struggles with sharp edges, thin structures, and surfaces seen at grazing angles. Much of the ongoing research in the field focuses on mitigating these problems, whether through better density control during optimization, post-processing cleanup, or alternative Gaussian formulations like 2D Gaussians that lie flat on surfaces rather than filling volumes.

Handling Moving Scenes

The original Gaussian splatting method assumes a static scene. Every photograph shows the same frozen moment, and the Gaussians represent geometry that does not change. Extending this to dynamic scenes, where objects move and deform over time, is one of the most active research frontiers.

One approach, known as 4D Gaussian Splatting, adds a time dimension. Rather than creating a separate static reconstruction for each video frame, 4D-GS learns a single representation that includes both the 3D Gaussians and a time-varying deformation model. A compact neural component predicts how each Gaussian should shift, rotate, or resize at any given timestamp, allowing the system to render novel views of the scene at arbitrary moments in time.5arXiv. 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering This keeps training and storage efficient compared to the brute-force alternative of rebuilding the entire scene frame by frame.

A harder variant of the problem involves large, fast-moving objects captured from a single monocular camera, like a phone video of someone dancing. Here, separating the object’s overall rigid motion from its surface deformation becomes important. Methods that decouple these two types of movement can maintain rendering speed while handling scenarios that trip up simpler approaches.6Proceedings of the AAAI Conference on Artificial Intelligence. Motion Decoupled 3D Gaussian Splatting for Dynamic Object Representation

Generating 3D From Text

Gaussian splatting has also found a role in generative AI pipelines that create 3D objects from text descriptions. Earlier text-to-3D methods typically used NeRF-based representations, inheriting all of NeRF’s rendering overhead. Switching the underlying 3D representation to Gaussians speeds up the generation loop considerably, since each candidate 3D object can be rendered quickly for evaluation during optimization.

One method called Gsgen uses a two-stage progressive optimization: first roughing out the geometry, then refining the appearance. The explicit nature of Gaussians, meaning each one has a concrete position and shape rather than being encoded implicitly inside a neural network, makes it straightforward to incorporate geometric priors that guide the optimization toward plausible shapes.7arXiv. Text-to-3D using Gaussian Splatting The results are not yet at the level of hand-crafted 3D models, but the field is moving fast, and Gaussian-based generation has become one of the standard approaches.

Extracting Meshes and Surfaces

One practical limitation of Gaussian splatting is that the output is not a traditional 3D mesh. Most downstream applications in games, film, engineering, and 3D printing expect triangle meshes with well-defined surfaces. A cloud of fuzzy ellipsoids does not slot neatly into those workflows.

Bridging this gap is an active area of research. One approach converts a trained Gaussian splatting scene into a dense point cloud by sampling points from each Gaussian, then applies standard surface reconstruction algorithms to produce a colored mesh, all without retraining the model.8arXiv. 3DGS-to-PC: Convert a 3D Gaussian Splatting Scene into a Dense Point Cloud or Mesh Another line of work tackles the problem from the other direction, using 2D Gaussians (flat discs rather than 3D ellipsoids) as a bridge between novel-view synthesis and learned geometric priors. These flat Gaussians naturally align with surfaces, making it easier to extract accurate geometry even from very sparse input views.9Proceedings of the AAAI Conference on Artificial Intelligence. MeshSplat: Generalizable Sparse-View Surface Reconstruction via Gaussian Splatting

The quality of extracted meshes still does not match dedicated multi-view stereo pipelines for precision applications, but for rapid prototyping, visualization, and content creation, the combination of fast Gaussian reconstruction followed by mesh conversion is becoming a viable workflow.

Photorealistic Avatars and Faces

Human faces and bodies are a natural testing ground for any 3D representation, since people are extremely sensitive to subtle errors in how skin, hair, and expressions look. Gaussian splatting has proven well-suited to avatar creation because each Gaussian can be “rigged” to a parametric face or body model, meaning the blobs move and deform according to a controllable skeleton underneath.

GaussianAvatars, for instance, attaches 3D Gaussians to a morphable face model, enabling photorealistic rendering of a person’s head with full control over expression, pose, and viewpoint. Expressions can be transferred from a driving video sequence, or adjusted manually by tweaking the underlying model parameters.10arXiv. GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians Taking this further, full talking-body avatars have been demonstrated that animate from speech audio combined with expression and body pose inputs, achieving frame rates above 150 fps on web-based rendering systems.11Computer Graphics Forum. THGS: Lifelike Talking Human Avatar Synthesis From Monocular Video Via 3D Gaussian Splatting

These avatar systems highlight one of Gaussian splatting’s underappreciated strengths: because the representation is explicit, each Gaussian can carry additional attributes beyond color and opacity. Attaching a Gaussian to a specific triangle on a face mesh, or tagging it with a semantic label, is conceptually straightforward in a way that would be awkward with a neural representation where everything is entangled inside network weights.

Running in a Web Browser

Getting Gaussian splatting to run outside of a desktop workstation environment is important for practical adoption. Several projects have ported the renderer to WebGPU, the next-generation graphics API available in modern browsers, aiming to let users view Gaussian-splatted scenes on phones, tablets, and laptops without installing any software.

This is harder than it sounds. WebGPU lacks certain low-level features that desktop GPU APIs provide, such as global atomic operations that the standard sorting step depends on. WebSplatter, one such browser-based renderer, works around this by replacing the global sort with a hierarchical approach and adding an opacity-aware culling stage that throws out nearly transparent Gaussians before they reach the rasterizer, reducing both overdraw and memory use. Across diverse hardware, it achieves speedups of roughly 1.2 to 4.5 times over previous web-based Gaussian splatting viewers.12arXiv. WebSplatter: Enabling Cross-Device Efficient Gaussian Splatting in Web Browsers via WebGPU

The web deployment angle matters because it lowers the barrier to entry for consumers. Real estate walkthroughs, product visualization, cultural heritage preservation, and casual social sharing of 3D captures all become more accessible when viewing requires nothing more than a browser tab. The gap between desktop and web performance is shrinking, though complex scenes with millions of Gaussians still push budget hardware to its limits.

Adding Meaning to Gaussians

A raw Gaussian splatting scene knows about geometry and color but has no concept of what anything is. A table, a chair, and the floor are all just blobs of different shapes and hues. Adding semantic understanding, the ability to label or query parts of the scene by meaning, opens up applications in robotics, augmented reality, and scene editing.

Semantic Gaussians is one approach that distills knowledge from 2D image-understanding models into the 3D Gaussian representation. Rather than training a new 3D model from scratch, it projects semantic features from pretrained image encoders onto existing Gaussians based on their spatial relationships, requiring no additional training.13arXiv. Semantic Gaussians: Open-Vocabulary Scene Understanding with 3D Gaussian Splatting The result is a scene where you can ask open-vocabulary questions, like “highlight all the books” or “select the lamp,” and get meaningful 3D responses.

This builds on the same explicit-representation advantage mentioned in the avatar context. Each Gaussian is a discrete object with properties you can read, write, and query. Adding a semantic feature vector to each Gaussian is just one more attribute alongside position, color, and opacity. In a neural representation, achieving the same thing typically requires retraining the network or running an expensive secondary model.

Scaling to Large Environments

Reconstructing a living room or a statue works well with a few million Gaussians. Scaling to large outdoor environments like city streets or driving corridors is more challenging. The sheer number of Gaussians required grows quickly, perspective distortion becomes more severe as camera distances vary widely, and lighting conditions may change across the scene.

One approach designed for autonomous driving scenarios uses a grouping strategy that partitions the scene to handle the perspective issues that arise in large-scale, complex environments. On a standard driving dataset, this method reached a PSNR of about 29 dB, indicating strong reconstruction fidelity for the task.14Proceedings of the AAAI Conference on Artificial Intelligence. EGSRAL: An Enhanced 3D Gaussian Splatting Based Renderer with Automated Labeling for Large-Scale Driving Scene Beyond driving, researchers are applying similar divide-and-conquer strategies to aerial mapping, urban reconstruction, and indoor environments that span many rooms.

Memory remains the primary bottleneck at scale. A single well-reconstructed room might use a few hundred megabytes of Gaussian data. A city block can balloon into gigabytes. Compression techniques, including vector quantization and pruning of redundant Gaussians, are being developed to bring storage requirements down without sacrificing visible quality, though the field has not yet converged on a standard compression format.

Where Gaussian Splatting Fits in the Broader Landscape

Gaussian splatting did not appear in a vacuum. It sits alongside NeRFs, traditional photogrammetry, LiDAR-based reconstruction, and conventional polygon-based 3D modeling. Each approach has trade-offs that make it better suited to certain tasks. Photogrammetry produces meshes directly but struggles with reflective and transparent surfaces. NeRFs handle view-dependent effects well but render slowly. LiDAR captures geometry precisely but lacks photorealistic appearance. Polygon modeling gives artists full control but requires significant manual effort.

Gaussian splatting’s niche is the intersection of photorealism, speed, and ease of capture. You point cameras at something, run an optimization for minutes to hours depending on scene complexity, and get a representation that renders in real time with quality that rivals or exceeds NeRFs. The trade-off is that the output is not a mesh (though conversion is improving), the memory footprint can be large, and certain scene types like transparent or highly reflective objects remain challenging.

The pace of research has been remarkable. In roughly two years since the original paper, the technique has been extended to dynamic scenes, avatar animation, text-to-3D generation, semantic scene understanding, web browsers, large-scale environments, and more. New papers appear almost daily on preprint servers. Whether Gaussian splatting becomes a lasting foundation for 3D graphics or gets supplanted by something newer, it has already shifted the field’s expectations for what real-time photorealistic rendering from photographs should look like.