How Neural Radiance Fields Turn 2D Photos Into 3D Scenes

Neural radiance fields, usually called NeRFs, are a way to build photorealistic 3D scenes from ordinary photographs using a neural network. Introduced in 2020, the core idea is surprisingly elegant: a small network learns to predict the color and density of every point in a 3D space, and from that single learned representation you can render the scene from any viewpoint you like. The technique landed like a bomb in the computer-vision community and has since spread into medicine, autonomous driving, virtual reality, and even audio simulation. What makes NeRF remarkable is less the raw visual quality, which keeps improving, and more what it replaced: decades of hand-crafted pipelines for 3D reconstruction, condensed into a single trainable function.

What a Neural Radiance Field Actually Does

At its heart, a NeRF is a function that takes in a position in 3D space and a viewing direction, then spits out two things: how dense (opaque) that point is, and what color light it emits toward the camera. The original paper described this as a fully connected neural network whose input is a continuous five-dimensional coordinate and whose output is volume density and view-dependent radiance at that location.1Communications of the ACM. NeRF “View-dependent” is the key phrase. A glossy surface looks different depending on where you stand. By conditioning the color output on viewing direction, NeRF can reproduce specular highlights and subtle sheen that shift as you move around an object.

To turn that function into an image, NeRF borrows from an old technique called volume rendering. For each pixel in the desired image, a virtual ray is cast through the scene. The network is queried at many sample points along that ray, and the resulting colors and densities are combined into a single pixel value using a weighted sum, with denser points contributing more. This rendering step is fully differentiable, meaning the whole system can be trained end to end with gradient-based optimization.2arXiv. Differentiable Rendering with Reparameterized Volume Sampling You feed in photos of a scene taken from known camera positions, compare the rendered pixels to the real ones, compute the error, and nudge the network’s weights to close the gap. After enough iterations, the network has internalized a complete 3D representation of the scene.

Volume rendering also solves a tricky geometry problem. Traditional 3D rendering needs explicit surfaces, triangles and meshes, and figuring out exactly where a ray hits a surface is hard to make differentiable. NeRF sidesteps this by treating the scene as a continuous cloud of density, leveraging a probabilistic notion of visibility rather than hard surface intersections.3arXiv. Volume Rendering Digest (for NeRF) The result is smoother gradients and more stable training.

The Speed Problem and How It Was Solved

The original NeRF was beautiful but slow. Training a single scene could take a day or more on a high-end GPU, and rendering a single frame took tens of seconds. Every pixel required hundreds of network evaluations along its ray, and the network, while small by deep-learning standards, was not small enough for real-time use. For researchers publishing papers, this was tolerable. For anyone wanting to use NeRF in a product, it was a dealbreaker.

The breakthrough came with a technique called Instant NGP (Neural Graphics Primitives), which replaced much of the neural network’s heavy lifting with a multiresolution hash table of trainable feature vectors. Instead of asking a large network to memorize every detail of the scene, the system stores learned features in a spatial lookup table and uses only a tiny network to decode them. This achieved a speedup of several orders of magnitude: training a high-quality scene dropped from hours to seconds, and rendering hit tens of milliseconds at full HD resolution.4ACM Transactions on Graphics. Instant neural graphics primitives with a multiresolution hash encoding The approach was trivial to parallelize on modern GPUs, and it reshaped expectations about what NeRF-based systems could do in practice.

Other researchers attacked the speed issue from the encoding side. Standard neural networks are sluggish at learning high-frequency details, the fine textures and sharp edges that make a scene look real. One approach, Quantized Fourier Features, encodes input coordinates in frequency-domain bins rather than spatial bins, which can yield smaller models, faster training, and better quality simultaneously.5arXiv. QFF: Quantized Fourier Features for Neural Field Representations These encoding tricks matter because they let you push quality up without pushing training time back into impractical territory.

Moving Beyond Static, Controlled Scenes

The original NeRF assumed a perfectly static scene photographed under consistent lighting. Point a phone camera at a coffee mug on your desk from 50 angles, hold the lighting steady, and NeRF will reconstruct the mug flawlessly. But the real world is not a controlled photo studio. People walk through the frame, the sun moves, and shadows shift between shots.

NeRF in the Wild tackled this head-on. It extended the framework to handle unstructured photo collections grabbed from the internet, photos of landmarks taken by tourists at different times of day, in different seasons, with pedestrians and vehicles cluttering the foreground. The method models variable illumination and transient occluders as part of the representation, allowing accurate reconstructions from messy, uncontrolled image sets.6arXiv. NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections If you have ever searched for photos of a famous building and found thousands of images taken under wildly different conditions, this is the kind of input the system was designed to digest.

Dynamic scenes present a different challenge. If the subject itself is moving, a single static NeRF cannot represent it. Non-Rigid NeRF addresses this by separating a dynamic scene into a canonical volume, the “default” shape, and a learned deformation field that bends rays to account for motion. This lets you reconstruct and re-render non-rigid objects like a person gesturing or fabric blowing in the wind, all from a single monocular video.7arXiv. Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synthesis of a Dynamic Scene From Monocular Video

Learning From a Handful of Photos, or Even One

A standard NeRF is trained from scratch for every new scene, and it needs dozens to hundreds of images to do so. That makes sense when you are digitizing a museum artifact with a turntable and a DSLR, but it falls apart when you have two snapshots of a room you want to model. The pixelNeRF framework addressed this by training a network across many scenes so it could learn a general prior about what 3D scenes look like. Given just one or a few images of a new scene, it predicts a neural radiance field in a single forward pass, no per-scene optimization required.8arXiv. pixelNeRF: Neural Radiance Fields from One or Few Images The quality is not as high as a scene-specific NeRF trained on a hundred images, but for many practical applications, speed and convenience matter more than last-mile perfection.

Text-to-3D generation pushed this idea even further. DreamFusion showed that you do not need any photos of a scene at all. Instead, it uses a pretrained 2D text-to-image diffusion model to guide the optimization of a randomly initialized NeRF. You type a text prompt, and the system optimizes a 3D model so that its renderings from random angles match what the diffusion model thinks that prompt should look like.9arXiv. DreamFusion: Text-to-3D using 2D Diffusion Early results tended toward over-saturated, blobby shapes. Follow-up work like ProlificDreamer improved fidelity and diversity by replacing the fixed optimization target with a variational framework, producing sharper geometry and more realistic textures.10NeurIPS Proceedings. ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation

Medical and Surgical Uses

NeRF’s ability to conjure a 3D model from limited views has obvious appeal in medicine, where you often cannot take as many images as you would like. One area seeing early results is endoscopic sinus surgery. Researchers applied NeRF to monocular endoscopic video, the kind of feed a surgeon already sees through the scope, and reconstructed the 3D surgical field in high fidelity. In cadaveric specimens, the reconstructions achieved submillimeter accuracy comparable to postoperative CT data, with mean errors well under a millimeter for key anatomical measurements.11PubMed Central. Neural Radiance Fields for 3D Reconstruction of Monocular Endoscopic Video in Sinus Surgery The practical promise is real-time spatial awareness during surgery without needing additional imaging hardware.

CT reconstruction is another frontier, though a harder one. Standard NeRF was built around visible-light photography, and X-ray imaging follows fundamentally different physics. Photons in visible light reflect off surfaces; X-rays pass through tissue and attenuate along the way. Directly applying NeRF to medical image reconstruction has had limited success precisely because of these differences in how photons travel and what prior information is available.12PubMed Central. TomoGRAF: An X-ray physics-driven generative radiance field framework for extremely sparse view CT reconstruction Researchers are addressing this by building X-ray-specific physics into the rendering model, but the work is still early. The potential payoff, reconstructing useful CT volumes from far fewer X-ray projections than a conventional scan requires, could reduce radiation exposure for patients.

Autonomous Driving and Large-Scale Environments

Self-driving car companies need massive amounts of training data showing realistic street scenes from every angle, under every lighting condition. Building that data with real cameras is expensive and limited; building it in a traditional simulator often looks fake enough to hurt model performance. NeRF-based simulation systems like S-NeRF++ train on real driving datasets and can then generate large numbers of realistic street scenes with flexible manipulation of camera angles, vehicles, and lighting.13PubMed. S-NeRF++: Autonomous Driving Simulation via Neural Reconstruction and Generation The system ingests noisy LiDAR data alongside camera images to improve depth accuracy, and it handles both the static background and moving foreground objects.

Scaling NeRF to large outdoor environments, entire neighborhoods or cityscapes rather than single objects, creates its own headaches. A single network struggles to hold the detail of a large scene. SCALAR-NeRF addresses this by training a coarse global model first and then partitioning the scene into smaller blocks, each handled by a dedicated local model. A shared decoder keeps the blocks consistent with each other, and the local outputs are fused for the final reconstruction.14arXiv. SCALAR-NeRF: SCAlable LARge-scale Neural Radiance Fields for Scene Reconstruction Aerial-NeRF takes a related but distinct approach for drone-captured footage, adapting spatial partitioning to different flight trajectories and using pose similarity rather than a separate network to decide which scene partition a new viewpoint falls into.15arXiv. Aerial-NeRF: Adaptive Spatial Partitioning and Sampling for Large-Scale Aerial Rendering

Where NeRF Still Struggles

Transparent and highly reflective objects remain a pain point. Standard NeRF assumes light travels in straight lines from the scene to the camera, which breaks down when a glass vase refracts light or a chrome faucet bounces it in unexpected directions. The rendered images of such objects tend to show ghosting, blurriness, or outright geometric errors. Neural Refractive-Reflective Fields attempt to fix this by modeling refraction and reflection explicitly, reconstructing the geometry of non-Lambertian objects with marching tetrahedra and then simulating the bent light paths using Fresnel terms.16arXiv. NeRRF: 3D Reconstruction and View Synthesis for Transparent and Specular Objects with Neural Refractive-Reflective Fields The results are promising, but the extra complexity and computation are significant, and fully general handling of mixed transparent-opaque scenes is still an open problem.

Relighting is a related challenge. A vanilla NeRF bakes the lighting conditions of the training photos into the representation. If you trained on photos taken at noon, you cannot re-render the scene under sunset light without additional work. Inverse rendering approaches like PBR-NeRF try to disentangle the scene into geometry, materials, and illumination so each can be edited independently.17arXiv. PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields Getting this decomposition right remains difficult because the same rendered pixel can be explained by many different combinations of material reflectance and light color, a classic ambiguity in computer vision.

Shrinking NeRFs for Deployment

A trained NeRF is a neural network, which means its “scene file” is a pile of floating-point weights. For a single scene, the original NeRF model is not enormous by modern standards, but it is still large enough to be impractical for streaming or mobile deployment, especially when you want to serve thousands of scenes. Research into compression has shown that combining low-rank optimization, network distillation, and weight quantization can reduce the model size to roughly 3% of the original, a dramatic shrink that makes NeRF-based light field compression viable for storage and transmission.18arXiv. Distilled Low Rank Neural Radiance Field with Quantization for Light Field Compression Compression at this level matters not just for bandwidth but for running NeRF on devices with limited memory, like AR headsets or smartphones.

Full-Body Avatars and Telepresence

NeRF’s photorealism has made it a natural fit for human avatars. AvatarReX builds NeRF-based full-body avatars from video data, with expressive control over body, hands, and face simultaneously, and achieves real-time animation and rendering. The approach models these body parts separately so that prior knowledge from standard body mesh templates can guide the reconstruction without locking the representation into rigid templates.19ACM Transactions on Graphics. AvatarReX: Real-time Expressive Full-body Avatars The practical implication for telepresence is a version of video calling where the remote person exists as a photorealistic 3D figure you can view from any angle, not a flat rectangle on a screen. The gap between “research demo” and “product you can use on a consumer headset” is still wide, mostly due to the capture setup needed, but it is narrowing fast.

Beyond Light: Audio-Visual Neural Fields

Sound behaves differently depending on the geometry and materials of a room. A hard-walled bathroom produces a very different echo from a carpeted living room. Researchers have extended the NeRF concept to model not just how a scene looks but how it sounds. One approach integrates audio propagation priors into NeRF, implicitly associating audio generation with the 3D geometry and material properties of the visual environment.20NeurIPS. Neural Radiance Fields for Real-World Audio-Visual Scene Synthesis A complementary method called Acoustic Volume Rendering adapts the volume rendering pipeline itself to model acoustic impulse responses, constructing an “impulse response field” that encodes wave propagation principles and can synthesize what a sound would sound like from a novel listener position in the scene.21NeurIPS Proceedings. Acoustic Volume Rendering The potential applications range from VR experiences with spatially accurate audio to architectural acoustics simulation.

NeRF and 3D Gaussian Splatting

If you follow 3D vision research, you have probably seen 3D Gaussian Splatting presented as NeRF’s successor. Gaussian Splatting represents a scene not as a continuous neural field but as a collection of small 3D Gaussian blobs, each with its own position, shape, color, and opacity. The rendering is done by “splatting” these blobs onto the image plane, which is extremely fast on a GPU and can easily reach real-time frame rates. A comparative review notes that NeRF tends to excel at high-quality rendering of fine geometric details, while Gaussian Splatting’s strength is high-speed real-time rendering.22Science and Technology of Engineering, Chemistry and Environmental Protection. A Review of Neural Radiance Fields and 3D Gaussian Splatting for 3D Reconstruction

In practice the two approaches are converging rather than one replacing the other. Many recent systems borrow ideas from both: Gaussian-based representations that are initialized or refined with NeRF-style optimization, or NeRF architectures that bake their output into splat-friendly formats for real-time viewing. The choice between them often comes down to the application. If you need to render a scene live on a headset, Gaussian Splatting’s speed is hard to beat. If you need the highest-fidelity reconstruction of a complex surface and can tolerate offline rendering, NeRF-style volume rendering still has the edge. And for tasks like inverse rendering, relighting, or physics-based simulation, NeRF’s continuous field representation is more naturally compatible with the underlying math.