What Is Inpainting? How AI Fills and Restores Images

Inpainting is the process of filling in missing or damaged parts of an image so that the result looks natural and seamless. The term originated in art conservation, where restorers would manually repaint lost portions of a painting, but over the past three decades the technique has moved almost entirely into the digital realm. Driven first by mathematical models and then by increasingly powerful neural networks, digital inpainting has evolved from simple texture-copying tricks into systems that can generate entirely new visual content that plausibly belongs in a scene.

From Brushstrokes to Algorithms

The earliest digital inpainting methods, developed in the late 1990s and early 2000s, borrowed ideas from physics and mathematics. One well-known family of techniques used partial differential equations to “diffuse” color and texture from the surrounding pixels into the missing region, much like heat spreading through a metal plate. These methods worked well for thin scratches and small gaps, smoothly blending neighboring pixel values inward. Another family, patch-based synthesis, worked differently: the algorithm searched the rest of the image for small patches that matched the texture along the border of the hole, then copied and stitched those patches in. Patch-based methods handled repetitive textures like brick walls and grass convincingly, and researchers refined them to run faster without sacrificing visual quality.

Both approaches shared a fundamental limitation. They could only recycle information already present in the image. If a large region was missing and nothing similar existed elsewhere in the photo, these tools produced blurry smears or obvious tiling artifacts. They had no understanding of what objects look like or how scenes are structured. A missing face, for instance, could not be reconstructed from a stretch of sky.

How Deep Learning Changed the Game

The turning point came around 2016, when researchers began training neural networks to predict what belonged inside a missing region by learning from millions of example images. A landmark approach called Context Encoders used a convolutional neural network conditioned on the surrounding image to generate plausible content for an arbitrary missing area. The key insight was combining a standard pixel-level reconstruction objective with an adversarial training signal, which pushed the network to produce sharp, realistic-looking fills rather than blurry averages.

1CVPR. Context Encoders: Feature Learning by Inpainting

This was a genuine paradigm shift. Instead of copying textures, the network was synthesizing new content that had never appeared in the image. It could hallucinate a plausible eye socket, invent a reasonable patch of foliage, or fabricate the curve of a missing building roofline, all because it had seen enough eyes, leaves, and buildings during training to know what they should look like. A broad survey of the field describes this transition as a move from low-level texture synthesis to high-level semantic generation.

2Expert Systems. Image Inpainting in 30 Years: A Survey

Subsequent architectures tackled a practical problem the early networks struggled with: handling irregular, free-form masks. When you erase an unwanted object from a photo, the hole it leaves is rarely a neat rectangle. A technique called gated convolution introduced a learned mechanism that lets the network figure out, at every pixel and every processing layer, how much to trust the incoming information versus how much to suppress it because it falls inside the masked area. This made results far more consistent for oddly shaped holes.

3arXiv. Free-Form Image Inpainting with Gated Convolution

Filling Large Holes and High-Resolution Images

Erasing a small blemish is one thing; removing an entire person from a crowded street scene is another. Large missing regions are harder because the network needs to understand the global structure of the image, not just the texture near the border. A method called LaMa addressed this by using fast Fourier convolutions, which give the network an effective receptive field spanning the entire image in a single operation. Combined with aggressive training on large, varied masks and a perceptual loss function tuned for broad spatial awareness, LaMa produced coherent fills even when half the image was missing.

4arXiv. Resolution-robust Large Mask Inpainting with Fourier Convolutions

Another line of work focuses explicitly on preserving structural information like edges and contours. One recent framework, called SAIN, splits the job into two stages: first, a network generates complete edge maps for the scene, providing a skeleton of the missing geometry; then a second network fills in textures and colors guided by those structural priors. This kind of architecture helps the model balance fine detail with global shape consistency, producing results where straight lines stay straight and curves maintain their arc even across wide gaps.

5Journal of King Saud University Computer and Information Sciences. SAIN: structure-aware image inpainting for large missing areas

Diffusion Models and Text-Guided Fills

The latest generation of inpainting tools rides the wave of diffusion models, the same technology behind popular image generators. These models work by learning to reverse a gradual noising process: given a noisy image, the model predicts a slightly less noisy version, step by step, until a clean result emerges. For inpainting, the known pixels anchor the process while the model freely generates content in the masked region.

What makes diffusion-based inpainting especially powerful is that it can be guided by text prompts. You might erase a parked car from a street photo and type “flower garden” to have the model fill the gap with plants instead of asphalt. A system called PILOT demonstrated how to optimize in the model’s internal representation space so that generated content stays faithful to a user’s text prompt while blending seamlessly with the surrounding image. Because it works as an optimization layer on top of existing pretrained models, it can plug into various tools and extensions without retraining from scratch.

6arXiv. Coherent and Multi-modality Image Inpainting via Latent Space Optimization

This text-guided capability is what powers the “generative fill” features now showing up in consumer photo editors. The underlying principle is the same: mask a region, optionally describe what you want there, and let the model generate it.

Measuring Whether the Fill Looks Good

Evaluating inpainting quality is surprisingly tricky. The most common metrics fall into three broad categories. Peak signal-to-noise ratio (PSNR) measures raw pixel-level accuracy against a known ground truth. Structural similarity (SSIM) focuses on perceived structural patterns like luminance and contrast. And learned perceptual similarity (LPIPS) uses a neural network to judge whether two image patches look similar the way a human would perceive them.

7arXiv. Assessing Image Inpainting via Re-Inpainting Self-Consistency Evaluation

Each metric captures something different, and none alone tells the full story. A fill can score well on PSNR by being pixel-accurate on average yet look obviously wrong to a human eye because it blurs over fine details. Conversely, a fill that invents plausible but technically “incorrect” content, like a slightly different flower arrangement, might score poorly on pixel metrics while looking perfectly natural. Researchers increasingly lean on perceptual metrics and human evaluation studies for this reason, though no single standard has emerged as definitive.

Medical Imaging

Inpainting has found a serious foothold in medical imaging, especially for dealing with metal artifacts in MRI scans. Dental implants, surgical screws, and other metallic objects create dark voids and bright streaks that obscure the surrounding tissue. One study tested several neural network architectures on this problem and found that gated convolution outperformed alternatives, achieving the lowest error rates and highest image quality scores in the regions affected by metal artifacts. The inpainted MRI images aligned well with corresponding CT scans that had undergone separate metal artifact reduction, suggesting the neural network was reconstructing anatomically plausible tissue rather than just generating convincing-looking filler.

8PubMed. Inpainting the metal artifact region in MRI images by using generative adversarial networks with gated convolution

Beyond artifact removal, medical inpainting serves other clinical purposes. Diffusion models in particular have shown strong results in reconstructing what healthy tissue should look like in a region affected by a tumor or lesion, which can help with anomaly detection and surgical planning. They are also used for data augmentation, generating synthetic training examples that help other diagnostic algorithms learn to spot diseases when real labeled data is scarce.

9arXiv. Diffusion Models in Medical Image Inpainting: Challenges, Solution Taxonomy, and Future Directions

Restoring Ancient Art

Art restorers face a problem that is almost the inverse of the medical case: the “missing data” is centuries of wear, water damage, and flaking pigment on irreplaceable frescoes. Training a neural network for this task is difficult because there are very few intact examples to learn from. A technique called Deep Image Prior sidesteps the data shortage by using an untrained network whose architecture itself acts as a regularizer. The network is optimized on a single damaged image, progressively learning to match the intact portions while generating plausible fills for the gaps. Compared with older variational and patch-based methods, this approach reduced visual artifacts and better captured non-local contextual information in highly damaged medieval paintings from chapels in the Mediterranean Alpine Arc.

10Heritage Science. Deep image prior inpainting of ancient frescoes in the Mediterranean Alpine arc

An important nuance here is that art historians do not want the algorithm to “invent” missing content the way a photo editor might. The goal is a plausible reconstruction that helps scholars understand the original composition, not a definitive claim about what the artist painted. Results are treated as hypotheses, and conservators can integrate information from multiple imaging modalities, such as infrared reflectography, to guide the process.

Satellite Imagery and Cloud Removal

If you have ever looked at satellite imagery on a mapping service, you have seen the problem: clouds get in the way. For earth scientists and urban planners who need clean views of the ground, cloud removal is effectively an inpainting task. The cloudy region is masked out, and an algorithm fills it in with what the ground surface should look like. Generative adversarial networks trained on paired cloudy and cloud-free images have shown strong results here, with one study achieving a PSNR of about 29.4 dB and an SSIM of roughly 0.97, a meaningful improvement over prior approaches tested on the same dataset.

11arXiv. Cloud Removal from Satellite Images

The challenge in this domain is that unlike a portrait or a street scene, satellite images contain information at scales ranging from individual buildings to entire landscapes, and the “correct” fill depends on season, lighting conditions, and land-use changes that may have occurred between the reference image and the cloudy capture. Temporal information from repeated satellite passes helps, but single-image methods are also valuable when multi-temporal data is unavailable.

Inpainting Beyond Flat Images

The concept of filling in missing information extends well beyond two-dimensional photos. In audio processing, inpainting refers to reconstructing missing or corrupted segments of a sound signal. The missing chunk might be a few milliseconds wiped out by a data transmission error or a longer gap left by removing an unwanted noise. Audio inpainting methods typically rely on sparse representations or autoregressive models that predict what the signal should sound like based on the surrounding context.

12Signal Processing. Algorithms for audio inpainting based on probabilistic nonnegative matrix factorization

Three-dimensional scenes present another frontier. Removing an object from a 3D scene, such as one represented by a neural radiance field or Gaussian splatting model, requires filling in the geometry and appearance from every possible viewing angle, not just one. A recent pipeline called GPGS tackles this by first completing the missing 3D point cloud geometry using a pretrained completion model, then refining the visual appearance of the filled region from multiple viewpoints to correct brightness shifts and texture misalignments. The system produces coherent results even in 360-degree panoramic scenes, where inconsistencies would be immediately obvious.

13Proceedings of the AAAI Conference on Artificial Intelligence. GPGS: Consistent 3D Object Removal via Geometry-Aware 3D Inpainting and Projected Image Refinement in 3D Gaussian Splatting

Detecting Inpainted Tampering

The same power that makes inpainting useful for legitimate editing also makes it a tool for deception. Someone can erase a person from a photograph, remove a watermark, or delete evidence of an event. This has spawned a parallel field: inpainting forensics, which aims to detect whether and where an image has been inpainted.

Deep learning-based forensic detectors analyze subtle statistical traces that inpainting leaves behind, such as inconsistencies in texture patterns, noise distributions, or compression artifacts. One approach feeds the image through a multi-task detection network that examines texture features and multi-scale patterns to flag tampered regions. This kind of detector can identify manipulations made by both traditional patch-based tools and modern neural network inpainting, and remains effective even after the image has been JPEG-compressed or rescaled.

14IETE Technical Review. Image Inpainting Detection Based on Multi-task Deep Learning Network

The cat-and-mouse dynamic here is real. As inpainting methods improve, the artifacts they leave become more subtle, forcing forensic tools to evolve in response. Older detectors that relied on hand-crafted features have given way to learned detectors that can adapt, but even these struggle when the inpainting quality is very high and the filled region is small. A deep learning forensic method demonstrated strong detection rates with low false positives, but the authors acknowledged that ongoing advances in generative models will keep raising the bar.

15Signal Processing: Image Communication. A deep learning approach to patch-based image inpainting forensics

Running Inpainting on Your Phone

Most state-of-the-art inpainting models are built for powerful GPUs in data centers. Running them on a phone or tablet in real time is a different engineering problem entirely. A project called RETHINED specifically targets this gap, proposing a lightweight architecture that combines a small convolutional network for recovering structure with a resolution-agnostic patch replacement step for texture. The result can inpaint ultra-high-resolution images in under 30 milliseconds on a variety of mobile devices, roughly a hundred times faster than existing high-quality methods.

16arXiv. RETHINED: A New Benchmark and Baseline for Real-Time High-Resolution Image Inpainting On Edge Devices

The tradeoff, predictably, is quality. A model that runs in 30 milliseconds on a phone cannot match the semantic understanding of a massive diffusion model running for seconds on a server. But for common use cases like removing a photobomber, erasing a blemish, or cleaning up a screenshot, the lightweight approach is often good enough, and the instant feedback makes the editing experience far more intuitive than waiting for a cloud round-trip.

Privacy, Consent, and the Ethics of Erasing

As inpainting tools become easier to use, the ethical questions become harder to ignore. Generative AI editing tools lower the barrier to modifying visual content, which is empowering when you are cleaning up your own vacation photos and troubling when someone else is altering images of you. A study exploring user attitudes toward AI-powered privacy tools found that participants saw generative AI as a promising way to lower editing barriers and enhance creative control, but also raised concerns about data usage, manipulation, and transparency. The study highlighted a tension between the person who uploads a photo and the people depicted in it, pointing to a need for shared consent mechanisms.

17Proceedings of the ACM on Human-Computer Interaction. Understanding User Needs and Attitudes for Privacy Protection Tools in Online Visual Content Sharing

Inpainting sits at a peculiar intersection of privacy and truth. The same tool that can protect someone’s identity by seamlessly removing their face from a shared photo can also erase evidence, fabricate scenes, or alter the historical record. Platform-level solutions, like embedding invisible metadata about which regions of an image were AI-generated, are still in early stages. For now, the technology’s capabilities have far outpaced the social and legal frameworks governing its use, and most people interacting with inpainted images have no reliable way to tell that anything was changed.