Few-shot learning is a branch of machine learning in which a model learns to handle new tasks or recognize new categories from just a handful of labeled examples, sometimes as few as one. Where conventional deep learning typically needs thousands or millions of training samples, few-shot methods aim to mimic something closer to how people learn: see a couple of examples of a new bird species or a new handwriting style, and generalize from there. The field has grown rapidly since the mid-2010s, driven by both standalone few-shot algorithms and, more recently, by the in-context learning abilities of large language models.
The Problem Few-Shot Learning Solves
Standard deep learning is data-hungry. Training a neural network to distinguish, say, a hundred species of butterflies may require thousands of labeled photographs per species. For many real-world problems, collecting that volume of data is expensive, slow, or outright impossible. Rare diseases may produce only a handful of confirmed diagnostic images per year. A factory might need to detect a new type of defect the day it first appears. A military analyst might encounter a vehicle variant never catalogued before. In all these cases, a system that can learn from five or ten examples rather than five thousand is not just convenient but necessary.
Few-shot learning formalizes this challenge. A model is evaluated on its ability to classify examples from categories it has never seen during training, given only a small “support set” of labeled samples per category. When the support set contains just one example per class, the task is called one-shot learning; when it contains zero and the model must rely entirely on a description, it is called zero-shot learning. The number of examples per class is called the “shot number,” and matching the shot number used during training with the one used during evaluation generally produces the best results.1arXiv. A Theoretical Analysis of the Number of Shots in Few-Shot Learning
How the Main Approaches Work
Researchers have attacked few-shot learning from several angles. The approaches differ in philosophy, but all share the goal of squeezing useful generalization out of minimal data.
Metric Learning
The simplest intuition is: if you can map every image (or text snippet, or audio clip) into a well-structured space, then classifying a new example becomes a matter of checking which known example it is closest to. Prototypical Networks, one of the most influential methods in this family, learn to embed each class’s few examples and compute a prototype, essentially the average position of those examples in the learned space. A new input is then assigned to whichever prototype is nearest.2arXiv. Prototypical Networks for Few-shot Learning The elegance of this approach is its simplicity: once the embedding space is good, classification requires no retraining at all, just a distance calculation.
Meta-Learning
Meta-learning, sometimes called “learning to learn,” takes a different angle. Rather than learning a fixed embedding, a meta-learner trains across many small tasks so that it develops initialization parameters or update rules that adapt quickly to any new task. The most widely cited method here is MAML (Model-Agnostic Meta-Learning), which trains a model’s parameters so that just a few gradient updates on a new task’s tiny dataset produce strong performance. Because MAML works with any model trained by gradient descent, it can be applied to classification, regression, and even reinforcement learning.3arXiv. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Transfer Learning and Fine-Tuning
A third family sidesteps the meta-learning machinery altogether. Instead of training specialized few-shot algorithms, you take a large model pre-trained on a broad dataset, then fine-tune it on the small support set. This sounds too simple to compete, and for a long time the field assumed meta-learning methods would dominate. But research comparing MAML and a similar technique called Reptile against straightforward fine-tuning found a surprising pattern: meta-learning methods excel when the new task closely resembles the training distribution, but when the new task comes from a different distribution, plain fine-tuning of a pre-trained network often works just as well or better. The pre-trained features turn out to be more diverse and discriminative than the specialized features learned by meta-learners.4Springer Link / Machine Learning. Understanding transfer learning and gradient-based meta-learning techniques This finding has had a lasting impact on how practitioners choose methods: if you expect your deployment tasks to look like your training tasks, meta-learning is a strong choice; if the world is unpredictable, a robust pre-trained backbone may be safer.
Few-Shot Learning Inside Large Language Models
When GPT-3 arrived in 2020, it popularized a form of few-shot learning that does not require any weight updates at all. You give the model a prompt containing a few input-output examples (the “shots”), then ask it to produce the output for a new input. The model never retrains; it infers the task from the pattern in the prompt. This is called in-context learning, and it has become the dominant way most people now encounter few-shot methods in practice.
In-context learning has proven particularly useful for low-resource languages, where labeled data is scarce by definition. A study across 25 low-resource and 7 higher-resource languages found that providing even a few semantically relevant in-context examples substantially improved a large language model’s understanding quality in the target language. The researchers also introduced a technique called query alignment, which closes the gap between the low-resource language and a high-resource language the model already handles well, outperforming earlier label-alignment strategies.5arXiv. LLMs Are Few-Shot In-Context Low-Resource Language Learners
This kind of in-context few-shot learning is now routine for tasks like text classification, summarization, translation, and question answering. But it has an important subtlety that catches people off guard.
When More Examples Hurt
A natural assumption is that giving a language model more in-context examples should always help. If five examples are good, fifteen should be better. Research on this question found the opposite can happen: loading too many domain-specific examples into a prompt can paradoxically degrade performance in certain models. By gradually increasing the number of carefully selected few-shot examples, researchers identified an optimal quantity for each model, beyond which accuracy actually dropped. A combined strategy of choosing the right examples and capping their count beat the state of the art while using fewer shots.6arXiv. The Few-shot Dilemma: Over-prompting Large Language Models
The reasons are not fully pinned down, but likely explanations include the model’s context window becoming cluttered, redundant examples introducing noise, and the model anchoring too heavily on surface patterns in the examples rather than the underlying task. For practitioners, the lesson is clear: treat the number of shots as a hyperparameter to tune, not a dial to crank up.
Applications in Medical Imaging
Healthcare is one of the most natural fits for few-shot learning because many conditions are rare, and expert-labeled data is both expensive and privacy-sensitive. A scoping review of the field describes few-shot learning as especially beneficial for medical imaging applications involving rare diseases or personalized medicine, where large labeled datasets simply cannot exist.7Discover Applied Sciences. The evolving landscape of few shot learning in medical image diagnosis a scoping review
A concrete example comes from ophthalmology. Standard deep learning systems trained on fundus photographs can detect common conditions like diabetic retinopathy, but they ignore rarer pathologies such as papilledema or anterior ischemic optic neuropathy because there are too few training images. A few-shot learning framework extended a convolutional neural network trained on common conditions with a probabilistic model for rare condition detection, allowing the system to flag conditions it had seen only a handful of times.8PubMed. Automatic detection of rare pathologies in fundus photographs using few-shot learning A later approach called dynamic feature splicing built on this idea for broader rare disease diagnosis, further demonstrating the practical significance of few-shot methods in settings where collecting more data is not a realistic option.9PubMed. Dynamic feature splicing for few-shot rare disease diagnosis
The pattern is consistent: wherever medicine encounters a long tail of rare conditions, few-shot learning offers a path forward that conventional data-hungry models cannot take.
Applications in Computer Vision Beyond Healthcare
Object detection, the task of not only classifying but locating objects within an image, has its own few-shot variant. Few-shot object detection asks a model to spot objects of a new category given just a handful of annotated examples. A survey of the field notes that few-shot learning’s core design goal, operating effectively in a scarce-data regime, maps directly onto this challenge.10ACM Computing Surveys. Few-Shot Object Detection: A Survey Recent work using domain adaptation strategies has pushed performance up by several percentage points on standard benchmarks.11Neural Processing Letters. Few-Shot Object Detection Based on Global Domain Adaptation Strategy
Practical scenarios abound: security cameras encountering a new type of vehicle, agricultural drones needing to spot a new pest, or warehouse robots asked to pick a product they have never handled. In each case, collecting and annotating thousands of images is too slow. A few-shot detector that generalizes from a dozen examples dramatically shortens the deployment timeline.
Robotics and Real-World Adaptation
Robots face a version of the few-shot problem in the physical world. A legged robot trained in simulation may encounter a missing leg, a steep slope, or an unexpected payload. Training a separate policy for every possible scenario is impractical. Meta-reinforcement learning addresses this by training a dynamics model that can be rapidly adapted online to the local context using only recent experience. One study demonstrated a real millirobot adapting in real time to a missing leg, novel terrains, miscalibrated sensors, and pulling payloads it had never encountered, all without retraining from scratch.12arXiv. Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning
What makes this particularly striking is the data efficiency: the meta-learned approach achieved performance comparable to model-free reinforcement learning methods while requiring roughly a thousand times less real-world experience. Instead of days of trial-and-error in the physical world, the robot needed only a few hours of equivalent interaction to adapt. For anyone working on deploying autonomous systems outside the lab, that difference is the gap between feasible and impossible.
Synthetic Data as a Crutch
One increasingly popular idea is to sidestep the data scarcity problem by generating synthetic training examples. If you only have five real images of a rare bird, maybe a generative model can produce fifty more. In principle, this augments the support set and gives the few-shot learner more to work with. In practice, the gap between real and synthetic data distributions often degrades performance rather than helping it. Models trained partly on synthetic samples can learn artifacts of the generator rather than genuine features of the target class.13arXiv. Provably Improving Generalization of Few-Shot Models with Synthetic Data
Researchers are actively working on methods to close this gap, including techniques that provably improve generalization when synthetic data is mixed in carefully. But the current state of the art is that naively adding generated images to a few-shot dataset is not a guaranteed win. The quality and realism of the synthetic samples matter enormously, and knowing when synthetic augmentation helps versus hurts is itself an open research question.
Parameter-Efficient Adaptation
A related trend is adapting very large pre-trained models to new few-shot tasks by tuning only a tiny fraction of their parameters. Rather than fine-tuning all billions of weights, methods like LoRA (Low-Rank Adaptation) insert small trainable modules into a frozen model. This approach has been applied not just to language models but also to image generation and restoration. One study took a 12-billion-parameter image editing model and, by fine-tuning only LoRA adapters on 32 to 128 paired images per task, turned a mediocre zero-shot editor into a competitive image restorer.14arXiv. Edit2Restore:Few-Shot Image Restoration via Parameter-Efficient Adaptation of Pre-trained Editing Models
The insight the researchers emphasized is worth quoting in spirit: what large pre-trained models lack is not capability but direction. They already contain rich representations of the world; they just need a nudge toward the specific task at hand. A few dozen examples and a lightweight adapter can provide that nudge without the computational expense of full retraining. This philosophy has become central to how few-shot learning is practiced in the era of foundation models.
The Data Contamination Problem
As large language models have become the default platform for few-shot learning, a serious methodological concern has emerged: how do we know the model is actually learning from the few examples in the prompt, rather than recalling the answer from its massive training data? This concern, known as task contamination, has been investigated systematically. One study found that language models performed surprisingly better on datasets released before their training data cutoff than on datasets released afterward. When evaluating on tasks with no possibility of contamination, the models rarely showed statistically significant improvements over simple baselines in either zero-shot or few-shot settings.15Proceedings of the AAAI Conference on Artificial Intelligence. Task Contamination: Language Models May Not Be Few-Shot Anymore
This finding is sobering. It suggests that some of the impressive few-shot benchmark numbers reported for large language models may be inflated by memorization rather than genuine generalization. A separate effort called Clean-Eval attempted to address this by paraphrasing contaminated evaluation data into semantically equivalent but superficially different forms, then re-evaluating models on the cleaned versions. Results across 20 benchmarks showed that cleaning the evaluation sets substantially changed the apparent performance of contaminated models.16Findings of the Association for Computational Linguistics: NAACL 2024. CLEAN–EVAL: Clean Evaluation on Contaminated Large Language Models
For anyone evaluating few-shot capabilities, the takeaway is to be skeptical of headline numbers, especially on well-known benchmarks that are likely present in training corpora. Truly novel tasks, or freshly created evaluation sets, provide a more honest measure of what a model can do from a few examples alone.
Where Standalone Few-Shot Methods Still Matter
Given the dominance of large foundation models, you might wonder whether the earlier generation of few-shot algorithms, the Prototypical Networks and MAMLs, still has a role to play. The answer is yes, for a few practical reasons. Large language and vision models require significant compute resources just for inference. A hospital running diagnostics on a local device, a drone operating without cloud connectivity, or an embedded sensor in a factory cannot afford to run a model with billions of parameters. Lightweight few-shot classifiers that operate in a learned metric space remain far more deployable in these settings.
There is also the question of domain specificity. Foundation models are trained on broad internet data, which may not cover specialized visual domains like metallurgical microscopy, satellite hyperspectral imagery, or niche biological specimens. In these areas, a purpose-built few-shot learner trained on domain-relevant meta-tasks can outperform a general-purpose giant. The field has not converged on one winner; it has stratified into regimes where different tools make sense.
Richer Supervision With Fewer Samples
One line of research explores whether giving the model more information per example, rather than more examples, can compensate for having few shots. Instead of labeling each training image with just a class name, you might provide per-sample annotations indicating which features are relevant to the classification. This “per-sample rich supervision” provides a denser learning signal from each of the few available examples, improving generalization error in both theoretical bounds and empirical benchmarks for tasks like scene classification and fine-grained recognition.17arXiv. Few-Shot Learning with Per-Sample Rich Supervision
The intuition is straightforward: if you can only show someone three photographs of a bird species, pointing out that the diagnostic feature is the shape of the beak rather than the background foliage makes each photograph far more informative. Translating that human insight into a form a model can use is the challenge, but early results suggest it is a productive direction, especially in domains where expert annotators are available even if example images are not.

