Foundation models are large-scale AI systems trained on massive, broad datasets that can then be adapted to perform a wide variety of tasks, from writing code to diagnosing medical images to folding proteins. The term was coined in a 2021 Stanford report to emphasize that these models serve as a shared base layer for AI development, much like a building’s foundation supports many different structures above it.1arXiv. On the Opportunities and Risks of Foundation Models GPT-4, Claude, Gemini, LLaMA, and DALL-E are all foundation models, though the concept extends well beyond chatbots into biology, robotics, and climate science. What makes them genuinely different from earlier AI is a shift in philosophy: instead of building a separate model for every task, you build one very capable general model and then steer it toward whatever you need.
What Makes a Model “Foundational”
Before foundation models, the standard approach in machine learning was to collect a labeled dataset for a specific task, train a model on that dataset, and deploy it for that one purpose. A spam filter learned to catch spam; a tumor detector learned to spot tumors. Each model started from scratch or nearly so, and the knowledge it gained was locked inside a narrow application. Foundation models break that pattern by training on enormous, diverse data first and specializing later.2PubMed. The new paradigm in machine learning – foundation models, large language models and beyond: a primer for physicians A single model trained on billions of web pages absorbs grammar, facts, reasoning patterns, and stylistic conventions all at once. That general knowledge becomes a resource that can be redirected toward medical questions, legal documents, or customer support without starting over.
Two ingredients make this possible. The first is scale: these models have billions or even trillions of learned parameters, and they consume datasets that can span a significant fraction of the publicly available internet. The second is the transformer architecture, a design introduced in 2017 that allows the model to weigh how every piece of an input relates to every other piece. Transformers turned out to be remarkably good at absorbing patterns from unstructured data, which is why nearly all current foundation models are built on some variant of them. Alternative architectures like state-space models have started to compete on efficiency, with designs like Mamba-2 running two to eight times faster on certain tasks while matching transformer-level performance, but transformers remain the dominant backbone for now.3arXiv. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Scaling Laws and Why Bigger Kept Working
One of the most consequential discoveries behind foundation models is that their performance improves in a surprisingly predictable way as you increase three things: the number of parameters, the size of the training dataset, and the amount of computing power used. These relationships follow power laws, meaning that plotted on a log scale, the improvement forms a nearly straight line across many orders of magnitude.4arXiv. Scaling Laws for Neural Language Models This finding, first systematically documented by OpenAI researchers in 2020, had a profound practical implication: if you had a fixed computing budget, you could predict roughly how well a model would perform at different sizes before actually training it. That predictability gave companies the confidence to spend hundreds of millions of dollars on training runs, because the returns were not a gamble but a reasonably foreseeable curve.
The same research showed that larger models are more sample-efficient. A bigger model extracts more useful knowledge from the same amount of data, which means the optimal strategy under a fixed compute budget is often to build a very large model and train it on a relatively modest amount of data rather than to train a smaller model to convergence. Theoretical work has since connected these empirical observations to deeper mathematical properties of the data, identifying distinct regimes where performance is limited by how much data the model has seen versus how many parameters it has to represent what it learned.5PubMed Central. Explaining neural scaling laws These scaling laws are not a universal guarantee, though. They describe average performance on broad benchmarks, and there are specific tasks where a smaller, purpose-built model still outperforms a much larger general one.
The Debate Over Emergent Abilities
As language models grew, researchers noticed something that looked dramatic: certain abilities seemed to appear suddenly at a particular scale. A model with 10 billion parameters might score near zero on a task like multi-step arithmetic, while a model with 100 billion parameters could handle it comfortably. These jumps were labeled “emergent abilities” and fueled a narrative that bigger models might develop qualitatively new capacities that smaller ones simply cannot possess.
That narrative has come under serious scrutiny. A 2023 study presented at NeurIPS argued that the apparent emergence is largely an artifact of how performance is measured. When researchers used metrics that are nonlinear or discontinuous, such as exact-match accuracy where anything less than a perfect answer scores zero, performance appeared to jump suddenly. When they switched to smoother metrics that give partial credit, the same models showed gradual, predictable improvement with no sudden transitions.6NeurIPS Proceedings. Are Emergent Abilities of Large Language Models a Mirage? The researchers concluded that what looked like emergent abilities may evaporate with different measurement choices and better statistical analysis. This does not mean that scale is unimportant, only that the story of abrupt phase transitions may have been overstated. Models do get meaningfully better with size; the improvement just appears to be smooth rather than a series of surprise breakthroughs.
How Foundation Models Get Customized
A foundation model straight out of pre-training is like a highly knowledgeable but unsteered conversationalist. It can complete text plausibly, but it has no clear objective beyond predicting the next word. Turning that raw capability into something useful involves several layers of customization.
Alignment Through Human Feedback
The most prominent customization step for chatbot-style models is alignment, the process of teaching the model to follow instructions, be helpful, and avoid harmful outputs. The original technique, reinforcement learning from human feedback (RLHF), works in stages: humans rank the model’s outputs by quality, a separate reward model learns to predict those rankings, and the foundation model is then fine-tuned to maximize the reward model’s scores. This process proved effective but is complicated and computationally expensive. A simpler alternative called Direct Preference Optimization, or DPO, showed that you can skip the reward model entirely by reformulating the problem as a straightforward classification task on the preference data. DPO achieves comparable or better alignment while being lighter to run.7NeurIPS Proceedings. Direct Preference Optimization: Your Language Model is Secretly a Reward Model Since DPO’s introduction, a wave of further alternatives has appeared, but a recent theoretical survey found that most of them reduce to variations along a few core design choices rather than being fundamentally different methods.8arXiv. From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models
Parameter-Efficient Fine-Tuning
Full fine-tuning of a foundation model means updating all of its billions of parameters on a new dataset, which requires enormous memory and compute. For most organizations, that is impractical. A family of techniques known as parameter-efficient fine-tuning sidesteps the problem by modifying only a tiny fraction of the model’s weights while freezing the rest. The most widely used of these is LoRA, which inserts small trainable matrices into each layer of the model. Because these matrices are low-rank, they add very few new parameters but can still redirect the model’s behavior substantially for a specialized task.9arXiv. LoRA: Low-Rank Adaptation of Large Language Models LoRA and its descendants have made it feasible for small teams and individual researchers to adapt state-of-the-art models on consumer hardware, which has been a significant democratizing force.10arXiv. Parameter-Efficient Fine-Tuning for Foundation Models
Beyond Text
Foundation models are not limited to language. The same basic recipe of large-scale pre-training followed by task-specific adaptation has been applied to images, audio, video, and various combinations of these modalities. Vision-language models, for instance, learn joint representations of images and text, allowing them to answer questions about photographs, generate images from descriptions, or match medical scans with clinical notes.11PubMed Central. Vision-language foundation models for medical imaging: a review of current practices and innovations In medical imaging specifically, these models can integrate radiology images with the text of clinical reports to assist with tasks like disease classification, image segmentation, and automated report generation.
Earth observation is another area where multimodal foundation models are gaining ground. Satellite imagery comes from many different sensors with different spectral bands and resolutions, which traditionally meant building separate models for each satellite platform. Newer architectures are designed to encode sensor parameters directly into the model, so a single foundation model can learn from imagery across different satellites at once.12German Conference on Pattern Recognition. SenPa-MAE: Sensor Parameter Aware Masked Autoencoder for Multi-Satellite Self-Supervised Pretraining The practical payoff is that a model pre-trained across multiple satellite platforms can then be adapted to tasks like crop monitoring or flood mapping with far less labeled data than a single-sensor model would need.
Foundation Models in Biology
Molecular biology was one of the first fields outside of language processing to benefit from the foundation model approach. Biological sequences, whether DNA, RNA, or protein chains, share a structural similarity with language: they are strings of symbols whose meaning depends on long-range patterns and context. Researchers have trained foundation models on vast databases of protein sequences, genomic data, and single-cell transcriptome measurements.13PubMed Central. Foundation models in molecular biology
The most famous example is AlphaFold, which used large-scale pre-training on known protein structures to predict how amino acid chains fold into three-dimensional shapes. That problem had resisted computational approaches for decades. Other models like DNABERT applied the pre-training-then-fine-tuning recipe to genomic sequences, learning transferable features that could be steered toward tasks as varied as predicting gene expression, identifying regulatory elements, and screening drug candidates.14National Science Review. Foundation models in bioinformatics – Section: EVOLUTION OF THE BIOINFORMATICS FM The biological domain is a useful illustration of why the foundation model paradigm matters: labeled data in biology is expensive and scarce (every label may require a wet-lab experiment), so a model that can learn general biological structure from unlabeled sequences and then adapt with very few labeled examples solves a genuine bottleneck.
Robots That Inherit Internet Knowledge
One of the more surprising extensions of foundation models is into physical robotics. Vision-Language-Action (VLA) models unify visual perception, language understanding, and motor control into a single system, aiming to build robots that can generalize across different tasks, objects, and environments.15arXiv. Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications The core idea is to start with a vision-language model that already understands the visual world and natural language from internet-scale training, then teach it to output robot actions alongside its usual text and image outputs.
Google’s RT-2 demonstrated this by expressing robot actions as text tokens and co-training a vision-language model on both web data and robotic trajectory data. The result was a robot that could follow novel instructions it had never been explicitly trained on, drawing on semantic knowledge absorbed from the web.16arXiv. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control More recent work has extended this to multiple robot platforms simultaneously, training a single policy that can control single-arm robots, dual-arm systems, and mobile manipulators.17arXiv. π₀: A Vision-Language-Action Flow Model for General Robot Control Robotics has historically struggled with the brittleness of hand-coded control; foundation models offer a path toward robots that handle novel situations by reasoning about them rather than by having every scenario pre-programmed.
Foundation Models as Autonomous Agents
Beyond generating text or controlling a robot arm, foundation models are increasingly used as the decision-making core of autonomous software agents. In this setup, the model interprets a user’s goal, breaks it into subtasks, decides which external tools to call (a calculator, a web browser, a database query, a code interpreter), evaluates the results, and adjusts its plan as needed. Research in this area examines both single-agent systems, where one model handles everything, and multi-agent systems, where multiple model instances divide labor and communicate with each other.18Springer Link / Artificial Intelligence Review. From language to action: a review of large language models as autonomous agents and tool users This agentic use case pushes foundation models from being passive responders into something closer to autonomous problem-solvers, though the reliability of their planning and tool use remains an active area of work.
Hallucination and the Limits of Confidence
The most discussed failure mode of foundation models is hallucination: generating content that sounds authoritative but is factually wrong. This is not a bug that can be patched easily, because it is rooted in how these models are trained. The training objective is to predict the next token in a sequence, optimizing for plausibility rather than truth. The model learns to produce text that looks statistically likely given what came before, and it has no built-in mechanism for distinguishing between a claim it has strong evidence for and one it is essentially fabricating. This creates what researchers describe as overconfidence and poorly calibrated uncertainty.19arXiv. Medical Hallucination in Foundation Models and Their Impact on Healthcare In low-stakes contexts, hallucinations are a nuisance. In medicine, law, or engineering, they can be dangerous.
Alignment techniques reduce but do not eliminate hallucination. A model trained with RLHF or DPO learns to decline some questions rather than guess, but it also learns to be confidently helpful, which can work against caution. Various mitigation strategies exist: retrieval-augmented generation (where the model checks an external knowledge source before answering), chain-of-thought prompting (which forces the model to show its reasoning), and human-in-the-loop verification. None of these are a complete fix. The honest state of the field is that hallucination remains a fundamental limitation, not merely an engineering problem waiting for a cleaner solution.
Looking Inside the Black Box
A related challenge is that nobody fully understands what happens inside a foundation model when it processes a prompt and generates a response. The field of mechanistic interpretability is trying to reverse-engineer the internal logic of these networks by identifying human-understandable circuits, algorithms, and causal structures within the model’s layers.20ACM Computing Surveys. Bridging the Black Box: A Survey on Mechanistic Interpretability in AI This work has produced some genuine insights. Researchers have identified specific components, such as “induction heads” in transformers, that appear to drive in-context learning, and have developed tools like sparse autoencoders that can decompose tangled network activations into distinct, interpretable features.21arXiv.org. Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning
But current interpretability tools cover only a small fraction of what these models are doing. The gap between what we can explain and what the models can do is enormous, and that gap creates real problems for deploying foundation models in high-stakes settings where regulators and users need to understand why a particular decision was made.
Energy, Data, and the Cost of Scale
Training a frontier foundation model consumes staggering amounts of electricity. The computational demands are large enough to raise genuine concerns about environmental impact, and the energy cost of training is compounded by the energy cost of inference, since every query to a deployed model also requires computation.22Discover Artificial Intelligence. Green AI: exploring carbon footprints, mitigation strategies, and trade offs in large language model training Some researchers have proposed decentralized training approaches that distribute the work across many smaller devices at the edge of the network rather than concentrating it in massive data centers, which could reduce both the environmental footprint and the concentration of power among a few large companies.23ACM SIGEnergy Energy Informatics Review. Towards Decentralized and Sustainable Foundation Model Training with the Edge
Data supply is another looming constraint. Foundation models have been trained on increasingly large slices of the internet, and there are real questions about whether high-quality text data will eventually run out. Synthetic data, generated by the models themselves to augment training sets, has emerged as one proposed solution, but it carries its own risks: models trained on their own outputs can amplify biases and degrade in subtle ways.24arXiv. Best Practices and Lessons Learned on Synthetic Data Getting synthetic data right requires careful filtering and quality control, and the field is still working out best practices.
Open Weights Versus Closed APIs
Foundation models exist on a spectrum from fully closed (accessible only through a company’s API, with no visibility into the model’s architecture or weights) to open-weight (the trained model is published for anyone to download, inspect, and modify). Models like Meta’s LLaMA family and Mistral sit on the open end; OpenAI’s GPT-4 and Anthropic’s Claude sit on the closed end. Open-weight models allow independent researchers to study and improve them, which accelerates scientific progress and enables uses that a company’s API might not support. But open weights also mean anyone can modify the model, removing safety guardrails or using it without oversight, and once released, the model cannot be recalled.25Transactions on Machine Learning Research. Open Technical Problems in Open-Weight AI Model Risk Management The tension between openness and control is one of the most actively debated policy questions in AI, with no consensus yet on where the line should be drawn as models become more capable.
Effects on Work and the Economy
Because foundation models can process and generate language, they touch a different slice of the labor market than previous waves of automation. Traditional automation displaced manual and routine physical tasks; foundation models disproportionately affect roles centered on information processing, administration, and managerial coordination.26Sustainable Development. Cognitive Automation and Sustainable Development: Global Task‐Level Evidence on Large Language Models and Labor Markets At the sector level, knowledge-intensive industries like finance, education, and professional services face higher exposure, while agriculture and manufacturing remain more insulated. One finding that complicates the optimistic narrative is that the potential productivity gains from this exposure are concentrated in areas that contribute significantly to economic value rather than to employment. In other words, the sectors where these models could boost output the most are not the same sectors that employ the most people, which creates a tension between aggregate efficiency gains and equitable distribution of those gains across workers.
The geographic dimension matters too. Countries whose economies rely heavily on administrative and clerical work face greater exposure than those with economies weighted toward physical production. This is the opposite pattern from earlier automation waves, which hit manufacturing-heavy economies hardest. For individual workers, the practical reality is that foundation models are not replacing entire jobs so much as reshaping which tasks within a job are done by a human versus delegated to a model. The workers who adapt fastest tend to be those who learn to use these tools as amplifiers for their own judgment rather than as replacements for it.

