Frontier, housed at Oak Ridge National Laboratory in Tennessee, became the world’s first public exascale supercomputer when it was inaugurated in 2022, capable of performing more than a quintillion calculations per second. That milestone, long pursued by the United States, China, Japan, and the European Union, represents more than a bragging-rights benchmark. The machine’s design choices in chips, cooling, power management, and software have shaped how researchers across dozens of scientific fields now approach problems that were previously too large to simulate.
What “Exascale” Actually Means
An exascale computer can sustain at least one exaflop of processing power, or 1018 floating-point operations per second. To put that in rough human terms, if every person on Earth performed one calculation per second, you would need more than a hundred million Earths running in parallel to match what Frontier does in one second. The system contains 9,408 CPUs, 37,632 GPUs, and a combined total of roughly 8.7 million cores working together.1arXiv. Challenges in Automatic Software Optimization: the Energy Efficiency Case Crossing the exascale threshold was not just about stacking more processors; it required a ground-up rethinking of how chips, memory, networking, cooling, and software work as a unified system.
The Hardware Inside Each Node
Frontier is built from HPE Cray EX cabinets, and its fundamental building block is the compute node. Each node pairs one 64-core AMD EPYC processor with four AMD Instinct MI250X graphics processing units.2DOE Data Explorer. Frontier (HPE Cray EX) Exascale Supercomputer at the Oak Ridge Leadership Computing Facility The CPU handles the kinds of branching, logic-heavy tasks that traditional processors excel at, while the GPUs do the heavy numerical lifting. Modern scientific simulation and AI training both lean heavily on the kind of massively parallel arithmetic that GPUs were originally designed for in video games but now dominate in high-performance computing.
The MI250X is itself a multi-chip module, meaning each physical GPU package actually contains two compute dies (AMD calls them Graphics Compute Dies, or GCDs). That detail matters for software developers because moving data between the two dies within a single GPU is faster than moving it between separate GPUs, and moving it between GPUs on the same node is faster than sending it across the network to a different node. Frontier’s programming model has to account for this hierarchy of communication speeds at every level.
With over 9,400 nodes, the system’s internal network is its own engineering feat. The Slingshot interconnect ties these nodes together in a topology designed to minimize the time data spends in transit. For problems that require billions of data points to be shared between processors mid-calculation, network performance can matter as much as raw chip speed.
Power Consumption and Energy Efficiency
Running 8.7 million cores simultaneously takes serious electricity. Frontier consumes about 21 megawatts during operation, roughly the power draw of a small town.3arXiv. Challenges in Automatic Software Optimization: the Energy Efficiency Case That sounds enormous, and it is, but the energy efficiency story is more nuanced than the raw wattage suggests. When Frontier first appeared on the Green500 list, which ranks supercomputers by performance per watt rather than raw speed, it took the top spot with 62.68 gigaflops per watt. The previous Green500 leader, Japan’s MN-3, had achieved 39.38 gigaflops per watt, so Frontier represented a substantial jump in efficiency even as it consumed far more total power.4arXiv. Challenges in Automatic Software Optimization: the Energy Efficiency Case
The efficiency gains come largely from the GPU-heavy architecture. GPUs perform more arithmetic operations per watt than CPUs for the kinds of parallel workloads that dominate scientific computing. Oak Ridge’s choice to pack four GPUs per node, with a single CPU acting more as an orchestrator, reflects a broader industry trend. The tradeoff is that software must be written or rewritten to run efficiently on GPUs, which is a nontrivial barrier for many legacy scientific codes that were designed for CPU-only machines.
Twenty-one megawatts translates to a substantial electricity bill and a significant carbon footprint, even before accounting for cooling overhead. The Tennessee Valley Authority supplies the facility’s power, and the site’s location was partly chosen for access to relatively affordable electricity. Still, the power question looms over the next generation of supercomputers. Machines aiming for performance beyond Frontier will need to find ways to hold power consumption steady or even reduce it, because doubling the wattage is not sustainable economically or environmentally.
Keeping It Cool
Twenty-one megawatts of electrical power becomes twenty-one megawatts of heat that has to go somewhere. Frontier uses a liquid cooling system rather than relying on air, because air cooling simply cannot remove heat fast enough at this density of computation. Warm liquid circulates through the cabinets, absorbs heat from the processors, and carries it to a cooling plant outside the building.
The cooling infrastructure is organized as multiple parallel subloops, each served by coolant distribution units. Figuring out how to allocate those units across subloops and how much flow to send through each one is a genuine optimization problem. Researchers at Oak Ridge have built simulation models to jointly determine how many cooling units each subloop needs, how to divide the flow, and how to adjust total flow rate and supply temperature over time while keeping every subloop within safe thermal limits.5arXiv. Co-Design Optimization for Data Center Cooling System via Digital Twin The fact that a digital twin of the cooling system itself requires sophisticated modeling gives a sense of the engineering complexity involved in simply keeping the machine from overheating.
Liquid cooling has become the default approach for the densest computing installations. The heat rejection problem scales with power consumption, and as future systems push past Frontier’s performance, the cooling infrastructure will need to evolve further. Some next-generation designs are exploring direct immersion cooling, where entire circuit boards are submerged in dielectric fluid, but Frontier’s warm-water approach represents the current mainstream for production exascale systems.
The File System That Feeds the Machine
A supercomputer this large generates and consumes staggering amounts of data. Frontier’s primary storage system is called Orion, and it runs on the Lustre parallel file system, which has been the workhorse of large-scale scientific computing for years.6ACM Transactions on Storage. Lustre Unveiled: Evolution, Design, Advancements, and Current Trends Lustre splits files across many storage servers simultaneously, so thousands of compute nodes can read and write data in parallel without creating a bottleneck at any single disk.
For many scientific workloads, the storage system is as critical as the processors. A climate simulation might produce terabytes of output per run. A genomics analysis might need to read enormous reference databases at high speed. If the file system cannot keep up with the rate at which the processors produce or request data, the CPUs and GPUs sit idle, wasting both time and electricity. Orion’s design has been studied to understand utilization patterns, performance characteristics, and usage trends at exascale, since Frontier is the first machine to push Lustre to this level of demand.7ACM Transactions on Storage. Lustre Unveiled: Evolution, Design, Advancements, and Current Trends
Beyond the parallel file system, Frontier also has a tiered storage architecture. Fast, node-local storage provides scratch space for intermediate results that do not need to persist, while the larger Lustre system handles longer-lived data. This tiered approach reduces network traffic and helps applications that do a lot of reading and writing mid-computation.
Scientific Applications Already Running on Frontier
The whole point of building an exascale supercomputer is to tackle scientific questions that were previously out of reach. Frontier’s user community spans physics, chemistry, biology, earth science, and engineering, with each field bringing problems that demand the machine’s scale.
In fusion energy research, scientists have used Frontier for gyrokinetic simulations that model how particles behave inside a fusion reactor’s plasma. These simulations track the turbulent motion of multiple ion species at extremely high resolution, the kind of calculation that is computationally brutal because it involves coupled physics across many spatial and temporal scales. Work on tungsten impurity transport in plasma, relevant to the design of future fusion reactors, has relied on Frontier for its most intensive computations.8IOP Publishing (Nuclear Fusion). Gyrokinetic prediction of core tungsten peaking in a WEST plasma with nitrogen impurities Understanding how tungsten migrates and concentrates in the reactor core is essential for keeping a fusion device running, since tungsten contamination can cool the plasma and kill the reaction.
Combustion science uses Frontier to simulate the chemistry inside engines and turbines at a level of detail that experiments alone cannot provide. One toolchain deployed on the machine, SUNDIALS, handles the large numbers of small, independent systems of chemical equations that arise when combustion is modeled with operator splitting. The same mathematical framework also serves cosmology simulations that model the evolution of the universe after the Big Bang.9The International Journal of High Performance Computing Applications. SUNDIALS time integrators for exascale applications with many independent systems of ordinary differential equations The fact that a single numerical library can serve both engine design and cosmology speaks to how exascale computing brings together domains that might seem unrelated but share the same mathematical structures.
Computational biophysics has also entered the exascale era with Frontier’s arrival. The combination of improved simulation software and hardware capable of sustaining exascale performance has opened new possibilities for modeling biological systems at atomic resolution over longer timescales than were previously feasible.10Biophysical Journal. Large-scale computing and the evolution of computational biophysics Drug design, protein folding, and the dynamics of cellular membranes are all areas where brute computational force translates directly into scientific insight, because the molecules involved are too complex and their behaviors too chaotic to model with shortcuts.
Training Large Language Models on a Supercomputer
Frontier was designed primarily for scientific simulation, but its GPU-rich architecture makes it naturally suited for training the large AI models that have reshaped the technology landscape since 2022. Researchers have explored using Frontier for training large language models with billions of parameters, and the results highlight both the machine’s strengths and the challenges unique to supercomputer-scale AI training.
The core difficulty is communication. When you spread a model with tens of billions of parameters across hundreds or thousands of GPUs, the processors spend a significant fraction of their time exchanging data rather than computing. Frontier’s internal network has varying bandwidths at different levels: communication between the two compute dies within a single MI250X GPU is fastest, transfer between separate GPUs on the same node is slower, and inter-node communication is slower still. Researchers have developed a three-level hierarchical partitioning strategy specifically tailored to this bandwidth hierarchy. For a 20-billion-parameter model, this approach yielded a 1.71 times increase in useful computation per GPU compared to a standard partitioning method, with a scaling efficiency of 0.94 across 384 compute dies.11arXiv. Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
Those numbers might sound abstract, but the practical implication is significant. A scaling efficiency of 0.94 means that adding more GPUs actually makes the training proportionally faster rather than wasting most of the added capacity on communication overhead. That is hard to achieve at this scale, and it depends on software that is carefully aware of the physical layout of the machine. The work also shows that government-funded supercomputers can play a role in AI development that is not entirely dominated by private companies with their own massive GPU clusters. Frontier offers something those private clusters often do not: open access for researchers who could never afford to rent equivalent hardware from a cloud provider.
How Frontier Compares to Commercial AI Clusters
The rise of enormous private GPU farms built by technology companies has changed the context in which Frontier operates. When Frontier debuted in 2022, its GPU count and aggregate performance were unmatched. Since then, several companies have built or announced clusters with comparable or larger numbers of GPUs, optimized specifically for AI training rather than general scientific simulation.
The differences matter. Commercial AI clusters tend to use NVIDIA GPUs with proprietary interconnects (like NVLink and InfiniBand) tuned for the specific communication patterns of deep learning. Frontier uses AMD GPUs and HPE’s Slingshot interconnect, which are designed for a broader set of workloads. For pure AI training throughput on the most common model architectures, the NVIDIA ecosystem currently has a software maturity advantage because most AI frameworks were developed on NVIDIA hardware first. Frontier’s researchers have had to do extra work porting and optimizing code for AMD’s ROCm software stack.
Where Frontier retains a clear advantage is in scientific workloads that mix traditional simulation with machine learning. A climate model that runs a physics simulation and then uses a neural network to correct its output in real time, for instance, benefits from a machine that can do both well on the same hardware. Commercial AI clusters are not designed for this kind of hybrid workload. Frontier’s architecture was explicitly built to serve it.
What Comes After Frontier
Oak Ridge is already preparing for Frontier’s successor. The U.S. Department of Energy’s next leadership computing systems will need to deliver significantly more performance while holding power consumption in check. The El Capitan system at Lawrence Livermore National Laboratory, which uses AMD’s newer Instinct MI300A accelerators that integrate CPU and GPU on the same chip package, represents one path forward. Integrating CPU and GPU memory eliminates some of the data-movement penalties that plague current architectures where the two are separate.
Europe and Japan are pursuing their own exascale programs as well. The JUPITER system in Germany and post-Fugaku planning in Japan reflect a global recognition that exascale computing has become infrastructure as fundamental to national scientific competitiveness as particle accelerators or space launch facilities. Frontier’s design choices, from its AMD chip selection to its liquid cooling approach to its Lustre-based storage, have influenced these efforts by demonstrating what works and what creates bottlenecks at exascale.
For Frontier itself, the near-term future involves expanding its scientific impact. Many of the codes now running on the machine were originally written for smaller systems and are still being optimized to use Frontier’s full capacity efficiently. The gap between a machine’s theoretical peak performance and what real applications actually achieve is always substantial, and closing that gap on a system this complex is a multi-year effort. Some of the most interesting science from Frontier will come not from its first year of operation but from the period when the software ecosystem has caught up to the hardware, and researchers can finally run the simulations they designed the machine to handle.

