What Is Federated Learning and How Does It Work?

Federated learning is a method of training machine-learning models across many devices or institutions without collecting everyone’s raw data in one place. Instead of uploading your photos, medical records, or text messages to a central server, each device trains a local copy of the model on its own data and sends only the learned updates back. A central server stitches those updates together into a single improved model, then pushes it back out. The data never leaves your device. This idea, first formalized by Google researchers in 2016, has since expanded into healthcare, finance, and smart-home applications, but it carries a set of engineering and privacy challenges that are far from solved.

How the Training Loop Works

Picture a hospital network where ten clinics each hold patient scans they cannot legally share. In a traditional setup, someone would need to pool all those scans in one data warehouse and train a model there. Federated learning flips the process. A coordinating server sends the same starting model to every clinic. Each clinic trains that model on its own scans for a few rounds, producing a set of weight updates. The clinics send those updates, not the scans, back to the server. The server averages the updates, producing a new global model, and the cycle repeats.

The most common averaging method is called Federated Averaging. It works by weighting each client’s contribution by the amount of data it trained on, so a clinic with 10,000 images has more pull than one with 500. This whole loop can run for hundreds of rounds before the model converges. The framework has been extended into three broad categories: horizontal federated learning, where different organizations hold data with the same features but different users; vertical federated learning, where organizations share users but hold different features about them; and federated transfer learning, which bridges gaps when both users and features differ across parties.1ACM Transactions on Intelligent Systems and Technology. Federated Machine Learning

Can It Match Centralized Training?

The obvious worry is that keeping data scattered will produce a worse model. Across a wide range of settings, though, federated learning achieves performance similar to centralized training, where all data sits in one place.2PubMed Central. A comprehensive experimental comparison between federated and centralized learning When accuracy gaps do appear, they tend to show up under specific stressful conditions, such as when each client’s data looks very different from every other client’s. One effective countermeasure is starting from a pre-trained model rather than training from scratch. Across multiple image-recognition tasks, pre-training closed the accuracy gap between federated and centralized approaches, especially in the hardest cases where each device’s data was highly unrepresentative of the overall distribution.3arXiv. On the Importance and Applicability of Pre-Training for Federated Learning

So the short version: if the data across clients is reasonably similar, you lose little by going federated. If it is wildly different, you lose more, but techniques like pre-training can recover most of the gap.

The Data Heterogeneity Problem

In textbooks, training data is assumed to be independent and identically distributed. In the real world, that assumption rarely holds. A dermatology clinic in northern Europe sees different skin conditions from one in Southeast Asia. A smartphone keyboard in Japan encounters different word patterns from one in Brazil. When each client’s data distribution is skewed in its own way, federated learning can slow down and produce a model that works well on average but poorly for any individual client.4arXiv. Non-IID data in Federated Learning: A Survey with Taxonomy, Metrics, Methods, Frameworks and Future Directions

This is arguably the single biggest open problem in the field. Skewed data does not just hurt the final model; it also discourages participation. If a hospital joins a federated network and the resulting model performs worse on its patients than its own local model did, it has no reason to keep contributing.5Future Generation Computer Systems. A state-of-the-art survey on solving non-IID data in Federated Learning Most proposed solutions attack the problem from the algorithm side, adjusting how updates are weighted, clustering similar clients together, or sharing small synthetic datasets that approximate each client’s distribution without revealing the real data. None of these approaches fully solve the problem yet, but they significantly narrow the gap.

Personalization Instead of One-Size-Fits-All

A related strategy is to abandon the goal of a single global model altogether and instead build personalized models for each client. In a personalized federated learning framework, every client maintains its own model while selectively absorbing useful knowledge from others. The APPLE framework, for instance, lets each client adaptively learn how much to borrow from every other client’s model, rather than blindly averaging everything.6PubMed Central. Adapt to Adaptation: Learning Personalization for Cross-Silo Federated Learning

More recent work takes this further by computing multidimensional similarity between clients and enabling selective sharing of only those model parameters that are relevant. In these systems, a client whose data distribution is unusual does not get dragged toward the average; it cherry-picks insights from whichever peers have the most useful knowledge to offer.7Information Sciences. FedPDA: Personalized federated learning based on attribute similarity migration Personalization is especially appealing in healthcare and mobile settings, where the differences between clients are not noise to be averaged out but genuine signal that each client’s model should respect.

Privacy Is the Selling Point, but It Has Limits

Federated learning’s headline promise is privacy: your data never leaves your device. That is true, but model updates themselves can leak information. Researchers have shown that under the right conditions, a malicious server or a curious participant can partially reconstruct training data from the gradients alone. If someone sends weight updates after training on a batch of chest X-rays, an attacker with enough computational resources may be able to infer what some of those X-rays looked like.

The standard defense is differential privacy, a mathematical framework that adds calibrated noise to the updates before they leave the client. This noise makes it provably difficult for anyone to reverse-engineer a specific person’s data. The trade-off is that more noise means more privacy but less model accuracy. In medical imaging experiments, federated learning with differential privacy achieved strong privacy bounds while maintaining performance comparable to centralized training and meaningfully better than each institution training alone.8PubMed Central. Defending against Reconstruction Attacks through Differentially Private Federated Learning for Classification of Heterogeneous Chest X-ray Data The key insight from privacy research is that the protections need to be local, obfuscating data before the server or any other participant can observe it, because limiting what a well-resourced adversary could reconstruct requires defense at the source.9arXiv. Protection Against Reconstruction and Its Applications in Private Federated Learning

Cryptographic approaches offer another layer. Homomorphic encryption allows a server to aggregate encrypted model updates without ever decrypting them. However, existing schemes that require all clients to share the same encryption key pair introduce a new vulnerability: if the shared key is compromised, everyone’s updates are exposed. Multi-key aggregation protocols address this by letting each client use its own key pair while still allowing the server to combine the encrypted updates.10Information Sciences. Secure and efficient multi-key aggregation for federated learning

Backdoor Attacks and Malicious Clients

Privacy is about protecting data from observers. Security is about protecting the model from participants. Because federated learning is distributed by design, a malicious client can poison the training process without anyone directly inspecting their data. The most studied form is the backdoor attack: a participant modifies their local data or tampering with their model updates so that the global model learns a hidden trigger. For example, a compromised client could train the model to misclassify any image containing a tiny watermark pattern, while performing normally on clean inputs.11PubMed Central. Federated Learning Backdoor Attack Based on Frequency Domain Injection

Defenses include robust aggregation methods that detect and down-weight suspicious updates, anomaly detection on gradient distributions, and requiring clients to prove properties of their data without revealing the data itself. The research community has been locked in an arms race between increasingly subtle attacks and increasingly sophisticated defenses. In safety-critical applications like autonomous driving or clinical diagnosis, this tension makes deploying federated learning more cautious than in, say, keyboard prediction.

When Devices Are Slow, Offline, or Underpowered

Not every device in a federated network is equal. A flagship smartphone trains faster than a budget smartwatch. A hospital with a GPU cluster computes updates in minutes; a rural clinic on a single laptop takes hours. If the server waits for every client to finish before moving to the next round, the slowest device (the “straggler”) sets the pace for everyone. This is the system heterogeneity problem, and it gets worse in networks with hundreds or thousands of participants.

Asynchronous approaches let the server proceed with whatever updates have arrived by a deadline, rather than waiting for all clients. The FedStrag framework, for instance, uses a time-bounded asynchronous paradigm that optimizes training even with stale updates from slow devices. In experiments run on Raspberry Pis with multiple straggler scenarios, it outperformed standard Federated Averaging in every case.12Digital Communications and Networks. FedStrag: Straggler-aware federated learning for low resource devices These solutions are especially relevant for Internet of Things networks, where devices range from powerful edge servers to tiny sensors with barely enough memory to hold a model.

Cutting Down Communication Costs

Even when every device is fast enough, the sheer volume of model updates flying back and forth can overwhelm a network. A modern deep-learning model has millions or billions of parameters. Sending the full set of updated weights from thousands of clients every round is expensive in bandwidth, time, and energy. This is why gradient compression has become a major research focus.

One approach is gradient sparsification: instead of sending every updated parameter, each client sends only the largest or most changed ones and zeros out the rest. Done naively, this degrades model quality because small but cumulatively important updates get dropped. Error feedback mechanisms track the dropped information and fold it back in during later rounds, recovering most of the lost accuracy while still achieving significant reductions in data transmission.13IEEE Transactions on Cognitive Communications and Networking. Gradient Compression and Correlation Driven Federated Learning for Wireless Traffic Prediction In bandwidth-constrained environments like cellular networks, these techniques can make federated learning practical where it otherwise would not be.

Dropping the Central Server Entirely

Standard federated learning still relies on a central server to coordinate rounds and aggregate updates. That server is a single point of failure and, in some settings, a trust bottleneck: whoever runs the server can see all the updates and potentially misbehave. Fully decentralized variants remove the server and let clients communicate directly with each other, peer-to-peer.

BrainTorrent, developed for medical applications, is one such framework. Each participating hospital interacts directly with the others, sharing model updates without routing through a central body.14arXiv. BrainTorrent: A Peer-to-Peer Environment for Decentralized Federated Learning Other decentralized designs use incentive mechanisms to keep participation fair. The Incentive-Aware Federated Bargaining framework, for example, distributes rewards based on each client’s actual contribution to the model, measured by a game-theoretic value metric. In experiments, it improved participation fairness by about 28% and reduced the time needed for the model to converge by roughly 35% compared to standard Federated Averaging.15PubMed Central. An incentive-aware federated bargaining approach for client selection in decentralized federated learning for IoT smart homes

The trade-off is complexity. Without a server orchestrating the rounds, clients need to agree on who to communicate with, when to exchange updates, and how to handle stragglers. Decentralized systems also face trickier security questions, because there is no single aggregation point where you can run anomaly detection on incoming updates.

Where Federated Learning Is Already Deployed

The most visible consumer deployment is on smartphone keyboards. Google’s Gboard was one of the first real-world federated learning applications. A recurrent neural network trained across millions of phones learns to predict the next word you will type, and it does so more accurately than a version trained on server-side data alone, because the on-device data better reflects how people actually write.16arXiv. Federated Learning for Mobile Keyboard Prediction A related line of work showed that federated learning can even pick up out-of-vocabulary words, expanding a keyboard’s dictionary with slang, new coinages, and proper nouns that users type frequently, without ever exporting that text to a server.17arXiv. Federated Learning Of Out-Of-Vocabulary Words

Healthcare is the field with perhaps the strongest motivation for federated learning, given strict patient-privacy regulations. In medical imaging studies, federated models trained across institutions have matched centralized performance with strong privacy guarantees, achieving this without any single hospital needing to share patient data outside its walls.18PubMed Central. Defending against Reconstruction Attacks through Differentially Private Federated Learning for Classification of Heterogeneous Chest X-ray Data Financial fraud detection is another active area, where banks want to collaborate on spotting suspicious patterns without exposing customer transaction data to competitors or regulators.

Fairness and Bias Across Clients

When clients contribute unequal amounts of data or represent different populations, the resulting model can develop biases. A skin-cancer classifier trained mostly on images from light-skinned populations will underperform on darker skin tones, and federated learning can amplify this if the contributing clinics are not diverse. Fairness in federated learning is complicated by the fact that heterogeneity shows up at multiple levels: differences in data volume, data distribution, device capability, and participation frequency can all skew the model’s predictions toward some groups and away from others.19Advanced Intelligent Systems. Fairness in Federated Learning: Trends, Challenges, and Opportunities

Proposed solutions range from re-weighting client contributions to ensure underrepresented groups have proportional influence, to post-hoc auditing of the global model for disparate performance across subpopulations. This is an area where the research is still catching up to the problem. Most fairness work in machine learning assumes you can inspect and re-balance the training data, which is exactly what federated learning is designed to prevent.

Federated Unlearning

Privacy regulations like the GDPR give people the right to have their data erased. In a centralized system, you delete the data and retrain the model. In federated learning, where the data was never centralized in the first place, the question becomes: how do you remove one client’s influence from a model that was trained across dozens of rounds with contributions from everyone?

One approach is to run a local “unlearning” process on the client being removed, producing a modified model that is then used as the starting point for a few additional federated rounds among the remaining clients. This avoids full retraining from scratch.20PubMed Central. Federated Unlearning: How to Efficiently Erase a Client in FL? A more exact solution partitions the training process into independent groups of lightweight modules. When a client requests erasure, the server simply deactivates the modules that the client contributed to, achieving instant removal without retraining.21arXiv. FedSGT: Exact Federated Unlearning via Sequential Group-based Training The distinction between approximate unlearning (which reduces but does not fully eliminate a client’s influence) and exact unlearning (which provably removes it) matters legally. Regulators have not yet clarified which standard they expect, but the technical community is preparing for the stricter one.

Fine-Tuning Large Language Models

The rise of large language models has raised a natural question: can federated learning be used to customize them? Training a model with billions of parameters from scratch across distributed devices is impractical given current communication and computation constraints. But fine-tuning a pre-trained model on domain-specific instructions is much more feasible.

Federated Instruction Tuning applies federated learning to the fine-tuning stage, letting organizations tailor a shared base model using their private instruction-response datasets without uploading those datasets anywhere.22arXiv. Towards Building the Federated GPT: Federated Instruction Tuning A law firm could fine-tune a language model on its internal case summaries while a medical group does the same with clinical notes, and both contribute to improving the base model without ever seeing each other’s data. This is still early-stage work, and the communication overhead of even fine-tuning-sized updates across many participants is significant, but it represents one of the more watched frontiers in the field.

The Energy Footprint

Training machine-learning models in massive data centers consumes a lot of electricity, and the carbon footprint of AI has become a growing concern. Federated learning, despite being slower to converge because of its distributed nature, can actually be greener than centralized alternatives. The energy used by many small devices spread across different power grids, some of which run on renewables, can compare favorably to the concentrated energy draw of a data-center GPU cluster.23arXiv. Can Federated Learning Save The Planet? The carbon math depends heavily on where the devices are located and what power sources they draw from, but the finding challenges the assumption that distributed training is necessarily less efficient. For organizations already motivated by privacy to adopt federated learning, the potential sustainability benefit is a welcome bonus rather than the primary justification, but it is worth knowing about.