Federated Learning
Federated learning trains a model using data held across devices or organizations, without collecting those raw examples into one central training dataset. Participants compute local learning updates, which are combined to improve a shared model.
Also known as: FL
A keyboard application can send a model to participating phones, train locally on each phone’s examples, and combine returned parameter updates. The next round starts from an improved shared model. This illustrates the coordinator-based pattern introduced by the federated averaging paper; it is not a claim about a particular keyboard’s deployment.
Keeping raw text on-device changes what is transferred, but updates can still reveal information. Privacy protections such as secure aggregation or differential privacy address additional risks; they are not automatic consequences of using the word “federated.” Participants can also submit harmful updates.
Device availability, communication cost and differences between participants’ data affect training. This matters when judging both quality and privacy: a model trained mostly by frequently connected devices may serve other users poorly. Federated learning is also different from querying several databases at inference time. It combines learning signals to train a model, rather than returning distributed query results.
Sources
- McMahan et al.: Communication-efficient learning from decentralized data — Introduces local training with aggregated updates and discusses communication limits and non-identically distributed participant data.
- Zhu et al.: Deep Leakage from Gradients — Demonstrates reconstruction of private training examples from shared gradients; keeping raw examples local does not guarantee privacy.
Go deeper
- TensorFlow Federated: Image classification tutorial docs
Simulate client datasets and follow a federated model through training and evaluation.