Communication-Efficient Federated Learning for Neural Machine Translation
Tanya Roosta, Peyman Passban, Ankit Chadha
Federated learningNeurIPS 2021 ENLSP Workshop · Amazon Alexa AI
PDF arXiv Publisher page BibTeX
To train translation models across separate data owners, the authors share only a few small added "Controller" layers, cutting exchanged parameters about six times with average translation quality slightly above the matched baseline.
Abstract
Training neural machine translation (NMT) models in federated learning (FL) settings could be inefficient both computationally and communication-wise, due to the large size of translation engines as well as the multiple rounds of updates required to train clients and a central server. In this paper, we explore how to efficiently build NMT models in an FL setup by proposing a novel solution. In order to reduce the communication overhead, out of all neural layers we only exchange what we term “Controller” layers. Controllers are a small number of additional neural components connected to our pre-trained architectures. These new components are placed in between original layers. They act as liaisons to communicate with the central server and learn minimal information that is sufficient enough to update clients. We evaluated the performance of our models on five datasets from different domains to translate from German into English. We noted that the models equipped with Controllers preform on par with those trained in a central and non-FL setting. In addition, we observed a substantial reduction in the communication traffic of the FL pipeline, which is a direct consequence of using Controllers. Based on our experiments, Controller-based models are ∼ 6 times less expensive than their other peers. This reduction is significantly important when we consider the number of parameters in large models and it becomes even more critical when such parameters need to be exchanged for multiple rounds in FL settings.
Why it matters
Federated training keeps each owner's data in place, but every round ships the whole model back and forth, which is costly for a translation engine with tens of millions of parameters. This work inserts two small trainable layers into the encoder and decoder and shares only those, holding average BLEU, a standard automatic translation score, about a point above a same-size baseline that shares everything, at roughly one sixth of the parameter traffic. Useful when bandwidth, not accuracy, is the binding constraint.
Key results
- Sharing and training only Controller layers (8E-8D/C-C (2-6)) reaches 29.46 average BLEU versus 28.40 for the matched federated baseline that shares and trains all layers (8E-8D/A-A).
- Controller-only exchange moves 15,240,704 parameters per round instead of the 94,079,328 of the 8E-8D/A-A baseline (which includes the 33,116,512-parameter embedding table), making Controller-based models about 6 times less expensive.
- Plain federated averaging over five clients (6E-6D/A-A) averages 33.83 BLEU, above the centralized mixed-domain Fine Tuning model at 32.00 and far above Chained Training at 21.83.
- Null or negativeThe Controller models do not actually match the non-FL centralized setting on average, scoring 29.46 BLEU against 32.00 for centralized Fine Tuning and 33.83 for the full-parameter federated model, so the abstract's "on par" wording overstates Table 1.
- Null or negativeGrowing the federated model from 6 to 8 encoder and decoder layers under the same 150K-step budget lowered average BLEU from 33.83 (6E-6D/A-A) to 28.40 (8E-8D/A-A), the opposite of the authors' stated expectation.
- Null or negativeReusing existing layers of the 6-layer model as Controllers is highly position-sensitive: average BLEU drops from 29.58 at positions (0-3) to 27.67 at (1-4), and the paper reports that placing a Controller between the embedding table and the first encoder layer stops the client from converging at all.
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper is about federated learning, neural machine translation, communication-efficient training, parameter-efficient fine-tuning, adapter layers, controller layers, transformer, fedavg, german-english translation, wmt newstest-14, opensubtitles, and non-iid data.
What this paper does not show
The Controller models lose to both comparators in our own tables, 29.46 BLEU against 33.83 for full-parameter federated training and 32.00 for centralized fine-tuning, so "on par" is generous. The roughly 6 times saving is a parameter count, not measured bytes on the wire or wall-clock time. No comparison against the standard communication-efficiency baselines such as sparsification or quantization, and the adapter comparison is qualitative with no numbers. Controllers are adapter-style layers under a new name, which the paper acknowledges.
How to cite
Tanya Roosta, Peyman Passban, and Ankit Chadha. 2021. Communication-Efficient Federated Learning for Neural Machine Translation. In NeurIPS 2021 Workshop on Efficient Natural Language and Speech Processing.
@inproceedings{roosta2021communication,
title = {Communication-Efficient Federated Learning for Neural Machine Translation},
author = {Roosta, Tanya and Passban, Peyman and Chadha, Ankit},
booktitle = {Proceedings of the 1st Workshop on Efficient Natural Language and Speech Processing (ENLSP), 35th Conference on Neural Information Processing Systems (NeurIPS 2021)},
year = {2021},
eprint = {2112.06135},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2112.06135}
}
Related work of mine
- Training Mixed-Domain Translation Models via Federated LearningNAACL 2022 (Main Conference)
- APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AIACL 2026 (Main Conference, Long Papers)
arXiv version. Cited 11 times as counted by Google Scholar on 2026-09-15.