Training Mixed-Domain Translation Models via Federated Learning
Peyman Passban, Tanya Roosta, Rahul Gupta, Ankit Chadha, Clement Chung
Federated learningNAACL 2022 (Main Conference), pages 2576–2586 · Amazon Alexa AI
PDF arXiv DOI Publisher page BibTeX
Five domain-specific translation models use federated learning to build a mixed-domain neural machine translation system without pooling their data, with a method for selecting impactful parameters to control communication bandwidth.
Research contributionIntroduced a federated approach to mixed-domain neural machine translation and a bandwidth-control method that selects impactful parameters during federated updates.
Abstract
Training mixed-domain translation models is a complex task that demands tailored architectures and costly data preparation techniques. In this work, we leverage federated learning (FL) in order to tackle the problem. Our investigation demonstrates that with slight modifications in the training process, neural machine translation (NMT) engines can be easily adapted when an FL-based aggregation is applied to fuse different domains. Experimental results also show that engines built via FL are able to perform on par with state-of-the-art baselines that rely on centralized training techniques. We evaluate our hypothesis in the presence of five datasets with different sizes, from different domains, to translate from German into English and discuss how FL and NMT can mutually benefit from each other. In addition to providing benchmarking results on the union of FL and NMT, we also propose a novel technique to dynamically control the communication bandwidth by selecting impactful parameters during FL updates. This is a significant achievement considering the large size of NMT engines that need to be exchanged between FL parties.
Why it matters
Teams often cannot pool translation data across customers, products, or hospitals, so each site ends up with a model that only handles its own text. This paper shows a simple weight-averaging setup reaches slightly better average quality than combining all the data centrally, using five German-to-English corpora. It also shows that sending only the parameters that barely moved between rounds cuts the upload in half and costs 1.55 BLEU, which gives a concrete knob for bandwidth-limited deployments.
Key results
- Across the five German-to-English test sets, the federated server averages 33.83 BLEU, which the paper reports as 1.83 points higher than the best centralized baseline (data combination).
- On the Ubuntu test set the federated server reaches 47.9 BLEU versus 35.61 for centralized data combination and 30.15 for the Ubuntu-only model.
- Null or negativeOn the OpenSubtitles test set the federated server scores only 19.17 BLEU, losing to centralized data combination at 21.82 and to the OpenSubtitles-only model at 23.58.
- A PHP client that runs one more round of local fine-tuning on top of the pushed server weights raises in-domain PHP BLEU from 37.32 to 45.07, but the paper notes it then loses its generalization to other domains.
- Null or negativeDynamic Pulling of the less active tensors (DPlc) halves the parameters exchanged per pull from 45,724,160 to 22,863,104, but average BLEU over the five domains falls from 33.83 to 32.28, a gap of 1.55 points against pulling every tensor.
- Null or negativeThe opposite selection rule, pulling the highly fluctuating tensors (DPgc), drops average BLEU to 24.49 and collapses PHP from 37.32 to 13.69, worse than random selection of the same number of parameters (35.33) and worse than DPlc (36.33).
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper studies federated learning for mixed-domain neural machine translation, domain adaptation without pooled data, communication-efficient federated learning, impactful parameter selection, FedAvg, cross-silo federated learning, Transformer NMT, German-to-English translation, BLEU, OpenSubtitles, and bandwidth reduction.
What this paper does not show
Preliminary work, and the paper says so. One language direction, one aggregation rule, single runs, BLEU only, no human evaluation. The federated server loses to centralized data combination on OpenSubtitles, 19.17 against 21.82. Both bandwidth-saving variants cost quality, 32.28 and 24.49 average against 33.83. The main comparator is our own group's prior paper, sharing three authors, so the headline head-to-head is largely against ourselves. No ablation shows how much of the gain comes simply from starting every client at the strong WMT model.
How to cite
Peyman Passban, Tanya Roosta, Rahul Gupta, Ankit Chadha, and Clement Chung. 2022. Training Mixed-Domain Translation Models via Federated Learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2576–2586.
@inproceedings{passban2022training,
title = {Training Mixed-Domain Translation Models via Federated Learning},
author = {Passban, Peyman and Roosta, Tanya and Gupta, Rahul and Chadha, Ankit and Chung, Clement},
booktitle = {Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies},
month = jul,
year = {2022},
address = {Seattle, United States},
publisher = {Association for Computational Linguistics},
pages = {2576--2586}
}
Related work of mine
- Communication-Efficient Federated Learning for Neural Machine TranslationNeurIPS 2021 ENLSP Workshop
- Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource LanguagesFindings of ACL 2023
arXiv version. Cited 27 times as counted by Google Scholar on 2026-09-15.