Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages
Shivanshu Gupta, Yoshitomo Matsubara, Ankit Chadha, Alessandro Moschitti
Web-scale QAFindings of ACL 2023, pages 14078–14092 · University of California Irvine; Amazon Alexa AI
PDF arXiv DOI Publisher page TyDi-AS2 (70,000+ questions, 8 languages) Xtr-WikiQA (9 languages) BibTeX
The paper trains answer-ranking models for languages with no labeled data by copying the soft scores of a strong English model, lifting Arabic accuracy on translated WikiQA from 64.2 to 76.3.
Abstract
While impressive performance has been achieved on the task of Answer Sentence Selection (AS2) for English, the same does not hold for languages that lack large labeled datasets. In this work, we propose Cross-Lingual Knowledge Distillation (CLKD) from a strong English AS2 teacher as a method to train AS2 models for low-resource languages in the tasks without the need of labeled data for the target language. To evaluate our method, we introduce 1) Xtr-WikiQA, a translation-based WikiQA dataset for 9 additional languages, and 2) TyDi-AS2, a multilingual AS2 dataset with over 70K questions spanning 8 typologically diverse languages. We conduct extensive experiments on Xtr-WikiQA and TyDi-AS2 with multiple teachers, diverse monolingual and multilingual pretrained language models (PLMs) as students, and both monolingual and multilingual training. The results demonstrate that CLKD either outperforms or rivals even supervised fine-tuning with the same amount of labeled data and a combination of machine translation and the teacher model. Our method can potentially enable stronger AS2 models for low-resource languages, while TyDi-AS2 can serve as the largest multilingual AS2 dataset for further studies in the research community.
Why it matters
Picking the sentence that answers a question works well in English because there is labeled data. For most languages there is not, and hand-labeling is expensive because one question can have hundreds of candidate sentences. This work shows a cheaper route: run an English model over machine-translated text and train the target-language model to copy its confidence scores. The authors also release two multilingual datasets, so others can test the idea rather than take it on trust.
Key results
- On Xtr-WikiQA with an XLM-R-Large student trained on a single target language, Arabic P@1 rises from 64.2 with supervised fine-tuning on gold labels to 76.3 with CLKD from either the ELECTRA-Large or RoBERTa-Large English teacher.
- On Xtr-WikiQA with a smaller mBERT student trained on a single language, Italian P@1 rises from 60.8 with supervised fine-tuning to 75.7 with CLKD using the RoBERTa-Large English teacher.
- On the Xtr-TyDi-AS2 translationese data with an XLM-R-Large student trained on all languages, Bengali P@1 goes from 62.8 with supervised fine-tuning to 67.2 with CLKD, above the 63.9 of the machine-translation-plus-English-teacher pipeline.
- On original-language TyDi-AS2 with an XLM-R-Large student trained on all languages, Korean P@1 goes from 80.5 with supervised fine-tuning to 83.4 with CLKD, above the machine-translation-plus-English-teacher pipeline at 77.8.
- Null or negativeIn that same strongest configuration, CLKD actually loses to supervised fine-tuning in five of seven languages: Japanese falls from 62.9 to 58.0, Finnish from 72.3 to 68.8, Bengali from 70.0 to 68.0, Swahili from 88.6 to 87.1, and Russian from 68.4 to 68.3.
- Null or negativeNo student ever catches the English teacher on Xtr-WikiQA: the best student result is 85.1 P@1 (Portuguese, XLM-R-Large, all-language CLKD with RoBERTa-Large) against teacher scores of 87.7 for ELECTRA-Large and 91.8 for RoBERTa-Large.
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper is about answer sentence selection, cross-lingual knowledge distillation, low-resource languages, knowledge distillation, multilingual question answering, wikiqa, tydi qa, xtr-wikiqa, tydi-as2, xlm-roberta, mbert, and machine translation for question answering.
What this paper does not show
The paper says the method beats supervised fine-tuning for all target languages. That is too strong. German and Swahili are losses in the tables, and with the largest student trained on all languages on original TyDi-AS2 data it loses in five of seven. No student ever reaches the English teacher. Three seeds were run but only averages are reported, so the sub-point gaps are not shown to be real. Both datasets were built with Amazon Translate by a team that is mostly Amazon, so the machine-translation side is not vendor-neutral.
How to cite
Shivanshu Gupta, Yoshitomo Matsubara, Ankit Chadha, and Alessandro Moschitti. 2023. Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages. In Findings of the Association for Computational Linguistics: ACL 2023, pages 14078–14092.
@inproceedings{gupta2023cross,
title = {Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages},
author = {Gupta, Shivanshu and Matsubara, Yoshitomo and Chadha, Ankit and Moschitti, Alessandro},
editor = {Rogers, Anna and Boyd-Graber, Jordan and Okazaki, Naoaki},
booktitle = {Findings of the Association for Computational Linguistics: ACL 2023},
month = jul,
year = {2023},
address = {Toronto, Canada},
publisher = {Association for Computational Linguistics},
pages = {14078--14092},
url = {https://aclanthology.org/2023.findings-acl.885/},
doi = {10.18653/v1/2023.findings-acl.885}
}
Related work of mine
- Question-Context Alignment and Answer-Context Dependencies for Effective Answer Sentence SelectionINTERSPEECH 2023
- WDRASS: A Web-scale Dataset for Document Retrieval and Answer Sentence SelectionCIKM 2022
arXiv version. Cited 15 times as counted by Google Scholar on 2026-09-15.