Question-Context Alignment and Answer-Context Dependencies for Effective Answer Sentence Selection
Minh Van Nguyen, Kishan KC, Toan Nguyen, Thien Huu Nguyen, Ankit Chadha, Thuy Vu
Web-scale QAINTERSPEECH 2023, pages 3437–3441 · University of Oregon; Amazon Alexa AI
PDF arXiv DOI Publisher page BibTeX
Instead of pasting neighboring sentences onto the question, this model matches question words to context words and builds a small sentence graph, raising top-1 answer accuracy on WikiQA from 68.09 to 74.16.
Abstract
Answer sentence selection (AS2) in open-domain question answering finds answer for a question by ranking candidate sentences extracted from web documents. Recent work exploits answer context, i.e., sentences around a candidate, by incorporating them as additional input string to the Transformer models to improve the correctness scoring. In this paper, we propose to improve the candidate scoring by explicitly incorporating the dependencies between question-context and answer-context into the final representation of a candidate. Specifically, we use Optimal Transport to compute the question-based dependencies among sentences in the passage where the answer is extracted from. We then represent these dependencies as edges in a graph and use Graph Convolutional Network to derive the representation of a candidate, a node in the graph. Our proposed model achieves significant improvements on popular AS2 benchmarks, i.e., WikiQA and WDRASS, obtaining new state-of-the-art on all benchmarks.
Why it matters
Voice assistants must pick one answer sentence out of retrieved web pages, and the sentences on either side of a candidate often decide whether it is correct. The common fix is to glue those neighbors into one long input, which also drags in irrelevant text. This work instead keeps only the context words that match the question and models how the three sentences depend on each other, gaining about six points of top-1 accuracy on WikiQA. On the larger WDRASS set the gain is smaller and MAP drops.
Key results
- On the combined WikiQA development and test sets with a non-finetuned RoBERTa-base encoder, CASSIE raises Precision-at-1 from 68.09 to 74.16 over the LOCT context-concatenation baseline.
- On the same non-finetuned WikiQA setting, MAP rises from 79.00 with LOCT (and 75.00 with TANDA) to 83.29 with CASSIE.
- With an ASNQ-finetuned RoBERTa-base encoder on WikiQA the margin narrows: P@1 goes from 81.31 to 83.77 and MAP from 88.00 to 89.28 over LOCT.
- Null or negativeOn the WDRASS test set CASSIE loses to TANDA on MAP, dropping from 63.5 to 61.8, even though P@1 improves from 54.6 to 55.9 and MRR from 64.3 to 69.7.
- A joint reranking variant that reorders TANDA's candidates recovers MAP on WDRASS from 63.5 to 64.2 with P@1 55.9 and MRR 65.0, but its MRR sits well below the single-candidate CASSIE score of 69.7.
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper is about answer sentence selection, as2, wikiqa, wdrass, optimal transport, graph convolutional network, open-domain question answering, answer reranking, roberta, tanda, mutual information neural estimation, and contextual answer selection.
What this paper does not show
There is no ablation anywhere in this paper, so the three components, the alignment, the sentence graph, and the mutual-information term, are never separated. The abstract claims state of the art on all benchmarks and that does not hold: our main model's WDRASS MAP falls below TANDA, 61.8 against 63.5, and the claim only works if the joint-reranking variant is counted. Three seeds, averages only, no significance tests. No latency measurement, which matters for a task that is latency-sensitive in production.
How to cite
Minh Van Nguyen, Kishan KC, Toan Nguyen, Thien Huu Nguyen, Ankit Chadha, and Thuy Vu. 2023. Question-Context Alignment and Answer-Context Dependencies for Effective Answer Sentence Selection. In INTERSPEECH 2023, pages 3437–3441.
@inproceedings{nguyen2023question,
author = {Minh Van Nguyen and Kishan KC and Toan Nguyen and Thien Huu Nguyen and Ankit Chadha and Thuy Vu},
title = {Question-Context Alignment and Answer-Context Dependencies for Effective Answer Sentence Selection},
booktitle = {Proc. INTERSPEECH 2023},
pages = {3437--3441},
year = {2023},
address = {Dublin, Ireland},
publisher = {ISCA},
doi = {10.21437/Interspeech.2023-2240}
}
Related work of mine
- WDRASS: A Web-scale Dataset for Document Retrieval and Answer Sentence SelectionCIKM 2022
- Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource LanguagesFindings of ACL 2023
arXiv version. Cited 3 times as counted by Google Scholar on 2026-09-15.