WDRASS: A Web-scale Dataset for Document Retrieval and Answer Sentence Selection
Zeyu Zhang, Thuy Vu, Sunil Gandhi, Ankit Chadha, Alessandro Moschitti
Web-scale QACIKM 2022, pages 4707–4711 · University of Arizona; Amazon Alexa AI
WDRASS introduced a web-scale open-domain question answering dataset with 64,000 questions and 800,000+ labeled passages and sentences drawn from 30 million documents.
Research contributionIntroduced a web-scale dataset that connects document retrieval, passage reranking, and answer sentence selection using complete sentence-level answers rather than short answer matching.
Abstract
Open-Domain Question Answering (ODQA) systems generate answers from relevant text returned by search engines, e.g., lexical features-based such as BM25, or embeddings-based such as dense passage retrieval (DPR). Few datasets are available for this task: they mainly focus on QA systems based on machine reading (MR) approach, and show problematic evaluation, mostly based on uncontextualized short answer matching. In this paper, we present WDRASS, a dataset for ODQA based on answer sentence selection (AS2) models, which consider sentences as candidate answers for QA systems. WDRASS consists of ∼64k questions and 800k+ labeled passages and sentences extracted from 30M documents. We evaluate the dataset by training models on it and comparing with the same models trained on Google NQ. Our experiments show that WDRASS significantly improves the performance of retrieval and reranking models, thus boosting the accuracy of downstream QA tasks. We believe our dataset can produce significant impact in advancing IR research.
Why it matters
Most question answering datasets call a passage relevant if it merely contains a short answer string, which mislabels passages and sidesteps questions that need a whole sentence, such as "what makes tides?". WDRASS instead asks human annotators whether a passage holds a complete, direct answer sentence, drawn from Common Crawl rather than Wikipedia alone. Teams building search-then-answer pipelines get training data whose notion of relevance matches what a user actually reads, plus an end-to-end evaluation judged by people instead of string matching.
Key results
- On the WDRASS test set in the Close setting (a 50-passage index per question), hit-rate at K=1 for dense passage retrieval rises from 0.318 with encoders trained on Google Natural Questions to 0.414 (41.4%) with encoders trained on WDRASS, about 10 absolute points.
- In the Open setting (a shared index of roughly 250K passages), hit-rate at K=50 rises from 0.777 for NQ-trained DPR to 0.903 for WDRASS-trained DPR, described in the text as more than 90% coverage versus only 77%.
- In the end-to-end human evaluation over 1,296 web-sampled questions against a ~130M-passage index, answer accuracy (H@1) rises from 0.372 for configuration C0 (NQ-trained DPR) to 0.447 for C3 (WDRASS-trained DPR plus WDRASS-trained passage reranker), an improvement the paper reports as 7.48%.
- Null or negativeBroken down by question type, the full WDRASS configuration C3 improves over the NQ baseline C0 by 8.21% on 585 factoid questions (H@1 0.439 to 0.521) and by only 6.89% on 711 non-factoid questions (H@1 0.316 to 0.385), the type the dataset was designed to help most.
- Null or negativeThe retrieval advantage is not durable at large K: in the Close setting both NQ-trained and WDRASS-trained DPR tie at 1.000 by K=50, and in the Open setting the gap narrows from 18.2 points at K=10 to 17.4 at K=25 and to 0.971 versus 0.991 at K=5,000.
- Null or negativeOn factoid questions at M=5 answer candidates, WDRASS-trained retrieval alone (C2, 0.619) is slightly below the NQ retrieval plus reranker configuration (C1, 0.622), so swapping in WDRASS encoders does not beat simply adding a reranker at that operating point.
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper introduces WDRASS for web-scale open-domain question answering and studies answer sentence selection, dense passage retrieval, passage reranking, non-factoid question answering, Common Crawl, Natural Questions, information retrieval datasets, transformer rerankers, TANDA, hit-rate at K, and web-scale question answering.
What this paper does not show
The retrieval comparison against Natural Questions is measured only on our own test set, and the paper says why: that test set is not a random sample of NQ data. There is no reverse evaluation on NQ, and no comparison against BM25, ColBERT, or any dense retriever other than DPR. The gap also closes to about 2 points by K equals 5,000. Quote the tables rather than the prose: Section 4.3 and Table 4 disagree on the accuracy gains, and the corpus size is stated four different ways across the paper.
How to cite
Zeyu Zhang, Thuy Vu, Sunil Gandhi, Ankit Chadha, and Alessandro Moschitti. 2022. WDRASS: A Web-scale Dataset for Document Retrieval and Answer Sentence Selection. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management (CIKM '22), pages 4707–4711.
@inproceedings{zhang2022wdrass,
author = {Zhang, Zeyu and Vu, Thuy and Gandhi, Sunil and Chadha, Ankit and Moschitti, Alessandro},
title = {{WDRASS}: A Web-scale Dataset for Document Retrieval and Answer Sentence Selection},
booktitle = {Proceedings of the 31st ACM International Conference on Information and Knowledge Management (CIKM '22)},
year = {2022},
pages = {4707--4711},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
location = {Atlanta, GA, USA},
isbn = {978-1-4503-9236-5},
doi = {10.1145/3511808.3557678},
url = {https://doi.org/10.1145/3511808.3557678}
}
Related work of mine
- Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource LanguagesFindings of ACL 2023
- Efficient Fine-Tuning Large Language Models for Knowledge-Aware Response PlanningECML PKDD 2023 (Research Track), LNAI 14170
Published version, licensed CC BY 4.0 by the authors. Cited 19 times as counted by Google Scholar on 2026-09-15.