AC

WDRASS: A Web-scale Dataset for Document Retrieval and Answer Sentence Selection

Web-scale QACIKM 2022, pages 4707–4711 · University of Arizona; Amazon Alexa AI

PDF DOI Publisher page BibTeX

In one sentence

WDRASS introduced a web-scale open-domain question answering dataset with 64,000 questions and 800,000+ labeled passages and sentences drawn from 30 million documents.

Research contributionIntroduced a web-scale dataset that connects document retrieval, passage reranking, and answer sentence selection using complete sentence-level answers rather than short answer matching.

Abstract

Open-Domain Question Answering (ODQA) systems generate answers from relevant text returned by search engines, e.g., lexical features-based such as BM25, or embeddings-based such as dense passage retrieval (DPR). Few datasets are available for this task: they mainly focus on QA systems based on machine reading (MR) approach, and show problematic evaluation, mostly based on uncontextualized short answer matching. In this paper, we present WDRASS, a dataset for ODQA based on answer sentence selection (AS2) models, which consider sentences as candidate answers for QA systems. WDRASS consists of ∼64k questions and 800k+ labeled passages and sentences extracted from 30M documents. We evaluate the dataset by training models on it and comparing with the same models trained on Google NQ. Our experiments show that WDRASS significantly improves the performance of retrieval and reranking models, thus boosting the accuracy of downstream QA tasks. We believe our dataset can produce significant impact in advancing IR research.

Why it matters

Most question answering datasets call a passage relevant if it merely contains a short answer string, which mislabels passages and sidesteps questions that need a whole sentence, such as "what makes tides?". WDRASS instead asks human annotators whether a passage holds a complete, direct answer sentence, drawn from Common Crawl rather than Wikipedia alone. Teams building search-then-answer pipelines get training data whose notion of relevance matches what a user actually reads, plus an end-to-end evaluation judged by people instead of string matching.

Key results

Every number above appears in the paper. Results that were null or went the wrong way are included and marked.

Web index retrieves candidates, annotators label answer sentences, then those labels propagate back to passages and docs.
Figure 1: ODQA pipeline for generating question-candidate pairs, which are labelled by expert annotators (upper/orange pipeline). The annotation is propagated back to the embedding passages and documents (lower/blue pipeline). Taken from the paper.

Topics

This paper introduces WDRASS for web-scale open-domain question answering and studies answer sentence selection, dense passage retrieval, passage reranking, non-factoid question answering, Common Crawl, Natural Questions, information retrieval datasets, transformer rerankers, TANDA, hit-rate at K, and web-scale question answering.

What this paper does not show

The retrieval comparison against Natural Questions is measured only on our own test set, and the paper says why: that test set is not a random sample of NQ data. There is no reverse evaluation on NQ, and no comparison against BM25, ColBERT, or any dense retriever other than DPR. The gap also closes to about 2 points by K equals 5,000. Quote the tables rather than the prose: Section 4.3 and Table 4 disagree on the accuracy gains, and the corpus size is stated four different ways across the paper.

How to cite

Zeyu Zhang, Thuy Vu, Sunil Gandhi, Ankit Chadha, and Alessandro Moschitti. 2022. WDRASS: A Web-scale Dataset for Document Retrieval and Answer Sentence Selection. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management (CIKM '22), pages 4707–4711.

@inproceedings{zhang2022wdrass,
  author    = {Zhang, Zeyu and Vu, Thuy and Gandhi, Sunil and Chadha, Ankit and Moschitti, Alessandro},
  title     = {{WDRASS}: A Web-scale Dataset for Document Retrieval and Answer Sentence Selection},
  booktitle = {Proceedings of the 31st ACM International Conference on Information and Knowledge Management (CIKM '22)},
  year      = {2022},
  pages     = {4707--4711},
  publisher = {Association for Computing Machinery},
  address    = {New York, NY, USA},
  location  = {Atlanta, GA, USA},
  isbn      = {978-1-4503-9236-5},
  doi       = {10.1145/3511808.3557678},
  url       = {https://doi.org/10.1145/3511808.3557678}
}

Related work of mine

Published version, licensed CC BY 4.0 by the authors. Cited 19 times as counted by Google Scholar on 2026-09-15.

← All publications