BERTQA -- Attention on Steroids
Ankit Chadha, Rewa Sood
Reading comprehensionarXiv preprint · Stanford University Preprint, not peer reviewed
Adds two extra attention layers to BERT so a question and its passage look at each other directly, which lifts SQuAD 2.0 dev F1 from 74.15 to 77.03 on the base model.
Abstract
In this work, we extend the Bidirectional Encoder Representations from Transformers (BERT) with an emphasis on directed coattention to obtain an improved F1 performance on the SQUAD2.0 dataset. The Transformer architecture on which BERT is based places hierarchical global attention on the concatenation of the context and query. Our additions to the BERT architecture augment this attention with a more focused context to query (C2Q) and query to context (Q2C) attention via a set of modified Transformer encoder units. In addition, we explore adding convolution based feature extraction within the coattention architecture to add localized information to self-attention. We found that coattention significantly improves the no answer F1 by 4 points in the base and 1 point in the large architecture. After adding skip connections the no answer F1 improved further without causing an additional loss in has answer F1. The addition of localized feature extraction added to attention produced an overall dev F1 of 77.03 in the base architecture. We applied our findings to the large BERT model which contains twice as many layers and further used our own augmented version of the SQUAD 2.0 dataset created by back translation, which we have named SQUAD 2.Q. Finally, we performed hyperparameter tuning and ensembled our best models for a final F1/EM of 82.317/79.442 (Attention on Steroids, PCE Test Leaderboard).
Why it matters
SQuAD 2.0 asks a model to say "no answer" when the passage does not contain one, and plain BERT is weak at that. The extra directed attention here mostly buys accuracy on the unanswerable questions, raising no-answer F1 from 68.21 to 72.54 on the base model, while costing accuracy on answerable ones. If you are tuning a reader for cases where refusing to answer matters, that trade is the useful signal. The back-translated SQuAD 2.Q data and code are public.
Key results
- On the SQuAD 2.0 dev set, adding localized convolutional feature extraction inside the coattention blocks raises overall F1 from 74.15 to 77.03 and EM from 71.09 to 74.37 over the BERT base baseline.
- Directed C2Q/Q2C coattention alone raises no-answer F1 on SQuAD 2.0 dev from 68.21 to 72.54 over BERT base, and the best base configuration reaches 73.83.
- Null or negativeThe gain is a trade, not a free win: has-answer F1 on SQuAD 2.0 dev falls from 80.62 for BERT base to 76.30 with C2Q/Q2C coattention, and the best base configuration only recovers to 78.47, still below the baseline.
- On BERT large, the three-model ensemble raises SQuAD 2.0 dev F1 from 80.58 to 83.42 and EM from 77.74 to 80.53, with no-answer F1 going from 77.08 to 85.56 but has-answer F1 dropping from 84.39 to 81.09; the best single model reaches 82.08 F1.
- Null or negativeIn the per-question-type error analysis, the model is worse than the baseline on "Which" questions, getting 69 wrong versus the baseline's 64, while on "What" questions it makes 30 fewer errors than the baseline's 776.
- The final ensemble reports F1/EM of 82.317/79.442 on the CS224n Pre-trained Contextual Embeddings test leaderboard, with roughly 300 GPU hours of total training.
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper is about question answering, squad 2.0, bert fine-tuning, directed coattention, context-to-query attention, unanswerable questions, no-answer f1, back-translation data augmentation, squad 2.q, bidaf, qanet, and model ensembling.
What this paper does not show
This is a Stanford CS224n course project report from March 2019, posted to arXiv nine months later. It is not a peer-reviewed paper and should not be read as one. Single runs on one dataset, no seeds or error bars, and no comparison against the 2019 leaders such as XLNet or RoBERTa. The conclusion claims a 3.5 point ensemble gain; the table shows 2.84. The headline test numbers are from the internal course leaderboard, not the public SQuAD 2.0 one. Answerable-question F1 never beats the baseline, so the gains come almost entirely from predicting "no answer" better.
How to cite
Ankit Chadha and Rewa Sood. 2019. BERTQA -- Attention on Steroids. arXiv:1912.10435.
@misc{chadha2019bertqa,
title = {{BERTQA} - Attention on Steroids},
author = {Ankit Chadha and Rewa Sood},
year = {2019},
eprint = {1912.10435},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/1912.10435},
note = {Stanford CS224n Pre-trained Contextual Embeddings course project report, dated March 17, 2019}
}
Related work of mine
- Deep Reinforced Self-Attention Masks for Abstractive Summarization (DR.SAS)arXiv preprint
- Question-Context Alignment and Answer-Context Dependencies for Effective Answer Sentence SelectionINTERSPEECH 2023
arXiv version. Cited 11 times as counted by Google Scholar on 2026-09-15.