AC

BERTQA -- Attention on Steroids

Reading comprehensionarXiv preprint · Stanford University Preprint, not peer reviewed

PDF arXiv Code BibTeX

In one sentence

Adds two extra attention layers to BERT so a question and its passage look at each other directly, which lifts SQuAD 2.0 dev F1 from 74.15 to 77.03 on the base model.

Abstract

In this work, we extend the Bidirectional Encoder Representations from Transformers (BERT) with an emphasis on directed coattention to obtain an improved F1 performance on the SQUAD2.0 dataset. The Transformer architecture on which BERT is based places hierarchical global attention on the concatenation of the context and query. Our additions to the BERT architecture augment this attention with a more focused context to query (C2Q) and query to context (Q2C) attention via a set of modified Transformer encoder units. In addition, we explore adding convolution based feature extraction within the coattention architecture to add localized information to self-attention. We found that coattention significantly improves the no answer F1 by 4 points in the base and 1 point in the large architecture. After adding skip connections the no answer F1 improved further without causing an additional loss in has answer F1. The addition of localized feature extraction added to attention produced an overall dev F1 of 77.03 in the base architecture. We applied our findings to the large BERT model which contains twice as many layers and further used our own augmented version of the SQUAD 2.0 dataset created by back translation, which we have named SQUAD 2.Q. Finally, we performed hyperparameter tuning and ensembled our best models for a final F1/EM of 82.317/79.442 (Attention on Steroids, PCE Test Leaderboard).

Why it matters

SQuAD 2.0 asks a model to say "no answer" when the passage does not contain one, and plain BERT is weak at that. The extra directed attention here mostly buys accuracy on the unanswerable questions, raising no-answer F1 from 68.21 to 72.54 on the base model, while costing accuracy on answerable ones. If you are tuning a reader for cases where refusing to answer matters, that trade is the useful signal. The back-translated SQuAD 2.Q data and code are public.

Key results

Every number above appears in the paper. Results that were null or went the wrong way are included and marked.

BERT embeddings are split into query and context, then passed through 7 directed C2Q/Q2C coattention blocks.
Figure 1: Proposed C2Q and Q2C directed coattention architecture Taken from the paper.

Topics

This paper is about question answering, squad 2.0, bert fine-tuning, directed coattention, context-to-query attention, unanswerable questions, no-answer f1, back-translation data augmentation, squad 2.q, bidaf, qanet, and model ensembling.

What this paper does not show

This is a Stanford CS224n course project report from March 2019, posted to arXiv nine months later. It is not a peer-reviewed paper and should not be read as one. Single runs on one dataset, no seeds or error bars, and no comparison against the 2019 leaders such as XLNet or RoBERTa. The conclusion claims a 3.5 point ensemble gain; the table shows 2.84. The headline test numbers are from the internal course leaderboard, not the public SQuAD 2.0 one. Answerable-question F1 never beats the baseline, so the gains come almost entirely from predicting "no answer" better.

How to cite

Ankit Chadha and Rewa Sood. 2019. BERTQA -- Attention on Steroids. arXiv:1912.10435.

@misc{chadha2019bertqa,
  title         = {{BERTQA} - Attention on Steroids},
  author        = {Ankit Chadha and Rewa Sood},
  year          = {2019},
  eprint        = {1912.10435},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/1912.10435},
  note          = {Stanford CS224n Pre-trained Contextual Embeddings course project report, dated March 17, 2019}
}

Related work of mine

arXiv version. Cited 11 times as counted by Google Scholar on 2026-09-15.

← All publications