Deep Reinforced Self-Attention Masks for Abstractive Summarization (DR.SAS)
Ankit Chadha, Mohamed Masoud
SummarizationarXiv preprint · Stanford University Preprint, not peer reviewed
Adds a reinforcement-learning agent that decides which input words a summarizer should ignore, raising the word-overlap score ROUGE-1 on CNN/Daily Mail news from 40.79 to 41.89 against fine-tuned UniLM.
Abstract
We present a novel architectural scheme to tackle the abstractive summarization problem based on the CNN/DM dataset [1] which fuses Reinforcement Learning (RL) with UniLM, which is a pre-trained Deep Learning Model, to solve various natural language tasks. We have tested the limits of learning fine-grained attention in Transfomers [3] to improve the summarization quality. UniLM applies attention to the entire token space in a global fashion. We propose DR.SAS which applies the Actor Critic (AC) algorithm [2] to learn a dynamic self-attention distribution over the tokens to reduce redundancy and generate factual and coherent summaries to improve the quality of summarization. After performing hyperparameter tuning, we achieved better ROUGE results compared to the baseline. Our model tends to be more extractive/factual yet coherent in details because of optimization over ROUGE rewards. We present detailed error analysis with examples of strengths and limitations of our model. Our codebase will be publicly available on our github [20]
Why it matters
Most summarization systems let the model attend to every word in the article. This work asks whether an agent can learn, per token, which words to hide, and reports a small gain: about one ROUGE-1 point over the same UniLM model fine-tuned normally. The practical lesson is mostly cautionary. Because the reward is ROUGE, an n-gram overlap score, the learned masks push the summaries toward copying source text, so gains in the metric can come with a shift toward extractive output rather than better writing.
Key results
- Raises ROUGE-1 F1 on the CNN/Daily Mail test set from 40.79 to 41.89 over the fine-tuned UniLM baseline (Table 1).
- Raises ROUGE-L F1 on CNN/Daily Mail from 38.10 to 39.28 over the fine-tuned UniLM baseline, the largest of the three ROUGE gains at +1.18.
- Null or negativeROUGE-2 F1 moves only from 19.01 to 19.22, a 0.21 point change that the authors themselves describe as a "marginal improvement" over fine-tuned UniLM, essentially a tie.
- Against the TextRank extractive baseline, DR.SAS lifts ROUGE-1 from 21.0 to 41.89 and ROUGE-L from 19.2 to 39.28, but almost all of that gap comes from UniLM itself rather than from the reinforcement learning.
- The only reported ablation-style number is a hyperparameter change: extending the maximum sequence length from 384 to 512 tokens gave a 0.1 ROUGE score improvement.
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper is about abstractive summarization, reinforcement learning for nlp, actor-critic, advantage actor-critic (a2c), self-attention masks, unilm, cnn/daily mail, rouge, policy gradient, transformer attention, textrank, and text summarization.
What this paper does not show
This started as a Stanford CS221 course project and was never peer reviewed. The gain is about one ROUGE point and the paper calls it marginal. Training was cut to 6 of 30 epochs at a quarter of the intended batch size, so the baseline is under-trained and 41.89 is not comparable to published CNN/DailyMail numbers. No human evaluation, no seeds, no significance tests. The claim that the summaries are more factual rests on two hand-picked articles with no factuality metric.
How to cite
Ankit Chadha and Mohamed Masoud. 2019. Deep Reinforced Self-Attention Masks for Abstractive Summarization (DR.SAS). arXiv:2001.00009.
@misc{chadha2019deep,
title = {Deep Reinforced Self-Attention Masks for Abstractive Summarization (DR.SAS)},
author = {Chadha, Ankit and Masoud, Mohamed},
year = {2019},
month = dec,
eprint = {2001.00009},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2001.00009},
note = {arXiv preprint arXiv:2001.00009}
}
Related work of mine
- BERTQA -- Attention on SteroidsarXiv preprint
- ACM -- Attribute Conditioning for Abstractive Multi Document SummarizationarXiv preprint
arXiv version.