AC

Deep Reinforced Self-Attention Masks for Abstractive Summarization (DR.SAS)

SummarizationarXiv preprint · Stanford University Preprint, not peer reviewed

PDF arXiv BibTeX

In one sentence

Adds a reinforcement-learning agent that decides which input words a summarizer should ignore, raising the word-overlap score ROUGE-1 on CNN/Daily Mail news from 40.79 to 41.89 against fine-tuned UniLM.

Abstract

We present a novel architectural scheme to tackle the abstractive summarization problem based on the CNN/DM dataset [1] which fuses Reinforcement Learning (RL) with UniLM, which is a pre-trained Deep Learning Model, to solve various natural language tasks. We have tested the limits of learning fine-grained attention in Transfomers [3] to improve the summarization quality. UniLM applies attention to the entire token space in a global fashion. We propose DR.SAS which applies the Actor Critic (AC) algorithm [2] to learn a dynamic self-attention distribution over the tokens to reduce redundancy and generate factual and coherent summaries to improve the quality of summarization. After performing hyperparameter tuning, we achieved better ROUGE results compared to the baseline. Our model tends to be more extractive/factual yet coherent in details because of optimization over ROUGE rewards. We present detailed error analysis with examples of strengths and limitations of our model. Our codebase will be publicly available on our github [20]

Why it matters

Most summarization systems let the model attend to every word in the article. This work asks whether an agent can learn, per token, which words to hide, and reports a small gain: about one ROUGE-1 point over the same UniLM model fine-tuned normally. The practical lesson is mostly cautionary. Because the reward is ROUGE, an n-gram overlap score, the learned masks push the summaries toward copying source text, so gains in the metric can come with a shift toward extractive output rather than better writing.

Key results

Every number above appears in the paper. Results that were null or went the wrong way are included and marked.

Actor-critic net reads UniLM attention scores, samples 0 or -10000 masks, and feeds them back into the encoder.
Figure 3. Encoder with AC Policy Learning for Self-Attention masks Taken from the paper.

Topics

This paper is about abstractive summarization, reinforcement learning for nlp, actor-critic, advantage actor-critic (a2c), self-attention masks, unilm, cnn/daily mail, rouge, policy gradient, transformer attention, textrank, and text summarization.

What this paper does not show

This started as a Stanford CS221 course project and was never peer reviewed. The gain is about one ROUGE point and the paper calls it marginal. Training was cut to 6 of 30 epochs at a quarter of the intended batch size, so the baseline is under-trained and 41.89 is not comparable to published CNN/DailyMail numbers. No human evaluation, no seeds, no significance tests. The claim that the summaries are more factual rests on two hand-picked articles with no factuality metric.

How to cite

Ankit Chadha and Mohamed Masoud. 2019. Deep Reinforced Self-Attention Masks for Abstractive Summarization (DR.SAS). arXiv:2001.00009.

@misc{chadha2019deep,
  title         = {Deep Reinforced Self-Attention Masks for Abstractive Summarization (DR.SAS)},
  author        = {Chadha, Ankit and Masoud, Mohamed},
  year          = {2019},
  month         = dec,
  eprint        = {2001.00009},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2001.00009},
  note          = {arXiv preprint arXiv:2001.00009}
}

Related work of mine

arXiv version.

← All publications