AC

ACM -- Attribute Conditioning for Abstractive Multi Document Summarization

SummarizationarXiv preprint · Amazon Alexa AI Preprint, not peer reviewed

PDF arXiv BibTeX

In one sentence

Adds a sentiment and polarity classifier to a news multi-document summarizer so it sticks to one viewpoint, raising the word-overlap ROUGE-1 score on MultiNews from 42.99 to 50.12 versus GraphSum.

Abstract

Abstractive multi document summarization has evolved as a task through the basic sequence to sequence approaches to transformer and graph based techniques. Each of these approaches has primarily focused on the issues of multi document information synthesis and attention based approaches to extract salient information. A challenge that arises with multi document summarization which is not prevalent in single document summarization is the need to effectively summarize multiple documents that might have conflicting polarity, sentiment or subjective information about a given topic. In this paper we propose ACM, attribute conditioned multi document summarization,a model that incorporates attribute conditioning modules in order to decouple conflicting information by conditioning for a certain attribute in the output summary. This approach shows strong gains in ROUGE score over baseline multi document summarization approaches and shows gains in fluency, informativeness and reduction in repetitiveness as shown through a human annotation analysis study.

Why it matters

When you summarize several news stories about the same event, the sources often disagree, and a summarizer trained only to compress text can blend opposing claims into one incoherent paragraph. This paper attaches a small sentiment and polarity classifier at three points, graph edge weights, training logits, and beam search, so the summarizer favors text that shares one viewpoint. The classifier is separate from the summarizer, so it can be bolted onto BART, BART with Longformer attention, or GraphSum without retraining the base model.

Key results

Every number above appears in the paper. Results that were null or went the wrong way are included and marked.

ACM inserts attribute conditioning modules into paragraph encoding and decoder output so summaries keep one viewpoint.
Figure 1: Attribute conditioned multi document summarization (ACM) model diagram. At each stage of the summarization process, ACM uses the attribute conditioning module to preserve viewpoint consistency in the output summary. As attribute future discriminators is a component of model evaluation, it is included separately in figure 3. Taken from the paper.

Topics

This paper is about multi-document summarization, abstractive summarization, attribute conditioning, controllable text generation, conditional language modeling, multinews, graphsum, bart longformer, xlnet sentiment classifier, rouge evaluation, future discriminators beam search, and viewpoint consistency.

What this paper does not show

Preprint, not peer reviewed, and its conclusion claims a new state of the art that no reviewer ever tested. No significance tests or seeds on any ROUGE number, so gains of 0.07 to 0.8 points cannot be separated from run-to-run noise. One dataset only. The human evaluation used 5 annotators over 182 summaries and the 78 percent figure is raw agreement, not chance-corrected. Table 5's best single-component variant beats the full system on two of three metrics, which the paper does not explain. The acronym in the title has nothing to do with the ACM.

How to cite

Aiswarya Sankar and Ankit Chadha. 2022. ACM -- Attribute Conditioning for Abstractive Multi Document Summarization. arXiv:2205.03978.

@misc{sankar2022acm,
  title         = {ACM -- Attribute Conditioning for Abstractive Multi Document Summarization},
  author        = {Sankar, Aiswarya and Chadha, Ankit},
  year          = {2022},
  month         = may,
  eprint        = {2205.03978},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2205.03978},
  note          = {arXiv preprint arXiv:2205.03978}
}

Related work of mine

arXiv version.

← All publications