ACM -- Attribute Conditioning for Abstractive Multi Document Summarization
Aiswarya Sankar, Ankit Chadha
SummarizationarXiv preprint · Amazon Alexa AI Preprint, not peer reviewed
Adds a sentiment and polarity classifier to a news multi-document summarizer so it sticks to one viewpoint, raising the word-overlap ROUGE-1 score on MultiNews from 42.99 to 50.12 versus GraphSum.
Abstract
Abstractive multi document summarization has evolved as a task through the basic sequence to sequence approaches to transformer and graph based techniques. Each of these approaches has primarily focused on the issues of multi document information synthesis and attention based approaches to extract salient information. A challenge that arises with multi document summarization which is not prevalent in single document summarization is the need to effectively summarize multiple documents that might have conflicting polarity, sentiment or subjective information about a given topic. In this paper we propose ACM, attribute conditioned multi document summarization,a model that incorporates attribute conditioning modules in order to decouple conflicting information by conditioning for a certain attribute in the output summary. This approach shows strong gains in ROUGE score over baseline multi document summarization approaches and shows gains in fluency, informativeness and reduction in repetitiveness as shown through a human annotation analysis study.
Why it matters
When you summarize several news stories about the same event, the sources often disagree, and a summarizer trained only to compress text can blend opposing claims into one incoherent paragraph. This paper attaches a small sentiment and polarity classifier at three points, graph edge weights, training logits, and beam search, so the summarizer favors text that shares one viewpoint. The classifier is separate from the summarizer, so it can be bolted onto BART, BART with Longformer attention, or GraphSum without retraining the base model.
Key results
- On MultiNews, ACM with the polarity attribute raises ROUGE-1 from 42.99 to 50.12, ROUGE-2 from 27.83 to 28.12, and ROUGE-L from 36.97 to 38.19 over the GraphSum baseline (Table 3).
- In a human study over 182 randomly selected MultiNews output summaries with 5 Amazon Mechanical Turk evaluators, ACM improves fluency from 3.48 to 3.91 and informativeness from 3.24 to 3.86, and lowers repetitiveness from 2.3 to 1.74, on a 1 to 5 scale, against GraphSum (Table 4).
- In the ablation on MultiNews, adding attribute future discriminators to BART+Longformer raises ROUGE-1 from 49.03 to 49.83 and ROUGE-2 from 19.04 to 19.85 over the plain BART+Longformer baseline (Table 5).
- Scoring generated summaries with the trained polarity classifier, attribute future discriminators reach mean 0.891 with standard deviation 0.103, versus 0.76 with 0.288 for graph conditional weighting and 0.82 with 0.274 for conditional training (Table 6).
- Null or negativeThe two attribute choices are effectively a tie: the paper states the sentiment and polarity modules "perform on par", with ACM w/ Sentiment at ROUGE-1 50.05 versus ACM w/ Polarity at 50.12, a gap of 0.07 with no significance test reported.
- Null or negativeThe ablation table contradicts the headline table: a single component, GraphSum w/ conditional training, reaches ROUGE-2 28.22 and ROUGE-L 38.61, above the full ACM w/ Polarity system's 28.12 and 38.19, so the complete model is not the paper's best number on two of three metrics; on BART, ROUGE-L barely moves at all, from 20.84 to 20.94 for both conditional training and attribute future discriminators, below the 21.01 from graph conditional weighting alone.
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper is about multi-document summarization, abstractive summarization, attribute conditioning, controllable text generation, conditional language modeling, multinews, graphsum, bart longformer, xlnet sentiment classifier, rouge evaluation, future discriminators beam search, and viewpoint consistency.
What this paper does not show
Preprint, not peer reviewed, and its conclusion claims a new state of the art that no reviewer ever tested. No significance tests or seeds on any ROUGE number, so gains of 0.07 to 0.8 points cannot be separated from run-to-run noise. One dataset only. The human evaluation used 5 annotators over 182 summaries and the 78 percent figure is raw agreement, not chance-corrected. Table 5's best single-component variant beats the full system on two of three metrics, which the paper does not explain. The acronym in the title has nothing to do with the ACM.
How to cite
Aiswarya Sankar and Ankit Chadha. 2022. ACM -- Attribute Conditioning for Abstractive Multi Document Summarization. arXiv:2205.03978.
@misc{sankar2022acm,
title = {ACM -- Attribute Conditioning for Abstractive Multi Document Summarization},
author = {Sankar, Aiswarya and Chadha, Ankit},
year = {2022},
month = may,
eprint = {2205.03978},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2205.03978},
note = {arXiv preprint arXiv:2205.03978}
}
Related work of mine
- Controlled Text Generation with Hidden Representation TransformationsFindings of ACL 2023
- Deep Reinforced Self-Attention Masks for Abstractive Summarization (DR.SAS)arXiv preprint
arXiv version.