Do Multi-Document Summarization Models Synthesize?

Jay DeYoung; Stephanie C. Martinez; Iain J. Marshall; Byron C. Wallace

Vol. 12 (2024)

TACL approved

Do Multi-Document Summarization Models Synthesize?

Published 2024-09-07

Jay DeYoung
Stephanie C. Martinez
Iain J. Marshall
Byron C. Wallace

Jay DeYoung
Northeastern University

Stephanie C. Martinez
Northeastern University

Iain J. Marshall
King's College London

Byron C. Wallace
Northeastern University

Abstract

synopsis should accurately synthesize inputs with respect to a key aspect, e.g., a synopsis of film reviews written about a particular movie should reflect the average critic consensus. As a more consequential example, narrative summaries that accompany biomedical systematic reviews of clinical trial results should accurately summarize the potentially conflicting results from individual trials. In this paper we ask: To what extent do modern multi-document summarization models implicitly perform this sort of synthesis? We run experiments over opinion and evidence synthesis datasets using a suite of summarization models, from fine-tuned transformers to GPT-4. We find that existing models partially perform synthesis, but imperfectly: even the best performing models are oversensitive to changes in input ordering and under-sensitive to changes in input compositions (e.g., ratio of positive to negative reviews). We propose a simple, general, effective method for improving model synthesis capabilities by generating an explicitly diverse set of candidate outputs, and then selecting from these the string best aligned with the expected aggregate measure for the inputs, or abstaining when the model produces no good candidate.

Article at MIT Press Presented at ACL 2024