Document Summarization with Conformal Importance Guarantees

Bruce Kuwahara (Signal 1 AI) · Chen-Yuan Lin (University of Toronto) · Xiao Shi Huang (Signal 1 AI) · Kin Kwan Leung (Layer 6 AI) · Jullian Yapeter (Signal 1 AI) · Ilya Stanevich (Signal 1 AI) · Felipe Perez (Layer 6) · Jesse Cresswell (Layer 6 AI at TD)
automatic summarizationblack-box llmscalibration setconformal importance summarizationconformal predictioncontrollable automatic summarizationcritical applicationsdistribution-free coverage guaranteesextractive document summarizationimportance-preserving summary generationinformation coverage ratemodel-agnosticrecall ratessentence-level importance scoresuser-specified coverage

Automatic summarization systems have advanced rapidly with large language models (LLMs), yet they still lack reliable guarantees on inclusion of critical content in high-stakes domains like healthcare, law, and finance. In this work, we introduce Conformal Importance Summarization, the first framework for importance-preserving summary generation which uses conformal prediction to provide rigorous, distribution-free coverage guarantees. By calibrating thresholds on sentence-level importance scores, we enable extractive document summarization with user-specified coverage and recall rates over critical content. Our method is model-agnostic, requires only a small calibration set, and seamlessly integrates with existing black-box LLMs. Experiments on established summarization benchmarks demonstrate that Conformal Importance Summarization achieves the theoretically assured information coverage rate. Our work suggests that Conformal Importance Summarization can be combined with existing techniques to achieve reliable, controllable automatic summarization, paving the way for safer deployment of AI summarization tools in critical applications. Code is available at github.com/layer6ai-labs/conformal-importance-summarization.