Abstract
Regulatory documents are extreme in length, structurally heterogeneous, and linguistically technical, yet they must be summarised with high precision because omissions or distortions can materially change compliance meaning. We introduce ReguSum, a dataset of U.S. agency regulatory documents from the Securities and Exchange Commission (SEC) and the Internal Revenue Service (IRS) (2020–2024), paired with agency-provided abstracts that we treat as gold reference summaries. Compared to widely used long-document summarisation benchmarks, ReguSum operates in a long-input with extreme compression ratios exceeding 200:1 that stresses content selection and length budgeting under practical context limits.
Using ReguSum, we evaluate three families of methods under a unif ied pipeline: (i) internal input structuring with length control (wholedocument truncation, dynamic chunking, and section-aware hierarchical summarisation); (ii) seven retrieval-augmented summarisation variants that differ in query formulation and context construction; and (iii) clustering-based semantic chunking with HDBSCAN, evaluated with global versus document-specific parameterisation and combined with hierarchical decoding. Overall, retrieval-augmented variants do not surpass strong internal-structuring baselines in this setting, whereas semantic chunking is most effective when paired with hierarchical summarisation.
Using ReguSum, we evaluate three families of methods under a unif ied pipeline: (i) internal input structuring with length control (wholedocument truncation, dynamic chunking, and section-aware hierarchical summarisation); (ii) seven retrieval-augmented summarisation variants that differ in query formulation and context construction; and (iii) clustering-based semantic chunking with HDBSCAN, evaluated with global versus document-specific parameterisation and combined with hierarchical decoding. Overall, retrieval-augmented variants do not surpass strong internal-structuring baselines in this setting, whereas semantic chunking is most effective when paired with hierarchical summarisation.
| Original language | English |
|---|---|
| Title of host publication | Natural Language Processing and Information Systems |
| Subtitle of host publication | 31st International Conference on Applications of Natural Language to Information Systems, NLDB 2026, Trondheim, Norway, June 17–19, 2026, Proceedings |
| Editors | Elena Cabrio, Eric Monteiro |
| Place of Publication | Cham |
| Publisher | Springer |
| Pages | 3-17 |
| Number of pages | 15 |
| Edition | 1st |
| ISBN (Electronic) | 9783032295323 |
| ISBN (Print) | 9783032295316 |
| DOIs | |
| Publication status | Published - 4 Jul 2026 |
| Event | The 31st Annual International Conference on Natural Language & Information Systems - Norwegian University of Science and Technology, Trondheim, Norway Duration: 17 Jun 2026 → 19 Jun 2026 https://www.ntnu.edu/nldb2026 |
Publication series
| Name | Lecture Notes in Computer Science |
|---|---|
| Publisher | Springer |
| Volume | 16696 |
| ISSN (Print) | 0302-9743 |
| ISSN (Electronic) | 1611-3349 |
Conference
| Conference | The 31st Annual International Conference on Natural Language & Information Systems |
|---|---|
| Abbreviated title | NLDB 2026 |
| Country/Territory | Norway |
| City | Trondheim |
| Period | 17/06/26 → 19/06/26 |
| Internet address |
Keywords
- Regulatorydocumentsummarisation
- Long-document summarisation
- Extreme compression
- Retrieval-augmented summarisation
- Hierarchical summarisation
- Semantic chunking
Fingerprint
Dive into the research topics of 'Summarising Regulations: an Empirical Study of Long-Document Summarisation Methods Under Extreme Compression'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver