paper-with-me

홈 › Papers

Summarising Historical Text in Modern Languages

2021-01-26 · EACL 2021 2 · Xutan Peng, Yi Zheng, Chenghua Lin, Advaith Siddharthan

We introduce the task of historical text summarisation, where documents in historical forms of a language are summarised in the corresponding modern language. This is a fundamentally important routine to historians and digital humanities researchers but has never been automated. We compile a high-quality gold-standard text summarisation dataset, which consists of historical German and Chinese news from hundreds of years ago summarised in modern German or Chinese. Based on cross-lingual transfer learning techniques, we propose a summarisation model that can be trained even with no cross-lingual (historical to modern) parallel data, and further benchmark it against state-of-the-art algorithms. We report automatic and human evaluations that distinguish the historic to modern language summarisation task from standard cross-lingual summarisation (i.e., modern to modern language), highlight the distinctness and value of our dataset, and demonstrate that our transfer learning approach outperforms standard cross-lingual benchmarks on this task.

📄 PDF Abstract BibTeX arXiv:2101.10759

Code (1)

Pzoom522/HistSumm 공식 구현

Tasks

Cross-Lingual TransferTransfer Learning

Similar Papers 제목 키워드 기반

Multilingual Event Extraction from Historical Newspaper Adverts

2023-05-18 · Nadav Borenstein, Natalia da Silva Perez, Isabelle Augenstein

NLP methods can aid historians in analyzing textual materials in greater volumes than manually feasible. Developing such methods poses substantial challenges though. First, acquiring large, annotated historical datasets …

Event ExtractionMachine Translation

A Neural Model for Part-of-Speech Tagging in Historical Texts

2016-12-01 · COLING 2016 12 · Christian Hardmeier

Historical texts are challenging for natural language processing because they differ linguistically from modern texts and because of their lack of orthographical and grammatical standardisation. We use a character-level …

Part-Of-Speech TaggingPOSPOS Tagging

CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition

2025-09-24 · Sina J. Semnani, Han Zhang, Xinyan He, Merve Tekgürler 외 arxiv

Accurate text recognition for historical documents can greatly advance the study and preservation of cultural heritage. Existing vision-language models (VLMs), however, are designed for modern, standardized texts and are…

Large Language Models for Summarizing Czech Historical Documents and Beyond

2025-08-14 · Václav Tran, Jakub Šmíd, Jiří Martínek, Ladislav Lenc 외 arxiv

Text summarization is the task of shortening a larger body of text into a concise version while retaining its essential meaning and key information. While summarization has been significantly explored in English and othe…

Text Summarization

The Classical Language Toolkit: An NLP Framework for Pre-Modern Languages

2021-08-01 · ACL 2021 5 · Kyle P. Johnson, Patrick J. Burns, John Stewart, Todd Cook 외

This paper announces version 1.0 of the Classical Language Toolkit (CLTK), an NLP framework for pre-modern languages. The vast majority of NLP, its algorithms and software, is created with assumptions particular to livin…

Diversity