paper-with-me

홈 › Papers

CLSE: Corpus of Linguistically Significant Entities

2022-11-04 · Aleksandr Chuklin, Justin Zhao, Mihir Kale

One of the biggest challenges of natural language generation (NLG) is the proper handling of named entities. Named entities are a common source of grammar mistakes such as wrong prepositions, wrong article handling, or incorrect entity inflection. Without factoring linguistic representation, such errors are often underrepresented when evaluating on a small set of arbitrarily picked argument values, or when translating a dataset from a linguistically simpler language, like English, to a linguistically complex language, like Russian. However, for some applications, broadly precise grammatical correctness is critical -- native speakers may find entity-related grammar errors silly, jarring, or even offensive. To enable the creation of more linguistically diverse NLG datasets, we release a Corpus of Linguistically Significant Entities (CLSE) annotated by linguist experts. The corpus includes 34 languages and covers 74 different semantic types to support various applications from airline ticketing to video games. To demonstrate one possible use of CLSE, we produce an augmented version of the Schema-Guided Dialog Dataset, SGD-CLSE. Using the CLSE's entities and a small number of human translations, we create a linguistically representative NLG evaluation benchmark in three languages: French (high-resource), Marathi (low-resource), and Russian (highly inflected language). We establish quality baselines for neural, template-based, and hybrid NLG systems and discuss the strengths and weaknesses of each approach.

📄 PDF Abstract BibTeX arXiv:2211.02423

Code (1)

google-research-datasets/clse 공식 구현

Tasks

nlg evaluationText Generation

Similar Papers 제목 키워드 기반

The SETimes.HR Linguistically Annotated Corpus of Croatian

2014-05-01 · LREC 2014 5 · {\v{Z}}eljko Agi{\'c}, Nikola Ljube{\v{s}}i{\'c}

We present SETimes.HR ― the first linguistically annotated corpus of Croatian that is freely available for all purposes. The corpus is built on top of the SETimes parallel corpus of nine Southeast European languages an…

AllBoundary DetectionDependency ParsingLemmatization+4

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models

2026-06-23 · Bin Chen, Yuxiang Cai, Yadan Luo, Yi Zhang 외 arxiv

Reducing visual token redundancy is critical for accelerating Multimodal Large Language Models (MLLMs) without degrading cross-modal reasoning performance. Existing token pruning methods typically rely on single-layer si…

The Annotation Guideline of LST20 Corpus

2020-08-12 · Prachya Boonkwan, Vorapon Luantangsrisuk, Sitthaa Phaholphinyo, Kanyanat Kriengket 외

This report presents the annotation guideline for LST20, a large-scale corpus with multiple layers of linguistic annotation for Thai language processing. Our guideline consists of five layers of linguistic annotation: wo…

POSPOS TaggingSentence

CLSEG: Contrastive Learning of Story Ending Generation

2022-02-18 · Yuqiang Xie, Yue Hu, Luxi Xing, Yunpeng Li 외

Story Ending Generation (SEG) is a challenging task in natural language generation. Recently, methods based on Pre-trained Language Models (PLM) have achieved great prosperity, which can produce fluent and coherent story…

Contrastive LearningText Generation

Collection and Annotation of the Romanian Legal Corpus

2020-05-01 · LREC 2020 5 · Dan Tufi{\textcommabelow{s}}, Maria Mitrofan, Vasile P{\u{a}}i{\textcommabelow{s}}, Radu Ion 외

We present the Romanian legislative corpus which is a valuable linguistic asset for the development of machine translation systems, especially for under-resourced languages. The knowledge that can be extracted from this …

Machine TranslationPOSTranslation