paper-with-me

Papers

TEIMMA: The First Content Reuse Annotator for Text, Images, and Math

2023-05-22 · Ankit Satpute, André Greiner-Petter, Moritz Schubotz, Norman Meuschke, Akiko Aizawa, Olaf Teschke, Bela Gipp

This demo paper presents the first tool to annotate the reuse of text, images, and mathematical formulae in a document pair -- TEIMMA. Annotating content reuse is particularly useful to develop plagiarism detection algorithms. Real-world content reuse is often obfuscated, which makes it challenging to identify such cases. TEIMMA allows entering the obfuscation type to enable novel classifications for confirmed cases of plagiarism. It enables recording different reuse types for text, images, and mathematical formulae in HTML and supports users by visualizing the content reuse in a document pair using similarity detection methods for text and math.

📄 PDF Abstract BibTeX arXiv:2305.13193

Code (1)

gipplab/teimma-reuse-annotator 공식 구현 pytorch

Tasks

Math

Similar Papers 제목 키워드 기반

Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization

2026-05-27 · Beiduo Chen, Pingjun Hong, Ziyun Zhang, Benjamin Roth 외 arxiv

Free-text explanations extend human label variation (HLV) beyond label disagreement by revealing the reasoning and preferences behind annotators' decisions. We study whether large language models (LLMs) can learn and rep…

Natural Language Inference

Quantifying Contextual Aspects of Inter-annotator Agreement in Intertextuality Research

2021-11-01 · EMNLP (LaTeCHCLfL, CLFL, LaTeCH) 2021 11 · Enrique Manjavacas Arevalo, Laurence Mellerin, Mike Kestemont

We report on an inter-annotator agreement experiment involving instances of text reuse focusing on the well-known case of biblical intertextuality in medieval literature. We target the application use case of literary sc…

BETA-Labeling for Multilingual Dataset Construction in Low-Resource IR

2026-02-16 · Md. Najib Hasan, Mst. Jannatun Ferdous Rain, Fyad Mohammed, Nazmul Siddique arxiv

IR in low-resource languages remains limited by the scarcity of high-quality, task-specific annotated datasets. Manual annotation is expensive and difficult to scale, while using large language models (LLMs) as automated…

Machine Translation

Modeling Ambiguity with Many Annotators and Self-Assessments of Annotator Certainty

2020-12-01 · COLING (LAW) 2020 12 · Melanie Andresen, Michael Vauth, Heike Zinsmeister

Most annotation efforts assume that annotators will agree on labels, if the annotation categories are well-defined and documented in annotation guidelines. However, this is not always true. For instance, content-related …

Sentence

Large Language Models for Propaganda Span Annotation

2023-11-16 · Maram Hasanain, Fatema Ahmad, Firoj Alam

The use of propagandistic techniques in online content has increased in recent years aiming to manipulate online audiences. Fine-grained propaganda detection and extraction of textual spans where propaganda techniques ar…

Propaganda detection