paper-with-me

Papers

LATA: A Tool for LLM-Assisted Translation Annotation

2026-02-11 · Baorong Huang, Ali Asiri arxiv

The construction of high-quality parallel corpora for translation research has increasingly evolved from simple sentence alignment to complex, multi-layered annotation tasks. This methodological shift presents significant challenges for structurally divergent language pairs, such as Arabic--English, where standard automated tools frequently fail to capture deep linguistic shifts or semantic nuances. This paper introduces a novel, LLM-assisted interactive tool designed to reduce the gap between scalable automation and the rigorous precision required for expert human judgment. Unlike traditional statistical aligners, our system employs a template-based Prompt Manager that leverages large language models (LLMs) for sentence segmentation and alignment under strict JSON output constraints. In this tool, automated preprocessing integrates into a human-in-the-loop workflow, allowing researchers to refine alignments and apply custom translation technique annotations through a stand-off architecture. By leveraging LLM-assisted processing, the tool balances annotation efficiency with the linguistic precision required to analyze complex translation phenomena in specialized domains.

📄 PDF Abstract BibTeX arXiv:2602.10454

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Translating the Untranslatable: An Operationalizable Ontology for Untranslatability

2026-06-15 · Jacob Bremerman, Brihi Joshi, Hirona Arai, Xiang Ren 외 arxiv

Untranslatability, cases where meaning cannot be directly preserved across languages, is well-studied in linguistics but underexplored in NLP. As machine translation (MT) systems improve on standard benchmarks, their lim…

Machine Translation

Steering AI-Driven Personalization of Scientific Text for General Audiences

2024-11-15 · Taewook Kim, Dhruv Agarwal, Jordan Ackerman, Manaswi Saha

Digital media platforms (e.g., social media, science blogs) offer opportunities to communicate scientific content to general audiences at scale. However, these audiences vary in their scientific expertise, literacy level…

Contemplata, a Free Platform for Constituency Treebank Annotation

2020-05-01 · LREC 2020 5 · Jakub Waszczuk, Ilaine Wang, Jean-Yves Antoine, Ana{\"\i}s Halftermeyer

This paper describes Contemplata, an annotation platform that offers a generic solution for treebank building as well as treebank enrichment with relations between syntactic nodes. Contemplata is dedicated to the annotat…

Cultural and Geographical Influences on Image Translatability of Words across Languages

2021-06-01 · NAACL 2021 4 · Nikzad Khani, Isidora Tourni, Mohammad Sadegh Rasooli, Chris Callison-Burch 외

Neural Machine Translation (NMT) models have been observed to produce poor translations when there are few/no parallel sentences to train the models. In the absence of parallel data, several approaches have turned to the…

Cultural Vocal Bursts Intensity PredictionLow Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine Translation+4

Automatic Input Rewriting Improves Translation with Large Language Models

2025-02-23 · Dayeon Ki, Marine Carpuat

Can we improve machine translation (MT) with LLMs by rewriting their inputs automatically? Users commonly rely on the intuition that well-written text is easier to translate when using off-the-shelf MT systems. LLMs can …

Machine TranslationText SimplificationTranslation