Papers Document Translation
“Document Translation” 태그가 달린 논문 47편 · 필터 해제
GRAFT: A Graph-based Flow-aware Agentic Framework for Document-level Machine Translation
Document level Machine Translation (DocMT) approaches often struggle with effectively capturing discourse level phenomena. Existing approaches rely on heuristic rules to segment documents into discourse units, which rare…
Document Level Machine TranslationDocument TranslationLarge Language ModelMachine Translation+1Agent Capability Negotiation and Binding Protocol (ACNBP)
As multi-agent systems evolve to encompass increasingly diverse and specialized agents, the challenge of enabling effective collaboration between heterogeneous agents has become paramount, with traditional agent communic…
Document TranslationCLIRudit: Cross-Lingual Information Retrieval of Scientific Documents
Cross-lingual information retrieval (CLIR) consists in finding relevant documents in a language that differs from the language of the queries. This paper presents CLIRudit, a new dataset created to evaluate cross-lingual…
BenchmarkingCross-Lingual Information RetrievalDocument TranslationInformation Retrieval+3Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
Large language models (LLMs) have demonstrated strong performance in sentence-level machine translation, but scaling to document-level translation remains challenging, particularly in modeling long-range dependencies and…
Document Level Machine TranslationDocument TranslationMachine TranslationSentence+1Quality-Aware Decoding: Unifying Quality Estimation and Decoding
An emerging research direction in NMT involves the use of Quality Estimation (QE) models, which have demonstrated high correlations with human judgment and can enhance translations through Quality-Aware Decoding. Althoug…
DecoderDocument TranslationNMTRe-Ranking+1Cross-Dialect Information Retrieval: Information Access in Low-Resource and High-Variance Languages
A large amount of local and culture-specific knowledge (e.g., people, traditions, food) can only be found in documents written in dialects. While there has been extensive research conducted on cross-lingual information r…
Cross-Lingual Information RetrievalCross-Lingual TransferDocument TranslationInformation Retrieval+2LLMs-in-the-Loop Part 2: Expert Small AI Models for Anonymization and De-identification of PHI Across Multiple Languages
The rise of chronic diseases and pandemics like COVID-19 has emphasized the need for effective patient data processing while ensuring privacy through anonymization and de-identification of protected health information (P…
De-identificationDocument TranslationNERRelation ExtractionAnalyzing Context Utilization of LLMs in Document-Level Translation
Large language models (LLM) are increasingly strong contenders in machine translation. We study document-level translation, where some words cannot be translated without context from outside the sentence. We investigate …
Document TranslationMachine TranslationSentenceTranslationDelTA: An Online Document-Level Translation Agent Based on Multi-Level Memory
Large language models (LLMs) have achieved reasonable quality improvements in machine translation (MT). However, most current research on MT-LLMs still faces significant challenges in maintaining translation consistency …
Document TranslationMachine TranslationProper NounSentence+1M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation
Document translation poses a challenge for Neural Machine Translation (NMT) systems. Most document-level NMT systems rely on meticulously curated sentence-level parallel data, assuming flawless extraction of text from do…
Document Level Machine TranslationDocument TranslationMachine TranslationNMT+4HLTCOE at TREC 2023 NeuCLIR Track
The HLTCOE team applied PLAID, an mT5 reranker, and document translation to the TREC 2023 NeuCLIR track. For PLAID we included a variety of models and training techniques -- the English model released with ColBERT v2, tr…
AllDocument TranslationNusaWrites: Constructing High-Quality Corpora for Underrepresented and Extremely Low-Resource Languages
Democratizing access to natural language processing (NLP) technology is crucial, especially for underrepresented and extremely low-resource languages. Previous research has focused on developing labeled and unlabeled cor…
DiversityDocument TranslationTranslationUnified Model Learning for Various Neural Machine Translation
Existing neural machine translation (NMT) studies mainly focus on developing dataset-specific models based on data from different tasks (e.g., document translation and chat translation). Although the dataset-specific mod…
Document TranslationMachine TranslationmodelNMT+2A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models
Machine Translation (MT) has greatly advanced over the years due to the developments in deep neural networks. However, the emergence of Large Language Models (LLMs) like GPT-4 and ChatGPT is introducing a new phase in th…
Document TranslationMachine TranslationPrivacy PreservingTranslationTransDocs: Optical Character Recognition with word to word translation
While OCR has been used in various applications, its output is not always accurate, leading to misfit words. This research work focuses on improving the optical character recognition (OCR) with ML techniques with integra…
Deep LearningDocument TranslationMachine TranslationOptical Character Recognition+3Extending English IR methods to multi-lingual IR
This paper describes our participation in the 2023 WSDM CUP - MIRACL challenge. Via a combination of i) document translation; ii) multilingual SPLADE and Contriever; and iii) multilingual RankT5 and many other models, we…
Document TranslationRerankingA Study on ReLU and Softmax in Transformer
The Transformer architecture consists of self-attention and feed-forward networks (FFNs) which can be viewed as key-value memories according to previous works. However, FFN and traditional memory utilize different activa…
Document TranslationModeling Context With Linear Attention for Scalable Document-Level Translation
Document-level machine translation leverages inter-sentence dependencies to produce more coherent and consistent translations. However, these models, predominantly based on transformers, are difficult to scale to long do…
Document Level Machine TranslationDocument TranslationInductive BiasMachine Translation+2NICT’s Submission to the WAT 2022 Structured Document Translation Task
We present our submission to the structured document translation task organized by WAT 2022. In structured document translation, the key challenge is the handling of inline tags, which annotate text. Specifically, the te…
Document TranslationNMTSentenceTAG+1Revamping Multilingual Agreement Bidirectionally via Switched Back-translation for Multilingual Neural Machine Translation
Despite the fact that multilingual agreement (MA) has shown its importance for multilingual neural machine translation (MNMT), current methodologies in the field have two shortages: (i) require parallel data between mult…
Document Level Machine TranslationDocument TranslationMachine TranslationTranslation