paper-with-me

홈 › Papers

Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows

2026-05-29 · Yuri Balashov, Rex VanHorn, Mingxi Xu, Austin Downes arxiv

Building on our previous work, this paper develops practical, low-barrier methods for freelance translators and smaller language service providers to evaluate translation technologies using rigorous yet accessible analytic methods. Here we address a high-stakes, specialized need: offline translation for confidentiality-sensitive domains in which privacy constraints preclude the use of cloud-based engines and commercial LLMs. We expand the Reeve Foundation Trilingual Corpus (RFTC) used in our previous work into a multilingual corpus (RFMC) by adding sentence-aligned German and Simplified Chinese reference translations. We then benchmark several locally runnable language models (via Ollama) across four language directions on 1000+ sentences selected from this corpus. We use consistent single-prompt calls without fine-tuning or domain adaptation, comparing local LLM outputs against commercial NMTs (DeepL, Baidu), a frontier LLM (GPT-5.2), and professional-grade local NMT systems (OPUS-CAT, NeuralDesktop, Promt). Automatic evaluation is conducted with MATEO. Results reveal substantial variation in local LLM performance across language directions and model sizes. The best local LLMs match or surpass local NMT systems and a frontier LLM, though they remain behind top commercial NMTs. These findings underscore the viability of carefully selected local LLM translation for privacy-constrained professionals and inform future research on model scaling and multilingual capability.

📄 PDF Abstract BibTeX arXiv:2605.31452

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Adaptation

Similar Papers 제목 키워드 기반

Translation Analytics for Freelancers: I. Introduction, Data Preparation, Baseline Evaluations

2025-04-20 · Yuri Balashov, Alex Balashov, Shiho Fukuda Koski

This is the first in a series of papers exploring the rapidly expanding new opportunities arising from recent progress in language technologies for individual translators and language service providers with modest resour…

Machine TranslationTranslation

AI and Jobs: Has the Inflection Point Arrived? Evidence from an Online Labor Platform

2023-12-07 · Dandan Qiao, Huaxia Rui, Qian Xiong

The emergence of Large Language Models (LLMs) has renewed the debate on the important issue of "technology displacement". While prior research has investigated the effect of information technology in general on human lab…

CUDABench: Benchmarking LLMs for Text-to-CUDA Generation

2026-02-13 · Jiace Zhu, Wentao Chen, Qi Fan, Zhixing Ren 외 arxiv

Recent studies have demonstrated the potential of Large Language Models (LLMs) in generating GPU Kernels. Current benchmarks focus on the translation of high-level languages into CUDA, overlooking the more general and ch…

"Be My Cheese?": Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMs

2026-02-04 · Madison Van Doren, Casey Ford, Jennifer Barajas, Riley VanMeter 외 arxiv

We present a large-scale human evaluation benchmark for assessing cultural localisation in machine translation produced by state-of-the-art multilingual large language models (LLMs). Existing MT benchmarks emphasise toke…

Machine Translation

Beyond Text: A Deep Dive into Large Language Models' Ability on Understanding Graph Data

2023-10-07 · Yuntong Hu, Zheng Zhang, Liang Zhao

Large language models (LLMs) have achieved impressive performance on many natural language processing tasks. However, their capabilities on graph-structured data remain relatively unexplored. In this paper, we conduct a …

Benchmarking