paper-with-me

홈 › Papers

When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation

2026-08-19 · Chenchen Mao, Hanjing Shi, Haiyan Jia, Emily Wegrzyn, Dominic DiFranzo arxiv

Readable AI output can leave an evaluability gap: even when the source is shown, an overall-quality judgment may not reflect what an output preserves. We investigated how source-text condition and output rendering relate to perceived translation quality, and how output and system appraisals relate to trust and stated disclosure willingness in a plain-text interface. A focal 2 * 2 comparison (N=306) using TransLingo examined simple generated narratives and complex literary-philosophical prose alongside LLM-generated readability-oriented outputs and researcher-revised fidelity-oriented outputs. A descriptive stimulus audit indicated greater source retention in fidelity-oriented outputs in both source-text conditions. Factorial analyses showed a significant rendering-by-source-text-condition interaction in perceived quality. Participants rated fidelity-oriented outputs higher than readability-oriented outputs for the simple narratives, whereas no reliable rendering difference emerged for the complex prose. A corresponding source-condition-dependent pattern was observed for perceived intelligence, agency-oriented anthropomorphic attribution, and task-performance trust. A separate theory-ordered appraisal-structure SEM characterized concurrent associations among perceived quality, perceived intelligence, agency-oriented anthropomorphic attribution, task-performance trust, and stated disclosure willingness across six domains, with task-performance trust as the proximal correlate of stated willingness. The observed rating pattern distinguishes source access from source evaluability: for the complex stimuli, displaying the source did not ensure that one overall-quality rating reflected differences in retained content. It also separates support for evaluating translation output from data-handling support for decisions about what personal text to entrust to a system.

📄 PDF Abstract BibTeX arXiv:2608.19083

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Theoretical Framework for Statistical Evaluability of Generative Models

2026-04-07 · Shashaank Aiyer, Yishay Mansour, Shay Moran, Han Shao arxiv

Statistical evaluation aims to estimate the generalization performance of a model using held-out i.i.d. test data sampled from the ground-truth distribution. In supervised learning settings such as classification, perfor…

Aggregating Incomplete Rankings

2024-02-26 · Yasunori Okumura

This study considers the method to derive a ranking of alternatives by aggregating the rankings submitted by several individuals who may not evaluate all of them. The collection of subsets of alternatives that individual…

The AI Evaluability Gap: The Missing Layer for Managing Risk and Sustaining Value

2026-06-19 · Vishal Srivastava, Tanmay Sah arxiv

Organizations deploying AI face two fundamental governance challenges: managing AI risk and sustaining AI value. Both depend on evidence whose sufficiency cannot be taken for granted. We call the shared underlying challe…

A Corpus of eRulemaking User Comments for Measuring Evaluability of Arguments

2018-05-01 · LREC 2018 5 · Joonsuk Park, Claire Cardie
Argument Mining

Making Knowledge Accessible: Divergent Readability-Accuracy Strategies of Mistral and QWen in Biomedical Text Simplification

2025-11-07 · P. Bilha Githinji, Aikaterini Melliou, Zeming Liang, Lian Zhang 외 arxiv

The growing public demand for accessible biomedical information calls for scalable text simplification. While large language models (LLMs) offer solutions, they too struggle with balancing improved readability against pr…

Text Simplification