Source-linked AI summary
When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation
Chenchen Mao, Hanjing Shi, Haiyan Jia, Emily Wegrzyn, Dominic DiFranzo
TL;DR
Readable AI translation can leave a gap between source access and evaluability of retained content. In a plain-text study comparing readability- and fidelity-oriented renderings across simple and complex sources, fidelity-oriented outputs received higher appraisals for simple sources but no reliable rendering difference appeared for complex sources.
Problem
The paper examines whether showing the source ensures that users can evaluate what an AI translation preserves.
Method
A plain-text translation study compared readability- and fidelity-oriented renderings across simple narratives and complex literary-philosophical prose.
Results
Fidelity-oriented rendering received higher quality, intelligence, agency-oriented attribution, and task-performance trust ratings for simple sources, but not complex sources.
Takeaways & Limitations
Source access and source evaluability are distinct: showing the source may not make retained content visible in overall judgments.
Takeaways & Limitations
Source complexity was represented by only two texts per level and was confounded with genre and provenance, limiting generalization beyond these stimuli.
Abstract
from arXiv · showhide
Readable AI output can leave an evaluability gap: even when the source is shown, an overall-quality judgment may not reflect what an output preserves. We investigated how source-text condition and output rendering relate to perceived translation quality, and how output and system appraisals relate to trust and stated disclosure willingness in a plain-text interface. A focal 2 * 2 comparison (N=306) using TransLingo examined simple generated narratives and complex literary-philosophical prose alongside LLM-generated readability-oriented outputs and researcher-revised fidelity-oriented outputs. A descriptive stimulus audit indicated greater source retention in fidelity-oriented outputs in both source-text conditions. Factorial analyses showed a significant rendering-by-source-text-condition interaction in perceived quality. Participants rated fidelity-oriented outputs higher than readability-oriented outputs for the simple narratives, whereas no reliable rendering difference emerged for the complex prose. A corresponding source-condition-dependent pattern was observed for perceived intelligence, agency-oriented anthropomorphic attribution, and task-performance trust. A separate theory-ordered appraisal-structure SEM characterized concurrent associations among perceived quality, perceived intelligence, agency-oriented anthropomorphic attribution, task-performance trust, and stated disclosure willingness across six domains, with task-performance trust as the proximal correlate of stated willingness. The observed rating pattern distinguishes source access from source evaluability: for the complex stimuli, displaying the source did not ensure that one overall-quality rating reflected differences in retained content. It also separates support for evaluating translation output from data-handling support for decisions about what personal text to entrust to a system.
1 Introduction
This study examines an evaluability gap in AI translation: readable outputs may weaken source faithfulness even when the source is visible. It distinguishes output evaluation from decisions about entrusting personal text to a system.
- Motivation: Readability-oriented generation may weaken adherence to the source, creating an evaluability gap between source visibility and recognizing retained or altered content.The gap arises because readability gains can come at the cost of faithfulness, while overall-quality judgments may not reflect source retention.
- Study design: The focal 2×2 comparison varied simple versus complex source text and readability-oriented versus fidelity-oriented output renderings in a plain-text interface.The study compared perceived quality and source-condition-dependent appraisals of intelligence, agency-oriented anthropomorphism, and task-performance trust.
- Findings: Greater source retention was reflected in overall-quality ratings for simple sources but not complex sources, and the source-dependent pattern extended to perceived intelligence, agency-oriented anthropomorphism, and task-performance trust.No reliable rendering contrast emerged for stated disclosure willingness in either source condition.
- Design implications: Output evaluation and information entrustment are distinct interface-design problems: rendering differences did not reliably change disclosure willingness despite appraisal differences.Model-implied rendering–disclosure associations were β= .050–.078 for simple sources and near zero for complex sources.
2 Related Work and Hypothesis Development
This section develops an evaluability problem: readable LLM outputs may diverge from source fidelity, making overall quality judgments theoretically ambiguous. It then proposes source-condition-sensitive quality comparisons and appraisal links from translation quality to intelligence, anthropomorphism, trust, and disclosure willingness.
- System Appraisals: The model hypothesizes that higher perceived translation quality is positively associated with perceived intelligence, which is also linked to agency-oriented anthropomorphic attribution and trust.Perceived quality concerns presented outputs, whereas perceived intelligence concerns broader system capabilities inferred from performance.
- Trust and Disclosure: The hypotheses further associate trust with stated disclosure willingness across six personal-information domains, including tastes and interests, attitudes and opinions, and work or studies.The proposed domains also include economic and social status, with H5a–H5f covering all six domains.
- Readability and Source Retention: LLM outputs can be clear and accessible while overgeneralizing source claims or deviating from underlying evidence, creating a potential trade-off between readability and source retention.This tension appears in scientific summarization and conversational question answering, where fluency can increase perceived trustworthiness despite reduced faithfulness.
- Quality Judgments: Two competing accounts predict opposite quality effects: processing fluency favors readability-oriented outputs, whereas source adequacy favors fidelity-oriented outputs.The section therefore poses RQ1a on perceived-quality differences between rendering types and RQ1b on whether those differences vary across source-text conditions.
- Source-Text Conditions: Source complexity may increase reliance on readability while making fidelity harder to recognize, motivating tests of whether rendering changes overall quality judgments across conditions.The study does not measure these mechanisms separately and instead examines the rendering–quality relationship across tested source-text conditions.
3 Methods
The study used a focal 2 × 2 TransLingo comparison of simple versus complex source texts and readability- versus fidelity-oriented renderings. Prepared-output audits and participant procedures evaluated readability, source retention, translation quality, and related appraisals while preserving important design limitations.
- Experimental design: The focal 2 × 2 design crossed simple and complex source-text conditions with readability-oriented and fidelity-oriented renderings of the same TransLingo output.The readability-oriented output was LLM-generated for ease of reading, whereas the fidelity-oriented output was researcher-revised toward closer source fidelity.
- Limitations: The source-text comparison was confounded with genre, topic, provenance, period style, and length, and therefore represented simple generated narrative versus complex literary-philosophical prose rather than pure linguistic complexity.The manuscript also notes that audit metrics described relative stimulus differences rather than independently validated translation adequacy.
- Stimulus audit: Readability-oriented outputs showed larger ΔFlesch gains than fidelity-oriented outputs for simple texts (14.08 vs. 4.83) and complex texts (70.43 vs. 56.40).Fidelity-oriented outputs retained more content words, had higher BERTScore, and received higher COMET scores at both complexity levels.
- Procedure: Participants were randomized to simple or complex texts within separate rendering periods, so source-text complexity was randomized within periods but rendering was compared between periods.Each participant evaluated two texts and then rated perceived complexity, translation quality, intelligence, anthropomorphism, trust, and willingness to disclose personal information.
4 Findings
The findings support distinct but closely related appraisal constructs and show that rendering effects depended on source complexity: fidelity-oriented outputs were preferred for simple sources but not complex ones. In the appraisal-structure model, task-performance trust was the proximal correlate of stated disclosure willingness across all six domains.
- Factorial condition analyses: For simple sources, fidelity-oriented outputs received higher perceived-quality ratings than readability-oriented outputs, but complex-source ratings did not differ reliably.Simple sources: ΔM = .584, 95% CI [.248, .921], p < .001, Hedges’ g = .567; complex sources: ΔM = .018, 95% CI [−.376, .413], p = .927, g = .014.
- Disclosure willingness: No reliable rendering-by-source-condition interaction was detected for stated disclosure willingness across the six domains.All interaction p-values were ≥ .076 and FDR-adjusted q-values were ≥ .360; simple-source contrasts did not remain significant after correction, and complex-source contrasts were nonsignificant.
- Appraisal-structure SEM: Perceived intelligence was positively associated with anthropomorphism and trust, while anthropomorphism was also positively associated with trust.The reported standardized associations were β = .650 for intelligence–anthropomorphism, β = .290 for intelligence–trust, and β = .671 for anthropomorphism–trust, all p < .001.
- Appraisal-structure SEM: Task-performance trust was positively associated with stated disclosure willingness in all six domains, with standardized coefficients ranging from β = .363 to β = .566.All six associations had p < .001 and remained significant after Benjamini–Hochberg FDR correction.
- Appraisal-structure SEM: Model-implied rendering associations with disclosure willingness were small for simple texts and negligible for complex texts.Simple-text indirect associations were β = .050–.078 with FDR-adjusted p = .004; complex-text associations were β = .009–.014, with confidence intervals including zero and FDR-adjusted p = .927.
5 Discussion
The discussion identifies an evaluability gap: source visibility did not ensure that holistic quality judgments reflected retained content for complex prose. It also distinguishes output evaluation from decisions about entrusting personal text.
- Evaluability gap: Fidelity-oriented outputs rated higher for simple narratives, but no rendering difference emerged for complex prose despite greater audited source similarity.The authors call this mismatch an evaluability gap because visible source evidence did not make retained-content differences apparent in overall-quality ratings.
- Evaluability gap: Showing the source provided evidence but did not ensure that holistic ratings reflected audited retention differences, possibly because difficult comparisons increased reliance on readability.The source-text conditions also differed in genre, provenance, period style, topic, and length, so matched stimuli and distinct measures are needed.
- System appraisals: For simple texts, fidelity-oriented outputs received higher perceived intelligence, agency-oriented attribution, and task-performance trust; complex texts showed no reliable appraisal differences.All three rendering-by-source-condition interactions were reliable.
- Appraisal structure: Perceived output quality was associated with perceived intelligence (𝛽= .770, 𝑝< .001), while intelligence and agency-oriented attribution were associated with task-performance trust.Their association with trust was 𝛽= .290 and 𝛽= .671, respectively; both 𝑝< .001, with HTMT = .871 despite evidence they remained distinct constructs.
- Disclosure willingness: Task-performance trust was the most direct correlate of stated disclosure willingness (𝛽= .363–.566, all FDR-adjusted 𝑝< .001).Adding direct agency-attribution paths did not improve fit: robust Δ𝜒2(6) = 3.67, 𝑝= .722.
- Design implications: Interfaces should support judging source retention separately from deciding whether to submit personal text, using alignment, omission indicators, distinct fidelity judgments, and source-based metrics.Task-performance trust concerns accuracy and reliability rather than privacy or data handling, motivating explicit information about retention, training use, and third-party sharing.
6 Limitations and Future Work
The study’s design limits generalization because participants could evaluate both English texts, complexity was confounded with stimulus properties, and data collection spanned changing AI familiarity. The appraisal and disclosure measures also limit causal and behavioral inference, motivating repeated-interaction and behavioral follow-up studies.
- Participant and evaluation constraints: English-speaking participants evaluated translations through back-translation, so the findings may not apply to users unable to evaluate the target language.Fidelity-oriented outputs received higher ratings only for simple sources, with no detected rendering difference for complex sources; this comparison should be tested directly in users who cannot evaluate the target language.
- Stimulus and rendering constraints: Source complexity used only two texts per level and was confounded with genre, provenance, period style, and length, so the interaction does not isolate responsible text properties.Future studies should vary complexity across more texts matched on these properties and replicate the comparison with independently generated system outputs.
- Temporal and familiarity constraints: Data collection spanned August 2023 to February 2025, prior AI familiarity was not measured, and rendering condition was not independently randomized across periods.Period-related variation is treated as a design boundary rather than a temporal finding.
- Appraisal-model constraints: Because appraisal constructs were collected in one post-survey and anthropomorphism and trust were closely related, common-method effects and unknown temporal ordering limit SEM interpretation.Repeated-interaction studies should examine appraisal changes after translation errors, correction, and recovery.
- Disclosure-behavior constraints: Stated willingness to disclose was measured instead of actual behavior, and topic-level intimacy, valence, and sensitivity perceptions were not measured.Future work should measure these perceptions and observe what information users actually choose to translate.
7 Conclusion
The study identifies an evaluability gap: showing the source did not ensure that users recognized what translation outputs preserved. Rendering affected quality-related appraisals but not stated disclosure willingness in the same way, distinguishing evaluation support from data-handling information.
- Conclusion: The findings distinguish source access from source evaluability because showing the source can leave users without enough support to recognize preserved output content.Thus, source display alone did not ensure evaluable differences in retained content for complex sources.
- Conclusion: Fidelity-oriented rendering received higher ratings for perceived quality, intelligence, agency-oriented attribution, and task-performance trust with simple sources, but not complex sources.This occurred despite greater source retention under fidelity-oriented rendering in both source-text conditions.
- Conclusion: No reliable rendering difference emerged in stated disclosure willingness under either source-text condition.The model-implied indirect association between rendering and disclosure was small for simple sources and negligible for complex sources.
- Conclusion: Greater task-performance trust was associated with greater disclosure willingness, making task-performance trust the proximal appraisal correlate of stated willingness.This association was clarified by the appraisal-structure model.
- Conclusion: Source-linked change support for translation evaluation and data-handling information for disclosure decisions should be evaluated as separate interventions.Translation quality and stated disclosure willingness did not respond to the rendering contrast in the same way.
A Measurement Items
This appendix reports the wording presented to participants.
- A Measurement Items: The appendix provides the wording shown to participants.
A.1 Perceived Intelligence
Perceived intelligence was assessed through four statements about TransLingo’s understanding, clarity, information processing, and ability to provide useful, accurate translations.
- Perceived Intelligence: Participants rated whether TransLingo could understand input text accurately.This item was labeled PI1.
- Perceived Intelligence: Participants rated whether TransLingo could provide translations in an understandable manner.This item was labeled PI2.
- Perceived Intelligence: Participants rated whether TransLingo could find and process necessary information for accurate translation.This item was labeled PI3.
- Perceived Intelligence: Participants rated whether TransLingo could provide a useful and accurate translation.This item was labeled PI4.
A.2 Anthropomorphism
Anthropomorphism was assessed through questions about TransLingo’s apparent intelligence, contextual understanding, anticipation, and translation planning. These items focused on human-like cognitive and decision-making capacities attributed to the machine translation system.
- A.2 Anthropomorphism: ANT1 asked how smart TransLingo seemed.
- A.2 Anthropomorphism: ANT2 assessed how well TransLingo could understand the text’s context.
- A.2 Anthropomorphism: ANT3 asked whether TransLingo could anticipate what would be written before it was written.
- A.2 Anthropomorphism: ANT4 assessed whether TransLingo could plan and choose the best translation.
A.3 Trust
Trust was assessed through participants’ judgments of TransLingo’s average accuracy, confidence in its translation quality, and willingness to rely on it across language scenarios and tasks.
- Trust: Trust items assessed perceived average translation accuracy and confidence in TransLingo’s translation quality during active use.These measures were captured by TR1 and TR4.
- Trust: Participants were asked about trusting TransLingo in busy, context-rich, real-time scenarios and simpler, less context-heavy scenarios.These scenario-specific items corresponded to TR2 and TR3.
- Trust: Additional items addressed trust in accurate translations for the next task and feeling at ease relying on TransLingo for communication needs.These measures were captured by TR5, TR6, and TR7.
A.4 Disclosure Willingness … A.4.6 Physical Appearance and Sex.
Disclosure willingness was assessed by asking how comfortable participants felt using TransLingo to translate and disclose personal information across seven domains. The domains ranged from tastes and opinions to work, status, relationships, and physical appearance and sex, using a four-point comfort scale.
- A.4 Disclosure Willingness: Participants were asked how comfortable they felt using TransLingo to translate and disclose each personal topic.The assessment introduced multiple topic-specific disclosure judgments.
- A.4.1 Tastes and Interests. •: Tastes and interests covered favorite foods, disliked entertainment, admired or disliked people and brands, and preferred or boring social gatherings.
- A.4.2 Attitudes and Opinions.: Attitudes and opinions included admired or disliked political and religious views, social progress, criticism of poverty and injustice, controversial choices, and skepticism about life choices.
- A.4.3 Work or Studies. •: Work or studies topics covered satisfaction, pressures, strengths, boredom, achievements, and frustration about work or study not being valued.
- A.4.4 Economic and Social Status. •: Economic and social status topics addressed wealth signals, immediate financial need, future financial optimism or pessimism, high-status connections, and feelings of inferiority.
- A.4.5 Interpersonal Relations and Self-Concept.: Interpersonal relations and self-concept included romantic experiences, homesickness, resentment toward friends, pride, and dislike of oneself.