Source-linked AI summary
Personas Differ from Native-Language Generation: Language Pathways Shape LLM Interpersonal Advice
Jinhee Won, Xinlan Emily Hu
TL;DR
The paper asks whether native-language generation followed by translation and native-speaker persona prompting produce equivalent multilingual interpersonal advice. Across 600 questions, 13 language conditions, and eight models, it finds systematic differences in linguistic style, behavioral scaffolding, and action recommendations, making elicitation strategy consequential.
Problem
It is unclear whether asking an LLM to respond as a native speaker is equivalent to generating advice in the target language and translating it into English.
Method
The study compares native-language generation followed by translation with native-speaker persona prompting across 600 advice questions, 13 language conditions, and eight language models.
Results
Native-speaker persona prompting and native-language generation systematically differ in linguistic style, behavioral scaffolding, and action recommendations; persona prompting often increases social cues while reducing concreteness and actionable guidance.
Takeaways & Limitations
Cross-lingual elicitation strategy is a substantive methodological choice, so multilingual advice systems should evaluate both linguistic and behavioral outcomes explicitly.
Takeaways & Limitations
The source dilemmas come from an English-language advice forum and may reflect Western, English-speaking assumptions rather than how people from different language communities would respond.
Abstract
from arXiv · showhide
LLMs are increasingly used for interpersonal advice and as tools for studying social behavior across languages and cultures. A common shortcut for eliciting language- or culture-related variation is to ask a model to answer as a native speaker. We test whether this native-speaker persona reproduces the outputs obtained when models instead generate advice in the target language and translate the response back into English. Using 600 interpersonal advice questions across 13 languages and eight LLMs, we compare native-language generation followed by translation (NL) with native-speaker persona prompting (NP), measuring linguistic style, behavioral scaffolding, and forced-choice action recommendations. We find that NP and NL are not interchangeable. Compared to NL, NP often increases lexical social cues, including affiliation and positive tone, while reducing qualities such as concreteness and social attunement; NP also provides less actionable scaffolding in open-ended advice. In forced-choice scenarios, NP changes which action the model selects, favoring confrontation over redirection, with effect sizes varying across languages, topics, and models. Our results show that cross-lingual elicitation strategy is a consequential methodological choice that can change both how advice is framed and which actions models recommend.
1 Introduction
Interpersonal advice links linguistic form with social meaning, yet multilingual evaluation and persona prompting leave open whether native-speaker prompting matches target-language generation. This study compares those pathways across languages and models to assess differences in advice framing and recommendations.
- Interpersonal advice requires attention to directness, politeness, hierarchy, and emotional expression, not only task accuracy.
- The study compares native-language generation followed by English translation with English responses prompted through a native-speaker persona.
- The analysis measures linguistic characteristics, behavioral guidance, and forced-choice recommendations among confrontation, disengagement, and redirection.
- Language pathway is presented as a consequential source of divergence in linguistic and behavioral outputs, varying by language, topic, and model.
2 Related Work
Prior multilingual NLP work emphasizes objective benchmarks, while related research shows that interpersonal advice and persona prompts involve broader social, emotional, and cultural variation.
- Multilingual evaluation has traditionally centered on classification, question answering, retrieval, translation, and structured prediction.
- Interpersonal advice quality depends on recommendation phrasing, tone, and perceived usefulness, alongside social reasoning.
- Persona prompts affect model outputs but capture only part of human variation and context.
3 Methodology
The study evaluates 600 socially situated English advice questions across 13 languages and eight models, comparing native-language generation with native-speaker persona prompting. It combines lexical, pragmatic, behavioral-scaffolding, and action-choice measures.
- The dataset contains 600 questions drawn from four interpersonal topics: work environment, friends, relationships, and family.
- The language set spans 13 languages, with English using original questions and other languages constructed through controlled translation.
- The study compares native-language generation, where answers are produced in the target language and translated, with English native-speaker persona prompting.
- Responses are evaluated using LIWC lexical features, LLM-based pragmatic annotations, behavior-change technique classifications, and forced-choice behavioral coding.
- Behavioral scaffolding measures guidance that helps recipients select, plan, carry out, and evaluate actions, complementing linguistic analysis.
- Matched prompt–model observations support NP–NL comparisons, with standardized continuous effects, action percentages, percentage-point differences, and false-discovery-rate control.
4 Linguistic Analysis
Native-language generation and native-speaker persona prompting produce distinct linguistic profiles in multilingual interpersonal advice. Persona prompting adds lexical social cues, whereas native-language generation more often preserves relational, pragmatic, and concrete detail, with differences varying by language.
- Overall Linguistic Patterns: NP increases affiliation, social references, prosocial language, and positive tone relative to NL.
- Overall Linguistic Patterns: NP is less formal, emotionally expressive, socially attuned, and concrete than NL.
- Overall Linguistic Patterns: NL provides more contextual and relational elaboration, including emotional nuance, relationship dynamics, and concrete interpersonal detail.
- Overall Linguistic Patterns: Additional lexical social cues in NP do not correspond to greater attention to recipients’ feelings, relational concerns, or specific circumstances.
- Language Variation: NP–NL differences vary systematically by language, with smaller social-attunement and concreteness gaps for Japanese, Korean, and Chinese than for several Romance and Germanic languages.
- Language Variation: Substituting NP for NL can widen or narrow observed cross-language differences and change conclusions about model behavior.
5 Behavioral Analysis
NP provides less behavioral scaffolding than NL and shifts forced-choice recommendations toward confrontation, with differences varying across topics and languages.
- Behavioral Guidance in Generated Advice: NP provides less behavioral scaffolding than NL across major behavior-guidance categories.The largest reductions appear in problem solving, action planning, consequence reasoning, feedback, and cognitive reframing.
- Behavioral Guidance in Generated Advice: These reductions suggest a strategy-level effect because they appear consistently across languages.
- Behavioral Guidance in Generated Advice: The loss of usable guidance is most pronounced for friendship and workplace questions, especially in problem solving, action planning, and cognitive reframing.
- Forced-Choice Action Preferences: NP and NL produce significantly different action distributions, with NL shifting toward redirection and NP selecting confrontation more often.
- Forced-Choice Action Preferences: Across most languages, NP selects confrontation more often than NL, while NL selects redirection more often.Persona prompting preserves some relative language-level structure but shifts absolute recommendations toward confrontation.
- Discussion: NP and NL are not interchangeable: NP can make advice more socially polished while reducing actionable scaffolding and changing recommended actions.
6 Robustness Analysis
Additional analyses tested sensitivity to persona wording, translation quality, and evaluator choice. The main NP–NL directional patterns persisted, although some effects varied in magnitude and significance.
- Overall Robustness: The main directional patterns were consistent across robustness checks, although some effects varied in magnitude and significance.
- Persona Wording: 14 of 17 linguistic dimensions retained the same NP–NL direction across three persona formulations.The three reversals were confined to nonsignificant contrasts, and all eight BCT dimensions remained negative under every formulation.
- Translation Quality: 59 of 60 LLM-annotated language–dimension contrasts retained their direction after excluding lower-similarity translations.
- Translation Quality: 131 of 144 LIWC contrasts retained their direction, while all 96 BCT contrasts remained negative.Twelve of the thirteen LIWC reversals involved near-zero effects.
- Translation Quality: Confrontation and redirection retained their directions in all 12 languages after translation-quality filtering.
- Evaluator Choice: The second evaluator agreed with the original judge on the direction of every aggregate effect across five linguistic-style dimensions.Across the broader evaluation, agreement was 12 of 13 effects; the sole reversal concerned directness, where the original estimate was null.
7 Conclusion
Across 600 advice questions, 13 language conditions, and eight models, NL and NP produced systematically different styles, behavioral scaffolding, and action recommendations. The findings identify prompting strategy as a substantive methodological choice for multilingual advice systems.
- Conclusion: NL and NP are not equivalent: they yield different linguistic styles, behavioral scaffolding, and action recommendations.
- Conclusion: Relative to NL, NP often amplifies lexical social cues while reducing concreteness, social attunement, emotional expressiveness, formality, and actionable scaffolding.
- Conclusion: NP more often shifts forced-choice recommendations from redirection toward confrontation.
- Conclusion: Prompting strategy is a substantive methodological choice, so multilingual advice systems should evaluate elicitation pathways and validate linguistic and behavioral outcomes.
Limitations
The study’s conclusions are constrained by the English-language source corpus, broad persona conditions, automated evaluation, and model composition. These boundaries limit claims about real language communities, cultural identities, evaluator validity, and generalizability.
- Corpus and cultural scope: The English-language advice forum may embed Western, English-speaking assumptions, so findings concern model responses to translated corpus questions rather than real language communities.The authors identify alignment with actual speakers as an open question.
- Persona scope: Native-speaker prompts collapse region, class, age, and gender, making them unsuitable for representing cultural identity.The authors recommend more specific regional, dialectal, and situational contexts.
- Evaluation: Automated lexical, LLM-based, and BCT annotations may miss nuance or reflect annotator-model biases.Future evaluation should include human raters from relevant linguistic and cultural backgrounds.
- Model generalizability: Differences in training data, alignment, and English-centered development may limit generalizability across models.The authors propose testing systems developed primarily in non-English contexts.
A.1 Advice Question Samples
The study builds a multilingual interpersonal-advice dataset and compares several response-generation pathways with automated linguistic, behavioral, and action evaluations. Translation candidates are selected for semantic similarity and back-translation stability before model responses are analyzed.
- Advice question samples: The dataset contains 600 English-language interpersonal advice questions spanning workplace, friendship, relationship, and family topics.Questions exclude country-specific tags to focus on broadly comparable dilemmas.
- Translation pipeline: Each non-English question is translated into 12 target languages using 32 candidates, back-translated four times, and scored against the English original.Selection favors high mean similarity and low back-translation variance.
- Response-generation prompts: The main comparison contrasts native-language generation followed by English translation (NL) with English native-speaker persona prompting (NP).Auxiliary TTG and OSP conditions help localize input-translation and target-language effects.
- Evaluation setup: Responses are evaluated for directness, formality, emotional expressiveness, social attunement, concreteness, behavioral scaffolding, and forced-choice action.The action labels are confrontation, disengagement, and redirection.
- Evaluation exclusions: 2.34% of forced-choice rows were excluded for labeling failures, while linguistic and BCT parsing failures accounted for 1.27% and 1.38% of rows.Most forced-choice failures were unlabeled paraphrases, including refusals, meta-commentary, or explanatory responses.
- Model variation: Model-level effects vary substantially, indicating that prompting-strategy effects differ across models.The paper links this variation to future study of training data, model origin, alignment, and multilingual coverage.
D.5 Linguistic Effects by Language Group
NP–NL differences are not uniform across language groups, features, or prompting formulations. Across the reported analyses, NP shifts linguistic framing and action recommendations while consistently providing less behavioral scaffolding than NL.
- Language-group effects: NP–NL linguistic gaps depend on both the target-language group and the measured feature rather than forming one stable transformation.Figure 12 aggregates responses from 600 questions and eight models.
- Baseline comparisons: NP shows the largest overall departure from the English baseline in composite BCT effects, while NL shifts forced-choice responses away from confrontation toward redirection relative to English.These results are reported in separate cross-strategy tables.
- Behavioral scaffolding: All NP formulations provide less behavioral scaffolding than NL, with six dimensions significantly lower across formulations.The shared significant dimensions include goal setting, action planning, consequence reasoning, self-monitoring, problem solving, and cognitive reframing.
- Forced-choice actions: All three NP formulations are more confrontation-oriented and less redirection-oriented than NL in forced-choice recommendations.Effect magnitudes and statistical significance remain sensitive to persona wording.
H Sensitivity to Translation Quality
Filtering for high-quality prompt translations leaves the principal NP–NL patterns largely unchanged. Linguistic contrasts are mostly directionally stable, every BCT contrast remains negative, and the confrontation-versus-redirection shift persists.
- Filtering procedure: The translation-quality filter retains 4,015 of 7,200 prompt–language pairs, or 55.8 percent, using mean back-translation similarity ≥0.95.Because retention differs by language, filtered NP–NL contrasts are estimated separately within each language.
- LLM-annotated features: 59 of 60 language–dimension LLM-annotated contrasts retain their direction after filtering, with a mean absolute change of approximately 0.03 z-score units.The largest change is 0.13, and NP remains generally more direct while NL remains more formal, emotionally expressive, socially attuned, and concrete.
- LIWC features: 131 of 144 LIWC language–dimension contrasts retain their direction, while most reversals involve estimates no larger than 0.09 in absolute value.The stronger lexical patterns remain more affiliation, positive tone, social references, and prosocial language under NP.
- Behavioral scaffolding: All 96 BCT language–dimension contrasts remain negative after filtering, with a mean absolute change of approximately 0.02 z-score units.The maximum change is 0.09, preserving lower behavioral scaffolding under NP in every language and dimension.
- Forced-choice actions: At the language level, confrontation and redirection contrasts retain their directions in all 12 languages, whereas smaller disengagement contrasts are less stable.The aggregate shift from redirection toward confrontation therefore persists under the quality filter.
I Cross-LLM-as-a-Judge Comparison
The principal NP–NL findings are largely robust to changing the automated evaluator, including when gpt-4o evaluates its own responses. Agreement remains strong across linguistic-style and BCT dimensions, with limited discrepancies.
- Scope of the robustness check: The robustness check applies only to linguistic-style and BCT dimensions because LIWC and forced-choice outcomes do not use an LLM evaluator.The comparison uses identical rubrics and matched significance-testing procedures across two evaluators.
- Linguistic-style robustness: 96.7% of language-by-dimension effects receive the same direction from both evaluators, with standardized effect sizes correlated at r = .89.Across generation models, evaluators also agree on 13 of 15 model-by-dimension effects, with effect sizes correlated at r = .90.
- BCT robustness: Seven of eight aggregate BCT effects have the same direction under both evaluators and indicate less behavioral scaffolding under NP than NL.The sole discrepancy is feedback: gpt-4o estimates a moderate NP–NL reduction, whereas claude-sonnet-5 estimates a near-null effect.
- Evaluator independence: For gpt-4o-generated responses, evaluators agree on 12 of 13 effects, with standardized effect sizes correlated at r = .77.The only directional difference is directness, estimated as near-null by the original evaluator and positive by claude-sonnet-5.
- Overall conclusion: Overall, the principal directional NP–NL findings persist across automated evaluators, despite a small number of dimension-specific differences.The evaluation covers five linguistic-style and eight BCT dimensions on a stratified subset spanning 100 prompts, 12 non-English languages, and three generation models.