Source-linked AI summary

CHARM: Character Hallucination for Multicultural Role Play Benchmark

Sunkyung Han, Nahyeon Park, Gaeun Seo, Seunghyun Yoon, JinYeong Bak

arXiv:2609.01352v1cs.CLcs.AI

TL;DR

Prior evaluations rarely distinguish recognizing a character’s knowledge boundary from complying with it. CHARM introduces a multicultural, abstention-enabled benchmark and two-stage evaluation, finding that hallucination is driven predominantly by compliance failures, often reflecting parametric overrides, with culturally patterned variation.

  • Problem

    Prior evaluations rarely distinguish failures to recognize a character’s knowledge boundary from failures to comply despite recognizing it.

  • Method

    CHARM benchmarks 40 characters from five cultural-linguistic regions across Temporal and Cross-Universe boundaries using abstention-enabled questions and separate awareness, compliance, and verification stages.

  • Results

    Across six LLMs, hallucination is driven primarily by compliance failures rather than boundary unawareness, with many cases verified as parametric overrides.

  • Takeaways & Limitations

    Role-playing agents should be evaluated for culturally grounded adherence to knowledge boundaries, not only for boundary recognition or factual accuracy.

  • Takeaways & Limitations

    CHARM samples five cultural-linguistic regions and does not represent the full diversity of global cultures and languages.

Abstract

from arXiv · show

Role-playing large language models (LLMs) are expected to adopt a character's style while also respecting that character's knowledge boundaries. Prior evaluations detect character hallucination but rarely distinguish whether errors arise from failure to recognize a boundary or from failure to comply despite recognition. We introduce CHARM, a multicultural benchmark of 40 real and fictional characters drawn from five cultural-linguistic regions, and validated by native reviewers. It probes two boundary types, Temporal (historical vs. modern) and Cross-Universe (entities outside a character's narrative or historical universe), using abstention-enabled multiple-choice questions. We propose a two-stage evaluation that separates Boundary-Awareness (explicit recognition that a query is out of scope) from Boundary-Compliance (abstention when answering concrete questions). Evaluations across six LLMs show that hallucination is driven predominantly by compliance failures. Models frequently acknowledge that a query lies outside the character's knowledge yet still provide factual, out-of-character answers. By re-posing the same questions to the target character, we confirm that a large fraction of these cases are verified parametric overrides; the model stores the relevant fact but fails to suppress it. We also observe systematic cultural variation in these failures, consistent with imbalances in how characters from different regions are represented in model knowledge.

1 Introduction

Role-playing LLMs must preserve characters’ knowledge boundaries, but prior evaluations rarely separate boundary recognition from compliance. CHARM addresses this gap with a multicultural benchmark and a two-stage diagnostic framework.

  • Role-playing LLMs must imitate a character’s style while respecting the character’s plausible knowledge boundary.
  • Existing evaluations rarely distinguish recognizing a knowledge boundary from suppressing out-of-character knowledge during answers.Models may possess relevant facts parametrically even when the simulated character should not.
  • CHARM evaluates 40 real and fictional characters from five cultural-linguistic regions using a multicultural benchmark.
  • The benchmark probes Temporal boundaries involving modern concepts and Cross-Universe boundaries involving entities outside a character’s narrative or historical universe.
  • CHARM separates Boundary-Awareness from Boundary-Compliance to identify whether hallucinations reflect boundary-recognition or compliance failures.Knowledge Verification further examines whether Compliance Gap cases arise from parametric override.
  • Across six LLMs, hallucination is driven more by compliance failures than by failures to recognize boundaries, with parametric override and cultural variation as diagnostic phenomena.

2 CHARM

CHARM constructs a validated benchmark around two knowledge-boundary types and evaluates them through awareness, compliance, and verification questions. Its design combines balanced character coverage with abstention-enabled multiple-choice testing.

  • Benchmark scope: CHARM evaluates knowledge-boundary violations across Temporal and Cross-Universe boundaries.Temporal items pair historical characters with modern concepts, while Cross-Universe items pair characters with entities outside their narrative or historical universe.
  • Evaluation stages: The benchmark uses Boundary-Awareness to measure explicit recognition and Boundary-Compliance to measure adherence in practice.
  • Characters: CHARM curates 40 characters from five cultural-linguistic regions, with eight characters per region balanced by reality status and historical period.
  • Temporal Questions: Temporal awareness questions ask whether historical characters recognize modern concepts, with the correct answer always “No”.Thirty modern concepts are instantiated through multiple templates to reduce format-based pattern matching.
  • Temporal Questions: Temporal compliance questions are five-choice MCQs whose correct answer is abstention, with four distractors designed for factual plausibility and structural diversity.
  • Cross-Universe Questions: Cross-Universe verification questions ask target-specific factual questions while prompting the model as the target entity, testing whether relevant facts remain in its parameters.
  • Validation: CHARM contains 680 awareness, 1,332 compliance, and 736 verification questions, validated by two native reviewers per region.

3 Experiments

The experiments evaluate six LLMs with independent awareness and compliance probes, matrix-based failure decomposition, parametric-override verification, and regional comparisons. Results show compliance failures dominate recognition failures, many correct non-abstentions reflect accessible parametric knowledge, and hallucination varies across cultural regions.

  • Experimental setup: Six LLMs independently answer Boundary-Awareness and Boundary-Compliance questions, with knowledge-verification questions used to test parametric overrides.Awareness requires recognizing that a query exceeds the character’s boundary, while compliance requires selecting abstention.
  • Evaluation metrics: The BA–BC matrix pairs awareness and compliance outcomes to distinguish compliance gaps from recognition failures among hallucination cases.Compliance Gap is (BA = True) ∧ (BC = False), whereas Recognition Failure is (BA = False) ∧ (BC = False); independent runs make this a cross-probe dissociation.
  • Main results: Compliance Gap rates consistently exceed Recognition Failure rates across models, indicating that hallucination more often reflects failure to comply with recognized boundaries.GPT-4o illustrates the dissociation with 91.3% BA accuracy, 72.1% Compliance Gap, and 8.9% Recognition Failure; Cross-Universe gaps also exceed Temporal gaps.
  • Parametric override analysis: For five of six models, 78–100% of factually correct Compliance Gap cases are confirmed as parametric overrides.Verification requires correct boundary recognition, a factually correct non-abstaining compliance answer, and a correct answer when responding as the target character.
  • Cultural patterns: Regional differences appear in Compliance Gap and parametric-override rates, with Western regions especially EN and Spain showing higher gap rates than Korea and Indonesia.The reported regional pattern is consistent with imbalances in how strongly characters from different cultural regions are represented in model knowledge.

4 Conclusion

CHARM introduces a two-stage framework for diagnosing character hallucination across Temporal and Cross-Universe boundaries. Across six LLMs, hallucination primarily reflects compliance failures, including verified parametric overrides, with variation across cultural contexts.

  • CHARM diagnoses character hallucination across Temporal and Cross-Universe knowledge boundaries using a two-stage evaluation framework.
  • Across six LLMs, hallucination stems primarily from compliance failures rather than boundary unawareness.
  • Many hallucinations are verified parametric overrides in which models fail to suppress accessible knowledge under role constraints.
  • Regional analyses show that this failure pattern varies across cultural contexts, motivating culturally grounded evaluation of knowledge-boundary adherence.

Limitations

The authors identify limitations concerning CHARM’s cultural coverage, multiple-choice format, and interpretation of parametric-override results. Future work will broaden regions and characters, support interactive evaluations, and investigate knowledge sources.

  • CHARM samples five cultural-linguistic regions and does not represent the full diversity of global cultures and languages.
  • Future work will expand regions and characters, support uncertainty and clarification, and apply provenance and attribution methods to investigate knowledge sources and causal mechanisms.
  • The benchmark uses abstention-enabled multiple-choice items, while future work will extend it to open-ended, interactive evaluations.
  • The parametric-override test demonstrates accessible knowledge under role prompts but does not prove its specific source.

Ethical Considerations

CHARM addresses cultural and ethical considerations through diverse character selection, exclusion of potentially problematic content, and native-speaker validation. The benchmark also acknowledges risks of stereotyping, misrepresentation, factual inaccuracies, and cultural misinterpretation.

  • CHARM includes real and fictional figures from diverse countries and eras, with characters selected across cultural-linguistic regions.
  • Culturally specific and historically sensitive content may reinforce stereotypes or misrepresent groups if cultural nuance is not adequately addressed.
  • Publicly available character information can still propagate factual inaccuracies or cultural misinterpretations through model evaluation.
  • The benchmark excludes potentially problematic content during construction and uses culturally and linguistically relevant annotators to validate accuracy and contextual appropriateness.

C Additional Results by Boundary Type

Table 6 reports the Compliance Gap rate separately for Cross-Universe and Temporal boundaries, enabling comparison of compliance failures by boundary type.

  • The Compliance Gap rate is reported separately for Cross-Universe and Temporal boundaries.

D Dataset Statistics

CHARM organizes character-boundary evaluation around two boundary types and three question stages, including a three-step verification of parametric overrides.

  • Evaluation stages: Boundary-Awareness asks whether the role character knows the target, whereas Boundary-Compliance tests whether the model abstains rather than selects a factual answer or distractor.FC-BC cases imply a compliance gap because the model recognizes the boundary but answers with factual knowledge.
  • Parametric override verification: The parametric override pipeline tests boundary awareness, factual compliance selection, and knowledge verification in sequence.An override requires BA=True, FC-BC=1, and FC-KVQ=1.
  • Parametric override verification: Knowledge Verification re-poses the factual content to the target character to assess whether the relevant knowledge is available parametrically.The denominator captures recognized-boundary cases with factual compliance answers; the numerator adds target-character factual verification to reduce chance correctness.
  • Boundary types: The benchmark covers Temporal and Cross-Universe boundaries, with Cross-Universe showing consistently higher dissociation than Temporal.Temporal and Cross-Universe gap rates are reported separately, and Cross-Universe questions concern entities outside the character’s narrative or historical universe.
  • Dataset organization: CHARM reports dataset statistics by region, boundary type, and question type, with examples and matrix breakdowns provided in Tables 7–9.BA uses binary Yes/No questions, while BC and KV use five-choice questions; BC includes abstention and distinguishes factual answers from unrelated distractors.

H Experiment Setting Detail

The experiments evaluate six LLMs under comparable prompting and independent single-pass runs for awareness, compliance, and knowledge verification.

  • Models: Six models are evaluated: three closed-source models and three open-source models.The models are GPT-4o, GPT-5.5, Gemini-3.5-Flash, Llama-3.1-8B-Instruct, Gemma-3-12B-IT, and Qwen3-8B.
  • Inference protocol: The evaluation uses comparable prompt formats and decoding parameters to support comparability and reproducibility.The common settings include temperature=0.0, top_p=0.95, and max_completion_tokens=256, with documented model-specific exceptions.
  • Inference protocol: Each model performs three independent inference runs for Explicit Awareness, Implicit Compliance, and Knowledge Verification questions.The runs are single-pass, with no fine-tuning or model modification.
  • Data processing: The study stores raw responses, parsed answers, and correctness labels, then computes accuracy by stage, boundary type, region, and model.Open-source models run locally, while closed-source models are accessed through APIs.

K Human Validation Detail

CHARM uses native regional reviewers and additional analyses to validate question quality, distractor plausibility, and model-performance comparisons.

  • Human validation: Two native reviewers from each region assess question appropriateness, answer correctness, and distractor quality.Flagged items are revised or removed when consensus cannot be reached.
  • Human validation: Validation checks criteria compliance, linguistic and cultural appropriateness, and semantic distinctiveness of distractors.Similar distractors are filtered and regenerated to prevent quality degradation.
  • Reporting: The benchmark includes model, licensing, and regional breakdowns, including Compliance Gap and parametric override rates.Table 17 reports regional means and standard deviations across six models; Table 18 provides the full model-by-region breakdown.
  • Reliability: Cohen’s κ indicates generally substantial inter-reviewer agreement across all regions.Agreement is reported by country in Table 13.
  • Sequential comparison: Under sequential BA→BC measurement, Cross-Universe BC accuracy rises from 10.4% to 80.6% for GPT-4o.The comparison uses 732 instances and is summarized in Table 15.
  • Distractor analysis: Close distractors consistently increase benchmark difficulty.Table 14 compares accuracy for high- and low-similarity questions and reports performance gaps.

M Sequential BA→BC Evaluation

The sequential BA→BC condition sharply improves abstention, but shared conversational context prevents attributing that improvement solely to genuine boundary compliance.

  • Evaluation design: The sequential condition asks GPT-4o to answer BA and then the corresponding BC question in the same conversation for 732 Cross-Universe instances.The main experiments measure BA and BC independently.
  • Results: BC accuracy increases from 10.4% to 80.6% under sequential rather than independent measurement.The preceding “No” response strongly influences subsequent refusal behavior.
  • Interpretation: The sequential setting cannot disentangle genuine boundary compliance from conversational consistency with the prior awareness response.The authors therefore do not attribute the improvement to boundary compliance alone.
  • Interpretation: Independent measurement avoids this consistency confound and provides a conservative lower bound on compliance capacity.The remaining gap under stricter measurement supports interpreting the dissociation as genuine.
  • Implications and caveat: Explicitly eliciting the character’s boundary may increase subsequent abstention, although further work is needed to separate genuine compliance from consistency bias.The limitation is specific to multi-turn, interactive settings.
  • Verification controls: Role removal and target-character verification converge, with No-role accuracy at 64.4–99.5% and KVQ accuracy at 50.5–100.0%.Requiring both verifications still yields 29.7–99.5%, above the 20% chance level.
  • Regional interpretation: Regional differences in Compliance Gap and override rates vary in magnitude across models and should be read as broad tendencies rather than fixed cultural properties.For example, override rates are Korea 68.9±30.6 and Spain 78.2±30.2.
Loading 2609.01352v1…