Source-linked AI summary
Toward Cultural Alignment: Human-Centered Evaluation of Multimodal AI Stories Across Five African Communities
Millicent Ochieng, Felermino D. M. A. Ali, Elizabeth A. Ankrah, Najeeb Gambo Abdulhamid, Migisha Boyd, Stephanie Nyairo, Mercy Muchai, Samuel Chege Maina, Aditya Vashistha, Anja Thieme, Jacki O'Neill
TL;DR
AI-generated multimodal stories can appear locally plausible while failing to fit communities’ lived practices, relationships, language, values, and visual expectations. The paper evaluates such stories with 19 representatives across five African communities and tests five multimodal LLM judges. Cultural alignment depends on contextual fit rather than recognizable markers alone, while judge reliability and calibration vary across communities, motivating community-calibrated evaluation with human review where needed.
Problem
AI-generated content may appear fluent and locally plausible while failing to reflect the lived practices, relationships, language, values, and visual expectations of represented communities.
Method
The study combines quantitative annotations, story-level scores, and qualitative focus groups from 19 culture representatives across five African communities, then evaluates five multimodal LLM judges against their judgments.
Results
Cultural alignment depends on contextual fit across five marker categories and eight misalignment mechanisms, while judge reliability and score calibration vary substantially across communities.
Takeaways & Limitations
Automated judges should be used only after community-specific validation, with community judgments establishing where automation can be trusted and human review remains necessary.
Takeaways & Limitations
The evaluation relies on 19 representatives who do not capture the full diversity of views within each community.
Abstract
from arXiv · showhide
In this paper, we examine how well AI-generated multimodal stories align with the lived practices, relationships, language, values, and visual expectations of the communities they represent. We conduct a community-grounded mixed-methods evaluation with 19 culture representatives across five African communities, combining quantitative annotations with qualitative focus group discussions. We find that cultural alignment depends not simply on recognizable cultural markers, but on how those markers fit social, linguistic, procedural, and visual context. From these evaluations, we develop a taxonomy of cultural alignment comprising five broader cultural marker categories and eight recurring mechanisms of misalignment. We additionally evaluate five multimodal LLM judges to examine whether automated evaluation can approximate community-grounded judgments at scale. Judge reliability and score calibration vary substantially across communities, with no single judge performing consistently across all five settings. These findings motivate community-calibrated evaluation pipelines in which automated judges are validated against community judgments to determine where they can be trusted and where human review remains necessary.
1 Introduction
The paper studies cultural alignment in AI-generated multimodal stories through community-grounded evaluation, asking whether recognizable cultural elements fit lived social, linguistic, procedural, and visual contexts. It also examines whether multimodal LLM judges can approximate these judgments across five African communities.
- Motivation: Community-grounded evaluation is difficult to scale because cultural judgments are complex, situated, and context-dependent.Outsiders may find forms of address, clothing, or food practices plausible even when community members experience them as inappropriate, foreign, generic, or incomplete.
- Approach: The study evaluates AI-generated multimodal stories with 19 culture representatives across Hausa, Kikuyu, Luo, AmaXhosa, and Xichangana communities.The evaluation combines quantitative annotations, story-level scores, and qualitative focus group discussions.
- Approach: The annotation platform captures frame-level text and image evidence, cultural marker categories, and overall story-level alignment judgments.Representatives identify influential text spans and image regions while evaluating how cultural elements shape alignment.
- Findings: Cultural alignment depends on contextual fit across five categories: referential, procedural, contextual, socio-geographic, and linguistic register markers.Examples include names and foods, preparation and exercise routines, settings, infrastructure, dialect, code-switching, and forms of address.
- Findings: Eight recurring mechanisms produce cultural misalignment: substitution, norm violation, omission, forced insertion, register conflation, cross-modal inconsistency, stereotyping, and hallucination.The paper also finds that LLM judges vary substantially across communities, with no single judge consistently matching community representatives.
2 Related Work
Prior research studies cultural knowledge and misalignment using demographic, value, norm, narrative, and visual proxies. This paper builds on that work while addressing the difficulty of evaluating culturally situated multimodal content at scale.
- Cultural alignment: Prior work finds that culture is difficult to operationalize through static demographic labels, national categories, or isolated value dimensions.Existing approaches include Hofstede’s dimensions, moral judgment datasets, and social etiquette norms.
- Cultural misalignment: Research on generated dialogue, narratives, and images documents culture-specific reasoning failures and forms of cultural misrepresentation.Examples include norm adherence and violation datasets, Western bias in stories about Arab cultures, and taxonomies for Indian stories.
- Scalable evaluation: LLM-as-judge methods offer scalable evaluation but remain sensitive to task framing, model bias, and subjective evaluation criteria.These sensitivities are especially consequential when judgments require lived and linguistic knowledge of a represented community.
3 Methodology
The study generates culturally situated multimodal stories and evaluates them through structured community annotations, focus groups, and multimodal LLM judges. The methodology captures both overall alignment and the specific textual, visual, and cultural evidence behind judgments.
- Study design: The study uses a mixed-methods pipeline spanning story generation, community-grounded cultural evaluation, automated validation, and characterization of misalignment mechanisms.Figure 2 presents the end-to-end workflow from generation through cultural evaluation.
- Story generation: Personas are generated through 19 sequential GPT-4.1 calls, conditioning each demographic or cultural attribute on previously generated attributes.The dependency-aware process is designed to maintain consistency across demographic, cultural, and regional attributes.
- Story generation: Each story contains four first-person text-image frames conditioned on a persona, a diabetes lifestyle question, a narrative arc, and one of three context settings.Stories address everyday diabetes lifestyle management rather than diagnosis or treatment.
- Story generation: The final dataset contains 199 multimodal stories across five communities, seven generation models, and three context settings.Images were produced through iterative editing with a fixed seed of 12345 and low editing strength of 0.3 to support continuity.
- Community evaluation: Culture representatives evaluate overall stories, individual frames, and the specific cultural markers influencing their judgments.The custom platform supports text-span highlighting, image-region selection, marker categorization, connectedness labels, influence ratings, comments, and story-level scores.
- Automated evaluation: Five multimodal LLM judges score the same stories with the community rubric and marker categories without community-specific calibration examples.The judges span proprietary and open-weight systems and evaluate text, images, and text-image coherence.
- Community evaluation: Nineteen participants from five African communities receive onboarding and practice training before individually evaluating 40 stories generated for their own communities.The study also conducts 15 focus group sessions, three per community, to explain agreement and disagreement in annotations.
- Taxonomy derivation: The taxonomy is derived through thematic analysis of focus group discussions and then examined in participant annotations.The resulting summary organizes cultural alignment categories derived from both focus groups and annotations.
4 Results
Community annotations show that cultural fit and disruption vary across modalities, communities, and marker categories, while misalignment arises through recurring contextual and representational failures. LLM judges also track community judgments unevenly and exhibit community-specific calibration problems.
- Cultural Marker Patterns Across Modalities: Text annotations were 86% connected (8,075/9,386), whereas image annotations were 74% connected (6,931/9,419), indicating more visible disruption in images.Not-connected annotations comprised 14% of text spans and 26% of image regions.
- Cultural Marker Patterns Across Modalities: Not-connected markers varied by community and category, with Hausa stories frequently showing Referential and Procedural disruptions in text.Figure 3 descriptively distributes not-connected markers by community, model, modality, and cultural marker category.
- Misalignment Mechanisms: Focus groups identified substitution, hallucination, forced insertion, norm violation, register conflation, cross-modal inconsistency, stereotyping, and omission as recurring misalignment mechanisms.Examples included replacing target-community markers, implausible cultural practices, generic narratives with inserted names, inappropriate dress or register, conflicting modalities, repeated tropes, and absent situated detail.
- LLM Judges Approximate Community Judgments Unevenly Across Cultures: LLM-judge correlations were very strong for Luo (r = 0.82–0.89), strong for Hausa (r = 0.67–0.80), moderate for Kikuyu (r = 0.50–0.64) and Xichangana (r = 0.39–0.51), and nonsignificant for AmaXhosa after correction.The table reports Pearson correlations between judge scores and mean culture-representative scores, with significance assessed using Bonferroni correction.
- LLM Judges Approximate Community Judgments Unevenly Across Cultures: Community-rating agreement ranged from good for Hausa (ICC(A,1)=0.81) to poor for Xichangana (ICC(A,1)=0.33).Agreement was moderate for Luo, AmaXhosa, and Kikuyu (ICC(A,1)=0.60–0.63).
- LLM Judges Approximate Community Judgments Unevenly Across Cultures: Judges overestimated Xichangana alignment by 35–45 points and some judges underestimated Kikuyu alignment by 14–16 points relative to culture representatives.For Xichangana, all paired t-tests had p < 0.001; Kikuyu biases were reported for Kimi and GPT-5.5.
5 Discussion
The discussion argues that cultural alignment depends on contextual fit rather than merely recognizable markers, and that automated judges require community-specific validation. Community judgments provide the reference needed to identify where automation is trustworthy and where human review remains necessary.
- Implications for Cultural Alignment Evaluation: Cultural alignment depends on whether names, foods, places, clothing, language, interactions, and settings fit the story context and represented community, not merely on marker presence.The taxonomy separates broader marker categories from the mechanisms that produce misalignment.
- Implications for Automated Evaluation: LLM judges vary substantially in reliability and calibration across communities, so no single judge consistently approximates community judgments across all five settings.Both correlation with community scores and score calibration matter because a judge can track differences while systematically shifting absolute scores.
- Implications for Automated Evaluation: Community-calibrated pipelines should validate automated judges against culture-representative reference judgments before relying on them for scalable evaluation.The taxonomy identifies cultural dimensions for judges to attend to, while its misalignment mechanisms make community feedback more actionable.
6 Conclusion
The paper presents a human-centered evaluation of multimodal AI stories across five African communities and shows that cultural alignment depends on contextual fit. It develops a taxonomy of cultural markers and misalignment mechanisms, and finds that judge reliability and calibration vary across communities.
- The study evaluates cultural alignment in AI-generated multimodal stories across five African communities using community judgments.
- Cultural alignment depends on how recognizable markers fit narrative, social, linguistic, procedural, and visual context.
- The authors develop a taxonomy containing five broader cultural marker categories and eight recurring mechanisms of misalignment.
- LLM-judge reliability and score calibration vary substantially across communities, with no single judge performing consistently across all five settings.
- The findings support community-calibrated evaluation in which community judgments ground assessment and automated judges are validated for trustworthiness.
Limitations
The study’s conclusions are bounded by a small and non-exhaustive participant sample, a focus on reader-independent alignment judgments, a diabetes-story domain, cross-community comparability concerns, and limited judge configurations.
- Participant scope: The 19 culture representatives provide situated knowledge but do not capture each community’s full diversity of views.Perspectives may vary by age, gender, region, language use, religion, class, and rural-urban experience.
- Outcome scope: The evaluation measures story-level and marker-level cultural alignment, not whether stories affect understanding, trust, behavior change, or lifestyle decisions.The authors call for future work connecting alignment judgments to reader-facing outcomes.
- Domain scope: The study focuses on everyday Type II diabetes lifestyle and behavior-change stories, so observed marker patterns, misalignment mechanisms, and judge behavior may differ elsewhere.The framework therefore requires validation across other generated-content domains and application settings.
- Comparability: Cross-community score comparisons require caution because Xichangana stories differed in language context and evaluation patterns and were systematically over-scored by judges.The authors caution that score differences should not be treated as direct measures of cultural-alignment difficulty.
- Judge scope: The judge analysis covers five multimodal judges and one fixed prompt, so reliability may change with other models, prompts, rubrics, or calibration methods.The authors recommend community-calibrated pipelines that identify where automation is trustworthy and where human review remains necessary.
Ethical Considerations
The study frames its stories as culturally contextualized lifestyle narratives rather than medical advice, and grounds evaluation in distinct communities and representative participation.
- Health-content boundary: The stories address healthy eating, physical activity, and everyday Type II diabetes management practices, without diagnoses, prescriptions, or medication recommendations.Cultural alignment is not evidence of medical correctness or safety; health-related content still requires medical review.
- Community context: The study treats Hausa, Kikuyu, Luo, AmaXhosa, and Xichangana as distinct evaluative contexts spanning different languages, geographies, histories, and social expectations.These differences include naming practices, foodways, dress norms, and everyday social expectations.
- Evaluation ethics: The annotation framework applies the same cultural marker categories across text and images, with representatives selecting influential spans or regions.This supports multimodal cultural evaluation while retaining community-based judgment.
C Annotation Interface and Evaluation Protocol
The study uses a custom platform and rubric to evaluate each text-image frame, cultural marker, and overall story, combining structured annotations with focus-group discussion.
- Annotation interface: The custom platform supports frame-by-frame evaluation, influential text-span and image-region identification, marker categorization, and overall story-level scoring.It was built because existing tools did not support this multimodal evaluation workflow.
- Marker evaluation: Annotators assign cultural marker categories to highlighted text spans or image regions and assess depth, authenticity, and paragraph-image alignment.They also penalize mismatched, generic, or culturally contradictory paragraph-image pairs.
E Additional Quantitative Results
Additional results quantify annotation volume, human agreement, and connected cultural markers, while providing span-level ground-truth procedures for evaluating automated judges.
- Annotation volume: 18,805 span-level annotations covered 199 stories, comprising 9,386 text spans and 9,419 image regions.Text annotations were 86% connected and 14% not connected; image annotations were 74% connected and 26% not connected.
- Deduplication: The annotation statistics distinguish connected from not-connected spans and report exact and partial deduplication based on identical text or token overlap above 50%.Image regions are separately clustered using bounding-box IoU greater than 0.3 across annotators.
- Annotation patterns: Visually salient cultural markers in images were more consistently recognized across annotators than textual markers after deduplication.The comparison is based on deduplicated annotation patterns across modalities.
- Ground truth: Ground truth for cultural-marker span evaluation aggregates representative annotations by majority vote before comparing LLM predictions against the resulting spans.The Hausa example illustrates overlapping dimensions such as names, dietary practices, local expressions, and occupation routines.
- Human agreement: Human annotators achieved span-level F1 of 0.36–0.47 and category agreement of 0.59–0.74.Annotators generally agreed on categories but differed more in span boundaries, establishing an empirical ceiling for automation.
E.1.2 LLM Judge Self-Consistency
Repeated evaluations show that the LLM judge produces relatively stable outputs, with consistency varying across languages. The evaluation compares predictions aggregated across runs with human-annotated ground truth.
- Evaluation setup: Inter-run agreement is measured pairwise across three identical runs.Table 16 reports the resulting agreement statistics.
- Self-consistency: Span F1 ranges from 0.68–0.82 and Category Agreement from 0.76–0.91 across repeated evaluations.These self-consistency scores substantially exceed human agreement, although Xichangana shows lower consistency and greater uncertainty.
- Evaluation setup: Predictions are aggregated across judges using majority voting and compared with human-annotated ground truth using NER-style metrics.Results are averaged across three runs.
E.1.4 Analysis
The analysis finds that LLM judge performance is constrained by mismatches in span boundaries, category assignments, and prediction density. Category-level agreement is stronger than exact span matching, while results vary across languages.
- Analysis: Strict matching yields F1 = 0.09, compared with type matching F1 = 0.46, showing stronger category recognition than boundary agreement.Partial matching reaches F1 = 0.22, indicating limited span overlap with human annotations.
- Span mismatch: Human annotators select short phrases, whereas the LLM judge often extracts longer sentence-level spans containing surrounding context.This span-granularity mismatch lowers strict and partial scores without necessarily indicating disagreement about cultural content.
- Span mismatch: For AmaXhosa, the model predicts 1,630 spans versus 781 human spans, while Luo and Kikuyu predictions are lower than human counts.The corresponding counts are 1,636 versus 2,453 for Luo and 1,479 versus 2,068 for Kikuyu.
- Category confusion: Type-level F1 ranges from 0.41–0.52, with category confusion especially pronounced for broad place and occupation categories.The model tends to over-predict broad categories at the expense of more specific ones.
- Ceiling effects: Human span-boundary agreement is F1 = 0.42, while model type-level F1 = 0.46 approaches human category agreement of 0.67.This comparison indicates that category-level performance is closer to human performance when boundary constraints are relaxed.
- Cross-language variation: Luo achieves the highest strict F1 at 0.107, whereas AmaXhosa has the lowest strict F1 at 0.072 but the highest type recall at 0.623.The passage attributes these differences potentially to annotation density, category distribution, and how discretely cultural markers are expressed.