Source-linked AI summary
Computational Hermeneutics: Evaluating generative AI as a cultural technology
Cody Kommers, Ruth Ahnert, Maria Antoniak, Emmanouil Benetos, Steve Benford, Mercedes Bunz, Baptiste Caramiaux, Shauna Concannon, Martin Disley, James Dobson, Yali Du, Edgar Duéñez-Guzmán, Kerry Francksen, Evelyn Gius, Jonathan W. Y. Gray, Ryan Heuser, Sarah Immel, Richard Jean So, Sang Leigh, Dalaki Livingston, Hoyt Long, Meredith Martin, Georgia Meyer, Daniela Mihai, Ashley Noel-Hirst, Kirsten Ostherr, Deven Parker, Yipeng Qin, Jessica Ratcliff, Emily Robinson, Karina Rodriguez, Adam Sobey, Ted Underwood, Aditya Vashistha, Matthew Wilkens, Youyou Wu, Yuan Zheng, Drew Hemment
TL;DR
Existing AI evaluation often treats culture as a measurable variable rather than a foundational part of GenAI’s operation, despite the open-ended cultural contexts in which these systems are developed and used. The paper develops computational hermeneutics as an interpretive framework centered on situatedness, plurality, and ambiguity, and proposes evaluation that examines systems, people, and cultural context. Its conclusion is a shift from standardized accuracy questions toward contextual questions about meaning.
Problem
AI evaluation commonly treats culture as a secondary variable and relies on standardized tasks, although GenAI systems operate across open-ended cultural contexts.
Method
The paper develops computational hermeneutics by applying humanities-based interpretation to GenAI systems, their architectures, interactions, and cultural outputs.
Results
The paper identifies situatedness, plurality, and ambiguity as inherent interpretive challenges and proposes iterative, people-inclusive, context-sensitive evaluation principles.
Takeaways & Limitations
GenAI evaluation should treat culture as foundational to system operation and assess contextual meaning rather than only standardized model output.
Takeaways & Limitations
The paper focuses on benchmark-based evaluation, while benchmarks can fail to meaningfully assess what they purport to measure because standardized metrics may function as inadequate proxies.
Abstract
from arXiv · showhide
Generative AI systems are increasingly recognized as cultural technologies, yet current evaluation frameworks often treat culture as a variable to be measured rather than fundamental to the system's operation. Drawing on hermeneutic theory from the humanities, we argue that GenAI systems function as "context machines" that must inherently address three interpretive challenges: situatedness (meaning only emerges in context), plurality (multiple valid interpretations coexist), and ambiguity (interpretations naturally conflict). We present computational hermeneutics as an emerging framework offering an interpretive account of what GenAI systems do, and how they might do it better. We offer three principles for hermeneutic evaluation -- that benchmarks should be iterative, not one-off; include people, not just machines; and measure cultural context, not just model output. This perspective offers a nascent paradigm for designing and evaluating contemporary AI systems: shifting from standardized questions about accuracy to contextual ones about meaning.
1 Introduction
GenAI systems are cultural technologies whose development, evaluation, and effects depend on culture and context. The paper argues that evaluation should treat culture as foundational to interpretation rather than as an optional variable measured through standardized tasks.
- GenAI systems depend on cultural norms, assumptions, meanings, practices, and social dynamics across their development, evaluation, and effects.
- AI evaluation often treats culture as a secondary variable, such as bias, a generalization constraint, an ethical parameter, or user-preference variability.
- Frontier models generate diverse cultural artifacts across open-ended contexts, making cultural considerations inseparable from their development and use.
- The paper reframes culture as a dynamic, contested space where meaning is made, challenging benchmarks built around universal tasks with convergent solutions.
- Quantitative proxy metrics can trivialize cultural activities and miss what makes outputs meaningful.
- Computational hermeneutics uses humanities-based interpretation to evaluate GenAI systems and identify when interpretations are legitimate.
- GenAI systems are interpretive processes whose data, architectures, and algorithms embed interpretive stakes, decisions, and processes without being equivalent to human interpretation.
- The paper identifies situatedness, plurality, and ambiguity as hermeneutic challenges and proposes iterative, people-inclusive, context-sensitive evaluation principles.
2 Computational Hermeneutics
Computational hermeneutics applies humanities theories of interpretation to GenAI, treating meaning as contextual, iterative, plural, and ambiguous. It therefore evaluates both models and interactions while shifting assessment from fixed correctness toward contextual legitimacy and appropriateness.
- Computational Hermeneutics: Hermeneutics studies how interpretations of cultural artifacts gain meaning and legitimacy within social or historical contexts.
- Computational Hermeneutics: The hermeneutic circle interprets specific parts and the whole iteratively, updating each understanding through the other.
- Computational Hermeneutics: Computational hermeneutics treats GenAI outputs as interpretive rather than binary right-or-wrong responses to cultural artifacts.
- Computational Hermeneutics: Judgments about cultural interpretations depend on the assumptions underlying the interpretive process.
- Computational Hermeneutics: Evaluation should examine both model-level architecture and context-specific dialogic interactions, distinguishing their general and particular effects.
- Hermeneutic Challenges for AI: Situatedness, plurality, and ambiguity make apparently peripheral model features significant interpretive choices rather than mere variables to optimize.
- Situatedness: Meaning depends on the historical or social context in which a cultural artifact is made, used, or perceived.
- Situatedness: Hermeneutic evaluation rejects a universal view from nowhere and requires the specific perspective offered by a model to be identified.
3 Generative AI systems as “Context Machines”
The paper frames GenAI systems as performing interpretation: they consolidate contextual cues, accommodate multiple possible meanings, and co-construct interpretations through human interaction. This interpretive capacity is shaped by both model mechanisms and the human choices surrounding system development and use.
- GenAI systems “do” interpretation as a fundamental capacity, although their interpretive processes differ from those of human interpreters.Interpretation occurs both internally within models and dialogically through interactions with people.
- GenAI systems function as “context machines” that use contextual cues to generate the next relevant token, pixel, or other value.Vector-space embeddings encode co-occurrence statistics and contextual information, whose probabilistic decoding accommodates multiple interpretations.
- Self-attention can be understood as a computational hermeneutic circle, iteratively updating token-level and sequence-level interpretations in relation to each other.The mechanism relates partial and holistic interpretations rather than treating tokens independently.
- GenAI systems co-construct interpretations with humans through interaction, so interpretive capacity arises from models together with the interfaces and interactions that frame them.This perspective does not treat AI as a substitute for human expertise or interpretation.
- Human choices shape GenAI interpretation through training data, objectives, reinforcement learning, prompting, and annotation practices.These choices encode particular goals, values, assumptions, and sometimes conflicting interpretations.
- GenAI systems also affect people by shaping metacognition, relational assumptions, thought-partner practices, explanations, and creative experiences.The collaboration therefore has effects in both directions, from people to machines and from machines to people.
4 Operationalizing Hermeneutics in AI
The paper proposes evaluating GenAI through contextual interpretation rather than standardized accuracy alone. Its three principles are iterative evaluation, human-inclusive assessment, and attention to cultural context in both interactions and outputs.
- Reframing benchmarks: Hermeneutic benchmarking shifts evaluation from standardized questions about accuracy toward contextual questions about meaning.The paper argues that cultural production varies too much across contexts for a comprehensive standardized task suite.
- Three principles: Hermeneutic benchmarks should be iterative, include people, and measure cultural context rather than relying only on isolated model outputs.These three principles are presented as ways to make AI benchmarks better reflect a hermeneutic lens on culture.
- Iterative evaluation: Evaluation should unfold across multiple prompts or exchanges because cultural outputs develop within evolving interpretive conversations.A single-prompt score can provide a limited and unreliable assessment.
- Iterative evaluation: Benchmarks should assess both holistic model capabilities and behavior within specific dialogic frames rather than only aggregate performance.Instance-by-instance evaluation helps address limits on generalizability associated with average metrics.
- People in evaluation: Human-AI collaboration should be evaluated directly because interpretive processes are bound up with the people using the systems.This includes examining interactive configurations and the dialogue through which interpretations are produced, not only final outputs.
- People in evaluation: Evaluating harms and cultural expectations increasingly involves human use and judgment, including benchmarks with more than 10,000 human annotations.These examples assess capabilities together with human interaction, norms, and lived experience.
- Cultural context: Cultural context should be evaluated as the medium through which performance emerges, including how and why responses become appropriate within particular cultural frameworks.Thin like/dislike or positive/negative signals cannot provide this contextual grounding.
- Cultural context: Context-sensitive evaluation can examine real-world use cases and distinguish cultural judgments by incorporating demographic or sociocultural markers.Examples include comparing Chinese and American viewer norms and organizing human feedback by demographic information.
5 Discussion
The discussion presents computational hermeneutics as a potential reframing of GenAI: culture is foundational to system operation, and systems become interpretive partners rather than merely answer generators. The paper focuses specifically on benchmark-based evaluation while positioning hermeneutics as relevant to broader technological and societal consequences.
- Computational hermeneutics reframes culture as foundational to GenAI operation and characterizes systems as interpretive partners engaging situatedness, plurality, and ambiguity.This shifts the conception of GenAI beyond answer generation toward participation in human meaning-making.
- The paper focuses on evaluating GenAI systems through benchmarks, while noting that hermeneutic perspectives could also inform debates about training-data culture.Benchmarking is presented as one possible lever for shaping AI metrics for success, not the only application of the framework.
- Computational hermeneutics treats AI as technology that both shapes and is shaped by cultural meaning, requiring attention to environmental and societal consequences.The discussion argues that powerful technological systems should not be considered only in isolation.