Source-linked AI summary
Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment
Jiangang Hao
TL;DR
Conversation-based assessment needs responses that remain semantically comparable, yet changing LLMs and adaptive context may alter generated replies. This study compares four LLMs across real collaborative-conversation messages with and without chat history, measuring reply similarity and alignment with human responses. Model choice and conversational context both affect response similarity, while generated replies are generally not highly similar to human replies.
Problem
LLM-based assessment requires sufficiently consistent responses for comparable measurement, but model evolution and conversational context may change generated-reply semantics.
Method
The study generated 100 responses per focal message from four LLMs under No History and With History conditions, then compared embedding-based semantic similarities and fitted mixed-effects models.
Results
Within-LLM similarity exceeded between-LLM similarity in both conditions, while chat history moderately changed reply content and its similarity effects differed across models.
Takeaways & Limitations
Prompting and conversational context alone may not preserve response consistency across evolving LLMs, motivating infrastructure and design strategies for stable, comparable assessment.
Abstract
from arXiv · showhide
This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.
1 Introduction
LLM-based agents may enable richer assessment of communication and collaboration, but their adaptive replies create tension between conversational flexibility and standardized assessment conditions. This study examines whether reply semantics remain consistent across models and contexts as LLMs evolve.
- Motivation: LLM-based agents can elicit richer evidence of communication and collaboration through natural, adaptive interactions.These skills are difficult to measure at scale with isolated-response formats because they emerge through situative, dynamic interactions.
- The adaptivity–standardization tension: The same learner input may receive semantically different replies because of the LLM, prompt, deployment setting, or prior chat history.Frequent model updates, releases, and retirements further complicate preservation of stable assessment conditions over time.
- Measurement implications: Uncontrolled reply variability can introduce construct-irrelevant variance that threatens validity, reliability, fairness, and comparability.Responses need not use identical wording, but should preserve the intended assessment function and comparable opportunities to demonstrate the target construct.
- Prior assessment practice: Traditional assessments likewise manage variability in task contexts, raters, and interaction processes while seeking comparable score interpretations.Examples include differing interviewer styles and follow-up prompts, or different writing prompts across forms and administrations.
- Design context: LLM-based systems replace much of traditional dialogue management with prompts and conversation history, enabling natural interactions but increasing response variability.Neuro-symbolic designs aim to combine LLM flexibility with greater control over relevant response consistency.
- Study focus: This study analyzes within-model similarity, human-reply alignment, and cross-model similarity across LLMs and conversational-context conditions.The results motivate infrastructure and design strategies for consistent and comparable conversation-based assessment as underlying models evolve.
2 Methods
The study uses real collaborative problem-solving conversations to compare LLM-generated replies across models and chat-history conditions. It measures semantic similarity among generated replies, their alignment with human replies, and the effects of model and context using exploratory comparisons and mixed-effects models.
- Data: The dataset contains dyadic online collaborative problem-solving science conversations collected through Amazon Mechanical Turk in 2018.Teams produced approximately 50–100 chat turns, and messages were segmented into early, middle, and late interaction stages.
- Data: The analysis selected 61 late-stage focal messages whose human replies were coded as highly relevant.The selection targeted messages with richer prior context and clearly interpretable human responses.
- Experimental conditions: Four LLMs generated 100 independent responses per focal message in each of two conditions: No History and With History.The models were GPT-4o mini, GPT-5.4, GPT-5.4 mini, and GPT-5.4 nano; the latter condition added all preceding conversational turns.
- Similarity measure: Cosine similarity between text-embedding-3-large vectors quantified semantic alignment among focal messages, human replies, and LLM-generated replies.Higher cosine-similarity values indicate greater semantic alignment.
- Exploratory analyses: The analyses compared within- and cross-LLM reply similarity, focal-message/reply similarity, and generated replies with versus without chat history.These comparisons tested whether model changes, semantically similar inputs, and prior context altered response comparability.
- Statistical analysis: Linear mixed-effects models estimated associations of LLM type, chat history, their interaction, and LLM–human similarity with mean similarity and similarity variability.A random intercept for focal message accounted for repeated observations across models and history conditions.
3 Results and Findings
Results show that semantic similarity varies with both the generating LLM and conversational context. Replies are more similar within the same model than across models, while chat history can weaken focal-message alignment and substantially change generated content.
- Within- and between-LLM similarity: 0.751–0.780 within-LLM versus 0.443–0.595 between-LLM similarity in the no-history condition, with within-model similarity consistently higher.With history, the corresponding ranges were 0.715–0.795 and 0.475–0.604.
- Within- and between-LLM similarity: GPT-5.4 had the highest within-model similarity under both prompting conditions, especially when conversational history was provided.Models in the GPT-5.4 family were also more similar to one another than to GPT-4o mini.
- Focal-message and reply similarity: Across all four LLMs, more similar focal messages produced more similar replies without history, whereas this relationship was generally weaker with history.For GPT-5.4 mini and GPT-5.4 nano, the with-history relationship remained below the no-history relationship at higher focal-message similarity.
- Cross-history similarity: Replies generated with history had median cross-history similarity of about 0.40–0.45, indicating that chat history meaningfully changed their semantic content.Figure 4 compared 100 no-history replies with 100 with-history replies for each focal message and model.
- Statistical analysis: Mixed-effects models found that history effects on mean pairwise similarity differed by model, while greater LLM-human similarity was associated with greater reply similarity.GPT-5.4 and GPT-5.4 mini were less negatively affected by history than GPT-4o mini, and GPT-5.4 showed a stronger tendency toward higher similarity with history.
- Statistical analysis: Variation in pairwise similarity also differed across models and conditions: GPT-5.4 showed higher variation than GPT-4o mini, whereas GPT-5.4 nano showed significantly lower variation.History increased variation for GPT-4o mini; higher LLM-human similarity was associated with lower variation.
4 Implications
The study finds that LLM-generated replies vary across models and conversational-context conditions, even for a single conversational turn. This variability makes stable, comparable conversation-based assessment an infrastructure and design challenge.
- The analysis compared within- and across-model reply similarity, alignment with human replies, and responses to semantically similar focal messages.These comparisons used embedding-based semantic similarity measures.
- LLM-generated reply semantics varied across models and between conditions with and without conversational history.Model-related differences remained evident under both prompting conditions.
- Prompting and conversational context alone were not sufficient to ensure highly similar response behavior when the underlying LLM changed.
- Maintaining consistent assessment interactions as LLMs evolve requires mechanisms to monitor, benchmark, and control response behavior across model transitions.The paper frames this as an infrastructure challenge rather than simply a prompt-engineering challenge.
- Response variability may affect the comparability of assessment conditions and the validity, reliability, fairness, and interpretability of score-based decisions.
Appendix A: System Instruction for LLM Reply Generation
The system instruction directs the LLM to generate concise, natural, accurate, helpful, safe, clear, and efficient replies to a teammate’s message. It also requires conversational continuity, calibrated uncertainty, and avoidance of fabricated or unsafe content.
- The LLM is instructed to generate a natural chat message based on the teammate’s message for an online science task.
- The instruction prioritizes accuracy, helpfulness, safety, clarity, and efficiency while favoring concise, typical human-human collaboration.
- The instruction requires explicit uncertainty and distinctions among facts, assumptions, and opinions when information is incomplete.
- The LLM must not hallucinate facts, citations, or capabilities, invent sources or events, pretend to perform unavailable actions, or produce unsafe, illegal, or privacy-violating content.
- The LLM should preserve conversational continuity and context, adapt explanations to the user’s expertise, and ask clarifying questions only when necessary.
- The output should avoid excessive verbosity, repetitive disclaimers, and unnecessary filler.