Source-linked AI summary
How You Ask Shapes What You Get: A Theory-Seeded Measurement of Articulation in Advice-Seeking LLM Conversations
Juneha Baek, Suhyeon Lee, Donghyuk Shin
TL;DR
Advice-seeking prompts can express similar needs with different levels of specificity, constraint, and context, but prior evaluation often treats that variation as noise. Using 16,447 prompts from three public chat corpora, the paper extracts articulation features, identifies recurring latent factors, and tests their relation to topics and model replies. Articulation is largely separable from topic, and roughly one in six WildChat prompts belongs to a long-form, information-poor style associated with shorter, vaguer answers and no clarifying question. The authors therefore propose articulation-stratified evaluation, while emphasizing that the observational data support association rather than causal claims.
Problem
The paper asks whether articulation variation in naturalistic advice-seeking chat forms a stable structure separable from topic and associated with model responses.
Method
The study extracts interpretable articulation features from 16,447 prompts across WildChat, LMSYS, and ShareChat, then learns factors and segments on WildChat for held-out replication.
Results
Roughly one in six WildChat prompts is long but information-poor; models give shorter, vaguer answers and no clarifying question, while articulation remains largely separable from topic.
Takeaways & Limitations
Benchmark evaluation should stratify on articulation rather than topic alone, using the extracted structure as a measurement instrument.
Takeaways & Limitations
Because chat logs do not reveal latent user intent or support controlled paraphrase comparisons, the response relationship is descriptive and associative rather than causal.
Abstract
from arXiv · showhide
Users articulate the same advice-seeking request in different ways: some specify detailed constraints, others gesture at a vague need. Prior work treats this variation as noise to be averaged away; we instead treat it as a stable, measurable structure in the input distribution. We ask whether articulation (how people ask) forms latent dimensions separable from topic (what they ask about), and whether it is associated with how language models respond. We extract interpretable features from 16,447 advice-seeking prompts pooled from public chat corpora (WildChat, LMSYS, and ShareChat) and recover a small set of latent articulation factors that replicate across train/test splits and across corpora. Because this structure is largely separable from topic, the populations it defines cut across topics and stay invisible to topic- or task-based evaluation. The factors define a handful of recurring articulation styles, one of which stands out: a long-form but information-poor style, roughly one in six prompts in the largest corpus, where models return shorter, vaguer answers and do not ask for clarification even though under-specification is exactly the condition that warrants it. The contrast holds within every topic group and length quintile, and is not under-specification alone -- a second, equally under-specified style does draw clarifying questions. Two independent human annotators reproduce this contrast. We argue that benchmarks should stratify on articulation, and we offer the extracted structure as a measurement instrument for doing so.
1 Introduction
The paper treats articulation variation in advice-seeking prompts as a stable structure to measure rather than noise, asking whether it is separable from topic and associated with model responses. It identifies a long-form, information-poor population that topic- and task-based evaluation can miss.
- The study asks whether naturalistic articulation variation forms a stable structure separable from topic.This question determines whether topic-based evaluation captures or overlooks articulation variation.
- 16,447 advice-seeking prompts from WildChat, LMSYS, and ShareChat reveal a stable, low-dimensional articulation structure that replicates across splits and corpora.
- Roughly one in six prompts form a long-form, information-poor population whose models produce shorter, vaguer answers and no clarifying question.The population cuts across topics, tasks, and length bands.
- A per-prompt under-specification score detects at-risk prompts, but population analysis is needed to estimate axes, replication, and real-traffic prevalence.One feature detects the prompts at AUC=0.91.
- The response link is treated as an association because chat logs do not reveal users’ unobserved preferences or support counterfactual comparisons.
2 Related Work
Prior research frames prompt variation as paraphrase noise, model behavior, ambiguity, or topic/task segmentation. This paper instead studies articulation as a measurable, potentially separable structure in naturalistic chat.
- Prior work often treats articulation variation as nuisance variation for paraphrase robustness or as an LLM grounding deficit.
- Research on ambiguity and underspecification largely evaluates whether individual prompts contain enough information for the model to answer appropriately.
- Most LLM-use segmentation organizes users by topic or task, whereas this study segments prompts by how users phrase their requests.
- The paper reports that articulation-based and topic-based partitions are largely separable.
- Consumer-preference theories inspire feature sources, but the paper does not treat them as constructs requiring validation.
3 Method
The method filters public chat logs for advice-seeking prompts, extracts interpretable articulation and response features, learns six factors and six segments on WildChat, and applies them transform-only to held-out corpora. It then compares articulation segments with topics and response outcomes.
- Data and filtering: The study defines advice-seeking as first-turn requests for help deciding about a user’s situation, excluding generic task, factual, and how-to prompts.
- Corpora: WildChat is the sole training corpus, while LMSYS and ShareChat receive the fitted model transform-only for replication.
- Data and filtering: 1,768,604 loaded English first-turn pairs yield 16,447 advice-seeking positives after a two-pass cascade and LLM intent classification.The cascade admits 535,612 candidates, or 30.3%, and its sampled false-negative rate is 0.4–1.6%.
- Feature extraction: LLM span tagging produces 30+ raw articulation features, including attribute type, hardness, specificity, crystallization, hedges, mood, and named alternatives.After preprocessing, 19 features enter factor analysis.
- Response analysis: Reply outcomes include recommendations, clarification questions, attribute coverage, specificity, hedging, direct answers, and diversity.
- Latent structure: Parallel analysis selects k=6 factors, and a Promax-rotated principal-axis model allows correlated articulation dimensions.
- Segmentation and tests: A six-component GMM defines articulation modes, whose separability from seven high-level topics is tested using dependence and within-topic reclustering.
4 Results
The analysis identifies six reproducible articulation factors and six soft-edged segments that are largely separable from topic. Segment membership covaries with response characteristics, especially for a long, information-poor style that receives neither a substantive answer nor clarification.
- 4.1 Factor structure: Six articulation factors capture specificity, density, named alternatives, search attributes, self-situation disclosure, and length versus imperative language.Top feature loadings range from λ=+0.72 to +1.02.
- 4.1 Factor stability: Four core factors show high train/test stability (ϕ ≥0.98) and cross-corpus stability (ϕ ≥0.97), whereas self-situation and length-imperative factors are less stable.F5 and F6 reach ϕ ≈0.90 train/test and ≈0.82 cross-corpus, and are partly dataset-specific.
- 4.2 Articulation segments: BIC selects six articulation modes, while bootstrap ARI=0.549 (95% CI [0.377, 0.926]) indicates moderate reproducibility with soft cluster boundaries.A stricter extractor raises ARI to 0.785.
- 4.2 Articulation segments: The six segments include baseline, open-asker, specifier, named-comparer, and an articulation-poverty style; s2 is largest at 26%.WildChat has the highest s4 share at 16%, while curated ShareChat platforms underrepresent s4 at 2–7%.
- 4.3 Topic separability: Topic association is weak (pooled Cramér’s V =0.243), and six of seven topics recover essentially the pooled articulation types.The exception is Personal Decisions, which reaches ARI=0.31, although stricter re-extraction raises it to 0.710.
- 4.4 Association with LLM response characteristics: The long, information-poor s4 style receives fewer direct answers and no meaningful clarification, unlike the similarly under-specified s5 style.s4 has 19% direct-answer rate versus 67% for the comparison, while s5 draws clarification at 11.9% versus 6.7% for s4.
5 Discussion
The discussion frames articulation as a reproducible input-side structure with practical implications for detection and benchmark coverage. It also highlights an equity concern while limiting claims because the segments are not linked to speaker identities.
- Articulation as a measurable, structured axis: ϕ ≥0.97 cross-corpus agreement on four core factors supports articulation as a stable, measurable structure rather than paraphrase noise.The authors recommend coverage across articulation segments instead of stochastic paraphrases.
- Detection and benchmark coverage: AUC=0.91 for specificity_mean shows that a single feature can detect the at-risk population without the full factor-analysis pipeline.The pipeline remains useful for establishing reproducibility, prevalence, and the distinction from plain ambiguity.
- From phenomenon to input-side structure: 16% of WildChat prompts belong to a long-form, information-poor population associated with shorter, vaguer answers and no clarifying question.This population cuts across topic, task, and surface length, making it relevant to benchmark coverage.
- Equity implications: Because the corpora lack demographic labels, the study cannot attribute segment-level disparities to protected attributes, despite consistency with prior sociolinguistic evidence.The authors identify this as a target for follow-up rather than an established attribution.
6 Conclusion
The conclusion presents articulation as a stable, low-dimensional property of advice prompts that is largely separable from topic. It identifies a previously obscured long-form, information-poor population and argues that evaluation should stratify on articulation.
- Articulation replicates across corpora and remains largely separable from topic, establishing it as structured input variation rather than noise.The extracted structure is released as a measurement instrument.
- Long-form, information-poor prompts receive shorter, vaguer answers and no clarifying question, making this population visible to evaluation.The conclusion contrasts articulation-based stratification with evaluation organized by topic alone.
7 Limitations
The limitations constrain the study’s causal, individual-level, measurement, selection, and generalization claims. Its evidence concerns a bounded English-language, first-turn dataset and a population-level partition whose boundaries are soft.
- The study estimates associations between encoded prompts and replies, not counterfactual effects of changing articulation while holding user intent fixed.Controlled multi-prompt elicitation would be required for the stronger causal claim.
- Stability is established for partitions and factor structures, not individual users, because the corpora do not track users across conversations.The relevant measures are bootstrap ARI and Tucker congruence.
- ∼0.93% admission by the advice-seeking cascade may favor explicitly articulated requests, making 16% s4 prevalence in WildChat a conservative lower bound.Under-articulated prompts are described as especially likely to be excluded by the filter.
- A single LLM family extracts both articulation features and response outcomes, creating a shared-method-variance concern despite human audits and manipulation checks.The paper reports four mitigations, including reply-only outcome extraction and external annotation.
- Cross-corpus congruence rules out corpus-specific, but not model-specific, extractor artifacts; stricter extraction sharpens the structure without establishing human validity.This correction narrows what replication evidence can support.
- Silhouette=0.19 and bootstrap ARI=0.55 indicate soft density modes rather than six discrete types, although the s4 finding survives without clustering.The study therefore does not defend the six segments as a typology.
- The claims are limited to English-language, first-turn, advice-seeking prompts across seven domains and exclude healthcare, non-English articulation, multi-turn conversations, and task-oriented prompts.Cross-corpus replication is claimed only within this subset.
- The analysis uses publicly released, de-identified corpora and characterizes articulation at the population level without identifying or profiling individual users.Example prompts receive minimal additional redaction where content could identify someone.
Use of AI Assistants
The study uses structured prompt features and a seven-group topic taxonomy to analyze advice-seeking prompts, with human annotation rules validating topic labels. AI assistants supported writing and code development, while empirical pipeline decisions remained with the authors.
- Use of AI Assistants: Claude assisted with drafting, editing, and code development, but authors made all final content, analysis, and claim decisions.
- Feature Extraction: The Stage 1 LLM extracts 30+ raw prompt features spanning spans, compositional ratios, categorical variables, continuous measures, and person distributions.
- Preprocessing: Compositional ratios are transformed into five ilr coordinates, counts are length-normalized, near-zero-variance columns are dropped, residual NaNs are mean-imputed, and features are z-standardized.
- Factor Analysis: After preprocessing, 19 features enter factor analysis, while beta and sentence_mood remain descriptive variables for known-groups validation.
- Topic Annotation: The topic taxonomy rolls 25 fine-grained categories into seven high-level groups, with annotators labeling 100 prompts blindly using substance-over-phrasing and dominant-intent rules.
- Topic Distribution: H2 Personal Decisions & Recommendations is the largest advice-seeking group, containing 8,730 prompts, or 53% of pooled positives.
C Segment Interpretations
The analysis identifies six articulation modes from a 19-feature factor model, each defined by distinctive combinations of specificity, density, alternatives, search hardness, self-situation, and length. These modes describe how users formulate advice requests independently of the substantive topic.
- Segment Interpretations: The six segments are baseline, open-asker, specifier, named-comparer, long-form/low-density, and vague/situated.
- Segment Interpretations: s2 specifier is the largest segment at 26%, with high specificity and detailed constraint articulation.
- Segment Interpretations: s3 named-comparer is jointly high on density, named alternatives, and search hardness, producing dense, brand-aware, specification-driven prompts.
- Segment Interpretations: s4 long-form/low-density comprises 16% of prompts and combines sparse intent spans with long declarative monologues, making it verbose but information-poor.
- Segment Interpretations: s5 vague/situated is under-specified with some self-disclosure and has the highest clarification rate among the six segments.
- Factor Structure: The Promax factor matrix contains 19 features across six factors: specificity/structure, density, named-alternative, search-attribute hardness, self-situation, and length versus imperative.
- Topic Separation: s4 appears in all seven topic groups, including H4 Technical/STEM and H5 Creative & Media, rather than concentrating in one subject area.
I Segment Regression Coefficients
Regression analyses compare articulation segments on response outcomes while controlling for topic, assistant model, prompt length, and source corpus. The strongest contrast is between s4, which receives shorter and less actionable responses without more clarification, and s5, which does prompt clarification despite comparable under-specification.
- Model Specification: Regression models include controls for topic, assistant model, log prompt length, and source corpus, using segment s0 as the reference category.
- s4 Effects: s4 responses are approximately 50% shorter than baseline, with substantially worse search coverage and a direct-answer logit of −1.48.
- s4 Effects: s4 receives approximately 1.9 fewer named recommendations than baseline and shows no significant increase in clarification, with p=0.11.
- Overall Results: All twelve modeled outcomes show jointly significant segment effects, although the table notes that significance is weak evidence at n=16,447.
- s5 Effects: s5 is the only segment triggering significantly more clarification, with a logit increase of +0.77∗∗∗, but its recommendation specificity remains lower by −0.44.
- s3 Effects: s3 produces 0.70∗∗∗ fewer options and a lower direct-answer rate, consistent with comparison-oriented responses to brand-listing prompts.
L.5 Matched-length analysis: density is not a length proxy
Matched-length analyses show that the s4 response deficit persists when prompt length is held fixed, supporting articulation density rather than verbosity as the relevant contrast.
- Matched-length results: 47.4pp: In Q1 prompts of 3–14 tokens, s4 received direct answers 25.2% of the time versus 72.6% for s0.The gap is comparable to the pooled effect, despite the shortest prompts being unable to form long monologues.
- Detection: Raw word count reached AUC=0.87 for s4 detection, while specificity_mean reached AUC=0.91.These detection results show that length and density correlate, but do not establish length as the operative cause.
- Topic validation: The topic tagger remains within the inter-annotator envelope, with κ=0.61 versus a human ceiling of κ=0.72.Within-class performance is strongest for H4 and H6 and weakest for the residual H7 “Other” class.
- Matched-length results: 29–48 percentage points: The s4 direct-answer deficit persisted in every token-count quintile.Search-coverage and option-count gaps showed the same pattern.
- Topic robustness: The contrast survives topic adjustment and appears in all seven topics, with direct-answer gaps ranging from −19pp to −54pp.Each topic-specific contrast was significant at p<0.002.
M Model-Selection Curves
The paper selects six factors and six articulation segments using separate model-selection procedures, while evaluating reproducibility, topic separability, and response associations.
- Factor selection: k=6: Parallel analysis retains six factors because the observed eigenvalue 1.042 exceeds both random benchmarks at k=6, while k=7 falls below them.The factor count is unchanged across standard-normal and column-permutation benchmarks.
- Segment selection: k=6: GMM BIC decreases through the allowed cap, whereas silhouette peaks at k=4.When the cap is lifted, BIC continues decreasing through k=10, so post-6 gains are treated as refinement rather than a clear new optimum.
- Estimands: Bootstrap ARI measures partition reproducibility between the observed segmentation and resampled partitions.The analysis also reports Cramér’s V and mutual information to assess segment–topic separability.
- Response associations: Response outcomes are regressed on segment indicators alongside topic, assistant model, source corpus, and log prompt length.This specification tests whether articulation segments covary with responses beyond these controls.
O Stricter Stage-1 Re-extraction and Calibration of the Congruence Statistic
A stricter Stage-1 extractor leaves the six-factor count intact and improves structural statistics, while null calibration and unrotated comparisons qualify how congruence should be interpreted.
- Structural robustness: Six factors remained, and every structural statistic improved under the stricter extraction standard.In the rerun, all six factors cleared 0.97 cross-corpus congruence, compared with approximately 0.82 for main-text F5 and F6.
- Re-extraction: 7.5%: The stricter extractor declined 1,241 prompts admitted by the original extraction before the pipeline was rerun end to end.Stages 3–6 were refit, including the factor model, GMM, and regressions.
- Null calibration: ϕ=0.382: Independently permuted features produced mean congruence 0.382, a 95th percentile of 0.668, and a maximum of 0.974.Observed values therefore clear the calibrated null, but the null has a long right tail.
- Alignment check: +0.001: Orthogonal Procrustes alignment increased half-sample congruence by only 0.001.Without rotation, congruence was 0.9880–0.9979 versus 0.9866–0.9987 with alignment.
- Comparability caveat: The rerun’s factor axes are not label-identical to the main-text solution, so its statistics establish rerun stability rather than unchanged individual factors.The discrepancy indicates that marginal axis identities remain less settled than the factor count.
P Pre-registered Construct Validation of the Factor Labels
Pre-registered prompt manipulations validate most factor labels against a measured extraction noise floor, while one label fails and its replacement remains provisional.
- Noise floor: Mean |∆Z| ranged from 0.029 to 0.135 across reruns, and 2 of 180 status decisions flipped.Effects should be interpreted against this nonzero temperature-0 extraction noise floor.
- Failed manipulations: Densify was invalid because it increased both intent spans (+2.05) and tokens (+10.0) in a ratio-based density construct.The manipulation therefore licenses no conclusion about the instrument.
- Failed manipulations: Soften lowered F3 by −0.79 even though crystallization increased by +1.53 and hard-constraint rate fell by −0.38.The intended manipulation worked, but the predicted factor response was wrong.
- Validation design and results: 20 of 24: Frozen known-groups criteria hold, with six passing manipulations moving their named factors at d=1.0 to 4.6.These shifts exceed the extractor’s test–retest noise floor by 9× to 49×.
- Label revision: F3 correlates most strongly with within-prompt specificity variance at r=+0.489, motivating a post-hoc mixture-of-precise-and-vague-requirements interpretation.The convergent manipulation needed to validate this replacement label has not been run.
- Reproduction: The pipeline is reproducible from public corpora, with documented code, automatic data download, and a smoke-test configuration.No new dataset is released because the claim concerns naturalistic traffic from public sources.