Source-linked AI summary

Reasoning Models Generate Societies of Thought

Junsol Kim, Shiyang Lai, Nino Scherrer, Blaise Agüera y Arcas, James Evans

arXiv:2601.10825v1cs.CLcs.CYcs.LG

TL;DR

The mechanisms behind sophisticated reasoning remain unclear beyond longer chains of thought. This paper analyzes reasoning traces and finds that reasoning models organize thought as internal societies whose conversational structure supports stronger problem solving.

  • Problem

    Whether reasoning models derive their advantage from perspective diversity and conversational organization, rather than computation alone, remains insufficiently understood.

  • Method

    The authors analyze reasoning traces across reasoning and instruction-tuned models using quantitative measures and mechanistic interpretability of implicit perspectives.

  • Results

    Accuracy doubles when a conversational-surprise feature signaling perspective shifts and contrast is amplified, while dialogue-based fine-tuning improves reasoning more than monologue-based fine-tuning.

  • Takeaways & Limitations

    Reasoning models appear to structure thought as societies of differentiated internal perspectives, with attention, role-taking, and conflict resolution coordinated during problem solving.

  • Takeaways & Limitations

    The comparison assumes that identical problems and correct answers isolate the effect of reasoning format from differences in task knowledge or solution exposure.

Abstract

from arXiv · show

Large language models have achieved remarkable capabilities across domains, yet mechanisms underlying sophisticated reasoning remain elusive. Recent reasoning models outperform comparable instruction-tuned models on complex cognitive tasks, attributed to extended computation through longer chains of thought. Here we show that enhanced reasoning emerges not from extended computation alone, but from simulating multi-agent-like interactions -- a society of thought -- which enables diversification and debate among internal cognitive perspectives characterized by distinct personality traits and domain expertise. Through quantitative analysis and mechanistic interpretability methods applied to reasoning traces, we find that reasoning models like DeepSeek-R1 and QwQ-32B exhibit much greater perspective diversity than instruction-tuned models, activating broader conflict between heterogeneous personality- and expertise-related features during reasoning. This multi-agent structure manifests in conversational behaviors, including question-answering, perspective shifts, and the reconciliation of conflicting views, and in socio-emotional roles that characterize sharp back-and-forth conversations, together accounting for the accuracy advantage in reasoning tasks. Controlled reinforcement learning experiments reveal that base models increase conversational behaviors when rewarded solely for reasoning accuracy, and fine-tuning models with conversational scaffolding accelerates reasoning improvement over base models. These findings indicate that the social organization of thought enables effective exploration of solution spaces. We suggest that reasoning models establish a computational parallel to collective intelligence in human groups, where diversity enables superior problem-solving when systematically structured, which suggests new opportunities for agent organization to harness the wisdom of crowds.

Results

Reasoning models exhibit more conversational, reciprocal, and perspective-diverse reasoning than instruction-tuned models, especially on difficult problems. These behaviors are linked to improved accuracy through cognitive strategies such as verification, backtracking, subgoal setting, and backward tracking.

  • Conversational behaviours: DeepSeek-R1 and QwQ-32B exhibit conversational behaviours far more frequently than instruction-tuned models.DeepSeek-R1 shows significantly more question–answering, perspective shifts, and reconciliation than comparison models.
  • Socio-emotional roles: DeepSeek-R1 and QwQ-32B display more reciprocal socio-emotional roles, combining asking and giving orientations, opinions, and suggestions with negative and positive roles.DeepSeek-R1 also shows higher Jaccard indices for co-occurring ask & give and positive & negative roles, indicating reciprocal coordination.
  • Task complexity: Conversational behaviours appear more frequently as DeepSeek-R1 tackles more complex problems, except for giving orientations and opinions.Problem complexity was assessed using Gemini-2.5-Pro judgments and error rates across conventional instruction-tuned models.
  • Mechanisms: Conversational behaviours and socio-emotional roles mediate reasoning models’ accuracy advantage by facilitating verification, backtracking, subgoal setting, and backward tracking.Conversational features also enhance reasoning directly by enabling more effective exploration of the solution space and indirectly by scaffolding systematic problem-solving strategies.
  • Causal steering: Positive steering (+10) doubles Countdown accuracy from 27.1% to 54.8%, whereas negative steering (−10) reduces accuracy to 23.8%.Positive steering simultaneously increases question-answering, perspective shifts, conflict of perspectives, and reconciliation; negative steering suppresses these behaviours.

Discussion

The discussion argues that reasoning models derive their advantage from emergent “societies of thought”: structured conversational interactions among diverse internal perspectives, rather than simply longer chains of thought. Reinforcement-learning and interpretability findings support conversational organization, perspective diversity, and coordinated conflict resolution as functional components of effective reasoning.

  • Core interpretation: Reasoning models exhibit “societies of thought” involving questions, alternative perspectives, conflict resolution, and coordinated socio-emotional roles, unlike non-reasoning models.These patterns rarely occur in non-reasoning models across model sizes of 671B, 70B, 32B, and 8B.
  • Functional role of conversation: Conversational behaviours and socio-emotional roles increase on more difficult problems and explain a substantial portion of reasoning models’ accuracy advantage.Steering experiments further link conversational markers, including perspective-shift signals, to reasoning performance.
  • Perspective diversity: Reasoning traces contain multiple implicit voices that differ systematically in personality traits and domain expertise.Mechanistic analyses find more diverse personality- and expertise-related features when models are steered toward conversational markers.
  • Reinforcement learning: Models fine-tuned on multi-agent dialogues reason more effectively than models trained only on correct monologue-like traces, indicating that conversational scaffolding—not initial correctness—provides the benefit.The experiments used relatively small 3B-parameter Qwen-2.5-3B and Llama-3.2-3B models.
  • Implications: As test-time computation expands, reasoning traces evolve from isolated monologues into structured dialogues whose effectiveness depends on coordinating attention, role-taking, and conflict resolution.The discussion frames this as “social scaling” and motivates agent architectures populated by diverse perspectives, personalities, and specialized expertise.

Methods

The study evaluates reasoning across 8,262 problems and six zero-shot models, measuring conversational, socio-emotional, and cognitive behaviors in generated reasoning traces. It also operationalizes role reciprocity, model effects, activation steering, and perspective attribution with automated annotation and statistical procedures.

  • Evaluation benchmark: The benchmark covers 8,262 reasoning problems spanning symbolic logic, mathematical and scientific reasoning, instruction following, and multi-agent inference.Tasks include BBH, GPQA, MATH (Hard), MMLU-Pro, and IFEval.
  • Model comparison: Responses are generated zero-shot from two reasoning models, DeepSeek-R1-0528 and QwQ-32B, and four instruction-tuned models.The instruction-tuned models are DeepSeek-V3-0324, Qwen-2.5-32B-Instruct, Llama-3.3-70B-Instruct, and Llama-3.1-8B-Instruct.
  • Conversational behaviours: An LLM-as-judge using Gemini-2.5-Pro counts question–answering, perspective shift, conflict of perspectives, and reconciliation in each reasoning trace.Each category is operationally defined with conversational examples, and counts are integers, including zero when absent.
  • Socio-emotional roles: Socio-emotional roles are annotated with Bales’ Interaction Process Analysis and aggregated into information-giving, information-asking, positive-emotional, and negative-emotional categories.The judge separately counts 12 operationally defined roles before aggregation.
  • Role reciprocity: Role reciprocity is measured with Jaccard indices for asking-versus-giving task roles and positive-versus-negative emotional roles within the same trace.The index is the proportion of traces containing both roles among traces containing either role.
  • Behavioral and mechanistic analyses: Model behavior is analyzed with categorical model indicators, cognitive-behavior annotations, activation steering across five retained strengths, and Hungarian-algorithm alignment for speaker attribution.The retained steering strengths are s∈{−10, −5, 0, 5, 10}; ±15 are excluded because excessive steering lowers accuracy.

Extended Data Figures

The extended data figures provide supporting analyses of reasoning-trace structure, simulated personas and socio-emotional roles, latent-voice identification, feature diversity, and reinforcement-learning effects of conversational scaffolding.

  • Simulated social structure: Extended Data Fig. 2 presents a chemistry reasoning trace as multi-turn dialogue among distinct cognitive personas, annotated for conversational behaviors and socio-emotional roles.It also reports Big Five personality profiles for five personas identified by LLM-as-judge on normalized 1–5 trait scales.
  • Simulated social structure: Extended Data Fig. 3 quantifies Bales’ 12 detailed socio-emotional roles in chain-of-thought reasoning and relates their presence to problem complexity.Complexity is measured using a 7-point Likert scale or error rates in non-reasoning models.
  • Measurement and mechanisms: Extended Data Fig. 5 validates LLM-as-judge identification of latent voices using 1,196 conversations, achieving Spearman’s ρ = 0.86 with p < 1 × 10−323.The benchmark hides speaker labels and concatenates each dialogue into one text block.
  • Measurement and mechanisms: Extended Data Fig. 6 estimates personality- and expertise-related diversity from SAE activations using coverage and entropy distributions with 95% confidence intervals.Median and interquartile-range markers summarize the distributions.
  • Reinforcement learning: Extended Data Fig. 7 shows that question-and-answering emerges first during reinforcement learning, while conflict and perspective shifts rise in parallel and reconciliation remains low.Extended Data Fig. 8 reports faster accuracy gains after multi-agent dialogue scaffolding than monologue-style scaffolding on Countdown for Qwen-2.5-3B and Llama-3.2-3B, although both eventually converge.

Extended Data Tables

The extended data include a reasoning trace before and after steering the conversational surprise feature at Layer 15, Feature 30939. The displayed +10 trace begins by framing an arithmetic puzzle and planning a step-by-step solution.

  • Steering reasoning traces: Extended Data Table 1 examines reasoning traces before and after steering the conversational surprise feature.The feature is identified as Layer 15, Feature 30939.
  • Steering reasoning traces: +10, the displayed trace frames a puzzle using 46, 54, 52, and 77 to create an equation equaling 75.The trace states that basic arithmetic operations may be used and each number only once.
  • Steering reasoning traces: The trace begins with explicit step-by-step problem solving, first listing the four numbers before considering arithmetic combinations.This excerpt is labeled “Steering Reasoning Trace Result.”

Supplementary Information · Supplementary Methods: Annotations Examples · DeepSeek-R1: Chemistry

The chemistry example illustrates DeepSeek-R1’s reasoning as a multi-perspective process, with specialized internal roles debating the reaction sequence and converging on a fused norbornane–anthracene-like product. The trace includes perspective shifts, conflicts, backtracking, and reconciliation while analyzing sequential Diels–Alder additions.

  • DeepSeek-R1: Chemistry: DeepSeek-R1 assigns the chemistry problem to distinct perspectives spanning planning, associative expertise, visualization, verification, and pragmatic decision-making.The listed roles differ in problem-solving approach and personality/expertise profiles, including Planner/Executor, Associative Expert, Visualizer/Builder, Critical Verifier, and Pragmatist/Strategist.
  • DeepSeek-R1: Chemistry: The initial annotated trace records question-answering, perspective shifting, conflict of perspectives, and subgoal setting during the synthesis analysis.Its conversational annotation is question_and_answering: 1, perspective_shift: 1, conflict_of_perspectives: 1, reconciliation: 0, while cognitive annotation gives subgoal_setting: 1.
  • DeepSeek-R1: Chemistry: The reasoning repeatedly revises uncertain interpretations of the dibromomethyl diene and sodium-iodide reaction, considering dehalogenation and o-quinodimethane formation.Several trace segments explicitly mark backtracking, conflict, or alternative interpretations before settling on an in-situ reactive diene.
  • DeepSeek-R1: Chemistry: The trace converges on o-xylylene as the diene and 7-(tert-butoxy)norbornadiene as the dienophile in the proposed Diels–Alder sequence.The analysis states that sodium iodide generates o-xylylene in situ and identifies the reaction pairing accordingly.
  • DeepSeek-R1: Chemistry: Two equivalents of the o-xylylene precursor are interpreted as enabling two additions, one across each norbornadiene double bond.The trace first places one addition at positions 2–3, then proposes a second at positions 5–6.
  • DeepSeek-R1: Chemistry: Later annotations show increased reconciliation and perspective shifting as the trace stabilizes its structural interpretation.One annotated segment records question_and_answering: 1, perspective_shift: 2, conflict_of_perspectives: 1, reconciliation: 1; another records perspective_shift: 4, conflict_of_perspectives: 3, reconciliation: 1.
  • DeepSeek-R1: Chemistry: The proposed product contains two benzene rings fused to the norbornane system, specifically at positions 2–3 and 5–6.The resulting structure is described as a symmetric, anthracene-like system fused to the central norbornane framework.

DeepSeek-R1: Creative Sentence Rewriting

DeepSeek-R1’s sentence-rewriting trace shows iterative synonym selection, structural alternatives, verification, and reconciliation while preserving the original meaning. Its reasoning alternates among conflicting perspectives before settling on a concise rewrite.

  • Conversational Behaviour: Conversational behavior rises from zero initially to question_and_answering: 1, perspective_shift: 1, conflict_of_perspectives: 1, and reconciliation: 1 in the final trace.Intermediate steps include perspective shifts, conflicts, question-answering, and reconciliation as candidate rewrites are compared.
  • Meaning Preservation: The trace evaluates synonyms and sentence structures while retaining the original meaning and avoiding redundant or newly introduced ideas.It considers replacements for “flung,” “hatred,” and “burning fire,” then rejects “deep-seated” because it adds an unsupported idea.
  • Cognitive Behaviour: Cognitive behavior includes subgoal_setting: 4 during lexical analysis, followed later by verification: 1 and backtracking: 1 when the candidate rewrite is checked.The trace verifies that “flames” appropriately denotes the visible part of a fire and backtracks from alternatives that alter meaning.
  • Final Rewrite: The selected rewrite is “I hurled my hatred into the flames,” replacing “flung” with a strong synonym and “burning fire” with “flames.”The trace explicitly presents this as the preferred rewrite after considering alternatives and checking semantic appropriateness.

DeepSeek-V3: Chemistry

The section presents a multistep synthesis and estimates that final product 4 has 8 chemically distinct hydrogen atoms. The reasoning attributes this estimate to symmetry in a rigid polycyclic structure and symmetric substitution patterns.

  • Reaction sequence: The problem asks for the number of chemically distinct hydrogen atoms in product 4 after a four-step reaction sequence.The sequence uses cycloheptadiene and dibromomethylcyclohexadiene with sodium iodide, followed by aqueous sulfuric acid, SO3/pyridine in DMSO, and heating at 150°C.
  • Product 4: The response characterizes product 4 as likely a symmetric, rigid polycyclic aromatic or conjugated system.It suggests the final heating step could involve elimination, rearrangement, or further cyclization.
  • Distinct hydrogens: The response estimates 8 chemically distinct hydrogens, selecting choice (B), because symmetric substitution patterns reduce the number of distinct hydrogen environments.It also notes that substituents may further reduce symmetry, while presenting 8 as a reasonable estimate.

DeepSeek-V3: Creative Sentence Rewriting

DeepSeek-V3 rewrites a sentence by exploring synonyms and alternative phrasings while preserving its meaning and emotional intensity. It selects stronger action and imagery after evaluating redundancy.

  • Method: The agent uses creative-writing and linguistic-analysis expertise to deconstruct the sentence, generate alternatives, evaluate redundancy, and synthesize a revision.Its criteria include emotional intensity and vivid imagery.
  • Method: It considers replacing “flung” with “hurled” or “tossed” and “hatred” with “anger” or “rage” while maintaining the original meaning.The alternatives are evaluated for phrasing and emotional intensity.
  • Result: After identifying “burning fire” as redundant, it chooses “hurled” and “flames” for stronger action and more vivid imagery.The reasoning simplifies the phrase because fire is inherently burning.

Reinforcement Learning Step 40 (Countdown)

At reinforcement learning step 40, the methodical perspective prioritizes simple, step-by-step arithmetic deduction. Its recorded conversational behavior is limited to question-and-answering, with no perspective shift, conflict, or reconciliation.

  • Perspective and expertise: The perspective is a methodical problem-solver that enumerates and evaluates arithmetic expressions for the numerical puzzle.It prioritizes simplicity and step-by-step logical deduction.
  • Conversational behavior: Conversational behavior consists solely of question-and-answering, with perspective shift, conflict of perspectives, and reconciliation all set to 0.The recorded values are question_and_answering: 1; perspective_shift: 0; conflict_of_perspectives: 0; reconciliation: 0.
  • Cognitive behavior: Cognitive behavior records verification, backtracking, subgoal setting, and backward chaining all at 0.No positive value is recorded for any of the four listed cognitive behaviors.

Reinforcement Learning Step 120 (Countdown)

The countdown trace uses arithmetic trial-and-error, repeatedly checking candidate combinations and revising its approach. After eight unsuccessful attempts, it concludes that reaching 14 may be impossible with the given numbers and operations.

  • Agent roles: The calculator-like agent systematically tests combinations of 3, 56, 66, and 44, using each number once while seeking 14.Its stated expertise is basic arithmetic and methodical problem-solving, with trial-and-error exploration of the solution space.
  • Conversational behavior: Conversational behavior accompanies the search through question-answering, perspective shifts, and conflicts of perspectives, while reconciliation remains absent.The recorded behavior repeatedly assigns nonzero values to these exploratory categories but gives reconciliation a value of 0.

Supplementary Methods: Behavioural Pathways Linking Reasoning Models to Accuracy Advantages

A structural equation model decomposes DeepSeek-R1 and QwQ-32B’s accuracy advantage over instruction-tuned models into social, cognitive, social–cognitive, and direct pathways. For DeepSeek-R1, the total accuracy effect is 0.26, with significant direct and social indirect effects but no statistically distinguishable cognitive indirect effect.

  • Structural equation model: The SEM treats model type as the treatment, behavioral mediators as pathways, and task accuracy as the outcome.It includes eight social and four cognitive behavior mediators while controlling for log-transformed reasoning time.
  • Mediation pathways: The model estimates social, cognitive, and social–cognitive pathways linking reasoning-model behaviors to accuracy.The social–cognitive pathway tests whether conversational behaviors facilitate cognitive strategies that then improve accuracy.
  • DeepSeek-R1 effects: 0.06 (p < 0.001) is DeepSeek-R1’s direct effect, representing variance unexplained by the measured mediators.The direct pathway is defined as β_D.
  • DeepSeek-R1 effects: 0.26 (p < 0.001) is the total effect of DeepSeek-R1 on accuracy.This estimate is reported in Extended Data Fig. 4a.
  • DeepSeek-R1 effects: 0.07 (p < 0.001) is DeepSeek-R1’s indirect effect through social behaviors, whereas the cognitive indirect effect is β = −0.00 (p > 0.05).The cognitive indirect effect is not statistically distinguishable from zero.

Supplementary Methods: Cross-domain Reasoning Transfer

Conversational scaffolding transferred beyond arithmetic: models fine-tuned on multi-agent Countdown dialogues developed factual reasoning faster during reinforcement learning on PolitiFact misinformation detection. This supports socially organized reasoning as a domain-general aid.

  • Cross-domain transfer: The study pretrained models with conversational scaffolding on Countdown before reinforcement learning on the distinct PolitiFact misinformation detection task.The conversational scaffold was generated through supervised fine-tuning data produced by GPT-4.1.
  • Cross-domain transfer: The PolitiFact corpus contains 23,299 fact-checked claims labeled True, Mostly True, Half True, Mostly False, False, or Pants on Fire.For evaluation, the labels were grouped into True, Half True, and False categories.
  • Cross-domain transfer: The comparison contrasted baseline models receiving reinforcement learning alone with models supervised-finetuned on correct multi-agent Countdown dialogues before reinforcement learning.The baselines were Llama-3.2-3B and Qwen-2.5-3B, and the conversational fine-tuning used Countdown rather than misinformation detection data.
  • Cross-domain transfer: Models fine-tuned with conversational scaffolding on Countdown achieved faster early-stage gains in factual reasoning accuracy on misinformation detection.The learning trajectories are illustrated in Supplementary Fig. 3.
  • Cross-domain transfer: Social interaction fine-tuning improved in-domain arithmetic reasoning and accelerated reasoning development in the different misinformation-detection domain.The authors interpret this cross-domain improvement as evidence for the generality of socially organized reasoning.

Supplementary Methods: Performance Comparison Between PPO and GRPO

In smaller-scale reasoning experiments, standard PPO achieved performance comparable to GRPO despite GRPO’s group-relative advantage normalization and potential variance reduction.

  • Algorithm comparison: PPO and GRPO achieved comparable performance after 250 training steps on Qwen-2.5-3B in the Countdown task.GRPO normalizes policy advantages within each mini-batch and uses a group-relative objective, whereas PPO does not.

Supplementary Methods: Replications on Llama-3.2-3B

The study tests generalizability by replicating its training pipeline with Llama-3.2-3B, comparing conversational and monologue-like reasoning supervision. It also applies PPO on Countdown arithmetic with a reward combining accuracy and format correctness.

  • Supervised fine-tuning: Llama-3.2-3B is supervised-fine-tuned on either conversational or monologue-like reasoning using standard next-token prediction loss.In the conversational condition, reasoning from multiple personas is concatenated into a single <think> </think> block.
  • Reinforcement learning: PPO training on Countdown arithmetic assigns reward R = 0.9 × Accuracy + 0.1 × Correct Format and proceeds for 250 steps.Accuracy and format are binary: correctness requires the reasoning trace to yield the right answer, while format requires reasoning and final-answer blocks with one equation-form answer.

Supplementary Methods: LLM-as-Judge prompts … Cognitive Behaviors

The supplementary methods use constrained LLM-as-Judge prompts to count conversational behaviors, socio-emotional interaction roles, problem difficulty, and cognitive reasoning behaviors from model-generated text. Each evaluator requires integer counts in a strict JSON schema, with zero used when a behavior is absent.

  • Conversational Behaviors: Conversational-behavior prompts count Question_and_Answering, Perspective_Shift, Conflict_of_Perspectives, and Reconciliation instances in chain-of-thought text.The categories capture posed-and-answered questions, transitions between viewpoints, disagreement or tension, and integration of conflicting views.
  • Conversational Behaviors: The conversational evaluator counts distinct occurrences per category and returns 0 when none are present.Its output must be a single valid JSON object with the four prescribed keys and integer values.
  • Socio-Emotional Roles: The socio-emotional evaluator applies Bales’ Interaction Process Analysis to count 12 interaction categories in chain-of-thought or group-interaction transcripts.The schema includes solidarity, tension release, agreement, suggestions, opinions, orientation, requests, disagreement, tension, and antagonism.
  • Problem Complexity: The problem-complexity prompt rates intrinsic zero-shot difficulty for a capable language model on a 1–7 scale.The evaluator maps 1 to very easy and 7 to very difficult, returning only {"difficulty": <integer from 1 to 7>}.
  • Cognitive Behaviors: The cognitive-behavior evaluator inspects model reasoning for Answer Verification, Backtracking, Subgoal Setting, and Backward-Chaining.Verification checks results against targets; backtracking abandons failed paths; subgoals decompose problems; backward-chaining works from targets toward initial problems.
  • Cognitive Behaviors: Cognitive-behavior outputs use verification_count, backtracking_count, subgoal_count, and backward_count, assigning 0 when the corresponding behavior is absent.The prompt requires one valid JSON object containing exactly these four integer keys.

Persona Identification Prompt … Generating Conversation-Like Reasoning Traces for Supervised Fine-Tuning

The section operationalizes reasoning traces as interactions among identifiable perspectives, whose personalities and expertise are inferred, segmented, and mapped to conversational or feature-level labels. It then contrasts monologue-like traces with supervised fine-tuning prompts that explicitly simulate a flexible collaborative group of thinkers.

  • Persona Identification Prompt: Perspectives are identified as distinct cognitive roles using transitions, role shifts, rhetorical changes, corrections, subproblem movement, domain knowledge, and personality traits.Each identifiable shift is treated as a boundary between perspectives.
  • Persona Identification Prompt: Each perspective receives BFI-10 personality answers and a concise profile of domain expertise in a single valid JSON object.The questionnaire requires one of five exact agreement strings for each item.
  • Persona Segmentation: The transcript is segmented sequentially into turns, preserving exact starting text, perspective indices, repeated non-consecutive appearances, and moderator or transition interjections.The required start_text is the exact first 10 words copied verbatim, including punctuation and capitalization.
  • Identifying Conversational Contexts: Conversational context is scored from 0 to 100, ranging from clearly a single-person thought to clearly a conversation or response to someone.The output must contain exactly one JSON object with an answer field.
  • Classifying Sparse Autoencoder (SAE) Personality Features: SAE features are scored for personality-trait relevance across behavioral and psychological patterns including sociability, agreeableness, conscientiousness, openness, emotional stability, confidence, empathy, risk, competitiveness, and curiosity.The scale runs from 0 for no personality relevance to 100 for complete personality relevance.
  • Classifying Sparse Autoencoder (SAE) Expertise Features: SAE features are separately scored for domain-expertise relevance, covering specialized professional or academic knowledge and terminology.The scale runs from 0 for no domain-expertise relevance to 100 for complete domain-expertise relevance.
  • Generating Monologue-Like Reasoning Traces for Supervised Fine-Tuning: Monologue-like supervised fine-tuning traces require step-by-step reasoning enclosed within <think> and </think> before the answer.The prompt explicitly prohibits answering directly without reasoning.
  • Generating Conversation-Like Reasoning Traces for Supervised Fine-Tuning: Conversation-like traces simulate a collaborative group of distinct personas in realistic back-and-forth problem solving, with flexible speaker order and repeated turns inside structured tags.The format defines persona blocks, thinker-specific utterance tags, a conversation container, and a group_solution container.

Supplementary Tables

The supplementary tables document comparative analyses of conversational and socio-emotional behaviors, task difficulty, SAE features, training settings, and generated fine-tuning data. Additional examples illustrate monologue- and conversation-style reasoning on mathematical and logical problems.

  • Comparative analyses: Supplementary Table 1 compares reasoning models with instruction-tuned counterparts on conversational behaviors, socio-emotional roles, and Jaccard indices.The reported comparisons use regression coefficients with 95% confidence intervals, t-statistics, degrees of freedom, and exact p-values.
  • Comparative analyses: The regression analyses include task fixed effects, control for log-transformed reasoning-trace length, and two-sided t-tests with task-level clustered standard errors.The comparisons are between DeepSeek-R1 and QWQ-32B and their respective instruction-tuned models, DeepSeek-V3 and Qwen-2.5-32B-IT.
  • Task difficulty: Supplementary Table 2 sorts the most and least challenging tasks by problem complexity measured by an LLM-as-judge.Cells report average problem complexity or average behavior frequency, with standard deviations in parentheses.
  • Training and examples: Supplementary Tables 7–8 report PPO training hyperparameters and monologue-style versus conversation-style fine-tuning data generated by Qwen-2.5-32B-IT.The examples include a GCD/LCM problem solved through conversational correction and an outfit-assignment problem answered as both monologue and conversation.
Loading 2601.10825v1…