Source-linked AI summary

ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning

Jiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li, Michael Qizhe Shieh, Yejin Choi, Ranjay Krishna, Yu Cheng

arXiv:2510.27492v3cs.CV

TL;DR

Multimodal reasoning lacks a clear recipe for meaningful text-image interleaving, especially when visual manipulation is required. ThinkMorph trains a unified model on approximately 24K complementary interleaved traces and achieves broad vision-centric gains while exhibiting emergent multimodal behaviors. These results support interleaved reasoning as a setting for studying generalizable multimodal problem solving and unified-model capabilities.

  • Problem

    Existing multimodal reasoning methods do not clearly establish how text and images can mutually advance reasoning, while text-only and prior interleaved approaches remain limited for visual manipulation and generalization.

  • Method

    ThinkMorph fine-tunes a unified model on approximately 24K high-quality traces in which textual and visual thoughts provide complementary, progressively advancing cues across four tasks.

  • Results

    ThinkMorph improves vision-centric performance by an average of 34.74% over its base model, generalizes to out-of-domain benchmarks, and exhibits unseen visual manipulation, autonomous mode switching, and diversified test-time reasoning.

  • Takeaways & Limitations

    Interleaved reasoning provides a controlled framework for studying when multimodal coordination helps and how unified models develop adaptive reasoning behaviors.

  • Takeaways & Limitations

    Prior unified interleaved reasoning can remain limited by simplistic text components that are isomorphic to generated images and generalize poorly beyond training domains.

Abstract

from arXiv · show

Multimodal reasoning requires iterative coordination between language and vision, yet it remains unclear what constitutes a meaningful interleaved chain of thought. We posit that text and image thoughts should function as complementary rather than isomorphic modalities that mutually advance reasoning. Guided by this principle, we build ThinkMorph, a unified model fine-tuned on approximately 24K high-quality interleaved reasoning traces spanning tasks with varying visual engagement. ThinkMorph learns to generate progressive text-image reasoning steps that concretely manipulate visual content while maintaining coherent verbal logic. It delivers large gains on vision-centric benchmarks (averaging 34.7 percent over the base model) and generalizes to out-of-domain tasks, matching or surpassing larger and proprietary VLMs. Beyond performance, ThinkMorph exhibits emergent multimodal intelligence, including unseen visual manipulation skills, adaptive switching between reasoning modes, and better test-time scaling through diversified multimodal thoughts. These findings suggest promising directions for characterizing the emergent capabilities of unified models for multimodal reasoning.

1 INTRODUCTION

ThinkMorph addresses limits of text-only and existing interleaved reasoning by treating language and images as complementary modalities. Its interleaved training improves vision-centric performance while exposing adaptive multimodal behaviors.

  • 1 INTRODUCTION: 34.74% average improvement over the base model across vision-centric tasks, including 85.84% on Spatial Navigation and 38.75% on Jigsaw Assembly.The model also generalizes to out-of-domain benchmarks and matches or surpasses larger systems.
  • 1 INTRODUCTION: ThinkMorph is fine-tuned on approximately 24K high-quality interleaved traces spanning tasks with different levels of visual engagement.The traces are designed so textual reasoning and visual manipulation progress together.
  • 1 INTRODUCTION: Interleaved reasoning treats text and images as complementary modalities that jointly advance reasoning, unlike approaches built around isomorphic or indirect representations.Existing tool-augmented methods depend on external visual modules, while prior unified approaches lack a generalizable recipe for mutual advancement.
  • 1 INTRODUCTION: ThinkMorph reveals unseen visual manipulations, autonomous switching between reasoning modes, and diversified multimodal exploration during test-time scaling.These behaviors are presented as emergent properties of the trained unified model.

2 THINKMORPH: INTERLEAVED CHAIN-OF-THOUGHT GENERALIZATION

ThinkMorph operationalizes interleaved reasoning as progressive sequences of textual and visual thoughts trained on verifiable multimodal traces. Its data and training design span four tasks with varied visual engagement and optimize both image and text generation.

  • 2 THINKMORPH: INTERLEAVED CHAIN-OF-THOUGHT GENERALIZATION: Interleaved thoughts combine generated text and image tokens in a sequence, with delimiter tokens controlling transitions between modalities.The model generates intermediate tokens conditioned on the question and prior thoughts.
  • 2 THINKMORPH: INTERLEAVED CHAIN-OF-THOUGHT GENERALIZATION: The training data make text and images complementary cues that progressively guide solutions rather than isomorphic representations.This design addresses the difficulty of externalizing and scaling meaningful visual reasoning.
  • 2 THINKMORPH: INTERLEAVED CHAIN-OF-THOUGHT GENERALIZATION: 24,990 questions span four tasks with different visual engagement, using custom synthesis for Jigsaw Assembly and Spatial Navigation and human-in-the-loop filtering for Visual Search and Chart Refocus.The dataset is designed around concrete, verifiable visual manipulations.
  • 2 THINKMORPH: INTERLEAVED CHAIN-OF-THOUGHT GENERALIZATION: Training optimizes separate image-token MSE and text-token negative log-likelihood objectives within the Bagel base-model implementation.Evaluation includes in-domain task benchmarks and additional out-of-domain settings.

3 WHEN DOES INTERLEAVING IMPROVE MULTIMODAL REASONING?

ThinkMorph helps most when reasoning requires sustained visual engagement and complementary text-image updates, while text-only reasoning can suffice when visual information is redundant. Its advantage therefore depends on task-specific visual grounding and manipulation demands.

  • 3 WHEN DOES INTERLEAVING IMPROVE MULTIMODAL REASONING?: 86.67% on Spatial Navigation represents an 85.84% improvement over the 0.83% base-model result, while vision-centric tasks average a 34.74% gain over the base model.Interleaved reasoning also surpasses the next-best reasoning mode by 5.33%.
  • 3 WHEN DOES INTERLEAVING IMPROVE MULTIMODAL REASONING?: Interleaved reasoning outperforms text-only and visual-only modes by 5.33% across reasoning-mode comparisons.The comparison is conducted across the evaluated tasks using the three reasoning modes.
  • 3 WHEN DOES INTERLEAVING IMPROVE MULTIMODAL REASONING?: On ChartQA, text-only reasoning exceeds interleaved reasoning by 1.88%, whereas interleaved reasoning exceeds text-only reasoning by 6.33% on out-of-domain MMVP.The contrast indicates that visual manipulation is supplementary for ChartQA but essential for MMVP.
  • 3 WHEN DOES INTERLEAVING IMPROVE MULTIMODAL REASONING?: Visual tokens provide task-specific evidence through rearranged pieces, route overlays, and bounding boxes that text alone cannot capture.Text-only reasoning suffices when additional visual information is redundant, but interleaving is crucial for precise visual grounding or manipulation.

4 EMERGENT PROPERTIES IN INTERLEAVED REASONING

ThinkMorph exhibits emergent multimodal behaviors: unseen visual manipulations, task-adaptive switching between reasoning modes, and stronger test-time scaling through diversified multimodal thoughts.

  • Unseen Visual Manipulations: ThinkMorph generates eight types of unseen visual manipulations, including zooming, inpainting, multi-box generation, motion forecasting, perspective transformation, and cropping.These operations can comprise up to 10% of visual operations on some benchmarks and directly support problem solving.
  • Unseen Visual Manipulations: Textual cues such as “examine closely” trigger zooming, whereas “restore” and “reconstruct” prompt inpainting in contextually appropriate ways.The passage attributes the capability to multimodal pretraining, with interleaved fine-tuning aligning it to reasoning steps.
  • Autonomous Mode Switching: 81.25% accuracy on switched instances exceeds interleaved reasoning’s 73.96%, while switching occurs in 5.3% of inference cases.The model switches from interleaved to text-only reasoning despite training exclusively on interleaved traces.
  • Autonomous Mode Switching: Mode switching follows visual complexity: unresolved fine-grained cues retain interleaved reasoning, whereas expressible initial visual information permits text-only reasoning.Switched samples use approximately 75% more tokens when forced to continue interleaved reasoning, such as 156 versus 89 tokens.
  • Better Test-Time Scaling: +8.0% on BLINK-J is the largest test-time scaling gain, with ThinkMorph improving from 65.33% to 73.33% while visual reasoning drops 2.0%.Across VSP, VStar, MMVP, and BLINK-J, interleaved reasoning gains are +5.2%, +1.0%, +0.7%, and +8.0%, respectively.
  • Better Test-Time Scaling: Diversified interleaved trajectories explore complementary multimodal solution subsets, making Best-of-N sampling more effective than unimodal reasoning.Interleaved reasoning spans both modalities, whereas unimodal chains remain confined to one representational space.

5 GENERALIZATION OF INTERLEAVED REASONING

Across broader vision-centric benchmarks, ThinkMorph generalizes beyond its training tasks, improving substantially over unified baselines and competing with much larger VLMs while showing task-dependent scaling behavior.

  • 5.1 RESULTS: ThinkMorph improves 20.74% over Bagel-7B across nine diverse tasks and exceeds Janus-Pro-7B and Chameleon-7B by 28.8%–42.7% on reported comparisons.On BLINK, ThinkMorph improves by 12.42%; the cited table provides the broader model comparison.
  • 5.1 RESULTS: ThinkMorph outperforms Qwen2.5VL-72B by 34% on VSP and 10.67% on BLINK-J, while surpassing InternVL3.5-38B on SAT.It also outperforms GPT-4o by 24.67% on SAT and matches Gemini 2.5 Flash at 80.33% on MMVP.
  • 5.2 ADDITIONAL ANALYSIS ON EMERGENT PROPERTIES: Unseen visual manipulations and adaptive mode switching remain consistent on broader out-of-domain tasks, while test-time scaling varies by task type.The passage introduces these as the first two properties persisting beyond the original evaluations.
  • 5.2 ADDITIONAL ANALYSIS ON EMERGENT PROPERTIES: Textual responses outperform interleaved responses by 9.75% on MMVP and 1.84% on VStar, but trail by 2.98% on BLINK-J.Across eight-response sampling, purely textual responses comprise 6.38% of MMVP, 8.64% of VStar, and 1.25% of BLINK-J responses.

6 RELATED WORK

Prior multimodal chain-of-thought methods use indirect tools or unified models with limited mutual advancement, while joint understanding and generation remain fragile despite evidence of synergy.

  • Multimodal Chain-of-Thought: Tool-augmented multimodal Chain-of-Thought approaches keep interleaving indirect and fragile, while unified approaches pursue tighter integration.The related work divides explicit multimodal Chain-of-Thought into these two broad lines.
  • Multimodal Understanding and Generation: Existing unified multimodal models often report that optimizing diffusion generation degrades understanding and learned representations, and vice versa.This makes joint training fragile and brittle.
  • Multimodal Understanding and Generation: MetaMorph shows that visual understanding and generation can be synergistic, with more data for either capability often benefiting both.Reasoning and understanding can also improve generative tasks according to the cited related work.

7 CONCLUSION

ThinkMorph uses interleaved fine-tuning to make text and images reinforce multimodal reasoning, producing broad gains and emergent behaviors that extend beyond explicit supervision.

  • 7 CONCLUSION: ThinkMorph enables text and images to reinforce each other through light interleaved fine-tuning, yielding gains on vision-centric benchmarks and competitiveness with larger proprietary systems.The conclusion frames this as generalizable multimodal reasoning.
  • 7 CONCLUSION: The model exhibits spontaneous visual manipulation, autonomous mode switching, and diversified exploration that improves test-time scaling.These behaviors are presented as emergent properties of unified multimodal reasoning.
  • 7 CONCLUSION: The paper identifies adaptive mode selection, stronger cross-modal alignment objectives, and coherent visual-text thought integration as directions for future work.These directions are proposed for nurturing interleaved reasoning behaviors.

B.1 TEST-TIME SCALING RESULTS

ThinkMorph compares text-only, visual-only, and interleaved reasoning modes, with test-time scaling results reported across these alternatives. Interleaved reasoning alternates text and image thoughts so the modalities can inform each other throughout reasoning.

  • B.1 TEST-TIME SCALING RESULTS: Test-time scaling is evaluated across text-only, visual-only, and interleaved reasoning modes.The section reports results through tables comparing these distinct thought-chain compositions.
  • B.1 TEST-TIME SCALING RESULTS: Text-only reasoning generates textual thought chains without intermediate image generation.
  • B.1 TEST-TIME SCALING RESULTS: Visual-only reasoning generates image-only thought chains followed by the final answer, without intervening text thoughts.
  • B.1 TEST-TIME SCALING RESULTS: Interleaved reasoning alternates text and image thoughts, allowing the two modalities to inform each other during reasoning.Text-only chains contain only textual tokens, whereas visual-only chains contain only image tokens before the final answer.
  • B.1 TEST-TIME SCALING RESULTS: Visual and interleaved reasoning use the same images, while interleaved reasoning adds complementary textual reasoning rather than duplicating visual content.

B.3 COMPUTATIONAL COST ANALYSIS

Interleaved reasoning costs more because it generates images, but it can achieve better performance-cost trade-offs than text-only reasoning and cannot be matched by simply adding text data.

  • B.3 COMPUTATIONAL COST ANALYSIS: Interleaved reasoning costs approximately 3× more than text at the same N because it generates images.
  • B.3 COMPUTATIONAL COST ANALYSIS: Interleaved N=4 reaches 70.0% on BLINK-J versus text N=8 at 68.0%, while using fewer total tokens.The reported token counts are 4,269 for interleaved reasoning versus 2,879 × 1 per sample for text reasoning.
  • B.3 COMPUTATIONAL COST ANALYSIS: Interleaved N=1 reaches 81.33% on MMVP, surpassing text N=8 at 80.33% with substantially lower cost.
  • B.3 COMPUTATIONAL COST ANALYSIS: Text-only training with 4× more samples does not close the interleaved gap, with VStar at 55.50% versus 63.87% and VSP at 47.17% versus 86.67%.The comparison is presented as evidence that the advantage comes from the multimodal search space rather than additional compute.
  • B.3 COMPUTATIONAL COST ANALYSIS: Interleaved reasoning produces richer and more diverse textual reasoning than text-only reasoning across the evaluated benchmarks.Table 7 reports text entropy and vocabulary diversity across reasoning modes.

B.4 MULTIMODAL TRAINING ENRICHES TEXT REPRESENTATIONS

Interleaved training enriches the model’s textual representations, producing higher entropy, greater vocabulary diversity, and more visually grounded language than text-only reasoning.

  • B.4 MULTIMODAL TRAINING ENRICHES TEXT REPRESENTATIONS: Unique vocabulary increases by 11.4% on VStar and 24.1% on Jigsaw under interleaved reasoning.The reported counts are 3,186 versus 2,860 on VStar and 3,795 versus 3,057 on Jigsaw.
  • B.4 MULTIMODAL TRAINING ENRICHES TEXT REPRESENTATIONS: Interleaved text has consistently higher entropy and richer vocabulary than text-only text across both benchmarks.
  • B.4 MULTIMODAL TRAINING ENRICHES TEXT REPRESENTATIONS: Interleaved reasoning uses more visual-relational terms, whereas text-only reasoning relies more on generic tokens.The paper links this pattern to more expressive and visually grounded textual reasoning, including when the model switches to text-only mode.

B.5 DETAILS ON QUESTION CONSTRUCTION AND FINETUNING DATA CURATION

The paper constructs diverse visual reasoning datasets and curates their supervision through task-specific generation, filtering, manual review, and answer-evaluation procedures.

  • B.5 DETAILS ON QUESTION CONSTRUCTION AND FINETUNING DATA CURATION: Jigsaw Assembly converts images into multiple-choice puzzles with two to four pieces across several grid configurations.The pipeline samples arrangement options, including the correct configuration, from images sourced across three datasets.
  • B.5 DETAILS ON QUESTION CONSTRUCTION AND FINETUNING DATA CURATION: Visual Search collects 144k problems and filters images whose target objects occupy 1%-30% of the image before manual review.The review identifies ambiguous phrasing, incorrect answers, and misplaced bounding boxes as quality issues.
  • B.5 DETAILS ON QUESTION CONSTRUCTION AND FINETUNING DATA CURATION: Spatial Navigation generates Frozen Lake problems on 3×3 to 6×6 grids and visualizes candidate paths with red lines and arrows.The pipeline prompts GPT-4.1 to describe maze layout information before generating navigation reasoning.
  • B.5 DETAILS ON QUESTION CONSTRUCTION AND FINETUNING DATA CURATION: Chart Refocus uses bar-chart questions filtered to require one highlighting or drawing operation, followed by manual review of 8.4k questions.
  • B.5 DETAILS ON QUESTION CONSTRUCTION AND FINETUNING DATA CURATION: Evaluation uses official or VLMEvalkit judging pipelines, with GPT-5 answer extraction for ChartQA responses.The ChartQA extraction prompt outputs only the final answer and removes units or extra symbols from numeric answers.

B.7 TRAINING AND INFERENCE DETAILS

ThinkMorph’s training and inference setup supports unified interleaved outputs, autonomous modality switching, and distinct decoding regimes for single-pass versus test-time scaling. The appendix also documents task-specific prompts that structure visual reasoning and answer generation.

  • Training and inference: Bagel-7B is trained on curated interleaved traces as unified autoregressive streams using two nodes with 16×A100 80GB GPUs.Other settings follow the official Bagel defaults unless specified in the hyperparameter table.
  • Training and inference: Two special tokens, <image start> and <image end>, enable autonomous switching between text reasoning and image generation.Text traces and final answers are additionally wrapped with dedicated markers.
  • Training and inference: Inference uses temperature=0 for single-pass runs and temperature=0.7 for test-time compute scaling, with max tokens=4096 in both settings.
  • Prompting: Task prompts require detailed, visually grounded reasoning while instructing the model to conceal provided answers and present conclusions as independent analysis.The prompts use visual descriptions, highlighted or edited images, and structured reasoning stages across chart, search, and other tasks.
Loading 2510.27492v3…