Source-linked AI summary

Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment

Zihao Wang, Xi Xiang, Yuwen Sun, Yingyu Li, Yabo Zhang, Yihan Zeng, Fan Li, Wangmeng Zuo

arXiv:2609.02573v1cs.CV

TL;DR

Existing multimodal evaluations often reduce interleaved text-image reasoning to static perception, leaving dynamic cross-modal evidence integration insufficiently assessed. TIC-Bench evaluates this capability across logical, temporal, and spatial associations, finding a substantial gap between current models and human experts, with the strongest model reaching 59.9% versus 91.7%.

  • Problem

    Existing evaluations often reduce interleaved text-image tasks to static perception, leaving dynamic cross-modal evidence integration insufficiently assessed.

  • Method

    TIC-Bench evaluates deeply interleaved text-image reasoning across three association domains and eight task types using complementary visual and textual clues.

  • Results

    59.9% overall accuracy for GPT-5.5 versus 91.7% for humans, with models especially challenged by convergent, parallel, and distributed-evidence reasoning.

  • Takeaways & Limitations

    TIC-Bench provides an analytical benchmark for comparing multimodal models’ ability to connect evidence distributed across interleaved images and text.

  • Takeaways & Limitations

    Existing models still struggle to reliably connect evidence distributed across multiple images and text segments, requiring stronger long-context multimodal reasoning.

Abstract

from arXiv · show

Current evaluations and training of multimodal models predominantly focus on multi-image tasks, largely overlooking interleaved text-image scenarios. In such multi-image tasks, text typically serves merely as task instructions, lacking deep semantic interaction with the visual content. In contrast, realworld applications like text-image co-creation, character tracking, and spatial reconstruction require constant interaction between text and images. Consequently, models must possess a deep understanding of these interleaved contexts. To bridge this gap, we introduce a novel benchmark, TIC-Bench (deeply interleaved Text-Image Contexts), designed to evaluate the capability of models to integrate text-image clues and recover the ground truth facts within deeply interleaved contexts. This benchmark encompasses three core domains: Logical, Temporal, and Spatial Association, which are further categorized into eight specific types, comprising a total of 2,280 questions. We evaluated 10 state-of-the-art MLLMs and observed a substantial performance gap compared to human experts, together with persistent difficulties in integrating evidence distributed across interleaved visual and textual inputs. Ultimately, this benchmark provides a valuable analytical tool for assessing and advancing the ability of multimodal models to effectively integrate text and image information in deeply interleaved contexts. TIC-Bench is publicly available at https://huggingface.co/datasets/pino10010/TIC-Bench

Introduction

TIC-Bench addresses the limitations of isolated text-image and traditional multi-image evaluations by testing whether MLLMs can integrate fragmented visual and textual clues across deeply interleaved contexts. The benchmark reveals substantial gaps from human performance and persistent difficulties in merging concurrent or multiple reasoning streams.

  • Motivation: TIC-Bench evaluates continuous binding, integration, and propagation of fragmented visual and textual clues across extended interleaved sequences.This targets reasoning over text-image contexts rather than perceiving each modality in isolation.
  • Motivation: Existing training corpora and evaluation paradigms largely use isolated text-image pairs or reduce interleaved tasks to static perception, limiting assessment of dynamic cross-modal reasoning.These paradigms deprive models of high-density interleaved data and fail to test evidence integration across sequences.
  • Benchmark: TIC-Bench contains 2,280 questions spanning Logical, Temporal, and Spatial Association domains and eight task types designed for deeply interleaved text-image reasoning.The benchmark minimizes answers based solely on parametric knowledge and emphasizes genuine cross-modal evidence integration.
  • Results: 59.9% overall accuracy for GPT-5.5 versus 91.7% for the human baseline demonstrates a substantial performance gap on TIC-Bench.The evaluation covered 10 state-of-the-art MLLMs across 15 inference settings.
  • Results: Models generally outperform on Cyclic and Linear logic relative to Convergent logic, while Parallel reasoning is most difficult in the Temporal Association domain.The findings identify merging multiple reasoning branches and maintaining concurrent evidence streams as persistent bottlenecks.

Related Work

Existing multimodal benchmarks have advanced evaluation across visual question answering, captioning, reasoning, and multi-image understanding, but often underrepresent relational reasoning and deep text-image interaction. TIC-Bench addresses this gap by systematically evaluating logical, temporal, and spatial associations in interleaved text-image contexts.

  • MLLM Reasoning Evaluation: Existing MLLMs have progressed on visual question answering, image captioning, and visual reasoning, yet their reliance on visual evidence remains unresolved.VisReason reports that current MLLMs still rely heavily on language priors rather than genuinely grounding reasoning in visual evidence.
  • Multi-Image Benchmarks: 11.5K college-level questions across six disciplines make MMMU broad, but its predominantly single-image format omits relational reasoning across multiple images.MMMU is contrasted with later benchmarks that explicitly target multi-image understanding.
  • Multi-Image Benchmarks: MMIU introduced systematic multi-image evaluation, while MMRB assessed spatial, temporal, and semantic reasoning across multiple images.These benchmarks nevertheless use text primarily as question descriptions rather than deeply interleaved evidence.
  • Interleaved Reasoning: MIR introduced Multi-image Interleaved Reasoning, but did not systematically categorize associations within interleaved text-image contexts.TIC-Bench extends this line of work by systematically evaluating logical, temporal, and spatial associations.

Benchmark

TIC-Bench evaluates joint cross-modal reasoning in deeply interleaved contexts, requiring models to align complementary textual and visual evidence across multi-step tasks. It covers logical, temporal, and spatial association with substantial multimodal scale and high interleaving density.

  • Benchmark formulation: Each instance interleaves visual and textual elements without requiring strict alternation, while requiring complementary evidence from both modalities.Solving tasks requires preserving cross-modal correspondence and integrating dispersed clues through multi-step reasoning.
  • Benchmark composition: TIC-Bench comprises Logical Association, Spatial Association, and Temporal Association subsets.These dimensions construct interleaved text-image evaluation scenarios.
  • Logical Association: Logical Association identifies targets through cross-image relation chains involving containment, proximity, or co-occurrence.Its structures are Linear Logic, which follows one chain; Cyclic Logic, which revisits an image; and Convergent Logic, which combines independent branches.
  • Temporal Association: Temporal Association combines comic-style scenes and interleaved descriptions, with images conveying identities and events while text conveys actions, causality, and plot progression.Sequential, Retrospective, and Parallel scenarios require timeline tracking, retrieval of earlier evidence, or integration of concurrent storylines.
  • Dataset scale: 2,280 questions and 45,776 image instances comprise the dataset, averaging 20.08 images per question and 1.86 references per image.The benchmark also averages 37.34 image references and 309.3 textual-context words per sample, exceeding comparison benchmarks on these density measures.

Experiments

Experiments evaluate ten open- and closed-source MLLMs against human experts using unified prompts and automatic judging with human adjudication for disagreements. Results show a substantial human–model performance gap, strong dependence on both modalities, and characteristic reasoning errors.

  • Evaluation setup: Ten MLLMs—five open-source and five closed-source—are evaluated, with open-source models tested in thinking and non-thinking modes and human experts providing a reference.Closed-source models use default inference settings, while all models receive unified prompts and consistent input formatting within each benchmark domain.
  • Overall human–model gap: 91.7% human accuracy exceeds GPT-5.5’s 59.9% and Gemini 3.1 Pro’s 59.0%, leaving a gap of more than 30 percentage points.Model performance also varies substantially across association types, unlike the consistently strong human performance.
  • Overall human–model gap: r = 0.81 for Parallel–Photo is the strongest cross-domain subtask correlation, followed by r = 0.74 for Parallel–Map.These correlations suggest a shared bottleneck in integrating evidence distributed across multimodal inputs for concurrent-event and spatial-relation reasoning.
  • Modality ablations: Full interleaved contexts achieve the highest overall accuracy for all four tested models, while removing either text or images causes a substantial performance drop.The ablations compare Full, Text-only, and Images-only settings for Gemma-4-31B-thinking, Qwen3.6-35B-A3B-thinking, GPT-5.5, and Gemini 3.1 Pro Preview.
  • Evaluation reliability: 88.6% sample-level agreement and Cohen’s κ = 0.772 characterize the automatic evaluation pipeline against human evaluation.The pipeline uses three independent judges and human adjudication when their judgments disagree, based on 500 human-scored Time QA responses.
  • Error analysis: Qwen’s primary error type is abstraction at 36.6%, followed by hallucination at 23.0% and reasoning errors at 21.9%.The analysis assigns each incorrect response one of seven mutually exclusive error labels.

Conclusion

The benchmark evaluates association reasoning across deeply interleaved text-image contexts in logical, temporal, and spatial domains. Results show a substantial model–human gap and persistent difficulty connecting distributed multimodal evidence, motivating stronger long-context capabilities.

  • Conclusion: The benchmark covers three domains and eight subtasks for association reasoning over interleaved text-image contexts.The domains are logical, temporal, and spatial associations.
  • Conclusion: Current models exhibit a substantial performance gap relative to human experts, while showing complementary strengths across tasks.Thinking mode improves performance in some cases but does not eliminate the broader gap.
  • Conclusion: Existing models still struggle to reliably connect evidence distributed across multiple images and text segments.This limitation persists despite improvements from thinking mode in some cases.
  • Conclusion: Future models need stronger abilities to manage long interleaved multimodal contexts, preserve relevant information, suppress distractions, and establish consistent cross-modal associations.These capabilities directly address the benchmark’s observed integration difficulties.

Supplementary Material Overview

The supplementary material documents dataset construction, prompt collection, evaluation protocols, model inference details, and additional analysis procedures. It covers the benchmark’s data sources, construction pipeline, evaluation conditions, processing settings, and error analysis.

  • Detailed Dataset Construction: Dataset construction describes the data sources and pipeline for the three benchmark domains.
  • Prompts: Prompts collects prompts for dataset construction, Temporal inference, and automatic evaluation.
  • Human Evaluation Protocol: The human evaluation protocol documents evaluator backgrounds, answering conditions, and accuracy computation.
  • Model Inference and Evaluation Details: Model inference and evaluation details report input assembly, image processing, output parsing, and judge settings.
  • Additional Analysis Protocols: Additional analysis protocols present the error analysis.

Detailed Dataset Construction

The dataset construction pipeline builds structured logical scenes, temporally ordered storyboards, and spatial crop-based instances, then assembles interleaved text-image questions with reference answers. Human review and source-image geometry help preserve observable evidence, event transitions, and unambiguous spatial relations.

  • Logical Association: 5–10 objects populate each structured visual node, with generated names, attributes, scene content, and image-generation descriptions organized under required field constraints.Qwen 3.7 Plus generates scene information, while Claude Opus 4.7 organizes it into structured visual-node descriptions.
  • Logical Association: Three logical tasks connect scenes through shared objects, attributes, and relations: Linear, Cyclic, and Convergent Logic.Linear Logic follows one image-to-image clue chain, Cyclic Logic revisits an earlier image, and Convergent Logic combines independent clue paths.
  • Temporal Association: Temporal instances use keyframes from six live-action productions and three animated series, arranged as comic-strip-like storyboards by human reviewers.Reviewers select frames preserving event transitions and character-state changes, revise descriptions and references, and interleave text supplying actions, causal links, and plot transitions.
  • Spatial Association: Spatial instances derive from COCO photographs and CVOGL aerial imagery, dividing each source into unrotated square crops and using overlap connectivity for reconstruction.Candidate target pairs come from different crops, while unsuitable or ambiguous pairs are rejected before directional relations are computed.
  • Spatial Association: Spatial prompts restrict models to locally visible content, while source-image coordinates determine neighboring crop relations and prevent cross-crop directional inference.The construction also uses class-wise non-maximum suppression to reduce duplicate landmarks across neighboring crops.

Prompts

This section collects the prompts used for dataset construction, Temporal inference, and automatic answer evaluation. The prompts cover Logical scene generation, question–answer construction, Temporal frame description, and identity-reference instruction.

  • Prompt Functions: The prompt set supports dataset construction, Temporal inference, and automatic answer evaluation.Prompt 1 addresses Logical scene generation, Prompt 2 converts question skeletons into fluent question–answer pairs, and Prompt 3 generates visually grounded Temporal-frame descriptions.
  • Prompt Functions: Prompt 1 specifies structured requirements for Logical scene generation.
  • Prompt Functions: Prompt 2 converts precomputed question skeletons into fluent question–answer pairs, while Prompt 3 generates visually grounded descriptions of Temporal frames.
  • Prompt Functions: Prompt 5 provides the identity-reference instruction used during Temporal inference.

Human Evaluation Protocol

The human baseline was produced by ten experienced undergraduate computer science evaluators answering all 2,280 questions under the same evidence and evaluation procedure as the models. Overall human accuracy was reported as the unweighted macro-average of Logical, Temporal, and Spatial domain accuracies.

  • Human Evaluation Protocol: 10 evaluators with prior multimodal research experience independently answered all 2,280 benchmark questions, yielding 22,800 responses.They received the same evidence as models, had no time limit, and could not use external tools.
  • Human Evaluation Protocol: Human answers were assessed using the same answer-evaluation procedure applied to model outputs.
  • Human Evaluation Protocol: Overall human accuracy was computed as the unweighted macro-average of Logical, Temporal, and Spatial domain-level accuracies.

Model Inference and Evaluation Details

The evaluation assembles multimodal inputs in their original interleaved order, with domain-specific prompt and ablation configurations. Images are standardized before submission, and model responses are judged against the question, context, and reference answer.

  • Prompting: Logical and Spatial inference use no task-specific system prompt, whereas Temporal inference requires identity identifiers for people and animals.The identity instruction applies only to Temporal instances.
  • Input construction: Inputs are dynamically assembled in original interleaved order into description and question components, with domain-specific images, textual clues, queries, and candidate options.Logical descriptions include scene and candidate-object images; spatial questions contain relation queries and candidate directions.
  • Input construction: Full input retains all interleaved content, Text-only removes image contents, and Images-only removes descriptive textual evidence while preserving answer-essential questions and options.Text-only preserves image-reference markers and nonanswer context; Images-only retains every available input image.
  • Image processing and output evaluation: Before API submission, images are converted to RGB, resized above 1,024 pixels with Lanczos interpolation, and JPEG-encoded at quality 95 as Base64.Aspect ratio is preserved during resizing.
  • Image processing and output evaluation: Each judge follows Prompt 6 evaluation instructions, with the default configuration using a 0–100 scale.The original model response is evaluated alongside the question, context, and reference answer.

Additional Analysis Protocols

The analysis protocols define seven mutually exclusive error labels by the earliest identifiable error stage and use constrained prompts for temporal captioning, spatial crop description, temporal reasoning, and automatic evaluation. Figure-based evidence illustrates how relational composition errors are distinguished from perception or abstraction errors.

  • Error annotation: Seven mutually exclusive labels classify each incorrect response by its earliest identifiable error stage: perception, abstraction, reasoning, understanding, hallucination, knowledge, or other.Perception concerns visual recognition, abstraction concerns mapping perceived evidence to the wrong entity, scene, or answer option, and reasoning concerns relational-composition failures.
  • Temporal frame captioning: Temporal frame captioning requires structured descriptions of scenes, entities, actions, spatial relations, and visually supported state changes relative to previous frames.The prompt also requires consistent entity names across frames and marks unavailable state changes as “None observable.”
  • Captioning constraints: Captioning protocols restrict outputs to visually supported information, prohibit inferred identities, intentions, causes, or unseen events, and require ambiguous details to be marked uncertain.These constraints apply alongside preserving consistent entity names across frames.
  • Reasoning-error evidence: A Figure 8 example classifies identifying target A northwest of target B as a reasoning-error case because source-image coordinates determine the Northwest relation.The cited evidence states that this classification is not perception or abstraction error.
  • Reasoning-error evidence: In Figure 9, reducing Northwest to West after discarding a northward offset is classified as reasoning error because the failure occurs during multi-step relational composition.The model correctly identifies the crops containing both targets and recovers the westward component before making the composition error.

Representative Benchmark Examples

The section presents human-validated representative examples spanning all eight task types, with complete visual inputs, textual contexts, questions, and reference answers. It also reports uneven correctness among core open-source models and illustrates logical, spatial, and temporal reasoning failures.

  • Representative examples: Eight human-validated examples cover the benchmark’s task types, each including the complete visual input, textual context, question, and reference answer.Figures 10–17 present examples for Linear, Cyclic, Convergent, Sequential, Retrospective, Parallel, Map, and Photo tasks.
  • Model performance: Three of five core non-thinking open-source models correctly answer each Logical and Temporal example, while Spatial examples represent middle difficulty.The examples are selected to illustrate both task coverage and difficulty distribution.
  • Representative errors: A Logical Association example shows unstable entity binding: the model follows most cross-scene relations but maps the final entity to the wrong option.The error arises from drift in the final entity-to-option mapping despite mostly correct relational tracking.
  • Representative errors: A Spatial Association example shows incomplete direction composition, preserving the westward component while losing the required northward component in a multi-hop relation.The response model is Qwen3.6-35B-A3B-thinking.
  • Representative errors: A Temporal Association example shows hallucination from an unsupported narrative trope and character relationship, producing a wrong option that all three judge models mark incorrect.The response model is Qwen3.6-35B-A3B-thinking.
Loading 2609.02573v1…