Source-linked AI summary
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
Zhekai Chen, Yuqing Wang, Manyuan Zhang, Xihui Liu
TL;DR
Multi-reference image generation remains largely unsolved because models struggle to reason over many visual references and structured training data and standardized evaluation are scarce. The paper introduces MacroData and MacroBench, finding consistent improvements in long-context multi-reference generation after training on MacroData.
Problem
Multi-reference image generation remains largely unsolved because existing data and benchmarks inadequately support reasoning over many interdependent visual references.
Method
The paper introduces MacroData, a 400K-sample dataset with up to 10 references across four dimensions, and MacroBench, a 4,000-sample standardized benchmark across task types and input scales.
Results
Training on MacroData yields consistent improvements in long-context multi-reference generation, while ablations provide guidelines for data construction, task synergy, and efficiency.
Takeaways & Limitations
MacroData and MacroBench provide a structured foundation for developing and evaluating multi-reference image generation across tasks and input scales.
Takeaways & Limitations
Performance still degrades with 6–10 input images, MacroBench covers limited predefined tasks, and fine-tuned models trail state-of-the-art closed-source models.
Abstract
from arXiv · showhide
Generating images conditioned on multiple visual references is critical for real-world applications such as multi-subject composition, narrative illustration, and novel view synthesis, yet current models suffer from severe performance degradation as the number of input references grows. We identify the root cause as a fundamental data bottleneck: existing datasets are dominated by single- or few-reference pairs and lack the structured, long-context supervision needed to learn dense inter-reference dependencies. To address this, we introduce MacroData, a large-scale dataset of 400K samples, each containing up to 10 reference images, systematically organized across four complementary dimensions -- Customization, Illustration, Spatial reasoning, and Temporal dynamics -- to provide comprehensive coverage of the multi-reference generation space. Recognizing the concurrent absence of standardized evaluation protocols, we further propose MacroBench, a benchmark of 4,000 samples that assesses generative coherence across graded task dimensions and input scales. Extensive experiments show that fine-tuning on MacroData yields substantial improvements in multi-reference generation, and ablation studies further reveal synergistic benefits of cross-task co-training and effective strategies for handling long-context complexity. The dataset and benchmark will be publicly released.
1 Introduction
Multi-reference image generation is limited by performance degradation with more inputs and a scarcity of structured training data and standardized evaluation. The paper introduces MacroData and MacroBench, then reports improvements from fine-tuning and identifies strategies for handling complex long-context references.
- Motivation: Real-world tasks such as multi-subject composition, narrative illustration, and novel view synthesis require reasoning over multiple visual references.Existing advances have primarily targeted single- and few-reference tasks.
- Motivation: Models struggle as reference count grows: OmniGen2 accepts at most five input images, while Bagel degrades severely beyond three references.The paper attributes this deficiency primarily to scarce high-quality, structured multi-reference training data.
- MacroData: MacroData contains 400K samples with up to 10 reference images, organized across Customization, Illustration, Spatial, and Temporal dimensions.This structure is intended to facilitate learning dense inter-reference dependencies.
- MacroBench: MacroBench provides 4,000 samples evaluating generative coherence across four task dimensions and graded input scales from 1 to 10 images.It addresses the lack of standardized evaluation beyond customization tasks with at most three inputs.
- Results: Fine-tuning open-source models on MacroData yields substantial multi-reference generation improvements and narrows the gap with closed-source models.The experiments include state-of-the-art open-source models such as Bagel.
- Results: Ablations identify cross-task co-training and token selection as effective strategies for enhancing long-context multi-reference generation.The studies examine data scaling and task synergies for handling complex multi-reference contexts.
2 Related Work
Prior in-context image generation research has explored autoregressive, hybrid, and diffusion architectures with specialized visual representations, while recent models improve understanding–generation integration and vision–language alignment. Existing datasets rely mainly on model distillation or real-world retrieval, and benchmarks commonly use LLM-as-Judge evaluation but cover only a few references and limited task dimensions.
- In-Context Image Generation Model: Recent in-context image generation models use autoregressive, hybrid, or diffusion-based architectures with specialized vision representations.Bagel separates understanding and generation tokens with a Mixture-of-Transformer design, while OmniGen2 co-trains diffusion models with LVLM hidden states for tighter vision-language alignment.
- In-Context Image Generation Dataset: Existing in-context generation datasets are constructed mainly through distillation from strong generative models or retrieval from real-world corpora.Echo4o and MICo synthesize identity-consistent pairs using closed-source models, whereas OpenSubject extracts and matches relevant images from web pages.
- In-Context Image Generation Benchmark: Existing benchmarks commonly use LLM-as-Judge evaluation but cover at most a few input images and omit spatial reasoning, temporal coherence, and systematic reference scaling.OmniContext uses GPT-4.1 to score prompt adherence and subject consistency.
3 MacroData
MacroData is a 400K-sample dataset for multi-reference image generation, supporting up to 10 inputs per sample and providing substantial long-context coverage. It is evenly organized across four tasks spanning customization, illustration, spatial reasoning, and temporal dynamics.
- Dataset scale and context: 400K samples support up to 10 reference inputs per sample, averaging 5.44 inputs and providing substantial coverage beyond four references.MacroData addresses the scarcity of many-to-one data and exceeds datasets lacking samples beyond five references.
- Task definition: 100K samples each are allocated to Customization, Illustration, Spatial, and Temporal tasks to maintain data balance.The four tasks cover subject customization, diverse interleaved-context topics, 3D consistency, and video dynamics.
- Customization: Customization covers five subject categories: human, object, scene, cloth, and style.This task targets varied subject-level customization scenarios.
- Illustration: Illustration features 100 diverse topics, including narratives and health, derived from interleaved contexts.The task is designed around diverse contextual content.
- Spatial and Temporal: Spatial focuses on 3D consistency for multiview objects and panoramas, while Temporal captures video dynamics across durations from 0 to 120+ seconds.These tasks target spatial coherence and temporal variation across visual references.
2) Illustration:
The Illustration subset combines heterogeneous visual sources with tailored preprocessing, then constructs coherent multi-reference samples through semantic filtering and VLM-based quality control. Its related spatial and temporal pipelines select plausible views or coherent keyframe sequences for target prediction.
- Source Collection and Preprocessing: Source data span humans, objects, scenes, clothing, and styles, with tailored keyframe sampling, face removal, style tagging, and aesthetic scoring.Sources include OpenSubject, MVImgNet, DL3DV, Vibrant Clothes Rental Dataset, and WikiArt.
- Sampling and Generation: Metadata combinations are iteratively resampled with LLMs, then VLM bidirectional assessments remove outputs that mismatch inputs or contradict prompts.This procedure targets both logical or spatial implausibility and failures of input-output or prompt-image fidelity.
- Illustration Subset: Web-crawled interleaved sequences provide anchor images selected for relevance to accompanying text and preceding images, while VLMs rewrite contexts and filter low-quality samples.Anchors become generation targets, and preceding context supplies the input conditions.
- Outside-in Objects: The outside-in spatial subtask uses canonical views of 3D objects, filters visually unsuitable renders in HSV space, and designates one view as the target.It is constructed from multi-view renderings in the G-buffer Objaverse dataset to capture central objects from surrounding external viewpoints.
- Video Clip Extraction: Temporal samples segment videos at shot boundaries, retain central keyframes, and group them into coherent sequences using DINOv2 visual similarity.This compresses redundancy while preserving key visual transitions within scene boundaries.
4 MacroBench
MacroBench evaluates multi-reference in-context generation across four task dimensions and input scales up to 10 images. Its 4,000 held-out pairs use task-specific 1–10 scoring and Gemini-3-Flash as the judge model.
- Benchmark scope: MacroBench covers Customization, Illustration, Spatial, and Temporal generation tasks with up to 10 input images.It is designed as a comprehensive benchmark for in-context generation across these four dimensions.
- Benchmark structure: The benchmark crosses four task categories with four input-count ranges: 1–3, 4–5, 6–7, and 8–10 images.This two-dimensional organization supports analysis across task complexity and input scale.
- Data curation: 4,000 held-out evaluation pairs are evenly allocated across four tasks and four image-count categories, with 250 pairs per task-category combination.The evaluation sources include metadata, documents, objects, panoramas, and videos, and are excluded from training.
- Judge model: Gemini-3-Flash is selected as the LLM judge because it provides reliable evaluations across task dimensions and input scales at reasonable cost.Commonly used judges such as GPT-4.1 show degraded quality with multiple references, especially on Spatial tasks requiring 3D reasoning.
- Evaluation metrics: Task-specific metrics scored from 1–10 measure consistency, instruction adherence, text alignment, view transformation, and content preservation.Customization uses ICS and PFS; Illustration uses TCS and ICS; Spatial uses VTS and CCS; the overall score averages across tasks.
5 Experiments
Experiments show that MacroData improves multi-reference generation across benchmarks, task dimensions, and input scales, while ablations identify effective data and long-context handling strategies. Fine-tuned models achieve strong overall and external-benchmark performance, with improved robustness as reference counts increase.
- MacroBench: MacroData fine-tuning outperforms all open-source baselines across MacroBench metrics, with Bagel scoring 5.71 versus 3.03 for base Bagel and ranking third overall.The fine-tuned model approaches Nano Banana Pro in Customization and surpasses it in Spatial tasks.
- MacroBench: Increasing inputs from 1–5 to 6–10 degrades performance generally, but MacroData improves robustness, mitigating Qwen’s Customization drop from 5.49 to 3.62 and Illustration drop from 2.50 to 1.55.MacroData also provides stable gains on challenging Spatial tasks where base models typically score below 1.0.
- OmniContext Benchmark: MacroData surpasses Echo4o on OmniContext, achieving 8.26 versus 8.09 despite targeting the broader multi-reference setting.The result supports the quality of the dataset’s data-collection pipeline on 1–3 image Customization tasks.
- Qualitative Results: Qualitative results show coherent multi-image feature integration in Customization, accurate target-view synthesis in Spatial tasks, and faithful generation in Temporal tasks.The method maintains strong Customization performance even with 10 inputs.
- Token Selection Strategies: Image-wise token selection underperforms the baseline, whereas block-wise and text-aligned selection outperform it by retaining crucial cross-reference information.Text-aligned selection remains effective even at low retention ratios, while image-wise filtering loses information by weakening cross-reference interactions.
6 Conclusion · Appendix
The paper addresses multi-reference image generation by introducing structured training data and standardized evaluation for reasoning over many visual references. It presents MacroData and MacroBench as complementary resources for improving and assessing coherent generation.
- 6 Conclusion: The paper targets coherent image generation from many visual references simultaneously.The challenge requires models to reason over multiple references while producing coherent outputs.
- 6 Conclusion: MacroData contains 400K samples with up to 10 reference images.The dataset is designed to address the scarcity of structured training data.
- 6 Conclusion: MacroData spans Customization, Illustration, Spatial, and Temporal dimensions.These four dimensions provide complementary coverage of the multi-reference generation setting.
- 6 Conclusion: MacroBench addresses the lack of standardized evaluation for multi-reference generation.The benchmark is introduced alongside MacroData to evaluate generation in this setting.
- 6 Conclusion: Together, MacroData and MacroBench target both training-data scarcity and evaluation gaps.The dataset supplies structured training data, while the benchmark provides standardized assessment.
- 6 Conclusion: The proposed resources are intended to support coherent generation under multi-reference conditions.Their design responds to the need for models that reason over many visual references simultaneously.
A Detailed Data Construction Pipeline … A.5 More Visualizations of MacroData
MacroData is constructed through large-scale source collection, model-assisted composition and curation, structured spatial and temporal processing, and visualized coverage across varied tasks and reference counts. The pipeline spans customization, illustration, spatial, and temporal subsets with explicitly defined data sources, sampling procedures, and view or sequence organization.
- A.1 Customization Subset: Customization source collection draws from millions of identities, videos, images, and artworks, with additional human, object, scene, and cloth data from Echo4o.The listed sources include over 2 million identities, 200,000 object videos, 10,000 scene videos, 50,000 clothing images, and 10,000 artworks.
- A.1 Customization Subset: Customization composition mixes human, object, scene, cloth, and style metadata at an 8:6:3:2:1 ratio before model-based evaluation, generation, and consistency assessment.Gemini-3-Flash evaluates combinations and bidirectional consistency, while Nano Banana Pro generates target images.
- A.2 Illustration Subset: Illustration construction reorganizes samples from 210 million interleaved image-text sequences using Qwen-3VL-8B and Gemini-3-Pro to identify anchors and synthesize concise context.The passage describes random sampling of candidate targets and preceding contexts followed by semantic re-evaluation and sample reorganization.
- A.3 Spatial Subset: Spatial object data uses a 10-category G-buffer Objaverse setup with multi-view rendering and canonical top, bottom, left, right, front, back, and diagonal perspectives.The source provides 24 views in one elevation range, 12 in another, plus top and bottom views; inputs are drawn from 15° or corresponding 30° rotation sets.
- A.3 Spatial Subset: Spatial inside-out scenes use Qwen-3VL-8B to classify filtered panoramas as indoor or outdoor, retain 10 canonical views, reverse horizontal ordering, and randomize camera yaw with pitch between −10° and 10°.The inside-out sequence is counter-clockwise, contrasting with the clockwise outside-in object sequence.
- A.4 Temporal Subset: Temporal data samples 1 million links from 10 million YouTube videos, then applies TransNetv2, DINOv2, and Gemini-3-Flash for segmentation, keyframe extraction, shot-boundary identification, and sequence description.The sampled subset is downloaded before clip segmentation and central keyframe extraction.
- A.5 More Visualizations of MacroData: Additional MacroData visualizations show varying reference counts across tasks, diverse customization inputs, broad illustration topics, three spatial categories, and temporal examples.The spatial categories are objects, indoor panorama, and outdoor panorama, with different target viewpoints for prediction.
B Benchmark … Correlation Metrics.
MacroBench specifies judge inputs, task-level aggregation rules, and an overall score across four tasks, then validates Gemini-3-Flash against human annotations using three correlation measures. Its validation set combines 160 model-generated samples with 120 ground-truth samples.
- B.1 Benchmark Prompt: For each query, the judge receives the task-specific instruction, all reference images, and the generated output; Spatial and Temporal tasks also provide the ground-truth target.These prompts are documented for every MacroBench task.
- B.2 Calculation Details: Raw metric scores lie in [0, 10], and each sample’s two task-specific scores are aggregated into one scalar using their geometric mean.The geometric mean penalizes severe failure on either dimension more strongly than the arithmetic mean.
- Single-score Aggregation.: The task metric pair for Customization is ICS × PFS.The metric-pair specification is part of the single-score aggregation procedure.
- Harmonic Mean for Customization ICS.: Customization ICS is computed per reference image and combined with a harmonic mean because fidelity to every individual reference subject is required.The harmonic mean suppresses the overall ICS when fidelity to any single reference is low.
- Overall MacroBench Score.: The overall MacroBench score is the arithmetic mean of per-task scores across the four tasks.
- B.3 Validation Consistency: Gemini-3-Flash is evaluated as the judge model through a human study measuring judge–human score consistency.
- Sample Construction.: 160 non-GT samples are created by sampling outputs from Nano Banana Pro and Bagel across four tasks and four image-count categories, with five outputs per model–task–category combination.The sampling uses fixed random seed 42, and professional annotators score each sample with the judge’s task-specific rubric.
- Correlation Metrics.: 120 GT samples come from Illustration, Spatial, and Temporal, with 10 items per task and image-count category; GT target images receive annotation score 10, while Customization is excluded.Agreement is reported with Pearson r, Spearman ρ, and Kendall τ under Overall and Human settings.
Results. … C.2 Evaluation Settings
The experiments validate Gemini-3-Flash as the evaluation judge through strong agreement with human judgments. Training and inference settings detail distributed fine-tuning, dynamic batching or resolution, and extended image embeddings for longer reference contexts.
- Results.: 0.821 Pearson correlation: Gemini-3-Flash agrees substantially more with human judgments than GPT-4.1 across Overall and Human settings.Gemini-3-Flash leads GPT-4.1 on all reported correlation metrics, supporting its selection as the judge model.
- C Experiments Details: The experiments section combines evaluation validation with implementation details for training and inference across the tested baselines.These settings cover judge selection, distributed optimization, context-length management, and multi-image embedding expansion.
- C.1 Training Settings: Training details cover dynamic resolution and the remaining model-specific hyperparameters.The appendix supplements the main text’s dynamic resolution strategy with detailed training configurations.
- C.1 Training Settings: BAGEL is fine-tuned with FSDP on 32 NVIDIA H800 GPUs across 4 nodes, using 32,768-token dynamic batching and a 2 × 10−5 learning rate.The ViT encoder remains frozen during training.
- OmniGen2 [48].: OmniGen2 uses DeepSpeed on 32 NVIDIA H800 GPUs across 4 nodes, with global batch size 64 and learning rate 8×10−7.Its baseline supports a maximum of 5 image embeddings, which are extended and normally initialized for more references.
- Qwen-Image-Edit-2511 [47].: Qwen-Image-Edit-2511 fine-tunes its DiT component with DeepSpeed on 128 NVIDIA H800 GPUs across 16 nodes at a 1 × 10−5 learning rate.The configuration uses 16 nodes and adapts the DiT component during fine-tuning.
- C.2 Evaluation Settings: Evaluation uses dynamic resolution for varying reference counts while otherwise following the baselines’ default inference settings.For OmniGen2, image embeddings expand from 5 to 10 with normal initialization, accommodating up to 10 input images.
C.3 Detailed Quantitative Results on MacroBench … Temporal
MacroBench results show that open-source models degrade sharply with increasing reference counts on progressive Customization, while non-progressive tasks remain comparatively stable. Qualitative results and failures further show strong multi-reference capabilities alongside persistent limitations in context retention, spatial reasoning, temporal consistency, and fine-detail perception.
- C.3 Detailed Quantitative Results on MacroBench: Qwen-Image-Edit-2511 drops from 8.27 with 1–3 references to 4.55 with 4–5 references on Customization.This illustrates the pronounced degradation observed in open-source models beyond three reference images.
- C.3 Detailed Quantitative Results on MacroBench: Existing open-source models and datasets struggle to provide robust inter-reference reasoning for long-context multi-reference generation, including on Customization.The passage specifically attributes MICo’s limitation to training data built through simple decomposition and recomposition.
- C.3 Detailed Quantitative Results on MacroBench: Customization metrics descend as input-image counts increase, whereas Illustration, Spatial, and Temporal performance remains comparatively stable for most baselines.The contrast distinguishes progressive Customization from non-progressive tasks across image-count categories.
- MacroBench Results of Different Tasks: Bagel fine-tuned on MacroData composes more than eight references into coherent Customization scenes and demonstrates 3D spatial understanding for novel viewpoints.Its Temporal outputs also capture transformation patterns across video frames and generate plausible future scenes.
- Token Selection: Text-aligned selection achieves competitive results while retaining only 30% of tokens, whereas block-wise selection retains high-quality generation at a 90% retention rate.Text-aligned selection can omit fine-grained details, while block-wise dropping can occasionally produce unnatural images.
- D Failure Cases: With too many inputs, Customization can forget references, omit identities, and mix associated attributes such as clothing and hair.The reported example generates three humans instead of the required four after one woman disappears.
- Customization: Illustration failures include retrieving accessories from earlier references and rendering semantically meaningful text.The model instead generates incorrect accessories, while text appears garbled because it was not specifically trained for text rendering.
- Spatial: Spatial failures reverse the required viewpoint shift, while Temporal failures break clothing consistency, scene structure, and scoreboard preservation in long contexts.These cases expose broader deficiencies in context retention, contextual consistency, 3D spatial reasoning, and fine-detail perception.
E Limitation … Evaluation Prompt for Temporal
MacroData still faces degradation with 6–10 references, while MacroBench remains preliminary; future work targets broader data, finer evaluation, and more efficient long-context modeling. The appendix provides task-specific evaluation prompts for customization, illustration, spatial, and temporal generation.
- E Limitation: Performance degrades with 6–10 input images, showing that highly complex, long-context visual dependencies remain challenging.MacroBench also covers a relatively limited range of predefined tasks and requires a more comprehensive evaluation framework.
- F Social Impact: Long-context multi-reference generation presents dual-use risks including deceptive content and unauthorized identity manipulation.MacroData construction relies on publicly available sources and standard permissible licenses to mitigate ethical and legal concerns.
- G Future Work: Future work will broaden MacroData to more general multi-image scenarios and samples with more reference images.The aim is to raise input capacity and narrow the performance gap with state-of-the-art closed-source models.
- G Future Work: Future work will refine MacroBench into a more granular evaluation framework with detailed scoring methodologies.The supplied passage identifies this refinement as a concurrent research direction.
- G Future Work: Future methods will improve dense multi-image utilization through specialized token representations and attention mechanisms optimized for in-context generation.These methods target both computational efficiency and generation performance, building on preliminary token-selection explorations.
- Evaluation Prompt for Customization: Customization evaluation uses a meticulous digital-art critic and quality-assurance specialist prompt.The appendix includes a corresponding evaluation prompt for the Customization task.
- Evaluation Prompt for Illustration: Illustration evaluation uses a visual-communication specialist and content-quality auditor prompt for diverse visual content.The appendix includes a corresponding evaluation prompt for the Illustration task.
- Evaluation Prompt for Spatial: Spatial evaluation uses a meticulous 3D quality-assurance specialist and digital-art critic prompt.The appendix includes a corresponding evaluation prompt for the Spatial task.