Source-linked AI summary
Show-o2: Improved Native Unified Multimodal Models
Jinheng Xie, Zhenheng Yang, Mike Zheng Shou
TL;DR
Existing multimodal systems motivate unified models that can understand and generate across text, images, and videos. Show-o2 combines 3D causal VAE representations with native autoregressive and flow-based heads and a two-stage training recipe, achieving state-of-the-art performance across diverse benchmarks. Its reported limitations include weak rendered-text generation and missing small-object details at limited image resolution.
Problem
The paper addresses the challenge of natively unifying multimodal understanding and generation across text, images, and videos.
Method
Show-o2 uses a 3D causal VAE with dual-path spatial-temporal fusion, autoregressive language modeling, flow matching for visual generation, and two-stage training.
Results
Show-o2 demonstrates state-of-the-art performance across multimodal understanding and visual generation benchmarks, surpassing existing methods across most metrics.
Takeaways & Limitations
The resulting model handles diverse multimodal understanding and image/video-generation tasks within a versatile unified model.
Takeaways & Limitations
The model is weak at rendering text in images and lacks small-object details because rendered-text data are limited and image resolution is constrained.
Abstract
from arXiv · showhide
This paper presents improved native unified multimodal models, \emph{i.e.,} Show-o2, that leverage autoregressive modeling and flow matching. Built upon a 3D causal variational autoencoder space, unified visual representations are constructed through a dual-path of spatial (-temporal) fusion, enabling scalability across image and video modalities while ensuring effective multimodal understanding and generation. Based on a language model, autoregressive modeling and flow matching are natively applied to the language head and flow head, respectively, to facilitate text token prediction and image/video generation. A two-stage training recipe is designed to effectively learn and scale to larger models. The resulting Show-o2 models demonstrate versatility in handling a wide range of multimodal understanding and generation tasks across diverse modalities, including text, images, and videos. Code and models are released at https://github.com/showlab/Show-o.
1 Introduction
Show-o2 is designed as a native unified multimodal model for understanding and generating text, images, and videos. It combines unified visual representations, autoregressive modeling, flow matching, and staged training to support this scope.
- 1 Introduction: Table 1 compares unified multimodal models by visual-representation design and unified modeling, distinguishing direct native decoding from external-decoder conditioning.
- 1 Introduction: Autoregressive modeling predicts text tokens through the language head, while flow matching generates images and videos through the flow head.
- 1 Introduction: Show-o2 surpasses existing methods across most metrics on multimodal understanding and visual generation benchmarks.
- 1 Introduction: Show-o2 targets native multimodal understanding and generation across interleaved text, images, and videos.
- 1 Introduction: A 3D causal VAE supports unified visual representations for images and videos by fusing semantic and low-level features through spatial-temporal processing.
- 1 Introduction: A two-stage training pipeline learns visual generation while retaining language knowledge and scaling to larger models without requiring a massive text corpus.
2 Related Work
Related work develops multimodal systems by combining language understanding with visual processing and generation. Unified multimodal models either integrate these objectives natively or assemble specialized components through adapters or learnable tokens.
- 2 Related Work: Large multimodal models commonly project visual-encoder features into an LLM embedding space, while encoder-free models align raw visual features directly.
- 2 Related Work: Diffusion-based generation uses text encoders and denoising networks, whereas autoregressive generation uses LLM-style architectures trained by next-token prediction.
- 2 Related Work: Native unified multimodal models combine multimodal understanding and generation objectives within one architecture through autoregressive modeling, diffusion modeling, or both.
- 2 Related Work: Assembled unified frameworks tune adapters or learnable tokens to connect specialized multimodal and visual generative models.
3 Methodology
Show-o2 builds unified visual representations for images and videos, then combines them with text in a language-model sequence using separate language and flow heads. A two-stage recipe trains these components for multimodal understanding and generation.
- Sequence Modeling: Text embeddings and unified visual representations are arranged in an interleaved sequence for a pretrained language model.The sequence can represent combinations of text, images, and videos using modality-specific boundary tokens.
- Unified Visual Representation: The framework uses a 3D causal VAE and dual-path spatial (-temporal) fusion to support unified image and video representations.Semantic layers extract high-level information, while a projector retains low-level visual information before fusion.
- Dedicated Heads: The language head uses autoregressive causal attention for text prediction, while the flow head uses flow matching for image and video generation.The flow head predicts visual-latent velocity, and omni-attention remains causal across the sequence while allowing full attention within unified visual representations.
- Two-Stage Training: Stage 1 trains the projector, spatial (-temporal) fusion, and flow head on image-text data while progressively adding interleaved and video-text pairs.The stage uses around 66M image-text pairs and trains with autoregressive modeling and flow matching.
- Two-Stage Training: Stage 2 tunes the full model with high-quality multimodal understanding, visual generation, and video understanding data.The recipe uses 9M understanding instruction examples, 16M visual generation examples, and 1.6M video understanding examples.
- Scaling Up: Scaling to the 7B model resumes a pretrained flow head and adds a lightweight MLP transformation to align hidden sizes.This approach is intended to enable quick adaptation and convergence of the larger model.
4 Experiments
Show-o2 is evaluated across multimodal understanding, image and video generation, and ablation studies. The results show strong benchmark performance, benefits from spatial (-temporal) fusion and stage-2 training, and remaining weaknesses in chart understanding.
- Multimodal understanding: Show-o2 consistently outperforms state-of-the-art models across many multimodal understanding metrics, including comparisons with larger models.The 1.5B variant leads on MME-p and MMU-val, while the 7B variant surpasses Janus-Pro and TokenFlow-XL on several metrics.
- Multimodal understanding: The model supports detailed image description, object counting, image-text recognition, bilingual English–Chinese question answering, and video understanding.Video understanding is evaluated after fine-tuning on 1.6M video samples together with 1.1M image-level samples.
- Image generation: Show-o2 surpasses most compared methods on GenEval and achieves the best overall DPG-Bench score while using 66M image-text pairs.The comparisons include TokenFlow-XL, Show-o, Emu3, Transfusion, Janus-Pro, SD3-Medium, and other generation models.
- Video generation: With 2B parameters, Show-o2 outperforms several video-generation models with more than 6B parameters and remains competitive with CogVideoX and Step-Video-T2V.The comparison covers both text-to-video and image-to-video generation.
- Ablation studies: Spatial (-temporal) fusion improves both multimodal understanding and generation metrics by combining semantic and low-level visual features.The pilot study reports improvements on MME-p, GQA, and FID-5K under the same training settings.
- Ablation studies: Stage-2 training significantly improves GenEval and DPG-Bench, while the two-stage recipe preserves language performance better than direct one-stage co-training.The models remain comparable to the original Qwen2.5-1.5B and Qwen2.5-7B Instruct models; chart understanding remains weaker than the LLaVA-OV-7B baseline.
5 Limitations and Broader Impacts
The model has difficulty rendering text and small objects in generated images, reflecting limited text-rich data and image resolution. Its text-and-image generation also carries misuse and intellectual-property risks.
- Limited text-rendering data contributes to poor text rendering in generated images.The authors report that relatively few training images contain rendered text and describe higher-resolution, text-rich data as a mitigation.
- Limited image resolution can leave small objects lacking detail.
- Generating text and images may enable misuse, including fake information or profiles.
- Large-scale training data containing celebrity and copyrighted content could result in intellectual-property infringement.
6 Conclusion
Show-o2 unifies multimodal understanding and generation across text, images, and videos by combining a 3D causal VAE, autoregressive modeling, and flow matching. A dual-path fusion mechanism and two-stage training support unified visual representations and versatile benchmark performance.
- Show-o2 targets scalable multimodal understanding and generation across image and video modalities.
- The model integrates a 3D causal VAE, autoregressive modeling, and flow matching.
- A dual-path spatial (-temporal) fusion mechanism constructs unified visual representations with high- and low-level features.
- Extensive experiments report state-of-the-art performance across various benchmarks.
A.1 More Qualitative Results
Figure 3 presents qualitative examples of text-to-video and image-to-video generation.
- The figure shows examples of text-to-video generation.
- The figure shows examples of image-to-video generation.
- These examples provide qualitative evidence for video-generation capabilities across two input conditions.
A.2 Text Prompts
The appendix provides the text prompts used for the image-generation examples in Figure 2. The prompts span portraits, wildlife, neon typography, and highly detailed photorealistic scenes.
- The appendix identifies these texts as the image-generation prompts used for Figure 2.
- The prompts include a highly detailed portrait of a mature man with graying hair and blue eyes.
- Another prompt describes a young woman in a winter portrait on a snowy Moscow street.
- A wildlife prompt specifies a colorful, photorealistic parrot in a tropical rainforest.
- A neon-scene prompt requests a dark room with a glowing sign spelling “SHOW O2.”