Source-linked AI summary
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Haiwen Diao, Penghao Wu, Hanming Deng, Jiahao Wang, Shihao Bai, Silei Wu, Weichen Fan, Wenjie Ye, Wenwen Tong, Xiangyu Fan, Yan Li, Yubo Wang, Zhijie Cao, Zhiqian Lin, Zhitao Yang, Zhongang Cai, Yuwei Niu, Yue Zhu, Bo Liu, Chengguang Lv, Haojia Yu, Haozhe Xie, Hongli Wang, Jianan Fan, Jiaqi Li, Jiefan Lu, Jingcheng Ni, Junxiang Xu, Kaihuan Liang, Lianqiang Shi, Linjun Dai, Linyan Wang, Oscar Qian, Peng Gao, Pengfei Liu, Qingping Sun, Rui Shen, Ruisi Wang, Shengnan Ma, Shuang Yang, Siyi Xie, Siying Li, Tianbo Zhong, Xiangli Kong, Xuanke Shi, Yang Gao, Yongqiang Yao, Yves Wang, Zhengqi Bai, Zhengyu Lin, Zixin Yin, Wenxiu Sun, Ruihao Gong, Quan Wang, Lewei Lu, Lei Yang, Ziwei Liu, Dahua Lin
TL;DR
Existing vision-language models typically separate understanding and generation, producing divergent representations and loosely integrated systems. SenseNova-U1 uses the NEO-unify architecture to unify both capabilities natively, rivaling top-tier understanding-only VLMs while achieving strong image generation and preliminary VLA and world-model performance.
Problem
Understanding and generation remain isolated in multimodal models, with distinct objectives, pipelines, and representations limiting their integration.
Method
SenseNova-U1 uses a native Mixture-of-Transformers architecture that processes multimodal inputs in one sequence under shared self-attention.
Results
SenseNova-U1 rivals top-tier understanding-only VLMs across diverse tasks while achieving strong X2I generation and preliminary VLA and world-model capabilities.
Takeaways & Limitations
The results support native unification of perception, reasoning, generation, and multimodal interaction within a shared architecture.
Abstract
from arXiv · showhide
Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fragmented architectures, cascaded pipelines, and misaligned representation spaces. We argue that this divide is not merely an engineering artifact, but a structural limitation that hinders the emergence of native multimodal intelligence. Hence, we introduce SenseNova-U1, a native unified multimodal paradigm built upon NEO-unify, in which understanding and generation evolve as synergistic views of a single underlying process. We launch two native unified variants, SenseNova-U1-8B-MoT and SenseNova-U1-A3B-MoT, built on dense (8B) and mixture-of-experts (30B-A3B) understanding baselines, respectively. Designed from first principles, they rival top-tier understanding-only VLMs across text understanding, vision-language perception, knowledge reasoning, agentic decision-making, and spatial intelligence. Meanwhile, they deliver strong semantic consistency and visual fidelity, excelling in conventional or knowledge-intensive any-to-image (X2I) synthesis, complex text-rich infographic generation, and interleaved vision-language generation, with or without think patterns. Beyond performance, we show detailed model design, data preprocessing, pre-/post-training, and inference strategies to support community research. Last but not least, preliminary evidence demonstrates that our models extend beyond perception and generation, performing strongly in vision-language-action (VLA) and world model (WM) scenarios. This points toward a broader roadmap where models do not translate between modalities, but think and act across them in a native manner. Multimodal AI is no longer about connecting separate systems, but about building a unified one and trusting the necessary capabilities to emerge from within.
1 Introduction
SenseNova-U1 addresses the separation of multimodal understanding and generation with a native unified architecture built on NEO-unify. Its dense and mixture-of-experts variants target strong understanding and emerging native action and world-modeling capabilities while supporting efficient multimodal scaling.
- Motivation: Understanding and generation have largely evolved in isolation because they rely on different system components, learning objectives, and training pipelines.Understanding typically uses pretrained vision encoders, whereas generation relies on latent variational autoencoders.
- Motivation: Existing native VLMs either discretize modalities into tokens, risking lossy non-linguistic representations, or follow another distinct architectural direction.The token-based approach enables cross-modal reasoning but can constrain high-level semantics and visual fidelity.
- Architecture: SenseNova-U1 uses NEO-unify to directly process pixels and words without pretrained vision encoders or deep decoder heads.This design aims to provide a unified architecture with concise and scalable training.
- Models and results: Two variants—SenseNova-U1-8B-MoT and SenseNova-U1-A3B-MoT—use dense 8B and mixture-of-experts 30B-A3B backbones with native MoT architectures.MoT enables efficient scaling while reducing interference across heterogeneous multimodal objectives.
- Broader capabilities: Preliminary experiments suggest native vision-language-action and world-modeling capabilities without external adapters or modular bridges.The broader direction is to learn perception, reasoning, and generation within one natively unified architecture.
2 Related Works
Prior multimodal systems evolved from vision-encoder–language-model coupling toward native backbones and unified architectures, but understanding and generation remain constrained by tokenization, modality conflicts, and decoupled pathways. NEO-unify advances continuous native modeling, providing the foundation for SenseNova-U1’s unified pixel–text paradigm.
- Vision–language models: VLMs advance multimodal understanding by coupling visual encoders with large language models through staged pretraining or joint optimization.These designs inherit pretrained semantic biases and introduce additional complexity and capacity trade-offs across components.
- Native multimodal backbones: Native multimodal backbones such as Fuyu and EVE remove visual encoders, while later methods address vision–language conflicts through distillation, data mixing, shared modules, and modality decomposition.NEO extends this direction by exploring native pixel–word modeling.
- Unified understanding and generation: Early unified systems including Show-o, Janus, OmniGen, and BAGEL combine perception and synthesis in shared backbones but retain different tokenizers, diffusion heads, or decoupled pathways.These design choices reflect a persistent mismatch between understanding and generation.
- Discrete and continuous native modeling: Discrete unified models achieve architectural unification through token-level autoregression but sacrifice visual fidelity and expressivity under discrete tokenization.Continuous native approaches instead pursue end-to-end modeling without explicit tokenizers or latent bottlenecks.
- SenseNova-U1 foundation: SenseNova-U1 builds on NEO-unify as a native paradigm operating directly on pixel and text inputs without separate visual or variational autoencoders.Its framework combines a near-lossless visual interface with a native Mixture-of-Transformers main architecture and supports perception, synthesis, and interleaved vision-language generation.
3 Methodology
SenseNova-U1 introduces a native end-to-end multimodal framework that represents understanding and generation within one shared architecture. Its methodology combines pixel-and-word processing, native mixture-of-transformers, resolution-adaptive generation, unified guidance and training, and optimized deployment.
- Unified Architecture: SenseNova-U1 operates directly on pixels and words, eliminating pretrained encoder priors and fixed-representation scaling limitations.The framework is designed from first principles as a native, unified, end-to-end system.
- Tokenization and Decoding: The patch encoder maps images or noise into 32 × 32 visual-patch tokens, while separate heads predict text vocabulary tokens or pixel patches directly.The generation head bypasses diffusion heads and VAE decoders, enabling end-to-end representation learning.
- Generation Dynamics: Resolution-adaptive noise scaling uses σR(H, W) = σ0√(N(H, W)/N0) to preserve approximately constant per-token noise energy across resolutions.The same scale initializes terminal training noise and the inference flow ODE, supporting a consistent SNR distribution.
- Unified Architecture: A native Mixture-of-Transformers backbone processes clean inputs and noise-conditioned inputs as one sequence under shared self-attention.This lets perception and synthesis interact natively at every layer.
- Training and Guidance: Unified classifier-free guidance drops text conditions with probability 10% and both text and image conditions with an additional probability of 10%.The best X2I settings are γ = 4 and γimg = 1, with stronger guidance primarily enforcing textual alignment.
- Inference System: Disaggregated deployment uses LIGHTLLM for understanding and orchestration and LIGHTX2V for image generation, exchanging state through pinned shared memory and optimized transfer kernels.This permits workload-specific parallelization and independent resource allocation while preserving a unified API.
4 Data Construction
SenseNova-U1’s data construction combines staged multimodal pre-training, mid-training, and supervised fine-tuning with balanced sampling, prompt augmentation, and multi-criteria quality control. Separate generation corpora cover text-to-image, editing, and interleaved vision-text capabilities through unified preprocessing and filtering pipelines.
- Pre-training Stage: Pre-training combines web text, image-text pairs, and interleaved multimodal documents, distributed across image-text pairs (32%), captions (17%), infographic understanding (14%), and pure text (37%).The pipeline applies cross-source deduplication, content and safety filtering, image quality filtering, and CLIP-ratio-balanced re-captioning.
- Mid-training Stage: Mid-training draws primarily from internal SenseNova V6.5 datasets spanning General (39.2%), Agent and Spatial (22.3%), Knowledge Reasoning (19.3%), and Pure Text (19.2%).Its corpus is curated through CLIP-based diversity sampling, attribute-profiled stratification, prompt augmentation, and automated checks for correctness, hallucination, and instruction following.
- Supervised Fine-Tuning: The SFT corpus uses capability-atomic supervision, with approximately 15% spatial intelligence, 13% general multimodal understanding, 12% reasoning, and 11% each general NLP and OCR/document analysis.Candidate refinement emphasizes visual fidelity, instruction clarity, response correctness, reasoning quality, safety, and increased sampling of difficult high-quality examples.
- Generation Corpus: The generation corpus spans Nature, Design, People, and Synthetic domains, preserving balanced coverage alongside a long tail of natural, synthetic, and text-rich content.Text-to-image data comprise Nature (∼40.5%), People (∼26.7%), and Design (∼20.7%), enriched with infographics, bilingual rendering, posters, charts, and cityscapes.
- Generation Corpus: Image-editing data cover diverse real-world content and operations, while interleaved data alternate text and images across Video, Lifestyle, Infographics, and Reasoning domains.Both generation capabilities use unified preprocessing and quality-control flows, with interleaved data additionally applying task-specific synthesis and post-processing.
5 Experiments · 5.1 Main Results
SenseNova-U1 performs strongly across multimodal, text, spatial, agentic, generation, editing, and interleaved-generation benchmarks. Its results show competitive understanding and generation, strong text-rich synthesis, and bidirectional interaction between understanding and generation.
- 5.1.1 Image Understanding: SenseNova-U1 consistently outperforms strong multimodal baselines in reasoning, text-rich understanding, and spatial intelligence, including without specialized reinforcement learning for understanding domains.The evaluation covers perception, multimodal reasoning, OCR, visual reasoning, and spatial benchmarks; the encoder-free architecture shows particular advantages in mathematical reasoning.
- 5.1.2 Text Understanding: SenseNova-U1 delivers strong instruction following, academic knowledge, professional reasoning, and agentic interaction, while its A3B variant approaches substantially larger reasoning-oriented models with fewer active parameters.It consistently outperforms the Qwen3.5 series on IFEval and IFBench, narrows gaps on MMLU-Pro and SuperGPQA, and shows reliable long-horizon tool use.
- 5.1.3 Image Generation: SenseNova-U1 remains highly competitive across dense-prompt, multilingual, and text-centric image-generation benchmarks, including against leading specialized systems.The results cover DPG-Bench, OneIG-Bench, TIIF-Bench, and LongText-Bench, with strengths in alignment, text understanding, and fine-grained instruction following.
- 5.1.3 Image Generation: 0.91 overall score is achieved by both SenseNova-U1-A3B-MoT and SenseNova-U1-8B-MoT on GenEval, surpassing Qwen-Image at 0.87, Lumina-DiMOO at 0.88, and BAGEL at 0.82.The models remain competitive across object co-occurrence, counting, color, position, and attribute binding.
- 5.1.3 Image Generation: 0.940 best average word accuracy is attained by SenseNova-U1-8B-MoT on CVTG-2K across settings with 2 to 5 text regions.This demonstrates accurate rendering under dense multi-region conditions.
- 5.1.3 Image Generation: 0.979 on LongText-Bench-EN and 0.962 on LongText-Bench-ZH are achieved by SenseNova-U1-8B-MoT, while the A3B variant reaches 0.950 and 0.955 despite incomplete convergence.These results indicate accurate, readable, and semantically faithful long-form text rendering in English and Chinese.
- 5.1.4 Image Editing: SenseNova-U1 performs competitively on general and reasoning-informed image editing, though specialized systems retain advantages in highly optimized editing workflows.On RISEBench, SenseNova-U1-A3B-MoT-SFT improves from 25.3 without CoT to 30.0 with CoT, reaching the best level among open-source models.
- 5.1.5 Interleaved Generation: SenseNova-U1 demonstrates reasoning within generation and bidirectional synergy, achieving strong results on OpenING, VBVR-Image, Uni-MMMU, and RealUnify.The 8B variant reaches a RealUnify overall average of 52.4, including 55.7 Avg-UEG and 47.5 Avg-GEU; it also attains 35.0 GaU on Uni-MMMU.
5.2 Ablation Studies
The ablations support NEO-unify’s encoder-free interface, MoT co-training, and data-scaling efficiency. Results show preservation of semantic and pixel-level information, effective capability co-evolution, and steadily improving generation quality and understanding–generation synergy with more data.
- Image Reconstruction: 31.56 PSNR and 0.85 SSIM show that NEO-unify (2B) retains high-level semantics and fine-grained visual details without pretrained vision encoders or latent autoencoders.The result follows only 90K pretraining steps on MS-COCO 2017 and approaches FLUX.1-dev VAE’s 31.56 PSNR and 0.93 SSIM.
- Image Editing: 3.32 ImgEdit demonstrates strong editing capability with a frozen understanding branch after 60K-step mixed training using public text-to-image and editing datasets.All conditional contexts route through the understanding branch, while the generation branch directly synthesizes target images.
- Understanding–Generation Co-training: Jointly optimizing pretrained dual branches keeps understanding stable and generation rapidly convergent, indicating effective co-evolution within the MoT backbone with minimal intrinsic conflict.This holds even with low data ratios and small understanding loss weights.
- Data Scaling: Data-scaling curves show steadily improving generation quality and understanding–generation synergy as training data scale increases.The training pipeline combines web-scale pretraining, mid-training, and supervised fine-tuning over diverse data spanning understanding and generation tasks.
5.3 Visualization Results
Qualitative visualizations show SenseNova-U1 operating across complex multimodal scenarios, including embodied action reasoning and world-model prediction. The examples indicate coherent temporal understanding, object-state and trajectory reasoning, and plausible visual state transitions from action instructions.
- Vision-Language-Action: SenseNova-U1 captures action-relevant visual dynamics across time in video-based robotic manipulation examples.Four uniformly sampled frames illustrate each manipulation process and its temporal progression.
- Vision-Language-Action: The model maintains coherent visual understanding while reasoning about object states and manipulation trajectories in embodied settings.
- World Modeling: SenseNova-U1 translates structured action instructions into plausible visual state transitions in world-modeling examples.Given an input image and action-oriented instruction, it predicts the corresponding visual outcome using the original prompts during inference.
6 Conclusion
The paper presents a native unified multimodal foundation model where understanding, generation, and reasoning emerge within one architecture. Across diverse tasks, it combines perception, semantic reasoning, high-fidelity generation, and interleaved multimodal interaction.
- Unified architecture: Understanding, generation, and reasoning emerge within a single native architecture rather than through coordination between separate systems.This design frames multimodal capabilities as arising from one shared model.
- Core capabilities: The model demonstrates strong vision-language perception, semantic reasoning, and high-fidelity generation across a broad range of tasks.These capabilities span both analytical and creative multimodal functions.
- Shared representation: Interleaved multimodal interaction further indicates that a shared representation can support analytical and creative intelligence simultaneously.The conclusion links unified representations with combined understanding and generation capabilities.
7 Contributors
The contributors are organized by role, with alphabetical ordering by first name within each category. Roles span sponsorship, project leadership, core contribution, contribution, and acknowledgement.
- The contributor list is organized by contribution role, with individuals alphabetized by first name within each category.
- Dahua Lin is credited as Project Sponsor and Advisor, while Haiwen Diao is Project Lead.
- Senior Project Leads are Lei Yang, Lewei Lu, Quan Wang, Ruihao Gong, Wenxiu Sun, and Ziwei Liu.
- The Core Contributor and Contributor categories include the listed project contributors, while an Acknowledgement category thanks additional supporters for their valuable support and contributions.