Source-linked AI summary
Phi-4-reasoning-vision-15B Technical Report
Jyoti Aneja, Michael Harrison, Neel Joshi, Tyler LaBonte, John Langford, Eduardo Salinas
TL;DR
The paper addresses the cost and latency of increasingly large vision-language models while seeking strong multimodal reasoning in a compact open-weight system. It combines a mid-fusion architecture, high-resolution visual processing, staged training, and curated multimodal data, reporting competitive accuracy with substantially lower compute and token use. The resulting model supports broad vision-language tasks alongside mathematical, scientific, and computer-use reasoning, while its learned mode switching remains imperfect.
Problem
Growing vision-language models increase training and inference cost and latency, limiting usability in resource-constrained or interactive settings.
Method
The paper develops a compact model using mid-fusion architecture, dynamic-resolution vision encoding, staged training, and mixed mathematics, computer-use, and multimodal instruction data.
Results
The model achieves competitive accuracy with much slower models requiring ten times or more compute time and tokens, while outperforming similarly fast models particularly in math and science reasoning.
Takeaways & Limitations
A single compact multimodal model can support broad vision-language tasks together with scientific, mathematical, and user-interface reasoning.
Takeaways & Limitations
The learned reasoning-mode boundary can be imprecise, and the optimal reasoning-to-non-reasoning data balance remains open across domains and deployment contexts.
Abstract
from arXiv · showhide
We present Phi-4-reasoning-vision-15B, a compact open-weight multimodal reasoning model, and share the motivations, design choices, experiments, and learnings that informed its development. Our goal is to contribute practical insight to the research community on building smaller, efficient multimodal reasoning models and to share the result of these learnings as an open-weight model that is good at common vision and language tasks and excels at scientific and mathematical reasoning and understanding user interfaces. Our contributions include demonstrating that careful architecture choices and rigorous data curation enable smaller, open-weight multimodal models to achieve competitive performance with significantly less training and inference-time compute and tokens. The most substantial improvements come from systematic filtering, error correction, and synthetic augmentation -- reinforcing that data quality remains the primary lever for model performance. Systematic ablations show that high-resolution, dynamic-resolution encoders yield consistent improvements, as accurate perception is a prerequisite for high-quality reasoning. Finally, a hybrid mix of reasoning and non-reasoning data with explicit mode tokens allows a single model to deliver fast direct answers for simpler tasks and chain-of-thought reasoning for complex problems.
1 Introduction
Phi-4-reasoning-vision-15B is presented as a compact open-weight model balancing multimodal reasoning capability with efficiency and modest training-data needs. It targets broad vision-language use while emphasizing mathematical, scientific, and user-interface reasoning.
- The model balances reasoning power, efficiency, and training data needs as a compact open-weight multimodal system.
- Its strongest stated capabilities include mathematics and science reasoning and understanding user interfaces.
- The report contributes development insights and an open-weight model intended to remain competitive with similarly sized models across general vision-language tasks.
- It is designed for broad vision-language tasks, including everyday activities such as interpreting receipts and garment care instructions.
- The model aims to reduce training and inference costs without relying on extremely large datasets, architectures, or excessive token generation.The authors report using 200 billion multimodal-data tokens, compared with more than 1 trillion for several recent multimodal models.
2 Architecture and Training
Phi-4-reasoning-vision-15B uses mid-fusion to balance multimodal expressivity with manageable compute, combining SigLIP-2 visual features with Phi-4-Reasoning. Its training uses high-resolution vision processing and a three-stage curriculum spanning alignment, instruction tuning, and long-context, multi-image, and safety data.
- Architecture: Mid-fusion is chosen over early fusion because it offers a practical trade-off between multimodal expressivity and compute, memory, and data requirements.Early fusion permits unrestricted cross-attention throughout the network but incurs heavier resource demands.
- Architecture: Mid-fusion projects SigLIP-2 image features into the language embedding space, interleaves visual soft tokens with text, and feeds them into Phi-4-Reasoning.This preserves pretrained unimodal components while enabling cross-modal reasoning.
- Vision Encoder and Image Processing: Dynamic-resolution encoders with many visual tokens perform uniformly well and best on high-resolution datasets.Using 3600 rather than 2048 maximum tokens substantially boosts high-resolution benchmarks, particularly ScreenSpot-Pro.
- Vision Encoder and Image Processing: Multi-crop with S2 outperforms standard multi-crop despite using fewer visual tokens, while dynamic resolution produces the most tokens on average.S2-based methods are constrained by original image resolution and often use about half their maximum token budget.
- Training Recipe: Training proceeds in three stages: MLP-only alignment, joint instruction tuning, and specialized long-context, multi-image, and safety training.Stage 1 freezes the vision encoder and language model; Stage 2 trains all components on diverse single-image tasks; Stage 3 adds specialized data.
- Training Recipe: Stage 2 mixes reasoning traces marked with <think> tokens and direct responses marked with <nothink> tokens.The bulk training stage covers visual question answering, mathematical and scientific reasoning, grounding, captioning, OCR, and computer use.
3 Training Data
The training data pipeline emphasizes systematic quality improvement, targeted augmentation, and experiments on data composition across mathematics, science, and computer-use tasks. Results suggest that targeted data and broader scale can improve multiple reasoning domains, while more extreme data ratios remain unresolved.
- Data sources: The final data mix combines filtered and improved open-source vision-language data with domain-specific and acquired datasets.The authors describe the open-source component as the overwhelming majority of the mix.
- Data quality: Quality control classifies samples by answer correctness, question quality, image quality, and formatting, then regenerates or verifies flawed answers and captions.Datasets with high percentages of wrong answers were excluded.
- Data augmentation: Synthetic augmentation expands domain coverage through image descriptions, reformatted instruction data, multi-image matching, and sequential “what’s changed?” examples.These transformations reuse images and records to support mathematical understanding, image attention, and real-time navigation.
- Data augmentation: Human prompt diversity is used to improve robustness beyond overly structured VQA prompts.
- Data proportions: Increasing math data by 3× while keeping computer-use data constant improves both math and computer-use benchmarks.The reported relationship appears in the Table 4 comparison.
- Data proportions: At the tested scale, a single model can show uniformly superior performance across multiple reasoning domains, but larger-scale and more extreme-ratio effects remain open questions.The experiments used a maximum mathematics-to-total-data ratio of 7.5%; ratios of 1% or less remain underexplored.
4 Mixed Non-Reasoning and Reasoning
The model mixes reasoning and non-reasoning multimodal data so it can choose between direct responses and longer reasoning paths. This design targets the latency costs of reasoning while preserving its benefits for complex mathematical and scientific tasks.
- Motivation: Reasoning adds compute and latency, while perception tasks may not benefit from it and mathematical or scientific tasks do.
- Training approach: Alternative training strategies trade off multimodal training burden, reasoning-data requirements, catastrophic forgetting, and reasoning-trace coverage.
- Training approach: The approach uses a reasoning-capable language model with mixed non-reasoning and reasoning multimodal training, allowing the model to learn when to reason.
- Model behavior: The model defaults to direct inference for perception-focused tasks and invokes longer reasoning paths for mathematics and science.Explicit <think> and <nothink> tokens support reasoning and direct-response modes.
- Model behavior: The model can interpret image sequences, including changes in Saturn’s rings across multiple frames.
- Limitations: The 20/80 reasoning-to-non-reasoning data split may not suit every domain or deployment context, and switching boundaries can remain imprecise.Users can override default behavior with explicit <think> or <nothink> prompts.
- Limitations: The mixed approach is presented as one point in the design space rather than a definitive solution for balancing latency, accuracy, and flexibility.
5 Applications
Phi-4-reasoning-vision-15B supports broad vision-language applications, with particular strengths in visual mathematical and scientific reasoning and graphical-user-interface interaction.
- The model handles general vision-language tasks including image description, visual question answering, sequence interpretation, object recognition, landmark recognition, and text transcription.
- Mathematical and scientific reasoning: It solves visually presented mathematical and scientific problems, including handwritten and diagram-based questions, while reasoning over quantitative information in documents and charts.
- Computer-use applications: The model supports GUI agents by interpreting screens, localizing interactive elements, and determining appropriate interactions across desktop, web, and mobile environments.Its high-resolution perception and fine-grained grounding support dense interfaces, while low inference-time needs suit latency-sensitive environments.
6 Evaluation
The evaluation measures accuracy and timing across vision-language benchmarks and finds that mixed reasoning generally outperforms forced reasoning modes while offering a favorable accuracy–cost trade-off.
- Accuracy was evaluated on benchmarks spanning visual question answering, chart understanding, hallucination, multimodal mathematics, OCR, and screen interaction.The study used Eureka ML Insights and VLMEvalKit for standardized analysis.
- Accuracy: The model’s default mixed-reasoning behavior has better average accuracy than forcing thinking or non-thinking modes.Forced thinking improves only MathVerse and MMMUVAL, while forced non-thinking improves ScreenSpotv2.
- Timing: Timing experiments sampled 100 examples from each of ChartQATEST, MathVistaMINI, MMMUVAL, and ScreenSpot using single-threaded batch-one H100 runs.Wall-clock latency and output token counts were measured for every model.
- Accuracy–cost trade-off: Compared with recent popular open-weight models, Phi-4-reasoning-vision-15B provides a favorable trade-off between accuracy and inference-time cost.Cost is assessed through inference-time compute and output tokens.
- Evaluation protocol: The reported benchmark numbers come from the authors’ own runs rather than quoted leaderboard results.The authors aimed for fair evaluations using recommended platforms, settings, and prompts.
7 Safety
Safety was treated as a core development consideration through safety-focused training data and quantitative and qualitative evaluation of harmful, misleading, and jailbreak-related behaviors.
- Safety training: The model was trained with public safety datasets and internally generated examples intended to elicit appropriate refusals.These signals target requests outside intended or acceptable use in alignment with Microsoft’s Responsible AI Principles.
- Safety training: Phase 3 incorporated responsible-AI data covering hateful-content detection, harmful-request refusal, and safe reasoning under adversarial prompts.Examples included Hateful Memes, VLGuard, Think-in-Safety, and WildGuard.
- Safety evaluation: Safety evaluation combined automated red teaming with qualitative assessment across disallowed content, copyright, intellectual property, jailbreak susceptibility, groundedness, and fabricated information.
8 Limitations
The model is competitive for its size and compute budget but remains bounded by larger proprietary models, imperfect mode switching, and difficulty with extremely detailed or nuanced images.
- Larger proprietary models outperform Phi-4-reasoning-vision-15B on broad, unconstrained vision-language and generalist multimodal tasks.The model is competitive with similarly sized open-weight models and achieves state-of-the-art accuracy relative to training and inference-time compute and tokens.
- The learned switch between reasoning and non-reasoning modes is not always optimal.Users can override the default behavior with <think> or <nothink> prompts.
- The model has limitations with extremely detailed or nuanced image understanding, so critical outputs involving fine-grained visual details should be verified.
9 Open Release and Community Engagement
The model is openly released through Microsoft Foundry and HuggingFace, with supporting examples, code, logs, and safety guidance. The release also includes permissively licensed artifacts intended to support community study and future development.
- Phi-4-reasoning-vision-15B is available on Microsoft Foundry and HuggingFace, with additional examples and details on GitHub.
- Additional guidance on using the model properly and safely is provided through its Model Card.
- The model is released under a permissive license with model weights, fine-tuning code, and benchmark logs.
- A portion of the training data is planned for release in the coming months.
- The release is intended to provide concrete artifacts that help close gaps in understanding how compact multimodal reasoning models can be built and studied.
10 Looking Forward
The authors position smaller vision–language models with selective, task-aware reasoning as a promising direction for practical and accessible multimodal systems. They share the model and its learnings to inform research on multimodal modeling, computer-using agents, and mathematical and scientific reasoning.
- Smaller vision–language models with selective, task-aware reasoning are presented as one promising direction for making multimodal systems more practical and accessible.
- The model and its learnings are intended to inform ongoing research in multimodal modeling, computer-using agents, and mathematical scientific reasoning.
- The authors invite critical evaluation, replication, and extension by the community.
A Open-Source Training Data
Table 7 identifies the open-source training data sources used across training stages 1–3.
- Table 7 lists open-source training data sources for stages 1–3.