Source-linked AI summary
Apple Intelligence Foundation Language Models: Tech Report 2025
Ethan Li, Anders Boesen Lindbo Larsen, Chen Zhang, Xiyou Zhou, Jun Qin, Dian Ang Yap, Narendran Raghavan, Xuankai Chang, Margit Bowler, Eray Yildiz, John Peebles, Hannah Gillis Coleman, Matteo Ronchi, Peter Gray, Keen You, Anthony Spalvieri-Kruse, Ruoming Pang, Reed Li, Yuli Yang, Emad Soroush, Zhiyun Lu, Crystal Xiao, Rong Situ, Jordan Huffaker, David Griffiths, Zaid Ahmed, Peng Zhang, Daniel Parilla, Asaf Liberman, Jennifer Mallalieu, Parsa Mazaheri, Qibin Chen, Manjot Bilkhu, Aonan Zhang, Eric Wang, Dave Nelson, Michael FitzMaurice, Thomas Voice, Jeremy Liu, Josh Shaffer, Shiwen Zhao, Prasanth Yadla, Farzin Rasteh, Pengsheng Guo, Arsalan Farooq, Jeremy Snow, Stephen Murphy, Tao Lei, Minsik Cho, George Horrell, Sam Dodge, Lindsay Hislop, Sumeet Singh, Alex Dombrowski, Aiswarya Raghavan, Sasha Sirovica, Mandana Saebi, Faye Lao, Max Lam, TJ Lu, Zhaoyang Xu, Karanjeet Singh, Marc Kirchner, David Mizrahi, Rajat Arora, Haotian Zhang, Henry Mason, Lawrence Zhou, Yi Hua, Ankur Jain, Felix Bai, Joseph Astrauskas, Floris Weers, Josh Gardner, Mira Chiang, Yi Zhang, Pulkit Agrawal, Tony Sun, Quentin Keunebroek, Matthew Hopkins, Bugu Wu, Tao Jia, Chen Chen, Xingyu Zhou, Nanzhu Wang, Peng Liu, Ruixuan Hou, Rene Rauch, Yuan Gao, Afshin Dehghan, Jonathan Janke, Zirui Wang, Cha Chen, Xiaoyi Ren, Feng Nan, Josh Elman, Dong Yin, Yusuf Goren, Jeff Lai, Yiran Fei, Syd Evans, Muyang Yu, Guoli Yin, Yi Qin, Erin Feldman, Isha Garg, Aparna Rajamani, Karla Vega, Walker Cheng, TJ Collins, Hans Han, Raul Rea Menacho, Simon Yeung, Sophy Lee, Phani Mutyala, Ying-Chang Cheng, Zhe Gan, Sprite Chu, Justin Lazarow, Alessandro Pappalardo, Federico Scozzafava, Jing Lu, Erik Daxberger, Laurent Duchesne, Jen Liu, David Güera, Stefano Ligas, Mary Beth Kery, Brent Ramerth, Ciro Sannino, Marcin Eichner, Haoshuo Huang, Rui Qian, Moritz Schwarzer-Becker, David Riazati, Mingfei Gao, Bailin Wang, Jack Cackler, Yang Lu, Ransen Niu, John Dennison, Guillaume Klein, Jeffrey Bigham, Deepak Gopinath, Navid Shiee, Darren Botten, Guillaume Tartavel, Alex Guillen Garcia, Sam Xu, Victoria MönchJuan Haladjian, Zi-Yi Dou, Matthias Paulik, Adolfo Lopez Mendez, Zhen Li, Hong-You Chen, Chao Jia, Dhaval Doshi, Zhengdong Zhang, Raunak Manjani, Aaron Franklin, Zhile Ren, David Chen, Artsiom Peshko, Nandhitha Raghuram, Hans Hao, Jiulong Shan, Kavya Nerella, Ramsey Tantawi, Vivek Kumar, Saiwen Wang, Brycen Wershing, Bhuwan Dhingra, Dhruti Shah, Ob Adaranijo, Xin Zheng, Tait Madsen, Hadas Kotek, Chang Liu, Yin Xia, Hanli Li, Suma Jayaram, Yanchao Sun, Ahmed Fakhry, Vasileios Saveris, Dustin Withers, Yanghao Li, Alp Aygar, Andres Romero Mier Y Teran, Kaiwei Huang, Mark Lee, Xiujun Li, Yuhong Li, Tyler Johnson, Jay Tang, Joseph Yitan Cheng, Futang Peng, Andrew Walkingshaw, Lucas Guibert, Abhishek Sharma, Cheng Shen, Piotr Maj, Yasutaka Tanaka, You-Cyuan Jhang, Vivian Ma, Tommi Vehvilainen, Kelvin Zou, Jeff Nichols, Matthew Lei, David Qiu, Yihao Qian, Gokul Santhanam, Wentao Wu, Yena Han, Dominik Moritz, Haijing Fu, Mingze Xu, Vivek Rathod, Jian Liu, Louis D'hauwe, Qin Ba, Haitian Sun, Haoran Yan, Philipp Dufter, Anh Nguyen, Yihao Feng, Emma Wang, Keyu He, Rahul Nair, Sanskruti Shah, Jiarui Lu, Patrick Sonnenberg, Jeremy Warner, Yuanzhi Li, Bowen Pan, Ziyi Zhong, Joe Zhou, Sam Davarnia, Olli Saarikivi, Irina Belousova, Rachel Burger, Shang-Chen Wu, Di Feng, Bas Straathof, James Chou, Yuanyang Zhang, Marco Zuliani, Eduardo Jimenez, Abhishek Sundararajan, Xianzhi Du, Chang Lan, Nilesh Shahdadpuri, Peter Grasch, Sergiu Sima, Josh Newnham, Varsha Paidi, Jianyu Wang, Kaelen Haag, Alex Braunstein, Daniele Molinari, Richard Wei, Brenda Yang, Nicholas Lusskin, Joanna Arreaza-Taylor, Meng Cao, Nicholas Seidl, Simon Wang, Jiaming Hu, Yiping Ma, Mengyu Li, Kieran Liu, Hang Su, Sachin Ravi, Chong Wang, Xin Wang, Kevin Smith, Haoxuan You, Binazir Karimzadeh, Rui Li, Jinhao Lei, Wei Fang, Alec Doane, Sam Wiseman, Ismael Fernandez, Jane Li, Andrew Hansen, Javier Movellan, Christopher Neubauer, Hanzhi Zhou, Chris Chaney, Nazir Kamaldin, Valentin Wolf, Fernando Bermúdez-Medina, Joris Pelemans, Peter Fu, Howard Xing, Xiang Kong, Wayne Shan, Gabriel Jacoby-Cooper, Dongcai Shen, Tom Gunter, Guillaume Seguin, Fangping Shi, Shiyu Li, Yang Xu, Areeba Kamal, Dan Masi, Saptarshi Guha, Qi Zhu, Jenna Thibodeau, Changyuan Zhang, Rebecca Callahan, Charles Maalouf, Wilson Tsao, Boyue Li, Qingqing Cao, Naomy Sabo, Cheng Leong, Yi Wang, Anupama Mann Anupama, Colorado Reed, Kenneth Jung, Zhifeng Chen, Mohana Prasad Sathya Moorthy, Yifei He, Erik Hornberger, Devi Krishna, Senyu Tong, Michael, Lee, David Haldimann, Yang Zhao, Bowen Zhang, Chang Gao, Chris Bartels, Sushma Rao, Nathalie Tran, Simon Lehnerer, Co Giang, Patrick Dong, Junting Pan, Biyao Wang, Dongxu Li, Mehrdad Farajtabar, Dongseong Hwang, Grace Duanmu, Eshan Verma, Sujeeth Reddy, Qi Shan, Hongbin Gao, Nan Du, Pragnya Sridhar, Forrest Huang, Yingbo Wang, Nikhil Bhendawade, Diane Zhu, Sai Aitharaju, Fred Hohman, Lauren Gardiner, Chung-Cheng Chiu, Yinfei Yang, Alper Kokmen, Frank Chu, Ke Ye, Kaan Elgin, Oron Levy, John Park, Donald Zhang, Eldon Schoop, Nina Wenzel, Michael Booker, Hyunjik Kim, Chinguun Erdenebileg, Nan Dun, Eric Liang Yang, Priyal Chhatrapati, Vishaal Mahtani, Haiming Gang, Kohen Chia, Deepa Seshadri, Donghan Yu, Yan Meng, Kelsey Peterson, Zhen Yang, Yongqiang Wang, Carina Peng, Doug Kang, Anuva Agarwal, Albert Antony, Juan Lao Tebar, Albin Madappally Jose, Regan Poston, Andy De Wang, Gerard Casamayor, Elmira Amirloo, Violet Yao, Wojciech Kryscinski, Kun Duan, Lezhi L
TL;DR
Apple addresses the need for capable, efficient foundation models across on-device and Private Cloud Compute settings. It develops complementary multilingual and multimodal models with specialized architectures and training pipelines, reports favorable benchmark comparisons, and exposes the on-device model through a developer framework.
Problem
Apple needs foundation models that support intelligent features across devices and services while meeting differing deployment and performance requirements.
Method
Apple develops an approximately 3B-parameter on-device model and a server model, refining them through training, reinforcement learning, and a Foundation Models framework for developer access.
Results
Apple’s on-device model performs favorably against larger InternVL and Qwen models and competitively against Gemma, while its server model outperforms Qwen-2.5-VL at less than half the inference FLOPS.
Takeaways & Limitations
The Foundation Models framework gives developers access to the on-device model for production-quality features such as summarization and entity extraction.
Takeaways & Limitations
Human graders’ preferences differ for about 20-30% of preference data, especially for subjective, difficult, or obscure prompts.
Abstract
from arXiv · showhide
We introduce two multilingual, multimodal foundation language models that power Apple Intelligence features across Apple devices and services: i a 3B-parameter on-device model optimized for Apple silicon through architectural innovations such as KV-cache sharing and 2-bit quantization-aware training; and ii a scalable server model built on a novel Parallel-Track Mixture-of-Experts PT-MoE transformer that combines track parallelism, mixture-of-experts sparse computation, and interleaved global-local attention to deliver high quality with competitive cost on Apple's Private Cloud Compute platform. Both models are trained on large-scale multilingual and multimodal datasets sourced via responsible web crawling, licensed corpora, and high-quality synthetic data, then further refined with supervised fine-tuning and reinforcement learning on a new asynchronous platform. The resulting models support several additional languages while understanding images and executing tool calls. In public benchmarks and human evaluations, both the server model and the on-device model match or surpass comparably sized open baselines. A new Swift-centric Foundation Models framework exposes guided generation, constrained tool calling, and LoRA adapter fine-tuning, allowing developers to integrate these capabilities with a few lines of code. The latest advancements in Apple Intelligence models are grounded in our Responsible AI approach with safeguards like content filtering and locale-specific evaluation, as well as our commitment to protecting our users' privacy with innovations like Private Cloud Compute.
1 Introduction
Apple developed new foundation models to power Apple Intelligence across its platforms, with a focus on broader capabilities, efficiency, and developer access.
- The models power intelligent features integrated across Apple platforms and experiences.
- The model family includes an approximately 3B-parameter on-device model optimized for Apple silicon and a scalable mixture-of-experts server model for Private Cloud Compute.
- The models improve tool use and reasoning, understand image and text inputs, and support 16 languages.
- The overview covers training data, architectures, training recipes, inference optimization, and evaluations against similar models.
2 Model Architectures
The models use complementary architectures: an efficient on-device design for low-latency inference and a scalable server design combining parallel tracks, MoE layers, and interleaved attention.
- The on-device model targets low-latency inference with minimal resource usage, while the server model targets high accuracy and scalability for complex tasks.
- On-Device Model: KV-cache sharing reduces on-device KV-cache memory usage by 37.5% and time-to-first-token by ∼37.5%.
- Server Model: With D = 4, track parallelism reduces synchronization overhead from 2L to L/D, a reduction of 87.5%.
- Server Model: The Parallel Track Transformer partitions layers into independently processed tracks, synchronizing at track-block boundaries to reduce overhead and improve latency without compromising model quality.
- Server Model: PT-MoE places local mixture-of-experts layers within track blocks so communication can overlap more effectively with computation.
- Server Model: Interleaved local and global attention supports long sequences while substantially reducing KV-cache size for long-context inference.
3 Data
Apple trains its models on diverse licensed, public, web-crawled, and synthetic data, with expanded multilingual and multimodal coverage supported by filtering and extraction pipelines.
- Training data comes from licensed publishers, public or open-source datasets, and publicly available information crawled by Applebot.
- Apple does not use users’ private personal data or interactions for foundation-model training and applies filters for personal information, profanity, and unsafe material.
- Applebot follows robots.txt protocols and provides publishers with controls over which pages it can access and how they are used.
- The pipeline expands general-domain, mathematical, programming, and multilingual content while using headless rendering and LLM-assisted extraction for complex web documents.
- The models use more than 10B high-quality image-text pairs and 175M interleaved image-text documents after quality and compliance filtering.
- Synthetic captioning produced over 5B image-caption pairs, while curated text-rich data includes PDFs, documents, infographics, tables, and charts.
4 Pre-training
Apple’s pre-training pipeline expands language and image capabilities through tokenizer growth, vision-encoder training, continued pre-training, and long-context adaptation.
- The pre-training recipe scales support for more languages and features requiring image understanding.
- Expanding the tokenizer from 100k to 150k tokens improves representation quality for additional languages with 50% more tokens.
- The vision encoder uses contrastive pre-training followed by joint training with an LLM decoder, starting from more than 6B image-text pairs.
- Text-only continued pre-training targets Math, Code, Knowledge, and Multilingual alignment using synthetic, high-quality organic, and bulk pre-training data.
- Multimodal adaptation improves visual understanding while avoiding regressions on text performance.
- Context lengthening trains on sequences up to 65K tokens while maintaining core capabilities.
5 Post-training
The post-training process expands multilingual and visual capabilities through supervised fine-tuning and reinforcement learning, using diverse synthetic and human-written data. A distributed asynchronous RL infrastructure improves training efficiency, while reward-aware prompt selection and multilingual data yield gains in automated and human evaluations.
- Training process: Post-training combines supervised fine-tuning and RLOO-based RLHF for both the on-device and server models.RLHF follows SFT, and Apple reports significant gains, especially on human preferences.
- Reinforcement learning infrastructure: Distributed asynchronous RL separates trajectory generators from policy updates, allowing generation and improvement to proceed simultaneously with independently optimized resources.The infrastructure supports diverse reward signals, including reward models, ground-truth verification, code execution, and LLM-as-a-judge.
- Reinforcement learning infrastructure: 37.5% fewer devices and 75% less compute time were used than in an earlier synchronous RL system, with similar performance.Trajectory generation and policy improvement overlap, while inference resources can scale independently for throughput.
- RLHF recipe: Cohesion-based prompt selection improved automated benchmarks and increased overall human satisfaction by 1.3–2.0% across locales.Reported automated gains include 4% in Arena Hard, 7% AlpacaEval win rate versus GPT4-Turbo, 10% in Agent Sandbox, 7% in GPQA, and 5% in Math500.
- Multilingual training: Multilingual data is included in both SFT and RLHF at an 80:20 English-to-multilingual sampling proportion, with RLHF producing a 16:9 win/loss rate over SFT in human evaluations.Both stages combine human-written and synthetic datasets.
6 Optimizations
The models use quantization, compression, and architectural techniques to reduce inference cost while preserving quality. The on-device model uses 2-bit quantization-aware training, while the server model uses ASTC compression with hardware decompression and LoRA-based quality recovery.
- Model compression: The on-device model is compressed to 2 bits-per-weight with quantization-aware training, while the server model reaches 3.56 bits-per-weight through ASTC post-training compression.Embedding tables use 4 bits per weight, and the KV-cache uses 8 bits per weight.
- On-device quantization: Quantization-aware training simulates quantization during training and uses a straight-through estimator to approximate gradients through rounding.The quantization range is defined using a scaling factor, zero point, and qmin/qmax bounds.
- On-device quantization: A learnable scaling factor adaptively tunes each weight tensor’s quantization range instead of deriving the scale conventionally from the weights alone.An iterative Newton-Raphson-inspired initialization estimates a clipping scalar to reduce outlier influence and stabilize 2-bit training.
- On-device quantization: A balanced 2-bit set {-1.5, -0.5, 0.5, 1.5} produces smoother training with fewer loss spikes than {-2, -1, 0, 1}.The authors also set weight decay to 0 to encourage use of the full quantization range.
- Server-model compression: ASTC compression reduces memory bandwidth and storage overhead without adding prompt-time inference latency because decompression occurs in hardware.The method uses 6×6 blocks encoded into 128-bit compressed values in HDR-ch mode.
- Quality recovery: LoRA adapters compensate for quantization and compression artifacts while keeping the core model weights frozen.The adapters are fine-tuned with the same data recipe as base-model training.
7 Foundation Models Framework
The Foundation Models framework gives developers Swift-centric access to the approximately 3B-parameter on-device model through guided generation, tool calling, and adapter customization.
- Guided generation: Guided generation converts Swift data structures into constrained model outputs, reducing manual format specification and string parsing.The framework uses the @Generable macro and optimized constrained decoding to produce instances of Swift structures.
- Tool calling: Tool calling guarantees structurally correct tool names and arguments by building on guided generation.Developers implement the Tool Swift protocol while the framework handles tool-call generation.
- Sessions: LanguageModelSession couples an append-only session to the model’s KV cache and supports streaming through snapshots.The session is designed to avoid unintended cache invalidation while partially generated content streams.
- Customization: Developers can train rank-32 LoRA adapters and optionally a draft model for specialized on-device use cases.Adapters are compatible with the framework, but each adapter supports only one specific model version and requires storage.
- Developer tooling: The framework integrates with Xcode through prompt-engineering playgrounds, an inference profiler, and iOS and visionOS simulator support.These tools accompany the Swift API for developing and testing on-device model features.
8 Evaluation
Apple evaluates its on-device and server models across language, reasoning, visual, human, and locale-specific tasks against external models. The reported comparisons show favorable or competitive performance against similarly sized systems, with larger models often remaining ahead.
- Pretraining benchmarks: The on-device model outperforms Qwen-2.5-3B, Gemma-3-4B, and Gemma-3n-E4B on MMLU/MMMLU, but trails Gemma-3n-E4B on MGSM.It also performs lower than the larger Qwen-3-4B model.
- Pretraining benchmarks: The server model slightly trails LLaMA 4 Scout and has a larger gap against Qwen-3-235B and GPT-4o.LLaMA 4 Scout has comparable total and active parameter counts.
- Optimization impact: Optimization evaluations measure quality before and after QAT, ASTC, and quality-recovery adapters alongside inference efficiency.The reported optimization goal is higher token throughput, lower latency, and reduced DRAM footprint relative to 16-bit weights.
- Evaluation design: Evaluation covers analytical reasoning, brainstorming, chat, coding, creative writing, extraction, mathematical reasoning, question answering, rewriting, summarization, and tool use.The evaluation suite also includes locale-specific assessments of native-sounding responses, excluding locale-agnostic domains such as math and coding.
- Image understanding: The on-device model performs favorably against larger InternVL-2.5-4B and Qwen-2.5-VL-3B models and competitively against Gemma-3-4B.The server model outperforms Qwen-2.5-VL-32B at less than half the inference FLOPS, while trailing larger LLaMA 4 Scout and GPT-4o.
9 Responsible AI
Apple frames its foundation models and features within Responsible AI principles covering user needs, representation, careful design, and privacy. It combines regional expertise, safety evaluation, guardrails, and user feedback to identify and mitigate risks.
- Responsible AI principles: Apple’s Responsible AI principles empower users, represent global users authentically, design with care, and protect privacy.The principles address user needs, stereotypes and bias, potential misuse or harm, and privacy-preserving infrastructure.
- Risk mitigation: Apple’s safety taxonomy is updated as risks evolve, including hallucinations and susceptibility to prompt injections.Newly discovered risks may lead to added override terms that trigger warnings or prevent generation.
- Safety evaluation: Safety evaluation combines internal and external human evaluation, auto-grading, external-model benchmarking, and targeted safety datasets before deployment.Both foundation models and individual features are assessed, including risks arising in sensitive app-specific content.
- Developer safeguards: The Foundation Models framework includes built-in safety guardrails and educational resources to help developers tailor AI safety to their apps.The resources include Generative AI Human Interface Guidelines for Responsible AI principles.
- Internationalization: Expanding language support requires broader safety representation across regions and cultures using representative data and legal, language, and cultural expertise.Locale expansion also incorporates local laws and regulations.
- Continuous improvement: User and developer feedback, evaluation data, and other metrics support continuous improvement of Apple Intelligence features and models.Examples include feedback on generated images and feedback submitted through Feedback Assistant.
10 Conclusion
Apple presents its language foundation models as more efficient and capable across many languages and software platforms. The Foundation Models framework gives developers direct access to the on-device model for integrating capabilities into apps.
- Conclusion: The models are intended to unlock helpful features across Apple software platforms and many languages.The conclusion emphasizes both increased efficiency and expanded capability.
- Conclusion: The Foundation Models framework provides developers direct access to Apple’s on-device language model with free inference and a few lines of code.Supported capabilities include text extraction and summarization.