Source-linked AI summary
Gemma 4 Technical Report
Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, Mayank Chaturvedi, Aditya Chawla, Victor Cotruta, Alice Coucke, Phil Culliton, Robert Dadashi, Lucas Dixon, Mohamed Elhawaty, Utku Evci, Clément Farabet, Johan Ferret, Filippo Galgani, Sertan Girgin, Jean-Bastien Grill, Maarten Grootendorst, Jiaxian Guo, Cassidy Hardin, Yanzhang He, Steven M. Hernandez, Omri Homburger, Léonard Hussenot, Juyeong Ji, Armand Joulin, Aishwarya Kamath, Parnian Kassraie, Olivier Lacombe, Preethi Lahoti, Gaël Liu, Gus Martins, Luciano Martins, Tatiana Matejovicova, Ramona Merhej, Nikola Momchev, Sneha Mondal, Ryan Mullins, Sindhu Raghuram Panyam, Shreya Pathak, Sarah Perrin, André Susano Pinto, Etienne Pot, Angéline Pouget, Alexandre Ramé, Sabela Ramos, Douglas Reid, David Rim, Morgane Rivière, Karsten Roth, Louis Rouillard, Omar Sanseviero, Pier Giuseppe Sessa, Shane Settle, Danila Sinopalnikov, Sara Smoot, Piotr Stanczyk, Andreas Steiner, Lawrence Stewart, Ilya Tolstikhin, Michael Tschannen, Anton Tsitsulin, Nino Vieillard, Renjie Wu, Pingmei Xu, Haichuan Yang, Edouard Yvinec, Biao Zhang, Li Zhang, Joe Zou, Nicolas Aagnes, Abdelrahman Abdelhamed, Jakub Adamek, Shivani Agrawal, Shubham Agrawal, Ibrahim Alabdulmohsin, Jean Baptiste Alayrac, Uri Alon, Chandramouli Amarnath, Ankesh Anand, Chrysovalantis Anastasiou, Setareh Ariafar, François-Xavier Aubet, Kyriakos Axiotis, Federico Barbero, Joelle Barral, Alexei Bendebury, Urs Bergmann, Stanley Bileschi, Kat Black, Mathieu Blondel, Sebastian Borgeaud, Arthur Bražinskas, Ryan Burnell, Robert Busa-Fekete, Mu Cai, Daniele Calandriello, Glenn Cameron, Charlotte Caucheteux, Rahma Chaabouni, Garima Chadha, Jetha Chan, Blake Jianhang Chen, Jesse Chen, Lin Chen, Xu Chen, Derek Cheng, Tzu-hsiang Chien, Nikolai Chinaev, Yi Chou, Zhaohui Chu, Benjamin Coleman, Pooja Consul, Sam Conway-Rahman, Scott Crowell, Dylan Cutler, Vivek Dani, Samira Daruki, Anil Das, Daniel Deutsch, Nishanth Dikkala, Li Ding, Qiuhan Ding, Shenil Dodhia, Konstantin Donhauser, Tulsee Doshi, Anca Dragan, Alex Druinsky, Sahil Dua, Zoltan Egyed, Danielle Eisenbud, Daniel Eppens, Cindy Fan, Bahare Fatemi, Yassir Fathullah, Vlad Feinberg, Milen Ferev, Sebastian Flennerhag, Takumi Fujimoto, João Gabriel Oliveira, Isaac Galatzer-Levy, João Gante, Simon Geisler, Soham Ghosal, Antonious M. Girgis, Tamara von Glehn, Alec Go, Alhaad Gokhale, Alex Grills, Yiming Gu, Mayank Gupta, Pramod Gupta, Guru Guruganesh, Raia Hadsell, Hamza Harkous, Jitendra Harlalka, Demis Hassabis, Anja Hauth, Joe Heyward, Arian Hosseini, Chih-Yang Hsia, I-Hung Hsu, Xiaopeng Huang, Yangsibo Huang, Kevin Hui, Adrian Hutter, Te I, Fotis Iliopoulos, Advait Jain, Ganesh Jawahar, Ziwei Ji, Qilin Jin, Melvin Johnson, Kandarp Joshi, Arun Kandoor, Wang-Cheng Kang, Koray Kavukcuoglu, Mehran Kazemi, Kathleen Kenealy, Amr Khalifa, Phoebe Kirk, Ivan Korotkov, Suraj Kothawade, Vitaly Kovalev, Neel Kovelamudi, Adam Kraft, Ravin Kumar, Vivek Kumar, Harish Kuppam, Justin Lannin, Chen-Yu Lee, Seungji Lee, Dmitry Lepikhin, Alon Levkovitch, Dongdong Li, Qiujia Li, Valentin Liévin, Ethan Lin, Ziqian Lin, Casper Liu, Tianlin Liu, Tianqi Liu, Xin Liu, Ivan Lobov, Mayank Lunayach, Min Ma, Gagan Madan, Andrii Maksai, Eric Malmi, Michal Matuszak, Daniel McDuff, Gaurav Menghani, Maciej Mikuła, Daniil Mirylenka, Karolis Misiunas, Vedant Misra, Andreea Mitran, Kareem Mohamed, Maksim Mukha, Eric Noland, James O'Donnell, Brendan O'Donoghue, Kate Olszewska, Bernett Orlando, Wanqiong Pan, Rina Panigrahy, Unnati Parekh, Nicolas Perez-Nieves, Chunjong Park, Eric Paskie, Liqian Peng, Bryce Petrini, Slav Petrov, Jonas Pfeiffer, Bilal Piot, Martyna Plomecka, Siim Poder, Octavio Ponce, Arijit Pramanik, David Racz, Anish Rajan, Michelle Ramanovich, Anand Rao, Marvin Ritter, Vitor Rodrigues, Evan Rosen, Mikołaj Rybiński, Noveen Sachdeva, Michaël E. Sander, Rohit Sathyanarayana, Sagar Savla, Samuel Schmidgall, Tal Schuster, George Scrivener, Benoit Seguin, Andrew Sellergren, Aliaksei Severyn, Izhak Shafran, Dhruv Shah, Bobak Shahriari, Yuan Shangguan, Ashish Shenoy, Pradeep Shenoy, Rakesh Shivanna, Pauline Sho, Lucas Spangher, Wojciech Stokowiec, Tim Strother, Yao Su, Yinghao Sun, Mukund Sundararajan, Andrea Tacchetti, Mor Hazan Taege, Pouya Tafti, Jean Tarbouriech, Chetan Tekur, Shantanu Thakoor, Rahul Thapa, Madeleine Traverse, Lenart Treven, Tao Tu, Chien Te Tung, Çağlar Ünlü, Petar Veličković, Malini Pooni Venkat, Sagar Gubbi Venkatesh, Vidya Venkiteswaran, Francesco Visin, Alex Vitvitskyi, Kiran Vodrahalli, Weiyi Wang, Xin Wang, Tris Warkentin, Jan Wassenberg, John Wieting, Cindy Wu, Lechao Xiao, Hao Xu, Yuhui Xu, Fuzhao Xue, Arun Yadav, Jun Yan, Antoine Yang, Lin Yang, Ming-Hsuan Yang, Ziyu Ying, Jae Hyeon Yoo, Morteza Zadimoghaddam, Sajjad Zafar, Fred Zhang, Jiageng Zhang, Jianyi Zhang, Xiaofan Zhang, Chao Zhao, David Zhou, Chen Zou
TL;DR
Gemma 4 addresses the need for open-weight models combining multimodal understanding, reasoning, and computational efficiency. It introduces dense and MoE architectures, thinking mode, efficiency optimizations, and an encoder-free 12B design, and reports improved benchmark performance with human-rated results comparable to larger open models.
Problem
Open-weight models need strong multimodal understanding, reasoning, and computational efficiency across varied hardware environments.
Method
Gemma 4 combines dense and MoE architectures with thinking mode, long-context and cache optimizations, QAT, MTP speculative-decoding drafters, and an encoder-free 12B multimodal architecture.
Results
Gemma 4 demonstrates a leap over Gemma 3 across benchmarks and performs comparably to significantly larger open models in human evaluations.
Takeaways & Limitations
Gemma 4 provides a scalable open-weight foundation for edge deployment, reasoning, and open research across text, image, and audio modalities.
Takeaways & Limitations
Gemma 4 can be misused to generate false or misleading text, requiring responsible-use guidance and technical limitations to mitigate malicious applications.
Abstract
from arXiv · showhide
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.
1. Introduction
Gemma 4 is an open-weight, natively multimodal model family spanning dense and MoE architectures, with designs targeting reasoning, memory, compute, and long-context efficiency. Its innovations include thinking mode, optimized attention and KV caching, quantization-aware training, speculative decoding, and an encoder-free 12B architecture.
- Model family: Gemma 4 spans dense 2.3B–31B models and a 3.8B-activated, 26B-total-parameter MoE model for varied hardware environments.The parameter-count table defines the MoE by active parameters and notes per-layer embeddings for E2B and E4B.
- Reasoning: Thinking mode generates a reasoning trace before the response to improve capabilities in reasoning-heavy domains such as mathematics and coding.
- Efficiency and multimodality: Gemma 4 adds MTP drafter heads for speculative decoding, QAT-trained quantized models, and a unified encoder-free 12B architecture that projects raw audio chunks and image patches into the LLM embedding space.These choices target decoding speed, parameter memory, latency, and memory fragmentation while preserving quality with minimal impact from quantization.
- Evaluation: Across text, image, and audio modalities, Gemma 4 performs comparably to larger frontier open models in comprehensive benchmarks and human evaluations.The models are released under an Apache 2.0 license.
2. Model Architecture
Gemma 4 combines dense and MoE decoder-only Transformers across model sizes with multimodal encoders, including an encoder-free 12B design. The architecture also targets long-context, memory, quantization, and decoding efficiency.
- Gemma 4 includes dense 2.3B, 4.5B, 12B, and 31B models plus a 26B MoE model with 3.8B activated parameters.
- Long-context efficiency: Long-context design uses local-to-global attention ratios, p-RoPE, key-value reuse, and KV-cache sharing to reduce the global KV cache by 37.5%.The stated ratios are 4-to-1 for E2B and 5-to-1 for the other models.
- Encoder-free architecture: The 12B model replaces separate vision and audio encoders with lightweight projections for raw image patches and audio chunks.Vision uses 48×48×3 RGB patches and a 35M-parameter matrix; audio uses 40ms chunks projected directly into the LLM embedding space.
- Quantization-Aware Training: Quantization-aware training reduces the 150M image encoder’s forward-pass memory from 400 MB to 200 MB and on-device latency by 44%.For the audio encoder, the on-disk footprint decreases from 390 MB to 87 MB.
- Compute efficiency: An autoregressive multi-token prediction drafter supports speculative decoding by generating future tokens while cross-attending to the main model’s key-value states.For E2B and E4B, clustered top-k projection reduces the final matrix multiplication from d×262,000 to d×4096 while preserving a similar acceptance rate.
3. Instruction Tuning
Gemma 4 instruction tuning follows Gemma 3’s post-training approach while adding thinking mode, which outputs a reasoning trace before the answer. Post-training data filtering targets safety, duplication, attribution, hedging, refusals, and factuality.
- Instruction-tuned models use a post-training approach similar to Gemma 3, with thinking mode as the significant difference.
- Thinking mode lets models output a reasoning trace before answering.
- Data filtering: Post-training data filtering removes personal information, unsafe or toxic outputs, mistaken self-identification, and duplicated examples.
- Data filtering: Examples encouraging in-context attribution, hedging, and refusals improve factuality metrics without degrading other metrics.
4. Evaluation of final models
Gemma 4 is evaluated across automated and human benchmarks spanning multiple domains and modalities. The results show stronger performance than Gemma 3 27B, competitive long-context and vision performance, and parity with much larger open models in human ratings.
- Gemma 4 final models are evaluated on automated benchmarks, human evaluations, and static benchmarks across varied domains.
- Human evaluation: Gemma 4 31B is the top open dense model in Arena, while Gemma 4 31B and 26B-A4B match much larger open models.Arena uses blind side-by-side evaluations by human raters and reports Elo scores.
- Static benchmarks: Gemma 4 31B is significantly better than similarly sized Gemma 3 27B across benchmarks, while E2B roughly matches it with 10x fewer parameters.
- Static benchmarks: E4B equals or outperforms Gemma 3 27B on all reported vision evaluations.
- Static benchmarks: Gemma 4 shows a leap in long-context capabilities over Gemma 3 27B, with E4B outperforming the older model.
5. Responsibility, Safety, Security
Gemma 4’s responsibility approach combines safety evaluation, data filtering, policy alignment, and deployment guidance for multimodal open models. Evaluations report improved content safety with low unjustified refusals, while the report acknowledges risks involving bias, misinformation, and privacy.
- Gemma 4 undergoes rigorous safety evaluations and is developed with responsibility and security as core priorities.
- Safety development accounts for expanded multimodal capabilities and the risk that malicious uses can cause individual and institutional harm.
- Safety Policies and Train-Time Mitigations: Training data is filtered for personal information and sensitive data, while post-training evaluations and train-time mitigations support alignment with safety policies.
- Safety Evaluations: Gemma 4 models significantly outperform Gemma 3 and 3n on content-safety improvements while keeping unjustified refusals low.
- Safety Evaluations: Across text-to-text and image-to-text modalities and all model sizes, testing without safety filters found minimal policy violations.
- Ethical Considerations and Risk Mitigation: The report identifies bias, misinformation, misuse, and privacy as risks requiring monitoring, developer education, safeguards, and privacy-preserving techniques.
6. Discussion and Conclusion
Gemma 4 combines open-weight multimodal dense and MoE models with thinking mode and an encoder-free architecture for raw audio and image patches. The models improve efficiency and benchmark performance while targeting edge deployment and open research.
- Gemma 4 is an open-weight multimodal family spanning dense and MoE architectures for varied hardware environments.
- Thinking mode generates reasoning traces before responses and is reported to improve overall performance.
- The unified encoder-free architecture processes raw audio and image patches.
- Long-context memory limitations are addressed through local-to-global attention ratios, positional encoding, and KV cache sharing.
- Gemma 4 shows a performance leap over Gemma 3 across benchmarks and comparable performance to significantly larger open models in human evaluations.
- The resulting models are presented as a scalable foundation for edge deployment and reasoning while supporting open research.
Core contributors
This section lists contributors to the Gemma 4 technical report. The supplied passages contain author names rather than technical findings.
- The contributor list includes Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, and Olivier Bachem.
- The contributor list continues with Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, and Mayank Chaturvedi.
- The listed contributors include Aditya Chawla, Victor Cotruta, Alice Coucke, Phil Culliton, and Robert Dadashi.
Contributors
This section provides additional contributor names for the Gemma 4 technical report. The supplied passages contain author lists rather than research claims.
- The contributor lists include Douglas Reid, David Rim, Morgane Rivière, Karsten Roth, and Omar Sanseviero.
- Additional listed contributors include Alexei Bendebury, Urs Bergmann, Stanley Bileschi, Kat Black, and Mathieu Blondel.
- The names include Nicolas Aagnes, Abdelrahman Abdelhamed, Jakub Adamek, Shivani Agrawal, and Shubham Agrawal.
- The contributor list includes Blake Jianhang Chen, Jesse Chen, Lin Chen, Xu Chen, and Derek Cheng.
- Additional contributors listed are Fotis Iliopoulos, Advait Jain, Ganesh Jawahar, Ziwei Ji, and Qilin Jin.
- The list also includes Anil Das, Daniel Deutsch, Nishanth Dikkala, Li Ding, and Qiuhan Ding.
Appendix
The appendix documents conversation formatting and vision processing for Gemma 4. It covers image resizing, vision architecture references, and benchmark reporting at a specified resolution.
- Vision: The vision appendix references the encoder architecture, image-resizing algorithm, and low-resolution vision benchmark scores.
- Vision: Images are resized to 2 × 4 pooled patches using 16-pixel patches, 3 × 3 pooling, and a maximum of 10 soft tokens.
- Vision: The resizing example processes 72 vision-encoder patches and produces 8 soft tokens for the LLM backbone.
- Conversation format: Conversation formatting includes a leading [BOS] token, an optional thinking-mode system marker, and documented function-calling syntax.
- Vision: Table 12 reports Gemma 4 vision benchmark performance at resolution N_max=280 with thinking enabled.