Source-linked AI summary
NVIDIA Nemotron 3: Efficient and Open Intelligence
NVIDIA, :, Aaron Blakeman, Aaron Grattafiori, Aarti Basant, Abhibha Gupta, Abhinav Khattar, Adi Renduchintala, Aditya Vavre, Akanksha Shukla, Akhiad Bercovich, Aleksander Ficek, Aleksandr Shaposhnikov, Alex Kondratenko, Alexander Bukharin, Alexandre Milesi, Ali Taghibakhshi, Alisa Liu, Amelia Barton, Ameya Sunil Mahabaleshwarkar, Amir Klein, Amit Zuker, Amnon Geifman, Amy Shen, Anahita Bhiwandiwalla, Andrew Tao, Anjulie Agrusa, Ankur Verma, Ann Guan, Anubhav Mandarwal, Arham Mehta, Ashwath Aithal, Ashwin Poojary, Asif Ahamed, Asit Mishra, Asma Kuriparambil Thekkumpate, Ayush Dattagupta, Banghua Zhu, Bardiya Sadeghi, Barnaby Simkin, Ben Lanir, Benedikt Schifferer, Besmira Nushi, Bilal Kartal, Bita Darvish Rouhani, Boris Ginsburg, Brandon Norick, Brandon Soubasis, Branislav Kisacanin, Brian Yu, Bryan Catanzaro, Carlo del Mundo, Chantal Hwang, Charles Wang, Cheng-Ping Hsieh, Chenghao Zhang, Chenhan Yu, Chetan Mungekar, Chintan Patel, Chris Alexiuk, Christopher Parisien, Collin Neale, Cyril Meurillon, Damon Mosk-Aoyama, Dan Su, Dane Corneil, Daniel Afrimi, Daniel Lo, Daniel Rohrer, Daniel Serebrenik, Daria Gitman, Daria Levy, Darko Stosic, David Mosallanezhad, Deepak Narayanan, Dhruv Nathawani, Dima Rekesh, Dina Yared, Divyanshu Kakwani, Dong Ahn, Duncan Riach, Dusan Stosic, Edgar Minasyan, Edward Lin, Eileen Long, Eileen Peters Long, Elad Segal, Elena Lantz, Ellie Evans, Elliott Ning, Eric Chung, Eric Harper, Eric Tramel, Erick Galinkin, Erik Pounds, Evan Briones, Evelina Bakhturina, Evgeny Tsykunov, Faisal Ladhak, Fay Wang, Fei Jia, Felipe Soares, Feng Chen, Ferenc Galko, Frank Sun, Frankie Siino, Gal Hubara Agam, Ganesh Ajjanagadde, Gantavya Bhatt, Gargi Prasad, George Armstrong, Gerald Shen, Gorkem Batmaz, Grigor Nalbandyan, Haifeng Qian, Harsh Sharma, Hayley Ross, Helen Ngo, Herbert Hum, Herman Sahota, Hexin Wang, Himanshu Soni, Hiren Upadhyay, Huizi Mao, Huy C Nguyen, Huy Q Nguyen, Iain Cunningham, Ido Galil, Ido Shahaf, Igor Gitman, Ilya Loshchilov, Itamar Schen, Itay Levy, Ivan Moshkov, Izik Golan, Izzy Putterman, Jan Kautz, Jane Polak Scowcroft, Jared Casper, Jatin Mitra, Jeffrey Glick, Jenny Chen, Jesse Oliver, Jian Zhang, Jiaqi Zeng, Jie Lou, Jimmy Zhang, Jinhang Choi, Jining Huang, Joey Conway, Joey Guman, John Kamalu, Johnny Greco, Jonathan Cohen, Joseph Jennings, Joyjit Daw, Julien Veron Vialard, Junkeun Yi, Jupinder Parmar, Kai Xu, Kan Zhu, Kari Briski, Katherine Cheung, Katherine Luna, Keith Wyss, Keshav Santhanam, Kevin Shih, Kezhi Kong, Khushi Bhardwaj, Kirthi Shankar, Krishna C. Puvvada, Krzysztof Pawelec, Kumar Anik, Lawrence McAfee, Laya Sleiman, Leon Derczynski, Li Ding, Lizzie Wei, Lucas Liebenwein, Luis Vega, Maanu Grover, Maarten Van Segbroeck, Maer Rodrigues de Melo, Mahdi Nazemi, Makesh Narsimhan Sreedhar, Manoj Kilaru, Maor Ashkenazi, Marc Romeijn, Marcin Chochowski, Mark Cai, Markus Kliegl, Maryam Moosaei, Matt Kulka, Matvei Novikov, Mehrzad Samadi, Melissa Corpuz, Mengru Wang, Meredith Price, Michael Andersch, Michael Boone, Michael Evans, Miguel Martinez, Mikail Khona, Mike Chrzanowski, Minseok Lee, Mohammad Dabbah, Mohammad Shoeybi, Mostofa Patwary, Nabin Mulepati, Najeeb Nabwani, Natalie Hereth, Nave Assaf, Negar Habibi, Neta Zmora, Netanel Haber, Nicola Sessions, Nidhi Bhatia, Nikhil Jukar, Nikki Pope, Nikolai Ludwig, Nima Tajbakhsh, Nir Ailon, Nirmal Juluru, Nishant Sharma, Oleksii Hrinchuk, Oleksii Kuchaiev, Olivier Delalleau, Oluwatobi Olabiyi, Omer Ullman Argov, Omri Puny, Oren Tropp, Ouye Xie, Parth Chadha, Pasha Shamis, Paul Gibbons, Pavlo Molchanov, Pawel Morkisz, Peter Dykas, Peter Jin, Pinky Xu, Piotr Januszewski, Pranav Prashant Thombre, Prasoon Varshney, Pritam Gundecha, Przemek Tredak, Qing Miao, Qiyu Wan, Rabeeh Karimi Mahabadi, Rachit Garg, Ran El-Yaniv, Ran Zilberstein, Rasoul Shafipour, Rich Harang, Rick Izzo, Rima Shahbazyan, Rishabh Garg, Ritika Borkar, Ritu Gala, Riyad Islam, Robert Hesse, Roger Waleffe, Rohit Watve, Roi Koren, Ruoxi Zhang, Russell Hewett, Russell J. Hewett, Ryan Prenger, Ryan Timbrook, Sadegh Mahdavi, Sahil Modi, Samuel Kriman, Sangkug Lim, Sanjay Kariyappa, Sanjeev Satheesh, Saori Kaji, Satish Pasumarthi, Saurav Muralidharan, Sean Narentharen, Sean Narenthiran, Seonmyeong Bak, Sergey Kashirsky, Seth Poulos, Shahar Mor, Shanmugam Ramasamy, Shantanu Acharya, Shaona Ghosh, Sharath Turuvekere Sreenivas, Shelby Thomas, Shiqing Fan, Shreya Gopal, Shrimai Prabhumoye, Shubham Pachori, Shubham Toshniwal, Shuoyang Ding, Siddharth Singh, Simeng Sun, Smita Ithape, Somshubra Majumdar, Soumye Singhal, Stas Sergienko, Stefania Alborghetti, Stephen Ge, Sugam Dipak Devare, Sumeet Kumar Barua, Suseella Panguluri, Suyog Gupta, Sweta Priyadarshi, Syeda Nahida Akter, Tan Bui, Teodor-Dumitru Ene, Terry Kong, Thanh Do, Tijmen Blankevoort, Tim Moon, Tom Balough, Tomer Asida, Tomer Bar Natan, Tomer Ronen, Tugrul Konuk, Twinkle Vashishth, Udi Karpas, Ushnish De, Vahid Noorozi, Vahid Noroozi, Venkat Srinivasan, Venmugil Elango, Victor Cui, Vijay Korthikanti, Vinay Rao, Vitaly Kurin, Vitaly Lavrukhin, Vladimir Anisimov, Wanli Jiang, Wasi Uddin Ahmad, Wei Du, Wei Ping, Wenfei Zhou, Will Jennings, William Zhang, Wojciech Prazuch, Xiaowei Ren, Yashaswi Karnati, Yejin Choi, Yev Meyer, Yi-Fu Wu, Yian Zhang, Yigong Qin, Ying Lin, Yonatan Geifman, Yonggan Fu, Yoshi Subara, Yoshi Suhara, Yubo Gao, Zach Moshe, Zhen Dong, Zhongbo Zhu, Zihan Liu, Zijia Chen, Zijie Yan
TL;DR
Nemotron 3 addresses the need for accurate, efficient open models for agentic AI applications. It combines a hybrid Mamba-Transformer MoE architecture with long-context support, multi-environment reinforcement learning, and specialized techniques for larger models. The family is reported to deliver leading accuracy with improved throughput and faster generation, while releasing substantial model and training resources openly.
Problem
Nemotron 3 targets the need for accurate, efficient open models that support agentic AI applications and long-context workloads.
Method
The family combines a hybrid Mamba-Transformer MoE architecture with multi-environment reinforcement learning, long-context support, and LatentMoE, NVFP4, and MTP for larger models.
Results
Nemotron 3 is reported to provide leading accuracy and throughput, including 3.3× higher throughput than Qwen3-30B-A3B for Nemotron Nano 30B-A3B and roughly 2.4% average benchmark improvement from MTP in an 8B active-parameter MoE ablation.
Takeaways & Limitations
Nemotron 3 is intended to support high-accuracy agentic applications, including multi-agent environments, long-context processing, and low-latency text generation.
Abstract
from arXiv · showhide
We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transformer architecture to provide best-in-class throughput and context lengths of up to 1M tokens. Super and Ultra models are trained with NVFP4 and incorporate LatentMoE, a novel approach that improves model quality. The two larger models also include MTP layers for faster text generation. All Nemotron 3 models are post-trained using multi-environment reinforcement learning enabling reasoning, multi-step tool use, and support granular reasoning budget control. Nano, the smallest model, outperforms comparable models in accuracy while remaining extremely cost-efficient for inference. Super is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Ultra, the largest model, provides state-of-the-art accuracy and reasoning performance. Nano is released together with its technical report and this white paper, while Super and Ultra will follow in the coming months. We will openly release the model weights, pre- and post-training software, recipes, and all data for which we hold redistribution rights.
1. Introduction
Nemotron 3 is an open family of Nano, Super, and Ultra models designed to combine leading accuracy with efficient inference for agentic AI applications. Its hybrid architecture, long-context support, reinforcement learning, specialized Super and Ultra techniques, and planned open releases define the family.
- Nemotron 3 is presented as an open family of Nano, Super, and Ultra models targeting efficient, accurate agentic AI applications.The planned release includes model weights, training software, recipes, and datasets.
- The family uses a hybrid Mamba-Transformer MoE architecture, supports context lengths up to 1M tokens, and provides inference-time reasoning budget control.These capabilities target efficient reasoning and long-context workloads.
- Nemotron 3 models use diverse reinforcement learning environments to improve accuracy across coding, mathematics, and agentic tool-use tasks.The environments cover a broad range of reasoning and agentic capabilities.
- Super and Ultra add NVFP4 training, LatentMoE, and MTP layers to improve model quality and accelerate long-form generation without sacrificing throughput or latency.MTP is described as providing faster generation and modest quality improvements.
- Nemotron 3 is positioned as an efficient open-model family with state-of-the-art accuracy on reasoning and ultra-long-context tasks.The cited figures describe accuracy and throughput advantages of the hybrid architecture.
2. Features and Technologies
Nemotron 3 combines hybrid Mamba-Transformer MoE layers, LatentMoE, MTP, and NVFP4 training to improve inference efficiency, accuracy, and long-context capability. The supplied results report higher throughput, broad MTP gains, stable NVFP4 performance, and stronger long-context robustness.
- 2.1. Hybrid MoE: 3.3× higher throughput: Nemotron 3 Nano 30B-A3B versus Qwen3-30B-A3B on common reasoning workloads.The hybrid architecture minimizes expensive self-attention layers and provides further speedups at longer sequences.
- 2.2. LatentMoE: Hardware-Aware Expert Design for Improved Accuracy per Byte: LatentMoE consistently outperforms Standard MoE across all evaluated downstream tasks at matched active and total parameter counts.The comparison uses models with 8B active and 73B total parameters trained for 1T tokens with identical hyperparameters.
- 2.3. Multi-Token Prediction (MTP): 2.4% average accuracy improvement: MTP across benchmarks spanning knowledge, code, commonsense, reading comprehension, and math.MTP adds minimal FLOPs while providing denser supervision and speculative-decoding benefits.
- 2.3. Multi-Token Prediction (MTP): 97% first-two-token acceptance: the lightweight MTP module in an 8B active MoE ablation supports low-latency speculative decoding.The benefit is particularly relevant to batch-size–1 and long-form generation scenarios.
- 2.4. NVFP4 Training: < 1% relative loss difference: Nano trained with NVFP4 versus BF16, decreasing to < 0.6% for the larger 8B-active MoE model.Downstream evaluations for the A8B model are comparable between BF16 and NVFP4 training.
3. Key Takeaways
Nemotron 3 is a family of open models—Nano, Super, and Ultra—designed for efficient, accurate agentic AI applications.
- Nemotron 3 comprises Nano, Super, and Ultra models for building high-accuracy agentic AI applications.