Source-linked AI summary
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aaron Blakeman, Aaron Thomas, Aastha Jhunjhunwala, Abhibha Gupta, Abhinav Khattar, Adam Rajfer, Adi Renduchintala, Adil Asif, Aditya Vavre, Adriana Flores Miranda, Ahmad Bilal, Aileen Zaman, Ajay Hotchandani, Akanksha Shukla, Akhiad Bercovich, Aleksander Ficek, Alex Gronskiy, Alex Kondratenko, Alex Steiner, Alex Ye, Alexander Bukharin, Alexandre Milesi, Ali Taghibakhshi, Alice Gatti, Alisa Liu, Alok Kumar, Amar Phanishayee, Ameya Sunil Mahabaleshwarkar, Amir Klein, Amit Zuker, Amnon Geifman, Anahita Bhiwandiwalla, Ananth Subramaniam, Andrea Santilli, Andrew Fulks, Andrew McHarg, Andrew Tao, Andrii Skliar, Anjulie Agrusa, Ankur Srivastava, Ankur Verma, Anna Shors, Anna Warno, Antoni-Joan Solergibert I Llaquet, Arham Mehta, Arkadiusz Nowaczynski, Arti Jain, Ashwath Aithal, Ashwin Poojary, Asif Ahamed, Asit Mishra, Asma Kuriparambil Thekkumpate, Atefeh Sohrabizadeh, Avinash Kaur, Avinash Vem, Ayush Dattagupta, Barath Subramaniam Anandan, Bardiya Sadeghi, Ben Lanir, Benedikt Schifferer, Besmira Nushi, Bilal Kartal, Bill Thiede, Bita Darvish Rouhani, Bo Deng, Bob Schatz, Boris Ginsburg, Boxin Wang, Brad Nemire, Brandon Norick, Brian Dang, Brian Westphal, Brian Yu, Brucek Khailany, Bryan Catanzaro, Carlo del Mundo, Caryln Aarish, Chankyu Lee, Chantal Hwang, Charbel Sakr, Charles Wang, Charlie Truong, Chen Cui, Cheng Cheng, Cheng-Ping Hsieh, Chenghao Zhang, Chenhui Deng, Chintan Patel, Chris Alexiuk, Christian Cosgrove, Christian Munley, Christine Harvey, Christopher Parisien, Chunyang Shen, Coco Li, Collin Neale, Cynthia Gao, Cyril Meurillon, Dan Gil, Dan Su, Dan Zhao, Dane Corneil, Daniel Afrimi, Daniel Egert, Daniel Korzekwa, Daniel Lo, Daniel Machlab, Daniel Serebrenik, Daniil Sorokin, Daria Gitman, Daria Levy, Darko Stosic, David Mosallanezhad, David Yu, Davit Karamyan, Deena Donia, Deep Debroy, Deepak Narayanan, Devin O'Kelly, Dheeraj Peri, Dhruv Nathawani, Di, Wu, Dima Rekesh, Divyanshu Kakwani, Donald Plummer, Dong Anh, Dongfeng Yu, Dongfu Jiang, Donnie Kim, Dorrin Poorkay, Duncan Riach, Dusan Stosic, Dustin VanStee, Eavan Meng, Edgar Minasyan, Edward Lin, Eileen Margaret Peters Long, Elad Sarafin, Elad Segal, Elena Lantz, Ellie Evans, Elliott Ning, Eric Chung, Eric Harper, Eric Pham-Hung, Eric Tramel, Eric Yang, Erick Galinkin, Erik Pounds, Erika Goncalves Goncalves, Evan Briones, Evan Wu, Evelina Bakhturina, Evgeny Tsykunov, Ewa Dobrowolska, Faisal Ladhak, Farzan Memarian, Fay Wang, Fei Jia, Felipe Soares, Felipe Vieira Frujeri, Feng Chen, Fengguang Lin, Ferenc Galko, Frank Sun, Frankie Siino, Frida Hou, Gal Hubara Agam, Gal Kaplun, Gantavya Bhatt, Gargi Prasad, Garvit Kulshreshtha, George Armstrong, Gerald Shen, Giulio Borghesi, Gordana Neskovic, Gorkem Batmaz, Grace Lam, Greg Mason, Greg Pauloski, Grigor Nalbandyan, Grzegorz Chlebus, Grzegorz Karch, Guan-Ting Liu, Guoming Zhang, Guyue Huang, Haggai Maron, Haifeng Qian, Haim Elisha, Haoxing Ren, Haran Kumar Shiv Kumar, Haribhau Hud, Harris Nover, Harrison Saturley Hall, Hayate Iso, Helen Ngo, Herbert Hum, Herman Sahota, Hexin Wang, Himanshu Soni, Hovhannes Tamoyan, Hua Li, Huanhuan Chen, Hui Li, Hui Wang, Huy Nguyen, Ian Chiles, Ido Galil, Ido Shahaf, Igor Gitman, Igor Shovkun, Ilya Loshchilov, Ingo Guehring, Itamar Schen, Itay Levy, Itay Neeman, Ivan Moshkov, Izik Golan, Izzy Putterman, Jaemin Choi, Jakub Slowikowski, Jan Kautz, Jane Polak Scowcroft, Jared Casper, Jatin Mitra, Jeffrey Glick, Jenny Chen, Jesse Oliver, Jiacheng Xu, Jiafan Zhu, Jialin Song, Jian Zhang, Jiantao Jiao, Jiaqi Zeng, Jie Lou, Jim King, Jimmy Zhang, Jingquan Wang, Jinhang Choi, Jinju Chu, Joey Conway, Joey Guman, Johan Jatko, Johannes Rausch, John Kamalu, John Roberts, Johnny Greco, Johnny Mensel, Jonah Alben, Jonas Yang, Jonathan Cohen, Jonathan Raiman, Joseph Jennings, Joshua Mabry, Joshua Pierce, Joyjit Daw, Julien Veron Vialard, Junkeun Yi, Jupinder Parmar, Kajal Jain, Kan Zhu, Kari Briski, Katherine Cheung, Katherine Luna, Keith Willowhawk, Keith Wyss, Keshav Santhanam, Kevin Shih, Kezhi Kong, Khanh Nguyen, Khushi Bhardwaj, Kirthi Shankar Sivamani, Konstantinos Krommydas, Krishna C. Puvvada, Krzysztof Pawelec, Kumar Anik, Kyle Keprios, Kylie Day, Lawrence McAfee, Leo Du, Leon Derczynski, Li Ding, Linda Liu, Lingjie Wu, Lior Kadoch, Lizzie Wei, Luis Vega, Luke Robison, Lun Su, Maarten Van Segbroeck, Maciej Jakub Mikulski, Maer Rodrigues de Melo, Magda Sypula, Mahan Fathi, Makesh Narsimhan Sreedhar, Makesh Tarun Chandran, Manoj Kilaru, Maor Ashkenazi, Marc Cuevas, Marc Romeijn, Marcin Chochowski, Mark Cai, Mark Mozolewski, Markus Kliegl, Marta Stepniewska-Dziubinska, Martyna Patelka, Mattei Machczynski, Matvei Novikov, Mauricio Ferrato, Maximilian Golub, Mehrzad Samadi, Melissa Corpuz, Mengru Wang, Mengxi Wu, Meredith Price, Meriem Boubdir, Micah Schaffer, Michael Andersch, Michael Boone, Michael Gschwind, Michael Lightstone, Michael Loh, Michal Bien, Michal Zawalski, Michelle Gill, Miguel Martinez, Mikail Khona, Mike Chrzanowski, Mike Houston, Mingyuan Ma, Minseok Lee, Mohamed Fawzy, Mohammad Dabbah, Mohammad Shoeybi, Mostofa Patwary, Nabin Mulepati, Najeeb Nabwani, Namit Dhameja, Narimane Hennouni, Natalie Hereth, Nathaniel Pinckney, Nave Algarici, Nave Assaf, Netanel Haber, Nicholas Knight, Nick Reamaroon, Nickson Quak, Nidhi Bhatia, Nikhil Desai, Nikolai Ludwig, Nima Tajbakhsh, Ning Xu, Nir Ailon, Nirmal Juluru, Nitin Nitin, Ofri Masad, Oleg Rybakov, Oleksii Hrinchuk, Oleksii Kuchaiev, Olivia Viessmann, Olivier Delalleau, Oluwatobi Olabiyi, Omer Ullman Argov, Omri Puny, Oren Tropp, Pablo Ribalta, Pallab Bhattacharya, Panos Lampropoulos, Parth Mannan, Pasha Shamis, Patrick Legresley, Paul Gibbons, Pavlo Molchanov, Pawel Morkisz, Peter Dykas, Peter Jin, Pierre-Yves Aquilanti, Pinky Xu, Piotr Januszewski, Piotr Laskiewicz, Pooya Jannaty, Prakash Gurumurthy, Pranav Prashant Thombre, Prasoon Varshney, Pritam Gundecha, Przemek Tredak, Puhui Meng, Qiyu Wan, Rabeeh Karimi Mahabadi, Rachel Oberman, Rachit Garg, Radha Sri-Tharan, Rahul Kandu, Rakshit Sanadhya, Ran El-Yaniv, Ran Zilberstein, Rasoul Shafipour, Ray Macalisang, Rayen Tian, Reka Kovacs, Renjie Pi, Rick Izzo, Rima Shahbazyan, Rishabh Garg, Rishi Puri, Rita Fernandes Neves, Ritchie Zhao, Ritika Borkar, Ritu Gala, Riyad Islam, Robert Clark, Robert Hesse, Robert Kirby, Roger Waleffe, Rohit Watve, Roi Koren, Ron Banner, Ruoxi Zhang, Russell J. Hewett, Ryan Prenger, Ryan Stewart, Ryota Egashira, Sadegh Mahdavi, Saee Paliwal, Sagar Singh, Sahil Modi, Salika Dave, Samantha Shinagawa, Samuel Kriman, Sandip Bhaskar, Sangkug Lym, Sanjay Kariyappa, Sanjeev Satheesh, Saran Vikas Murari, Satish Pasumarthi, Saurabh Mishra, Saurav Muralidharan, Scott Hara, Sean Narentharen, Selvaraj Anandaraj, Seonjin Na, Seonmeyong Bak, Seonmyeong Bak, Sepehr Sameni, Seph Mard, Serge Panev, Seth Henneman, Seth Poulos, Shahar Mor, Shantanu Acharya, Shaona Ghosh, Sharath Turuvekere Sreenivas, Sharon Mendelson, Shaun Kotek, Shawn Wang, Shay Aharon, Shaya Gharghabi, Sheng-Chieh Lin, Shi Chen, Shiqing Fan, Shirish Baskaran, Shreya Gopa, Shrimai Prabhumoye, Shubham Pachori, Shubham Toshniwal, Shuoyang Ding, Shwetha Krishnamurthy, Siddharth Singh, Simeng Sun, Sirshak Das, Sivakumar Arayandi Thottakara, Smita Ithape, Somshubra Majumdar, Soumye Singhal, Sri Harsha Singudasu, Sridhar Bhuvanapalli, Srimukh Veccham, Stas Sergienko, Stefania Alborghetti, Stephen Ge, Su Rong, Sugam Dipak Devare, Sukrit Rao, Sumeet Kumar Barua, Sungsoo Ha, Sunny Gai, Suriya Gunasekar, Suseella Panguluri, Suyog Gupta, Sviataslau Hinzburh, Sweta Priyadarshi, Syeda Nahida Akter, Talor Abramovich, Tan Bui, Tanay Varshney, Tatevik Ter-Hovhannisyan, Teodor-Dumitru Ene, Terry Kong, Thanh Do, Tianhe Zhang, Tiffany Moore, Tijmen Blankevoort, Tim Moon, Tiyasa Mitra, Tom Balough, Tomasz Grzegorzek, Tomasz Hliwiak, Tomer Asida, Tomer Bar Natan, Tomer Keren, Tomer Ronen, Tony Salim, Tony Wang, Traian Rebedea, Tugrul Konuk, Twinkle Vashishth, Udi Karpas, Ushnish De, Vahid Noorozi, Venkat Srinivasan, Venmugil Elango, Vibhor Agrawal, Victor Cui, Vijay Korthikanti, Vikas Mehta, Vinay Rao, Virginia Wu, Vitaly Kurin, Vitaly Lavrukhin, Vladimir Anisimov, Vu Pham, Wanli Jiang, Wasi Uddin Ahmad, Wataru Ishihara, Wei Du, Wei Ping, Weiheng Chai, Wenliang Dai, Wesley Helmholz, Will Jennings, Will Zhu, Wojciech Prazuch, Xiaowei Ren, Xiwen Yu, Yan Breek, Yang Chen, Yang Yu, Yangyi Chen, Yaniv Galron, Yashaswi Karnati, Yejin Choi, Yev Meyer, Yi-Fu Wu, Yian Zhang, Ying Lin, Yonatan Geifman, Yonggan Fu, Youngeun Kwon, Yu Yao, Yugi Guvvla, Yuki Huang, Yunsheng Liu, Zach Moshe, Zachary Newell, Zhilin Wang, Zhiyu Li, Zhongbo Zhu, Zhuolin Yang, Zihan Liu, Zijie Yan, Zsolt-Alon Wertheimer
TL;DR
Long-running autonomous agents need efficient inference without sacrificing accuracy. Nemotron 3 Ultra combines a Mixture-of-Experts hybrid Mamba-Attention architecture with SFT, RL, and MOPD post-training, achieving 5x higher inference throughput than other state-of-the-art open LLMs with on-par accuracy.
Problem
Long-running autonomous agents require inference that is both fast and accurate, motivating improvements along the inference-throughput-to-accuracy frontier.
Method
Nemotron 3 Ultra combines a Mixture-of-Experts hybrid Mamba-Attention architecture with SFT, RL, and MOPD post-training.
Results
5x higher inference throughput than other state-of-the-art open LLMs is achieved while maintaining on-par accuracy.
Takeaways & Limitations
The model and its pre-trained, post-trained, and quantized checkpoints, along with training data, are open-sourced on HuggingFace.
Takeaways & Limitations
MOPD’s effectiveness is limited on long-horizon agentic tasks because most experiments use single-turn rollouts, leaving end-to-end rollouts open for exploration.
Abstract
from arXiv · showhide
We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 trillion text tokens, then extended the context length to 1M tokens, and post-trained using Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD). Nemotron 3 Ultra is our most capable model yet, employing multiple key technologies - LatentMoE, Multi Token Prediction (MTP), NVFP4 pre-training, multi-environment RLVR, MOPD, and reasoning budget control. Nemotron 3 Ultra achieves up to ~6x higher inference throughput as compared to state-of-the-art publicly available LLMs while attaining on-par accuracy. The state-of-the-art accuracy, high inference throughput, and 1M token context length make Nemotron 3 Ultra ideal for long-running autonomous agentic tasks. We open-source the base, post-trained, and quantized checkpoints, along with the training data and recipe on HuggingFace.
NVIDIA · 1. Introduction
Nemotron 3 Ultra is presented as the largest and most capable Nemotron 3 model, combining a Mixture-of-Experts hybrid Mamba-Attention architecture with NVFP4 pre-training and agent-focused post-training. NVIDIA releases its base, post-trained, and quantized checkpoints together with training recipes, data, and RL environments.
- NVIDIA: Nemotron 3 Ultra is presented as the largest and most capable model in the Nemotron 3 family.
- 1. Introduction: Its Mixture-of-Experts hybrid Mamba-Attention architecture targets improved inference-throughput-to-accuracy tradeoffs for long-running agents.
- 1. Introduction: Nemotron 3 Ultra 550B-A55B Base was pretrained with NVFP4 pre-training, LatentMoE, and Multi Token Prediction.
- 1. Introduction: 20 trillion text tokens were used for NVFP4 pretraining under a Warmup-Stable-Decay learning-rate schedule.Pretraining used 15 trillion tokens in a diversity- and broad-coverage-focused first phase, followed by 5 trillion tokens focused on high-quality data.
- 1. Introduction: The agent-focused post-training pipeline combines curated SFT, unified RLVR across reasoning and agentic environments, specialized teachers, and Multi-teacher On-Policy Distillation.It targets long-horizon reasoning, tool use, and autonomous task completion.
- Data: NVIDIA also open-sources the training recipes, data, and RL environments alongside the checkpoints.The accompanying materials include fresh GitHub code data, synthetic legal datasets, specialized pretraining data, and post-training datasets.
2. Pretraining
Nemotron 3 Ultra’s pretraining combines a 550B-parameter hybrid Mamba-Attention MoE architecture, NVFP4 training, extended synthetic and domain-specific data, and stability interventions. Ablations show these data additions improve several benchmark accuracies, while training analyses identify and mitigate divergence and routing instabilities.
- Architecture: Nemotron 3 Ultra scales a hybrid Mamba-Attention MoE to 550B total and 55B active parameters per token, using LatentMoE and two shared-parameter MTP heads.The MTP heads support autoregressive drafting for inference acceleration.
- NVFP4 pretraining: The model uses an NVFP4 pretraining recipe with E2M1, two-dimensional weight block quantization, Random Hadamard input transforms, and stochastic gradient rounding.The final 15% of the network and selected projections, MTP, and embedding layers were retained in higher precision.
- Pretraining data: Benchmark-oriented synthetic Q&A data improved MMLU-Pro from 64.8 to 66.6, average code from 73.2 to 75.1, commonsense understanding from 72.9 to 74.5, and GPQA from 30.8 to 41.9.Average math remained stable, changing from 87.6 to 87.9, in a 100B-token continued-pretraining ablation.
- Pretraining data: Fact-seeking data improved SimpleQA from 40.24 to 50.16, while legal-specific data increased average LegalBench accuracy from 64.6 to 74.7.The SimpleQA evaluation converted questions to multiple-choice format, so its scores are not directly comparable to original SimpleQA scores.
- Training stability: Training divergences occurred around 8T and 16T tokens; reducing output-layer gradient accumulation precision contributed to the first, while learning-rate annealing after a 15T-token rollback mitigated the second.The practical pretraining horizon was reduced to 20T tokens.
3. Post-Training
Nemotron 3 Ultra uses a substantially redesigned post-training pipeline that begins with general supervised fine-tuning and adds Multi-teacher On-Policy Distillation (MOPD) for broad capability acquisition and targeted specialization. The section covers student preparation, specialized teacher training, iterative MOPD optimization, and MTP Boosting.
- Pipeline overview: The redesigned pipeline starts with general Supervised Fine-tuning and augments consecutive reinforcement learning with Multi-teacher On-Policy Distillation.MOPD is intended to support both broad capability acquisition and targeted specialization.
- Student preparation: Student preparation comprises Supervised Fine-tuning, RLVR, and MOPD warmup.The section organizes the student-preparation process around these three stages.
- Teacher training and optimization: The section next presents training specialized teacher models before detailing iterative MOPD optimization.This sequence follows the student-preparation stages in the described post-training organization.
- MTP Boosting: The post-training procedures conclude with MTP Boosting.The passage identifies MTP Boosting as part of the iterative optimization procedures.
MOPD … 3.2. Reinforcement Learning
Nemotron 3 Ultra’s post-training combines two-stage SFT with long-context, efficiency-control, safety, multilingual, and agentic data, followed by unified RLVR across diverse environments. The pipeline retains MTP during SFT, uses length-aware data packing, and trains RL with asynchronous GRPO-style optimization, large batches, and multiple rollouts.
- MTP Boosting: Two MTP layers remain active during SFT with a per-token auxiliary-loss scaling factor of 0.1.The shared-weight MTP objective is retained during post-training.
- 3.1.1. Data: SFT data target 512K long-context abilities, reasoning efficiency and control, robust safety, search, terminal use, conversational tools, software issue resolution, math, science, chat, code, CUDA, RTL, and multilingual tasks.The data include GPT-OSS-120B medium-effort reasoning samples, randomly truncated reasoning budgets, a 45K safety blend, and diverse synthetic agent trajectories.
- 3.1.1. Data: Search training combines 4–8-hop Wikidata-based trajectories, over 97K offline OpenResearcher trajectories, and challenging BrowseComp samples requiring 50–100 searches.Teachers include MiniMax 2.1, gpt-oss-120b, MiniMax 2.5, and GLM 5.1 across the search datasets.
- 3.1.1. Data: The multilingual pipeline processes full JSON objects end-to-end, addressing quality issues identified in an earlier line-by-line translation pipeline.The multilingual mixture combines sentence-level parallel corpora with synthetic English SFT translations in math, code, and science.
- 3.1.2. Data Packing: Length-aware best-fit packing combines conversations up to the maximum context length while using fixed-size in-memory pools, round-robin source interleaving, and in-pack deduplication.A final shuffle mixes completed packs across the data distribution.
- 3.2. Reinforcement Learning: Unified RLVR spans terminal use, productivity, software engineering, search, tool-calling, math, code, STEM, safety, chat, instruction following, long-context QA, reasoning, structured outputs, and general usability.The training data are refreshed from recent collections and reward-profiled before training.
- 3.2. Reinforcement Learning: RL uses a Gaussian-based mixture and curriculum, asynchronous GRPO with stability optimizations, global batch size 8192, and 16 rollouts per sample.The procedure also incorporates improvements to the training infrastructure described in Section 3.6.
3.3. MOPD
MOPD distills dense, domain-specific teacher signals into an RLVR student through asynchronous, iterative on-policy training. It improves performance broadly, especially on agentic tasks, while exposing limitations from student–teacher trajectory overlap and long-horizon training efficiency.
- Teacher specialization: More than ten specialized teachers provide domain-specific supervision to address diluted learning signals in mixed-environment RLVR.Teachers are optimized for individual capability areas, including agentic safety and STEM reasoning.
- Training procedure: MOPD asynchronously pipelines student rollouts, teacher scoring, and optimization across multiple iterative cycles.The RLVR student generates rollouts across domains, receives dense rewards from corresponding teachers, and supplies updated checkpoints for later teacher training.
- Results: MOPD improves over the RLVR student across the evaluation suite, with strong recovery on agentic and instruction-following or factuality benchmarks.Recovery rate is defined as (MOPD2 − RLVR)/(Teacher − RLVR), measuring the fraction of the teacher–student gap closed by MOPD.
- Limitations: Gains are smaller on self-contained reasoning benchmarks, especially HLE, because the student has not directly seen the additional reasoning data used to improve the teacher.The authors interpret this as a limitation of on-policy distillation rather than a failure of the teacher.
- Limitations: Long-horizon agentic and reasoning mixtures create substantial training inefficiency because rollout times can differ dramatically.The paper identifies sophisticated infrastructure and careful asynchronous algorithm design as necessary to balance efficiency and accuracy.
3.4. MTP Boosting
MTP Boosting addresses train–inference mismatch in Nemotron 3 Ultra’s native speculative-decoding head by exposing it to inference-like hidden-state noise and distilling the backbone’s token distribution. On SPEED-Bench, it consistently improves acceptance length and speculative-decoding speedup over base MTP.
- MTP mechanism: The MTP head predicts multiple future tokens, whose drafts are verified against the backbone so several tokens can be accepted per verification step.A shared head is applied recursively across several MTP steps, increasing the draft horizon without additional heads.
- Train-Inference Mismatch: Teacher-forced MTP training differs from autoregressive inference, causing later draft positions to condition on increasingly noisy mixtures of target-model and generated hidden states.This distribution mismatch degrades acceptance at deeper draft positions.
- Training Procedure: MTP Boosting samples MTP inputs from hidden states produced at earlier MTP steps, exposing training to inference-like noise and improving robustness at longer draft lengths.The procedure continues training from the MOPD checkpoint.
- Data: The head is trained for 12K steps on on-policy rollouts from general-purpose and agentic seed datasets, using temperature 1, global batch size 64, and sequences capped at 8K tokens.The loss is accumulated over assistant responses.
- Loss: The boosting objective uses temperature-scaled forward-KL against backbone logits, disabling gold-token cross-entropy so the head matches the backbone’s full distribution.The loss is defined over assistant-token positions and MTP generation steps.
- Results: MTP-Boosting consistently increases average acceptance length over base MTP, improving relative speculative-decoding speedup by 3.15% on summarization tasks and 5.82% on coding tasks.The results use draft length 7 on the SPEED-Bench qualitative split; main values use greedy decoding, with parenthesized values using temperature sampling at temp = 1.
3.5. Reasoning Efficiency and Control · 3.6. Infrastructure
Nemotron 3 Ultra provides controllable reasoning modes and inference-time budgets spanning accuracy–efficiency trade-offs, while infrastructure optimizations accelerate rollout generation and support resilient RL/MOPD training. Its production system combines speculative Multi-Token Prediction with large-scale cluster orchestration and instrumentation.
- 3.5. Reasoning Efficiency and Control: Three reasoning modes—reasoning-off, regular, and medium-effort—combine with inference-time budget control to cover accuracy–efficiency trade-offs.These controls complement task-level limits such as turn-limits in agentic applications.
- 3.5. Reasoning Efficiency and Control: 2.5% of RLVR prompts use medium-effort reasoning across math, STEM, and coding, with length-based reward adjustments.The medium-effort mode is introduced during SFT and optimized during RLVR; its effects generalize beyond those domains.
- 3.5. Reasoning Efficiency and Control: 2.5X fewer tokens are used by medium-effort reasoning than regular reasoning on average, at approximately 7% lower accuracy.Figure 11 compares these modes using Artificial Analysis Intelligence Index V4 and relative verbosity across 10 tasks.
- 3.6.1. Accelerating Rollout Generation with Multi-Token Prediction: MTP proposes k candidate tokens recurrently and lets the base model verify them in one forward pass, committing accepted tokens without extra sequential decoding.This speculative decoding accelerates rollout generation during RL and MOPD.
- 3.6.1. Accelerating Rollout Generation with Multi-Token Prediction: 1.46× faster rollout generation is achieved with MTP at k=5 versus the k=0 no-MTP baseline.The sweep evaluates k∈{0, 3, 5, 7}, and the benefit is concentrated in the slowest long-tail generations.
- 3.6.1. Accelerating Rollout Generation with Multi-Token Prediction: Long-tail generations benefit most from MTP because they emit more tokens and decode at lower concurrency near the end of the batch.These factors make speculative decoding particularly effective for the slowest generations.
- 3.6.2. Scaling RL Infrastructure: The production RL cluster uses NVIDIA GB200 nodes, Slurm orchestration, and co-located CPUs for sandbox execution.Generation and sandbox/tool-calling failures account for approximately 92% of observed RL software failures.
- 3.6.2. Scaling RL Infrastructure: High resiliency required sustained engineering through systematic instrumentation and infrastructure optimizations reviewed in the paper.The infrastructure section summarizes key optimizations and their impact.
Ray GCS Scalability and Slurm Launch Overheads … Multi-Node vLLM Operational Stability
The paper scales heterogeneous RL infrastructure by addressing Ray and Slurm startup bottlenecks, enforcing topology- and NUMA-aware placement, reducing checkpoint and JIT delays, and hardening multi-node vLLM initialization. These changes improve throughput and startup reliability across large GB200 deployments.
- Ray GCS Scalability and Slurm Launch Overheads: At 3K+ GPU scale, Ray’s single-threaded GCS was overwhelmed by actor registrations, causing 25-49-minute startups and thundering-herd failures; converting actors to tasks eliminated 40% of registrations.The heterogeneous RL job also required separate Slurm srun invocations for training, generation, environment, and judge roles, with each request processed serially by slurmctld.
- Topology-Aware NVLink Domain Placement: 20% end-to-end throughput improvement on GB200 came from domain-aware rank assignment that co-located each Expert Parallelism group within one NVLink domain, keeping MoE all-to-all traffic on NVLink.ClusterUUID-based custom resources, topology-ranked bundle ordering, and Ray-controlled GPU mapping prevented EP groups from spanning racks and using InfiniBand.
- NUMA Binding for Policy and vLLM Workers: 10% end-to-end throughput improvement on GB200 came from binding policy and vLLM workers to the CPU socket local to their assigned GPUs.Local NUMA placement keeps optimizer offloading, tokenization, preprocessing, and pinned-memory allocations on local DRAM through the socket-local C2C path.
- JIT Cache and Initialization: 99% reduction in JIT initialization time reduced the warm-cache init phase from 38.8 minutes to 0.4 minutes at 1K GPU scale.Persistent shared caches, node-local extraction into /tmp, and container-baked FlashInfer cubins avoided redundant cold-start compilation and shared-filesystem contention.
- Multi-Node vLLM Operational Stability: Multi-node vLLM startup failures arose from ABI mismatches, divergent subprocess environments, and JIT kernels incompatible with NCCL’s multi-node NVLink memory registration.Each vLLM data-parallel leader spawned EngineCore subprocesses that independently called ray.init(), adding GCS connections during startup.
- Multi-Node vLLM Operational Stability: Operational fixes unified GPU-kernel dependencies, propagated environment variables to spawned subprocesses, and disabled affected NVLink memory registration until an upstream FlashInfer fix.These changes directly addressed library incompatibility, environment divergence, and collective-communication hangs during distributed initialization.
- Multi-Node vLLM Operational Stability: vLLM health checks, TP-NCCL communication validation, RPC timeouts, graceful shutdown, and orphan-process cleanup improved stability across multi-node jobs.The safeguards detect invalid device or communication states and prevent stalled workers or residual processes from destabilizing subsequent jobs.
Container and Storage I/O · 3.7. Post-trained Model Evaluations
The paper addresses large-scale container and cache I/O bottlenecks with local caching and asymmetric persistence paths, then evaluates Nemotron 3 Ultra across broad agentic, reasoning, conversational, long-context, and multilingual suites. Results show strong, competitive post-trained performance, including held-out agentic validation, tool-integrated reasoning, million-token context support, and test-time scaling for Olympiad mathematics.
- Container and Storage I/O: ∼44 GB container images generated tens of terabytes of concurrent startup reads, causing I/O errors or extraction stalls exceeding 12 minutes.Normal extraction took ∼2 to 3 minutes, and a single slow node delayed Ray initialization.
- Container and Storage I/O: Container caching reuses locally extracted squashfs images on warm nodes, effectively eliminating their shared-storage read load.Enroot’s local squashfs cache supports reuse across subsequent jobs.
- Container and Storage I/O: Node-local JIT writes and single-sidecar archival replace simultaneous shared-storage writes, while startup restores caches through one sequential read per cache type.Because nodes compile identical kernels, only one node’s cache needs persistence.
- 3.7.1. Evaluation Setup: Evaluations span agentic capabilities, reasoning and knowledge, conversation and instruction following, long-context understanding, and multilingual performance under standardized settings.Results were collected through the Nemo Evaluator SDK and three main harnesses: Nemo Gym, Nemo Skills, and Harbor.
- 3.7.2. Evaluation Results: Nemotron 3 Ultra performs strongly across terminal execution, productivity, software engineering, real-world agents, deep research, and conversational tool use while remaining competitive with leading open models.The comparison includes models with substantially larger total parameter counts.
- 3.7.2. Evaluation Results: PinchBench and ProfBench serve as held-out generalization gates, evaluated once after finalization to test transfer to unseen agentic settings.They were excluded from training monitoring, checkpoint selection, and other development decisions.
- 3.7.2. Evaluation Results: 570.0 on IOI 2025 and 92.3 on IMOAnswerBench with tools demonstrate strong competitive-programming and tool-integrated mathematical reasoning, while 78.7 is the highest non-hallucination score on AA-Omniscience.The model also supports contexts of up to 1M tokens and performs competitively on million-token evaluation suites.
- 3.7.3. Test-time Scaling in Math Olympiad Problems: A generate–verify–refine pipeline evaluates Nemotron 3 Ultra with high-compute search-based test-time scaling on IMO-ProofBench Advanced, IMO 2025, Putnam 2025, and USAMO 2026.Accuracy is reported with corresponding graded scores in Table 11, and refinement rounds serve as a proxy for compute.
4. Quantization
Nemotron 3 Ultra is post-training-quantized for efficient Blackwell inference using a mixed-precision NVFP4 recipe selected through BPE sweeps and benchmark evaluation. The chosen 5.03-BPE configuration recovers long-context performance while preserving accuracy and deployment efficiency.
- Quantization approach: Post-training quantization maps the Ultra checkpoint to NVFP4 for efficient inference on NVIDIA Blackwell GPUs, beginning with a heuristic mixed per-layer precision recipe.The recipe is refined using Model-Optimizer sensitivity analysis, BPE ablations, and FP4 weight-quantization algorithm comparisons.
- BPE selection: The BPE sweep evaluates averaged pass@1 scores across seven benchmarks covering coding, scientific reasoning, instruction following, knowledge, and long-context reasoning.Evaluations are served through the Nemo Evaluator SDK on vLLM v0.20.0, with increased repeats to suppress run-to-run variance.
- BPE selection: 5.03 BPE (NVFP4 with mixed-FP8) is selected because it is the smallest tested budget that recovers long-context performance without measurable higher-precision gains.AA-LCR improves by +2.4 points from 4.85 to 5.03 BPE, then plateaus at 64.2–65.0 through 7.19 BPE; other capabilities remain flat within run-to-run noise.
- FP4 scale selection: At 5.03 or 5.43 BPE, MSE weight scaling slightly outperforms max scaling, while Four-Over-Six improves the balanced 5.03-BPE setting but degrades substantially at 4.85 BPE.Four-Over-Six reduces median relative MSE of routed-expert weight reconstruction by 16.4% versus the cited baseline.
- Final recipe: The final checkpoint operates at 5.03 BPE and combines NVFP4 routed-expert GEMMs with FP8 GEMMs for shared experts and Mamba linear layers.The combined recipe is designed to reduce naive NVFP4 accuracy loss while preserving runtime efficiency for deployment.
- Mamba cache quantization: Block-scaled INT8 quantization with stochastic rounding largely preserves FP32-cache accuracy and verbosity, whereas FP8 E4M3 degrades accuracy for Mamba caches.The authors identify mantissa precision and stochastic rounding as key ingredients for preserving accuracy after Mamba cache quantization.
5. Inference
Nemotron 3 Ultra’s inference design combines LatentMoE, hybrid Mamba-2 with sparse global Attention anchors, and Multi-Token Prediction to improve scaling, KV-cache efficiency, and speculative decoding. Performance depends on workload balance and batch size, motivating specialized serving and communication strategies.
- Inference-aware architecture: LatentMoE trades hidden-dimension width for more routed experts at fixed inference cost, while hybrid Mamba-2 provides sub-quadratic prefill scaling and bounded decode KV-cache size.The architecture also uses sparse global Attention anchors and Multi-Token Prediction for native speculative decoding.
- Speculative decoding: Speculative-decoding settings depend on batch size: high draft lengths favor latency at small batches, whereas lower lengths or disabling MTP often favor throughput at large batches.Weight reads dominate per-pass cost at small batches; compute and verification overhead dominate at large batches.
- Speculative decoding: SSM-state snapshots at every draft step enable rollback after rejected tokens, and coarser snapshots support longer-horizon state recovery.Unlike Attention KV truncation, the Mamba SSM state is a single fixed-size entry per sequence overwritten at every token.
- Workload-dependent throughput: 1.6× relative throughput over Qwen-3.5 on decode-heavy 8K/64K workloads, while trailing Qwen-3.5 on prefill-heavy 50K/2K workloads.Figure 15 normalizes both settings to GLM-5.1 and attributes the contrast to active-parameter versus total-weight-I/O behavior.
- Speculative decoding: 2.89× speedup over the no-MTP baseline occurs at draft length DL=6 on a single-user 10K/16K workload before verification overhead causes throughput to decline.The measurement uses one GB200 node with TP=4 and acceptance lengths from SPEED-Bench.
- Serving infrastructure: 15% to 20% of runtime is spent on routed-expert all-to-all communication in prefill-heavy 50K/2K workloads, motivating a less wasteful backend than activation AllGather plus output ReduceScatter.The default vLLM scheme replicates every token on every rank during AllGather.
6. Conclusion
The conclusion presents Nemotron 3 Ultra as a 550-billion-total, 55-billion-active-parameter MoE Hybrid Mamba-Attention model that combines efficient inference with on-par accuracy. It was trained on 20 trillion text tokens, post-trained with SFT, RL, and MOPD, and released with checkpoints and training data.
- Model and training: Nemotron 3 Ultra has 550 billion total and 55 billion active parameters.It uses a Mixture-of-Experts Hybrid Mamba-Attention architecture with LatentMoE and MTP.
- Efficiency and accuracy: 5x higher inference throughput than other state-of-the-art open LLMs is achieved with on-par accuracy.The architecture incorporates LatentMoE and Multi Token Prediction for inference and accuracy.
- Model and training: 20 trillion text tokens were used for pre-training, followed by SFT, RL, and MOPD post-training.The training pipeline includes supervised fine-tuning, reinforcement learning, and Multi-teacher On-Policy Distillation.
- Release: Pre-trained, post-trained, and quantized checkpoints, along with the training data, are open-sourced.The release covers the Nemotron 3 Ultra model artifacts and associated training data.
A. Post-training Evaluations · A.1. Benchmark Details
The benchmark details specify evaluation settings for TauBench V3, ProfBench, BrowseComp, and the Vals.ai Finance Agent Benchmark. These settings include user simulation, tools, task domains, evidence handling, and evaluation sample selection.
- A.1. Benchmark Details: TauBench V3 adds an additional prompt to the user simulation across all domains to mitigate premature termination.Banking uses the terminal_use setting, which lets the agent search a knowledge base through a terminal tool.
- A.1. Benchmark Details: TauBench V3 evaluates DeepSeek-V4 models under max reasoning with GPT-5.2 as the user simulator at low reasoning effort.Results are averaged across 8 trials.
- A.1. Benchmark Details: ProfBench evaluates deep research capability across Finance MBA, Consulting MBA, Chemistry PhD scientific research, and Physics PhD scientific research.Tasks reflect real-world professional workflows and are judged using rubric criteria annotated by domain professionals.
- A.1. Benchmark Details: ProfBench evaluations enable both search and browse tools so models can identify relevant context.
- A.1. Benchmark Details: BrowseComp uses a custom agentic search harness with Tavily search and browsing, terminal access, and a per-task disk workspace for retrieved content.The model receives metadata and snippets, then selectively inspects saved pages with shell commands such as grep, head, and sed.
- A.1. Benchmark Details: The Vals.ai Finance Agent Benchmark tests entry-level financial analyst tasks across nine categories, ranging from retrieval to financial modeling and market analysis.The evaluation uses 200 questions: 50 publicly available validation samples plus 150 additional samples from the privately licensed validation set.
A.2. Harness Robustness
Harness Robustness strengthens post-training by spanning five task-distribution verticals and training each distribution under at least two harnesses. This avoids reliance on a single harness, improving generalization and robustness across dynamic real-world execution contexts.
- Task distributions: Harness Robustness is a key component of Nemotron-3 Ultra’s post-training stage, organizing tasks into five verticals.The verticals include terminal use and software engineering, existing-repository bug fixing, office and general productivity, and general or multi-domain knowledge tasks.
- Multi-harness training: Each task distribution is trained under at least two harnesses rather than a single harness.This design targets robustness across varying execution setups.
- Robustness outcome: Training across multiple harnesses improves generalization and robustness in dynamic execution contexts used in real-world settings.Nemotron 3 Ultra’s performance across different harnesses is shown in Figure 17.