Source-linked AI summary
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aakshita Chandiramani, Aaron Blakeman, Abdullahi Olaoye, Abhibha Gupta, Abhilash Somasamudramath, Abhinav Khattar, Adeola Adesoba, Adi Renduchintala, Adil Asif, Aditya Agrawal, Aditya Vavre, Ahmad Kiswani, Aishwarya Padmakumar, Ajay Hotchandani, Akanksha Shukla, Akhiad Bercovich, Aleksander Ficek, Aleksandr Shaposhnikov, Alex Gronskiy, Alex Kondratenko, Alex Neefus, Alex Steiner, Alex Yang, Alexander Bukharin, Alexander Young, Ali Hatamizadeh, Ali Taghibakhshi, Alina Galiautdinova, Alisa Liu, Alok Kumar, Ameya Sunil Mahabaleshwarkar, Amir Klein, Amit Zuker, Amnon Geifman, Anahita Bhiwandiwalla, Ananth Subramaniam, Andrew Tao, Anjaney Shrivastava, Anjulie Agrusa, Ankur Srivastava, Ankur Verma, Ann Guan, Anna Shors, Annamalai Chockalingam, Anubhav Mandarwal, Aparnaa Ramani, Arham Mehta, Arti Jain, Arun Venkatesan, Asha Anoosheh, Ashwath Aithal, Ashwin Poojary, Asif Ahamed, Asit Mishra, Asli Sabanci Demiroz, Asma Kuriparambil Thekkumpate, Atefeh Sohrabizadeh, Avinash Kaur, Ayush Dattagupta, Barath Subramaniam Anandan, Bardiya Sadeghi, Barnaby Simkin, Ben Lanir, Benedikt Schifferer, Benjamin Chislett, Besmira Nushi, Bilal Kartal, Bill Thiede, Bita Darvish Rouhani, Bobby Chen, Boris Ginsburg, Brandon Norick, Branislav Kisacanin, Brian Yu, Bryan Catanzaro, Buvaneswari Mani, Carlo del Mundo, Chankyu Lee, Chanran Kim, Chantal Hwang, Chao Ni, Charles Wang, Charlie Truong, Cheng-Ping Hsieh, Chenhan Yu, Chenjie Luo, Cherie Wang, Chetan Mungekar, Chintan Patel, Chris Alexiuk, Chris Holguin, Chris Wing, Christian Munley, Christopher Parisien, Chuck Desai, Chunyang Sheng, Collin Neale, Cyril Meurillon, Dakshi Kumar, Dan Gil, Dan Su, Dane Corneil, Daniel Afrimi, Daniel Burkhardt Eliuth Triana, Daniel Egert, Daniel Fatade, Daniel Lo, Daniel Rohrer, Daniel Serebrenik, Daniil Sorokin, Daria Gitman, Daria Levy, Darko Stosic, David Edelsohn, David Messina, David Mosallanezhad, David Tamok, Deena Donia, Deepak Narayanan, Devin O'Kelly, Dheeraj Peri, Dhruv Nathawani, Di Wu, Dima Rekesh, Dina Yared, Divyanshu Kakwani, Dmitry Konyagin Brandon Tuttle, Dong Ahn, Dongfu Jiang, Dorrin Poorkay, Douglas O'Flaherty, Duncan Riach, Dusan Stosic, Dustin Van Stee, Edgar Minasyan, Edward Lin, Eileen Peters Long, Elad Segal, Elena Lantz, Elena Lewis, Ellie Evans, Elliott Ning, Eric Chung, Eric Harper, Eric Pham-Hung, Eric W. Tramel, Erick Galinkin, Erik Pounds, Esti Etrog, Evan Briones, Evan Wu, Evelina Bakhturina, Evgeny Tsykunov, Ewa Dobrowolska, Farshad Saberi Movahed, Farzan Memarian, Fay Wang, Fei Jia, Felipe Soares, Felipe Vieira Frujeri, Feng Chen, Fengguang Lin, Ferenc Galko, Fortuna Zhang, Frankie Siino, Frida Hou, Gantavya Bhatt, Gargi Prasad, Geethapriya Venkataramani, Geetika Gupta, George Armstrong, Gerald Shen, Giulio Borghesi, Gordana Neskovic, Gorkem Batmaz, Grace Lam, Grace Wu, Greg Pauloski, Greyson Davis, Grigor Nalbandyan, Guoming Zhang, Guy Farber, Guyue Huang, Haifeng Qian, Haran Kumar Shiv Kumar, Harry Kim, Harsh Sharma, Hayate Iso, Hayley Ross, Herbert Hum, Herman Sahota, Hexin Wang, Himanshu Soni, Hiren Upadhyay, Huy Nguyen, Iain Cunningham, Ido Galil, Ido Shahaf, Igino Padovani, Igor Gitman, Igor Shovkun, Ikroop Dhillon, Ilya Loshchilov, Ingrid Kelly, Itamar Schen, Itay Levy, Ivan Moshkov, Izik Golan, Izzy Putterman, Jain Tu, Jan Baczek, Jan Kautz, Jane Polak Scowcroft, Janica Rosenberg, Jared Casper, Jarrod Pflum, Jason Grant, Jason Sewall, Jatin Mitra, Jeffrey Glick, Jenny Chen, Jesse Oliver, Jiacheng Xu, Jiafan Zhu, Jialin Song, Jian Zhang, Jiaqi Zeng, Jie Lou, Jill Milton, Jim Chow, Jimmy Zhang, Jinhang Choi, Jining Huang, Jocelyn Huang, Joel Caruso, Joey Conway, Joey Guman, Johan Jatko, John Kamalu, Johnny Greco, Jonathan Cohen, Jonathan Raiman, Joseph Jennings, Joyjit Daw, Juan Yu, Julio Tapia, Junkeun Yi, Jupinder Parmar, Jyothi Achar, Kari Briski, Kartik Mattoo, Katherine Cheung, Katherine Luna, Keith Wyss, Kevin Shih, Kezhi Kong, Khanh Nguyen, Khushi Bhardwaj, Kirill Buryak, Kirthi Shankar Sivamani, Konstantinos Krommydas, Kris Murphy, Krishna C. Puvvada, Krzysztof Pawelec, Kumar Anik, Laikh Tewari, Laya Sleiman, Leo Du, Leon Derczynski, Li Ding, Lilach Ilan, Lingjie Wu, Lizzie Wei, Luis Vega, Lun Su, Maarten Van Segbroeck, Maer Rodrigues de Melo, Magaret Zhang, Mahan Fathi, Makesh Narsimhan Sreedhar, Makesh Sreedhar, Makesh Tarun Chandran, Manuel Reyes Gomez, Maor Ashkenazi, Marc Cuevas, Marc Romeijn, Margaret Zhang, Mark Cai, Mark Gabel, Markus Kliegl, Martyna Patelka, Maryam Moosaei, Matthew Varacalli, Matvei Novikov, Mauricio Ferrato, Mehrzad Samadi, Melissa Corpuz, Meng Xin, Mengdi Wang, Mengru Wang, Meredith Price, Micah Schaffer, Michael Andersch, Michael Boone, Michael Evans, Michael Z Wang, Miguel Martinez, Mikail Khona, Mike Chrzanowski, Mike Hollinger, Mingyuan Ma, Minseok Lee, Mohammad Dabbah, Mohammad Shoeybi, Mostofa Patwary, Nabin Mulepati, Nader Khalil, Najeeb Nabwani, Nancy Agarwal, Nanthini Balasubramaniam, Narimane Hennouni, Narsi Kodukula, Natalie Hereth, Nathaniel Pinckney, Nave Assaf, Negar Habibi, Nestor Qin, Neta Zmora, Netanel Haber, Nick Reamaroon, Nickson Quak, Nidhi Bhatia, Nikhil Jukar, Nikki Pope, Nikolai Ludwig, Nima Tajbakhsh, Nir Ailon, Nirmal Juluru, Nirmalya De, Nowel Pitt, Oleg Rybakov, Oleksii Hrinchuk, Oleksii Kuchaiev, Olivier Delalleau, Oluwatobi Olabiyi, Omer Ullman Argov, Omri Almog, Omri Puny, Oren Tropp, Otavio Padovani, Ouye Xie, Parth Chadha, Pasha Shamis, Paul Gibbons, Pavlo Molchanov, Peter Belcak, Peter Jin, Pinky Xu, Piotr Januszewski, Pooya Jannaty, Prachi Shevate, Pradeep Thalasta, Pranav Prashant Thombre, Prasoon Varshney, Prerana Gambhir, Pritam Gundecha, Przemek Tredak, Qing Miao, Qiyu Wan, Quan Tran Minh, Rabeeh Karimi Mahabadi, Rachel Oberman, Rachit Garg, Rahul Kandu, Raina Zhong, Ran El-Yaniv, Ran Zilberstein, Rasoul Shafipour, Renee Yao, Renjie Pi, Richard Mazzarese, Richard Wang, Rick Izzo, Ridhima Singla, Rima Shahbazyan, Rishabh Garg, Ritika Borkar, Ritu Gala, Riyad Islam, Robert Clark, Robert Hesse, Roger Waleffe, Rohit Varma Kalidindi, Rohit Watve, Roi Koren, Ron Fan, Ruchika Kharwar, Ruisi Cai, Ruoxi Zhang, Russell J. Hewett, Ryan Prenger, Ryan Timbrook, Ryota Egashira, Sadegh Mahdavi, Sagar Singh Ashutosh Joshi, Sahil Modi, Samuel Kriman, Sandeep Pombra, Sanjay Kariyappa, Sanjeev Satheesh, Santiago Pombo, Saori Kaji, Satish Pasumarthi, Saurav Mishra, Saurav Muralidharan, Scott Hara, Sean Narenthiran, Sebastian Rogawski, Seonjin Na, Seonmyeong Bak, Sepehr Sameni, Seth Poulos, Shahar Mor, Shantanu Acharya, Shaona Ghosh Adam Lord, Sharath Turuvekere Sreenivas, Shaun Kotek, Shaya Gharghabi, Shelby Thomas, Sheng-Chieh Lin, Shibani Likhite, Shiqing Fan, Shiyang Chen, Shreya Gopal, Shrimai Prabhumoye, Shubham Pachori, Shubham Toshniwal, Shuo Zhang, Shuoyang Ding, Shyam Renjith, Shyamala Prayaga, Siddhartha Jain, Simeng Sun, Sirisha Rella, Sirshak Das, Smita Ithape, Sneha Harishchandra S, Somshubra Majumdar, Soumye Singhal, Sri Harsha Singudasu, Sriharsha Niverty, Stas Sergienko, Stefana Gloginic, Stefania Alborghetti, Stephen Ge, Stephen McCullough, Sugam Dipak Devare, Suguna Varshini Velury, Sukrit Rao, Sumeet Kumar Barua, Sunny Gai, Suseella Panguluri, Sushil Koundinyan, Swathi Patnam, Sweta Priyadarshi, Swetha Bhendigeri, Syeda Nahida Akter, Sylendran Arunagiri, Tailling Yuan, Talor Abramovich, Tan Bui, Tan Yu, Terry Kong, Thanh Do, Thomas Gburek, Thorgane Marques, Tiffany Moore, Tijmen Blankevoort, Tim Moon, Timothy Ma, Tiyasa Mitra, Tomasz Grzegorzek, Tomer Asida, Tomer Bar Natan, Tomer Keren, Tomer Ronen, Traian Rebedea, Trenton Starkey, Tugrul Konuk, Twinkle Vashishth, Tyler Condensa, Udi Karpas, Ushnish De, Vahid Noorozi, Vahid Noroozi, Vanshil Atul Shah, Veena Vaidyanathan, Venkat Srinivasan, Venmugil Elango, Victor Cui, Vijay Korthikanti, Vikas Mehta, Virginia Adams, Virginia Wu, Vitaly Kurin, Vitaly Lavrukhin, Vladimir Anisimov, Wan Seo, Wanli Jiang, Wasi Uddin Ahmad, Wei Du, Wei Ping, Wei-Ming Chen, Wendy Quan, Wenliang Dai, Wenwen Gao, Will Jennings, William Zhang, Xiaowei Ren, Xiaowen Xin, Xin Li, Yang Yu, Yangyi Chen, Yaniv Galron, Yashaswi Karnati, Yejin Choi, Yev Meyer, Yi-Fu Wu, Yian Zhang, Ying Lin, Yonatan Geifman, Yonggan Fu, Yoshi Suhara, Youngeun Kwon, Yuan Zhang, Yuki Huang, Zach Moshe, Zhilin Wang, Zhiyu Cheng, Zhongbo Zhu, Zhuolin Yang, Zihan Liu, Zijia Chen, Zijie Yan, Zuhair Ahmed
TL;DR
Large language models need higher capability without proportionally increasing inference cost, especially for long-context and agentic tasks. Nemotron 3 Super addresses this with a hybrid Mamba-Attention MoE model using LatentMoE, MTP, NVFP4 pre-training, and agent-focused post-training, achieving comparable or higher accuracy with substantially higher throughput. The authors also release the model’s training assets and checkpoints openly.
Problem
Existing MoE designs emphasize accuracy per FLOP but give less consideration to online latency, memory bandwidth, and communication constraints.
Method
Nemotron 3 Super combines a hybrid Mamba-Attention MoE architecture with LatentMoE, MTP layers, NVFP4 pre-training, and post-training across diverse reinforcement-learning environments.
Results
Comparable or higher accuracy was achieved with up to 2.2× higher throughput than GPT-OSS-120B, while the model was pre-trained on 25 trillion tokens.
Takeaways & Limitations
The released base, post-trained, and quantized checkpoints provide an openly available model designed for efficient long-context and agentic inference.
Takeaways & Limitations
Merge experiments explored only one merge schedule and fixed checkpoint granularity, leaving alternative longer-horizon merging strategies untested.
Abstract
from arXiv · showhide
We describe the pre-training, post-training, and quantization of Nemotron 3 Super, a 120 billion (active 12 billion) parameter hybrid Mamba-Attention Mixture-of-Experts model. Nemotron 3 Super is the first model in the Nemotron 3 family to 1) be pre-trained in NVFP4, 2) leverage LatentMoE, a new Mixture-of-Experts architecture that optimizes for both accuracy per FLOP and accuracy per parameter, and 3) include MTP layers for inference acceleration through native speculative decoding. We pre-trained Nemotron 3 Super on 25 trillion tokens followed by post-training using supervised fine tuning (SFT) and reinforcement learning (RL). The final model supports up to 1M context length and achieves comparable accuracy on common benchmarks, while also achieving up to 2.2x and 7.5x higher inference throughput compared to GPT-OSS-120B and Qwen3.5-122B, respectively. Nemotron 3 Super datasets, along with the base, post-trained, and quantized checkpoints, are open-sourced on HuggingFace.
1. Introduction
Nemotron 3 Super combines a 12 billion active, 120 billion total parameter MoE hybrid Mamba-Attention architecture with extensive pre-training and agent-focused post-training. It targets comparable benchmark accuracy with substantially higher inference throughput and releases its recipe, datasets, and checkpoints openly.
- Nemotron 3 Super is a 12 billion active, 120 billion total parameter MoE hybrid Mamba-Attention model.
- Up to 2.2× and 7.5× higher inference throughput were achieved than GPT-OSS-120B and Qwen3.5-122B, respectively, at 8k input and 64k output lengths.The figure reports comparable accuracies across popular benchmarks and measures throughput on B200 GPUs using vLLM and TRT-LLM.
- 25 trillion text tokens were used for pre-training, split between broad-coverage data and high-quality data focused on benchmark accuracy.The first phase used 20 trillion tokens and the second used 5 trillion tokens.
- The model combines scaled agentic training data, expanded reinforcement-learning environments, and asynchronous infrastructure for long-horizon tool-using tasks.This recipe improved software engineering, terminal use, and general tool-use benchmarks over Nemotron 3 Nano.
- The authors publicly share the training recipe, datasets, and multiple base, post-trained, and quantized model checkpoints.The report also organizes its technical coverage around pre-training, post-training, and quantization.
2. Pretraining
Nemotron 3 Super combines hybrid Mamba-Attention modeling, sparse LatentMoE scaling, and MTP-based speculative decoding to improve accuracy and inference efficiency. Its pretraining uses NVFP4 at 25T tokens, while experiments examine data additions, checkpoint merging, and quantization-related training behavior.
- Model architecture: The model combines hybrid Mamba-Attention layers with sparse LatentMoE scaling and MTP-based inference acceleration.Its architecture targets long-context performance and deployment efficiency through linear-time Mamba blocks, sparse expert capacity, and attention anchors.
- Multi-Token Prediction: 3.45 tokens per verification step is Nemotron 3 Super’s highest overall average acceptance length on SPEED-Bench with draft length 7.It outperforms DeepSeek-R1 across domains and remains competitive with Qwen3-Next.
- Multi-Token Prediction: Acceptance declines with draft depth, but Nemotron 3 Super maintains higher acceptance than DeepSeek-R1 and closely tracks or exceeds Qwen3-Next across most positions.The advantage becomes more pronounced at draft indices 4–7, indicating more stable longer recursive rollouts.
- Multi-Token Prediction: MTP shifts the throughput–latency frontier by delivering higher aggregate TPS at a given median user latency than MTP-disabled decoding.The improvement comes from increasing draft depth from D=1 to D=3 through native speculative decoding.
- NVFP4 pretraining: 7% of parameters had zero-valued weight gradients by the end of pretraining, associated with NVFP4 underflow in low-norm expert-layer channels.Switching a partially trained model to BF16 restored zero-gradient counts to baseline, while BF16 retained small gradients that NVFP4 quantized to zero.
- Pretraining data: 0.2B-token specialized datasets improved HumanEval, MBPP, and CRUXEval-O by 1–2 points when added to the final 100B tokens of pretraining.A separate 1B-token synthetic MCQ ablation improved MMLU from 77.22 to 77.51, MATH Level 5 from 78.55 to 79.05, AIME-2024 from 53.3 to 56.7, and MBPP from 74.8 to 75.2.
3. Post-Training
Nemotron 3 Super’s post-training uses a two-stage SFT procedure followed by multi-stage RL and MTP healing, with expanded agentic data and infrastructure. The resulting training recipe broadens agentic coverage, supports multiple reasoning modes, and scales data and environments substantially.
- Pipeline: Post-training follows SFT, three RL stages—RLVR, SWE-RL, and RLHF—and a final MTP-healing phase.The pipeline overview places SFT first, followed by RLVR, SWE-RL, RLHF, and MTP healing.
- Reinforcement learning: Large-scale asynchronous RL trains across 21 diverse environments and long-horizon software-engineering tasks.The infrastructure supports training on thousands of GPUs, covering varied interaction scenarios and multi-step tasks.
- Supervised fine-tuning: SFT expands agentic datasets and interaction scenarios while adding low-effort reasoning mode and preserving the existing chat template.A single-stage SFT degraded long-input-short-output performance, motivating the two-stage procedure.
- Supervised fine-tuning: The two-stage SFT objective first averages loss over output tokens, then weights conversations equally after per-conversation normalization.Stage 2 reduces the dominance of long outputs by normalizing each conversation by its own output-token count.
- Training data: The post-training corpus scales to 279,116 conversations across 838 domains, compared with 15,588 conversations across 5 domains for Nemotron 3 Nano.The data pipeline uses several teacher models and covers diverse agentic capabilities.
- Reasoning control: Nemotron 3 Super supports reasoning-off, regular, and low-effort modes, which can combine with inference-time budget control.These controls span different accuracy-efficiency trade-offs across application scenarios.
3.2. Reinforcement Learning
Reinforcement learning is organized into distinct stages for broad verifiable-reward training, long-horizon software engineering, instruction quality, and MTP recovery. The unified RLVR stage uses many environments simultaneously because single-environment training caused regressions elsewhere.
- RL stages: The RL pipeline comprises multi-environment RLVR, a separate SWE-RL stage, RLHF, and MTP healing.SWE-RL is isolated because its rollouts are slower and typically require longer contexts; MTP healing freezes the remaining weights.
- RL stages: Unified RLVR jointly optimizes the full environment mixture, while SWE-RL handles slower long-context software-engineering rollouts separately.RLHF is applied afterward to improve instruction-following behavior, robustness, and interaction quality.
- RLVR: Training on all environments simultaneously yields stable gains, whereas single-environment training causes severe regressions on other benchmarks.The RLVR strategy scales the number of environments relative to Nemotron 3 Nano.
- RLVR: RLVR spans 21 environments covering math, code, STEM, safety, chat, instruction following, long context, puzzles, and agentic tasks.The mixture filters consistently correct SFT prompts and orders remaining samples with a difficulty-based curriculum.
- Reasoning control: Low-effort RL prompts receive rewards based on both correctness and generated-token count, initially comprising 2% of RL prompts before being reduced to 1%.The low-effort mix begins with math, STEM QA, and competitive coding, then narrows to math and STEM QA.
RLVR Data
The RLVR data and infrastructure combine 21 environments and 37 datasets with asynchronous, multi-environment training for diverse reasoning and agentic tasks. PivotRL reuses expert SFT trajectories at uncertain assistant turns to improve agentic RL efficiency without the stated OOD degradation of SFT.
- RLVR data: The RL program uses 21 environments and 37 RL datasets spanning reasoning, math, code, STEM, safety, instruction following, long context, and agentic tool use.The environments include formal proof verification, jailbreak robustness, and conversational tool-use or terminal-use tasks.
- SWE-RL: SWE-RL launches containerized repositories through OpenHands and evaluates generated patches against ground-truth tests for binary rewards.OpenCode and Codex agent classes vary tool formats while reusing the OpenHands harness.
- PivotRL: PivotRL reuses offline SFT expert trajectories and trains on uncertain assistant turns using domain-appropriate rewards that credit similar actions.The paper reports improved agentic RL efficiency without the OOD degradation issues associated with SFT.
- Infrastructure: Asynchronous RL decouples generation from training, trading off rollout on-policy-ness for improved training efficiency.Training and generation are placed on separate workers, simplifying deployment and memory management.
3.3. Post-trained Model Evaluations
Nemotron 3 Super is evaluated across general reasoning, agentic, instruction-following, long-context, and multilingual benchmarks against GPT-OSS-120B and Qwen-3.5-122B-A10B. It is broadly competitive, stronger across agentic tasks against GPT-OSS 120B, and matches or exceeds baselines on multilingual benchmarks.
- Evaluation scope: The evaluation suite covers general knowledge, reasoning, agentic capability, instruction following, long context, and multilingual performance.Comparisons use officially reported values when available and otherwise follow the Nemotron 3 Nano evaluation procedure.
- Reasoning capabilities: Across reasoning benchmarks, Nemotron 3 Super is competitive with GPT-OSS-120B but lags behind Qwen-3.5-122B slightly.The reported reasoning tasks include AIME 25, HMMT Feb 25, GPQA, LiveCodeBench v5, SciCode, and HLE.
- Agentic capabilities: Across agentic benchmarks, Nemotron 3 Super outperforms or is competitive with GPT-OSS 120B and is competitive with Qwen 3.5 122B on some harnesses.Agentic evaluations include TerminalBench, SWE-Bench, TauBench V2, and BrowseComp.
- Chat and instruction following: Across chat and instruction-following benchmarks, Nemotron 3 Super is competitive with the baseline models.The evaluation includes IFBench, Multi-Challenge, and Arena-Hard V2.
- Multilingual capabilities: Nemotron 3 Super matches or outperforms the baseline models on both multilingual benchmarks.Multilingual evaluation uses MMLU-ProX and WMT24++ en→xx.
4. Quantization For Inference
The paper develops FP8 and NVFP4 deployment checkpoints with mixed-precision post-training quantization, including specialized handling for Mamba state caches. The final recipe preserves accuracy while targeting efficient inference on Hopper and Blackwell GPUs.
- Deployment checkpoints: FP8 (W8A8) targets Hopper, while NVFP4 (W4A4) targets Blackwell for efficient deployment.The checkpoints quantize weights and activations for hardware-specific inference.
- FP8 calibration: The FP8 checkpoint uses a 256-sample, 65536-context calibration subset and retains the KV cache in FP8 while storing the Mamba state cache in FP16.MoE GEMMs and Mamba Linear layers are quantized for the FP8 checkpoint.
- NVFP4 mixed precision: The NVFP4 recipe combines MSE-minimizing weight scales, dynamic max-based activation scales, and selective promotion of sensitive layers.AutoQuantize allocates higher precision layerwise to address sensitivity while preserving efficiency.
- Quantization outcome: 99.8% median accuracy relative to the BF16 baseline was achieved after a mixed-precision PTQ process completed in less than 2 hours on one 8-GPU B200 node.The process used 512 SFT samples at sequence length 4096 and retained near-FP4 performance.
- SSM cache: Up to 40% greater verbosity occurred when the SSM cache was directly cast to FP16 with W8A8, and up to 37% occurred even with BF16 weights and activations.These effects motivated stochastic rounding before FP16 casting.
- SSM cache: FP16 with Philox<5> stochastic rounding maintained accuracy and verbosity similar to the FP32 baseline and was selected for the SSM cache.The choice balances cache efficiency, statistical quality, and pseudorandom-number-generation overhead.
5. Conclusion
Nemotron 3 Super combines a sparse hybrid Mamba-Attention MoE architecture, LatentMoE, MTP layers, NVFP4 pre-training, and post-training on diverse RL environments. Its optimized FP8 and NVFP4 checkpoints deliver higher inference throughput while maintaining model accuracy, and the checkpoints are released on HuggingFace.
- Model: Nemotron 3 Super has 12B active and 120B total parameters in a hybrid Mamba-Attention MoE architecture.The model is designed for strong agentic capabilities.
- Architecture: LatentMoE improves accuracy, while MTP layers accelerate inference through speculative decoding.These are core architectural components of the model.
- Training: The model was pretrained on 25 trillion text tokens with low-precision NVFP4 and then post-trained on diverse RL environments.The conclusion describes the full training sequence.
- Results: Up to 2.2× higher throughput than GPT-OSS-120B was achieved while maintaining higher accuracy across a wide range of tasks.The result applies to the reported optimized model checkpoints.
- Release: Pre-trained, post-trained, and quantized Nemotron 3 Super checkpoints are released on HuggingFace.The release also includes the associated model artifacts described in the conclusion.
A. Per-Benchmark Merge Evaluation
Figure 17 compares per-benchmark accuracy for trained checkpoints with the best offline checkpoint merge across the full 25T-token training run. The comparison spans general knowledge, code generation, mathematical reasoning, and commonsense understanding.
- Evaluation setup: Figure 17 compares trained checkpoints against the best offline checkpoint merge across the full 25T-token training run.The figure reports per-benchmark accuracy.
- Benchmark coverage: The 12 benchmarks cover general knowledge, code generation, mathematical reasoning, and commonsense understanding.Named benchmark families include MMLU-Pro and MMLU, HumanEval and MBPP, GSM8K and MATH-500, and RACE, ARC-Challenge, HellaSwag, and Wino-.
B.1. PTQ Algorithm Ablation
The PTQ ablation evaluates accuracy across alternative quantization algorithms under a common NVFP4 setup. Except for the final classification and attention linear layers, the linear layers are quantized to NVFP4.
- Ablation: The ablation compares evaluation accuracy for various PTQ algorithms.The results are presented as a PTQ algorithm ablation.
- Experimental setup: All linear layers except the final classification layer and attention linear layers are quantized to NVFP4.This defines the shared quantization configuration for the experiments.
B.2. AutoQuantize Algorithm
AutoQuantize measures quantization sensitivity and performance cost per operator, then selects formats through a constrained optimization under a deployment budget.
- AutoQuantize measures each operator’s sensitivity at its immediate or closest available output using a second-order metric.For linear layers, the measurement point is the linear-layer output.
- The sensitivity metric compares quantized outputs with BF16 outputs using a local Hessian approximation.The Hessian is approximated diagonally and estimated empirically with the diagonal Fisher information matrix.
- AutoQuantize defines each quantization choice’s performance cost and solves a constrained optimization over operator formats.The optimization selects formats subject to a total deployment cost budget B.
B.2.1. Deployment-Restriction-Aware Search
AutoQuantize incorporates deployment restrictions by coupling operators that runtimes require to share formats, while measuring and costing each coupled group jointly.
- Linear layer fusion: Fused Q, K, and V projections within each layer share one quantization format and are modeled as one decision variable.Their sensitivity and cost are aggregated for the fused QKV projection.
- Linear layer fusion: The fused-output sensitivity remains additive across Q, K, and V under the stated second-order approximation and additive branch contributions.This is the assumption supporting the aggregated fused-operator formulation.
- MoE layer constraints: Sparse experts within each MoE layer share one quantization format, with each expert’s up_proj and down_proj assigned jointly.The coupling reflects quantized MoE API restrictions.
- MoE layer constraints: Sensitivity for a coupled MoE sparse-expert group is measured at the MoE block output, capturing the combined contribution of its sparse experts.Deployment cost is defined as the sum over sparse experts.
- MoE layer constraints: Latent projection layers and shared experts remain outside the sparse-expert coupling constraint and may use different quantization formats.