Source-linked AI summary
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aaron Blakeman, Aaron Grattafiori, Aarti Basant, Abhibha Gupta, Abhinav Khattar, Adi Renduchintala, Aditya Vavre, Akanksha Shukla, Akhiad Bercovich, Aleksander Ficek, Aleksandr Shaposhnikov, Alex Kondratenko, Alexander Bukharin, Alexandre Milesi, Ali Taghibakhshi, Alisa Liu, Amelia Barton, Ameya Sunil Mahabaleshwarkar, Amir Klein, Amit Zuker, Amnon Geifman, Amy Shen, Anahita Bhiwandiwalla, Andrew Tao, Ann Guan, Anubhav Mandarwal, Arham Mehta, Ashwath Aithal, Ashwin Poojary, Asif Ahamed, Asma Kuriparambil Thekkumpate, Ayush Dattagupta, Banghua Zhu, Bardiya Sadeghi, Barnaby Simkin, Ben Lanir, Benedikt Schifferer, Besmira Nushi, Bilal Kartal, Bita Darvish Rouhani, Boris Ginsburg, Brandon Norick, Brandon Soubasis, Branislav Kisacanin, Brian Yu, Bryan Catanzaro, Carlo del Mundo, Chantal Hwang, Charles Wang, Cheng-Ping Hsieh, Chenghao Zhang, Chenhan Yu, Chetan Mungekar, Chintan Patel, Chris Alexiuk, Christopher Parisien, Collin Neale, Damon Mosk-Aoyama, Dan Su, Dane Corneil, Daniel Afrimi, Daniel Rohrer, Daniel Serebrenik, Daria Gitman, Daria Levy, Darko Stosic, David Mosallanezhad, Deepak Narayanan, Dhruv Nathawani, Dima Rekesh, Dina Yared, Divyanshu Kakwani, Dong Ahn, Duncan Riach, Dusan Stosic, Edgar Minasyan, Edward Lin, Eileen Long, Eileen Peters Long, Elena Lantz, Ellie Evans, Elliott Ning, Eric Chung, Eric Harper, Eric Tramel, Erick Galinkin, Erik Pounds, Evan Briones, Evelina Bakhturina, Faisal Ladhak, Fay Wang, Fei Jia, Felipe Soares, Feng Chen, Ferenc Galko, Frankie Siino, Gal Hubara Agam, Ganesh Ajjanagadde, Gantavya Bhatt, Gargi Prasad, George Armstrong, Gerald Shen, Gorkem Batmaz, Grigor Nalbandyan, Haifeng Qian, Harsh Sharma, Hayley Ross, Helen Ngo, Herman Sahota, Hexin Wang, Himanshu Soni, Hiren Upadhyay, Huizi Mao, Huy C Nguyen, Huy Q Nguyen, Iain Cunningham, Ido Shahaf, Igor Gitman, Ilya Loshchilov, Ivan Moshkov, Izzy Putterman, Jan Kautz, Jane Polak Scowcroft, Jared Casper, Jatin Mitra, Jeffrey Glick, Jenny Chen, Jesse Oliver, Jian Zhang, Jiaqi Zeng, Jie Lou, Jimmy Zhang, Jining Huang, Joey Conway, Joey Guman, John Kamalu, Johnny Greco, Jonathan Cohen, Joseph Jennings, Joyjit Daw, Julien Veron Vialard, Junkeun Yi, Jupinder Parmar, Kai Xu, Kan Zhu, Kari Briski, Katherine Cheung, Katherine Luna, Keshav Santhanam, Kevin Shih, Kezhi Kong, Khushi Bhardwaj, Krishna C. Puvvada, Krzysztof Pawelec, Kumar Anik, Lawrence McAfee, Laya Sleiman, Leon Derczynski, Li Ding, Lucas Liebenwein, Luis Vega, Maanu Grover, Maarten Van Segbroeck, Maer Rodrigues de Melo, Makesh Narsimhan Sreedhar, Manoj Kilaru, Maor Ashkenazi, Marc Romeijn, Mark Cai, Markus Kliegl, Maryam Moosaei, Matvei Novikov, Mehrzad Samadi, Melissa Corpuz, Mengru Wang, Meredith Price, Michael Boone, Michael Evans, Miguel Martinez, Mike Chrzanowski, Mohammad Shoeybi, Mostofa Patwary, Nabin Mulepati, Natalie Hereth, Nave Assaf, Negar Habibi, Neta Zmora, Netanel Haber, Nicola Sessions, Nidhi Bhatia, Nikhil Jukar, Nikki Pope, Nikolai Ludwig, Nima Tajbakhsh, Nirmal Juluru, Oleksii Hrinchuk, Oleksii Kuchaiev, Olivier Delalleau, Oluwatobi Olabiyi, Omer Ullman Argov, Ouye Xie, Parth Chadha, Pasha Shamis, Pavlo Molchanov, Pawel Morkisz, Peter Dykas, Peter Jin, Pinky Xu, Piotr Januszewski, Pranav Prashant Thombre, Prasoon Varshney, Pritam Gundecha, Qing Miao, Rabeeh Karimi Mahabadi, Ran El-Yaniv, Ran Zilberstein, Rasoul Shafipour, Rich Harang, Rick Izzo, Rima Shahbazyan, Rishabh Garg, Ritika Borkar, Ritu Gala, Riyad Islam, Roger Waleffe, Rohit Watve, Roi Koren, Ruoxi Zhang, Russell J. Hewett, Ryan Prenger, Ryan Timbrook, Sadegh Mahdavi, Sahil Modi, Samuel Kriman, Sanjay Kariyappa, Sanjeev Satheesh, Saori Kaji, Satish Pasumarthi, Sean Narentharen, Sean Narenthiran, Seonmyeong Bak, Sergey Kashirsky, Seth Poulos, Shahar Mor, Shanmugam Ramasamy, Shantanu Acharya, Shaona Ghosh, Sharath Turuvekere Sreenivas, Shelby Thomas, Shiqing Fan, Shreya Gopal, Shrimai Prabhumoye, Shubham Pachori, Shubham Toshniwal, Shuoyang Ding, Siddharth Singh, Simeng Sun, Smita Ithape, Somshubra Majumdar, Soumye Singhal, Stefania Alborghetti, Stephen Ge, Sugam Dipak Devare, Sumeet Kumar Barua, Suseella Panguluri, Suyog Gupta, Sweta Priyadarshi, Syeda Nahida Akter, Tan Bui, Teodor-Dumitru Ene, Terry Kong, Thanh Do, Tijmen Blankevoort, Tom Balough, Tomer Asida, Tomer Bar Natan, Tugrul Konuk, Twinkle Vashishth, Udi Karpas, Ushnish De, Vahid Noorozi, Vahid Noroozi, Venkat Srinivasan, Venmugil Elango, Vijay Korthikanti, Vitaly Kurin, Vitaly Lavrukhin, Wanli Jiang, Wasi Uddin Ahmad, Wei Du, Wei Ping, Wenfei Zhou, Will Jennings, William Zhang, Wojciech Prazuch, Xiaowei Ren, Yashaswi Karnati, Yejin Choi, Yev Meyer, Yi-Fu Wu, Yian Zhang, Ying Lin, Yonatan Geifman, Yonggan Fu, Yoshi Subara, Yoshi Suhara, Yubo Gao, Zach Moshe, Zhen Dong, Zihan Liu, Zijia Chen, Zijie Yan
TL;DR
Nemotron 3 Nano addresses efficient, capable open-model design for agentic reasoning across chat, reasoning, and tool-use settings. It combines a sparse MoE hybrid Mamba-Transformer architecture with 25T-token pretraining and SFT plus large-scale reinforcement learning. The model reports competitive accuracy, up to 3.3× higher inference throughput, and support for context lengths up to 1M tokens, with weights and training resources released openly.
Problem
The paper studies how to build an open model with agentic reasoning capabilities while improving the inference-throughput-to-accuracy trade-off.
Method
The model uses a sparse MoE hybrid Mamba-Transformer architecture, 25T-token pretraining, supervised fine-tuning, multi-environment RLVR, and RLHF.
Results
Nemotron 3 Nano achieves better or on-par accuracy than competitive models, up to 3.3× higher inference throughput, and context lengths up to 1M tokens.
Takeaways & Limitations
The released model and training resources provide an open system for agentic reasoning with competitive accuracy, efficient inference, and long-context support.
Takeaways & Limitations
Some comparison benchmark values come from official reports, ArtificialAnalysis, or scores computed using official protocols when direct results were unavailable.
Abstract
from arXiv · showhide
We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and large-scale RL on diverse environments. Nemotron 3 Nano achieves better accuracy than our previous generation Nemotron 2 Nano while activating less than half of the parameters per forward pass. It achieves up to 3.3x higher inference throughput than similarly-sized open models like GPT-OSS-20B and Qwen3-30B-A3B-Thinking-2507, while also being more accurate on popular benchmarks. Nemotron 3 Nano demonstrates enhanced agentic, reasoning, and chat abilities and supports context lengths up to 1M tokens. We release both our pretrained Nemotron 3 Nano 30B-A3B Base and post-trained Nemotron 3 Nano 30B-A3B checkpoints on Hugging Face.
1. Introduction
Nemotron 3 Nano is an open MoE hybrid Mamba-Transformer model designed for agentic reasoning, chat, and efficient inference. It combines large-scale pretraining and multi-stage post-training with released weights, recipes, code, and data.
- Model and contributions: Nemotron 3 Nano combines Mamba-2, grouped-query attention, and sparse MoE layers, activating 6 of 128 experts per forward pass.The model has 31.6B total parameters, with sparse activation intended to improve the inference-throughput-to-accuracy frontier.
- Results: 3.3× and 2.2× faster inference throughput were measured against Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B, respectively, on 8K-input/16K-output generation.Measurements used a single H200 GPU and the better result from vLLM or TRT-LLM for each model.
- Results: Nemotron 3 Nano supports context lengths up to 1M tokens and outperforms the cited comparison models on RULER across different context lengths.GPT-OSS-20B supports 128K tokens, limiting direct 1M-context comparison.
- Model and contributions: 25 trillion tokens spanning 15 categories were used to pretrain the base model across two phases.The first phase used 23.5T diverse tokens, followed by 1.5T high-quality tokens.
- Post-training: Post-training uses SFT, multi-environment RLVR, and RLHF to develop chat, agentic, reasoning, controllable-budget, and tool-integrated capabilities.RLVR trains on all environments simultaneously, while SFT includes diverse chat, agentic, and reasoning traces.
- Open release: The release includes base and final model weights, training recipes, code, and most of the training data.The paper also lists released Nemotron pretraining, SFT, and RL datasets.
2. Pretraining
Pretraining expands Nemotron 3 Nano’s corpus with curated, synthetic, translated, code, STEM, and cross-domain data. The architecture and data pipeline target broader coverage and more complex reasoning and coding examples.
- Architecture: Nemotron 3 Nano Base replaces standard FFN layers with sparse MoE layers in a hybrid Mamba-Transformer architecture.The model contains 31.6B total parameters and 3.2B active parameters per forward pass, excluding embeddings.
- Corpus construction: Code pretraining data combines filtered Common Crawl pages, curated GitHub repositories, synthetic code documents, rewriting, and Python-to-C++ transpilation.The pipeline preserves code and technical elements while filtering boilerplate, duplicates, and low-relevance documents.
- Corpus construction: 2.5T new tokens were curated or generated from Common Crawl data.Sources include recent snapshots, synthetic rephrasing, and translation into English from other languages.
- Specialized data: Synthetic specialized data covers STEM reasoning, mathematical textbooks, scientific coding, and advanced computational coding problems.Scientific coding examples include code-embedded articles and graduate- or research-level Python problems.
- Specialized data: InfiniByte cross-breeds datasets across coding, mathematics, physics, chemistry, and other sciences to generate cross-domain programming problems.The approach targets questions at domain boundaries and produces more diverse and complex code data.
3. Post-Training
Post-training combines improved supervised fine-tuning with large-scale multi-environment RLVR and RLHF. The resulting methodology targets agentic, reasoning, chat, tool-use, multilingual, safety, and long-context behavior.
- Training strategy: Nemotron 3 Nano’s post-training uses SFT, multi-environment RLVR, and RLHF as complementary stages.RLVR trains on all environments simultaneously, while RLHF uses a generative reward model.
- SFT: SFT data emphasizes diverse multi-step and multi-turn agentic tasks alongside chat and reasoning traces.The stage also supports reasoning-budget control, reasoning on/off control, and tool-integrated reasoning.
- SFT: The chat template supports reasoning and non-reasoning modes, preserving reasoning across multi-step calls while dropping previous-turn reasoning after a new user message.Tool calling uses XML-style special tags.
3.2. Multi environment Reinforcement Learning from Verifiable Rewards
Nemotron 3 Nano uses unified multi-environment RLVR with curriculum sampling to improve reasoning and agentic capabilities across domains. This training combines verifiable task environments with RLHF-oriented reward modeling and length control, producing stable gains while reducing verbosity.
- Multi-environment RLVR: Unified RLVR training across all environments produces stable gains across benchmarks, whereas single-environment training can cause unrecoverable degradation elsewhere.The model undergoes two RLVR stages: after SFT and after RLHF.
- RL environments: The RL curriculum spans competition math, competition coding, question answering, structured outputs, instruction following, long context, and agentic tool use.These environments include automatically verified tool calls, schema constraints, difficult STEM questions, and multi-document long-context tasks.
- Curriculum training: The curriculum filters out tasks already solved by the SFT checkpoint and dynamically shifts training toward harder cases.Task difficulty is controlled through Gaussian sampling whose target pass-rate mean decreases over training.
- Curriculum training: Curriculum sampling preserves domain ratios and stable learning across domains, while random sampling biases training toward easier tasks.The curriculum therefore supports learning challenging tasks without abandoning domain diversity.
- RLHF and length control: Group Relative Length Control applies group-relative penalties and quality-gated conciseness bonuses, reducing verbosity by 30% without sacrificing accuracy.The method uses reasoning and answer length coefficients of 0.5.
3.4. Post-trained Model Evaluations
Nemotron 3 Nano is evaluated across reasoning, coding, agentic, instruction-following, long-context, and multilingual benchmarks, where it shows especially strong agentic, chat, and long-context performance. It also achieves higher inference throughput than GPT-OSS-20B and Qwen3-30B-A3B-Thinking-2507.
- Evaluation coverage: Evaluations cover mathematical and scientific reasoning, coding, agentic tool use, instruction following, long-context understanding, and multilingual capability.
- Nemotron 3 Nano surpasses both GPT-OSS 20B and Qwen3-30B-A3B-Thinking-2507 across all evaluated categories.
- Evaluation results: On reasoning benchmarks, Nemotron 3 Nano surpasses Qwen3 and is competitive with GPT-OSS.
- Evaluation results: In agentic, chat, and long-context categories, Nemotron 3 Nano significantly outperforms both comparison models.
4. Quantization
The quantization study uses post-training FP8 quantization with selective retention of sensitive components in BF16. This design preserves accuracy while improving throughput, especially through FP8 KV-cache quantization and larger batch sizes.
- Selective quantization: Self-attention layers and preceding Mamba layers are retained in BF16, while other components are quantized to FP8.The model weights, activations, and KV cache use FP8, while Conv1D in Mamba layers remains BF16.
- Accuracy: Approximately 99% median accuracy recovery is achieved by Nemotron 3 Nano FP8 compared with BF16.
- Ablation design: The ablation varies attention-layer, Mamba-layer, and KV-cache precision to characterize accuracy–efficiency trade-offs.
- Throughput: FP8 KV-cache quantization significantly improves throughput by enabling larger batch sizes.
- Measurement setup: Accuracy recovery and throughput improvements are normalized relative to the Nemotron 3 Nano BF16 checkpoint, whose baseline is 100%.The benchmark uses a single H100 with ISL/OSL=8K/16K and maximum batch size for each configuration.
5. Conclusion
Nemotron 3 Nano is presented as an open, efficient MoE hybrid Mamba-Transformer model for agentic reasoning. It combines competitive accuracy, up to 3.3× higher inference throughput, 1M-token context support, and released training resources.
- Nemotron 3 Nano achieves better or on-par accuracy than competitive models while providing up to 3.3× higher inference throughput.
- The model supports context lengths of up to 1M tokens.
- The authors release base and final model weights on Hugging Face together with the training recipe, data, and code.
A. Base model evaluations
The released improved base checkpoint is compared with the pre-alignment base and Qwen3 across several benchmark categories. It improves average Code and General Knowledge performance over Qwen3, widens the Math lead, and has a slight Long Context regression from the earlier checkpoint.
- The improved checkpoint surpasses Qwen3 in average Code and General Knowledge performance, unlike the pre-alignment base.
- The improved checkpoint’s lead in Math tasks significantly widens relative to the pre-alignment base.
- Multilingual benchmarks: The improved checkpoint surpasses Qwen3 on MGSM, while Qwen3 retains a lead on MMLU Global Lite.
- Long Context: A slight Long Context performance drop occurs relative to the pre-alignment base, although the improved checkpoint maintains a commanding margin over Qwen3.
B. MMLU-redux evaluation
The MMLU-redux evaluation introduces chain-of-thought and tweaked variants to assess reasoning and reduce concerns about benchmark saturation. Enabling CoT substantially improves the model’s accuracy, particularly on STEM subjects.
- MMLU-redux design: MMLU-redux includes CoT and Tweak variants designed to better test step-by-step reasoning and performance on modified examples.The CoT variant uses five subject exemplars with detailed solutions, while Tweak changes numerical values and equations while preserving underlying skills.
- MMLU-redux results: +5.27 average accuracy improvement from enabling CoT for our model, compared with +0.79 for Qwen.The evaluation reports especially strong gains on STEM subjects.
- MMLU-redux results: 64.00 to 77.00 accuracy on Professional Accounting, a +13.00 improvement under the Other category.The passage attributes this gain to the task’s reliance on calculation skills.
- MMLU-redux results: On MMLU-redux Tweak, our model’s STEM score increases by 5.31, while Qwen’s decreases marginally by 0.83.Both models gain across non-STEM categories, according to the reported evaluation.
C. DPO for Reducing Tool Hallucination
The study tests whether lightweight DPO can reduce hallucinated tool calls while improving reasoning accuracy. Across evaluated benchmarks, small-scale DPO produces consistent gains with minimal training cost.
- Motivation and setup: DPO targets tool hallucinations, defined as tool invocations when no tools are declared in the system message.The evaluation treats outputs such as Python execution requests, search invocations, or tool-specific API formats as hallucinations.
- DPO data construction: 2,000 reasoning tasks combine 1,000 mathematics problems and 1,000 STEM multiple-choice questions across No-Tools, With-Tools, and Hallucination-Penalty settings.The settings separate pure reasoning, tool-assisted reasoning, and calibration of tool usage.
- Training and results: 50 training steps with a 3e-6 learning rate yield consistent improvements across all evaluated benchmarks.The lightweight setup is designed to provide preference-learning signal while minimally perturbing the SFT model.
- Benchmark results: 80.88% to 84.58%: AIME25 accuracy improves while hallucination falls from 1.25% to 0%.The AIME25 result reports complete elimination of spurious tool invocation in this setting.
- Benchmark results: 65.15% to 69.19%: GPQA accuracy rises as hallucination drops from 8.33% to 0.7%.The passage characterizes GPQA as a more challenging setting with higher baseline hallucination.
- Overall finding: Overall, minimal DPO reduces hallucinated tool usage while improving reasoning accuracy and providing a complementary signal to RL-based alignment.The reported improvements are described as occurring with negligible computational cost.
D. Safety Preference Data
The safety preference dataset pairs safe, unsafe, and over-refusal responses to train reward models for robust safety alignment. Its construction uses multiple open-source models and filtering to diversify candidate responses.
- Data sources: RLHF reward-model data uses the same underlying datasets as the SFT safety subset, preserving a similar starting seed-prompt distribution.Response generation additionally addresses over-refusals and harmful engagements through rejected responses.
- Preference construction: Harmful prompts pair safe chosen responses with unsafe rejected responses generated through jailbreak templates or direct prompting with moderation filtering.The two generation methods target unsafe completions for the rejected set.
- Preference construction: Safe prompts pair safe chosen responses with over-refusal rejected responses.These pairs distinguish desirable safety behavior from unnecessary refusal behavior.
- Purpose: The resulting preference pairs support reward-model training for robust safety alignment and mitigation of over-refusal behaviors.Harmful prompts use safe-versus-unsafe labels, while safe prompts use safe-versus-over-refusal labels.
- Diversity and selection: Five open-source models generate candidate responses, after which filtering retains safe chosen and unsafe or over-refusal rejected candidates.One chosen and rejected pair is randomly selected per prompt from the filtered candidates.
E. Prompt Sensitivity Analysis
The analysis measures how benchmark accuracy changes under varied prompt formulations. Nemotron 3 Nano is evaluated with multiple prompts and seeds to reduce dependence on any single wording or formatting choice.
- Prompt design: Prompt variations change wording, instruction granularity, problem placement, and answer formatting for each dataset.The design covers minimal versus detailed instructions and placement before, within, or after the prompt.
- Metric: Mean accuracy is computed across eight seeds for each prompt, with the standard deviation of prompt averages used as prompt sensitivity.Lower sensitivity indicates less variation across prompt formulations.
- Results: Below 1: Nemotron 3 Nano's sensitivity scores are reported across all datasets.Prompt sensitivity results are presented in Table 8, which compares Nemotron 3 Nano with Qwen3-30B-A3B-Thinking-2507 and GPT-OSS 20B.