Source-linked AI summary

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

NVIDIA, :, Amala Sanjay Deshmukh, Kateryna Chumachenko, Tuomas Rintamaki, Matthieu Le, Tyler Poon, Danial Mohseni Taheri, Ilia Karmanov, Guilin Liu, Jarno Seppanen, Arushi Goel, Mike Ranzinger, Greg Heinrich, Guo Chen, Lukas Voegtle, Philipp Fischer, Timo Roman, Karan Sapra, Collin McCarthy, Shaokun Zhang, Fuxiao Liu, Hanrong Ye, Yi Dong, Mingjie Liu, Yifan Peng, Piotr Zelasko, Zhehuai Chen, Nithin Rao Koluguri, Nune Tadevosyan, Lilit Grigoryan, Ehsan Hosseini Asl, Pritam Biswas, Leili Tavabi, Yuanhang Su, Zhiding Yu, Peter Jin, Alexandre Milesi, Netanel Haber, Yao Xu, Sarah Amiraslani, Nabin Mulepati, Eric Tramel, Jaehun Jung, Ximing Lu, Brandon Cui, Jin Xu, Zhiqi Li, Shihao Wang, Yuanguo Kuang, Shaokun Zhang, Huck Yang, Boyi Li, Hongxu Yin, Song Han, Bilal Kartal, Pavlo Molchanov, Adi Renduchintala, Charles Wang, David Mosallanezhad, Soumye Singhal, Luis Vega, Katherine Cheung, Sreyan Ghosh, Yian Zhang, Alexander Bukharin, Venkat Srinivasan, Johnny Greco, Andre Manoel, Maarten Van Segbroeck, Suseella Panguliri, Rohit Watve, Divyanshu Kakwani, Shubham Pachori, Jeffrey Glick, Radha Sri-Tharan, Aileen Zaman, Khanh Nguyen, Shi Chen, Jiaheng Fang, Qing Miao, Wenfei Zhou, Yu Wang, Zaid Pervaiz Bhat, Varun Praveen, Arihant Jain, Ramanathan Arunachalam, Tomasz Kornuta, Ashton Sharabiani, Amy Shen, Wei Huang, Yi-Fu Wu, Ali Roshan Ghias, Huiying Li, Brian Yu, Nima Tajbakhsh, Chen Cui, Wenwen Gao, Li Ding, Terry Kong, Manoj Kilaru, Anahita Bhiwandiwalla, Marek Wawrzos, Daniel Korzekwa, Pablo Ribalta, Grzegorz Chlebus, Besmira Nushi, Ewa Dobrowolska, Maciej Jakub Mikulski, Kunal Dhawan, Steve Huang, Jagadeesh Balam, Yongqiang Wang, Nikolay Karpov, Valentin Mendelev, George Zelenfroynd, Meline Mkrtchyan, Qing Miao, Omri Almog, Bhavesh Pawar, Rameshwar Shivbhakta, Sudeep Sabnis, Ashrton Sharabiani, Negar Habibi, Geethapriya Venkataramani, Pamela Peng, Prerit Rodney, Serge Panev, Richard Mazzarese, Nicky Liu, Michael Fukuyama, Andrii Skliar, Roger Waleffe, Duncan Riach, Yunheng Zou, Jian Hu, Hao Zhang, Binfeng Xu, Yuhao Yang, Zuhair Ahmed, Alexandre Milesi, Carlo del Mundo, Chad Voegele, Zhiyu Cheng, Nave Assaf, Andrii Skliar, Daniel Afrimi, Natan Bagrov, Ran Zilberstein, Ofri Masad, Eugene Khvedchenia, Natan Bagrov, Borys Tymchenko, Tomer Asida, Daniel Afrimi, Parth Mannan, Victor Cui, Michael Evans, Katherine Luna, Jie Lou, Pinky Xu, Guyue Huang, Negar Habibi, Michael Boone, Pradeep Thalasta, Adeola Adesoba, Dina Yared, Christopher Parisien, Leon Derczynski, Shaona Ghosh, Wes Feely, Micah Schaffer, Radha Sri-Tharan, Jeffrey Glick, Barnaby Simkin, George Zelenfroynd, Tomasz Grzegorzek, Rishabh Garg, Aastha Jhunjhunwala, Sergei Kolchenko, Farzan Memarian, Haran Kumar, Shiv Kumar, Isabel Hulseman, Anjali Shah, Kari Briski, Padmavathy Subramanian, Joey Conway, Udi Karpas, Jane Polak Scowcroft, Annie Surla, Shilpa Ammireddy, Ellie Evans, Jesse Oliver, Tom Balough, Chia-Chih Chen, Sandip Bhaskar, Alejandra Rico, Bardiya Sadeghi, Seph Mard, Katherine Cheung, Meredith Price, Laya Sleiman, Saori Kaji, Wesley Helmholz, Wendy Quan, Michael Lightstone, Jonathan Cohen, Jian Zhang, Oleksii Kuchaiev, Boris Ginsburg, Jan Kautz, Eileen Long, Mohammad Shoeybi, Mostofa Patwary, Oluwatobi Olabiyi, Andrew Tao, Bryan Catanzaro, Udi Karpas

arXiv:2604.24954v2cs.LGcs.AIcs.CV

TL;DR

Cross-modal understanding requires sophisticated reasoning across heterogeneous inputs, yet efficient omni-modal models remain challenging to develop. Nemotron 3 Nano Omni addresses this with native audio support, multimodal token reduction, and a unified training recipe, achieving consistent gains across modalities and substantially higher inference throughput.

  • Problem

    Cross-modal understanding remains challenging because extending reasoning across multiple modalities requires sophisticated cross-modal reasoning.

  • Method

    Nemotron 3 Nano Omni combines a Nemotron 3 Nano 30B-A3B MoE backbone, native audio support, multimodal token reduction, and multi-stage cross-modal training.

  • Results

    Across broad evaluations, Nemotron 3 Nano Omni delivers consistent gains over Nemotron Nano V2 VL across modalities and 3× higher single-stream output token throughput than Qwen3-Omni.

  • Takeaways & Limitations

    Released checkpoints, training data, and code make the model available for further multimodal research and development.

Abstract

from arXiv · show

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its predecessor, Nemotron Nano V2 VL, across all modalities, enabled by advances in architecture, training data and recipes. In particular, Nemotron 3 delivers leading results in real-world document understanding, long audio-video comprehension, and agentic computer use. Built on the highly efficient Nemotron 3 Nano 30B-A3B backbone, Nemotron 3 Nano Omni further incorporates innovative multimodal token-reduction techniques to deliver substantially lower inference latency and higher throughput than other models of similar size. We are releasing model checkpoints in BF16, FP8, and FP4 formats, along with portions of the training data and codebase to facilitate further research and development.

NVIDIA · 1. Introduction

Nemotron 3 Nano Omni is an efficient omni-modal model built on the Nemotron 3 Nano 30B-A3B backbone, with native audio support and improved multimodal reasoning. Its architectural and training advances yield strong benchmark performance, higher inference throughput, lower latency, and released model and research resources.

  • 1. Introduction: Nemotron 3 Nano Omni combines the Nemotron 3 Nano 30B-A3B language backbone with C-RADIOv4-H1 and Parakeet-TDT-0.6B-v22 encoders.It extends the Nemotron multimodal family with native audio support and improved reasoning capability across supported modalities.
  • 1. Introduction: The model replaces the dense Nemotron Nano V2 12B hybrid backbone with a 30B-A3B Mixture-of-Experts hybrid backbone for more efficient long-sequence processing and higher throughput.It also introduces native audio inputs alongside text, images, and video, plus dynamic image resolution instead of tiling-based processing.
  • 1. Introduction: A multi-stage training strategy progressively introduces modalities and scales context length to preserve text reasoning, mitigate catastrophic forgetting, and stabilize cross-modal alignment.The approach addresses modality alignment, training stability, and data balancing across heterogeneous sources.
  • 1. Introduction: Nemotron 3 Nano Omni achieves substantial gains over Nemotron Nano V2 VL and leading results in document understanding, audio-visual reasoning, and audio benchmarks.The cited leaderboards include OCRBench-V2, MMLongBench-DOC, VoiceBench, WorldSense, and DailyOmni.
  • 1. Introduction: 3× higher single-stream output token throughput than Qwen3-Omni and 9× higher output token throughput per GPU at a fixed interactivity target are achieved on NVIDIA B200.Compared with Nemotron Nano V2 VL, the model provides 3× higher throughput at the same interactivity target and 2× higher single-stream output token throughput.
  • 1. Introduction: Model checkpoints are released on HuggingFace in BF16, FP8, and NVFP4 formats.The listed checkpoints are Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16, -FP8, and -NVFP4.
  • 1. Introduction: The release includes part of the training datasets, data-generation pipelines, Megatron-Bridge training code, and a NeMo RL guide.Nemotron-Image-Training-v3 contains approximately 6.9 million training samples.
  • 1. Introduction: The paper proceeds from model architecture to training pipeline and datasets, then evaluation results across all modalities.These topics are covered in Sections 2, 3, and 4, respectively.

2. Model Architecture

Nemotron 3 Nano Omni uses an encoder-projector-decoder architecture that combines modality-specific vision and audio encoders with the Nemotron 3 Nano 30B-A3B language model. Dynamic-resolution vision processing, temporal compression, and audio tokenization support efficient multimodal sequence construction and joint temporal reasoning.

  • Overall architecture: The model combines vision and audio encoders with the Nemotron 3 Nano 30B-A3B language model through MLP projectors.The overall design is encoder-projector-decoder.
  • Vision processing: Dynamic-resolution processing preserves native image aspect ratios and replaces the predecessor’s tiling strategy.Images use a variable number of 16 × 16 patches, constrained to 1,024–13,312 visual tokens before projection.
  • Vision processing: Pixel shuffle provides 4× image token downsampling, while Conv3D compresses every two video frames into one.The architecture also optionally uses Efficient Video Sampling for higher video throughput.
  • Audio and sequence construction: Audio is resampled to 16 kHz mono and encoded at approximately 12.5 tokens per second after ∼8× temporal downsampling.Audio streams are segmented into 30-second clips of approximately 375 tokens, and multimodal tokens are interleaved temporally for joint reasoning.

3. Training Recipe & Datasets

Nemotron 3 Nano Omni uses a staged training strategy for heterogeneous multimodal encoders: supervised fine-tuning aligns modalities and expands capabilities, followed by reinforcement learning to refine reasoning and safety.

  • Staged Training Strategy: The recipe first applies SFT for modality alignment, multimodal instruction-following, and context extension, then uses RL to refine reasoning and safety.This staged progression is illustrated in Figure 2.

3.1. SFT

The SFT pipeline uses seven progressive stages that introduce modalities, unfreeze components, and extend context length to promote stable cross-modal alignment and long-context understanding. It culminates in jointly trained omni-modal stages and a 262,144-token long-context stage focused on documents and multimodal reasoning.

  • SFT pipeline: Seven progressive SFT stages introduce modalities and increase context length to stabilize cross-modal alignment, mitigate catastrophic forgetting, and improve multimodal understanding.The curriculum progresses from vision and audio projector alignment to joint omni-modal training and long-context adaptation.
  • Vision SFT: 9.35 million vision–text samples (∼15.5B tokens) train the vision projector across captioning, grounding, OCR, document, GUI, and visual question-answering tasks.Only the vision MLP projector is trained initially, with a maximum context length of 16384 while other components remain frozen.
  • Vision-language SFT: 86.3M multimodal samples (∼214.8B tokens) support joint vision–language fine-tuning with improved text reasoning, relabeled data, reasoning traces, broader domains, and multiple languages.The dataset uses public, internally curated, and synthetic data with filtering and distillation to improve correctness and coverage.
  • Audio SFT: 59.2M audio samples (∼11.4B tokens) warm up the audio projector, after which the audio encoder and projector train on ASR, sound, music, and speech understanding.The initial audio stage freezes the LLM, vision encoder, and Parakeet-TDT encoder; the subsequent stage adds captions, multiple-choice questions, open-ended questions, and reasoning traces.
  • Omni-modal SFT: All-parameter omni-modal SFT combines vision, text, safety, video, audio-video QA and captioning, ASR, and audio reasoning, with dominant Stage 4 sources of 30.6B vision, 9.7B audio, and 6.3B short-video tokens.The 48k stage extends context to 49,152 tokens and rebalances training toward medium and long video, omni-modal data, and reasoning.
  • Ultra-long-context SFT: 262,144-token SFT uses ∼34.0B tokens from long-context text and vision data to improve reasoning over 10-to-100-plus-page documents, including text, charts, and complex tables.The audio encoder and projector are frozen during this stage to focus capacity on long-context text and document understanding.

3.2. Reinforcement Learning

The model undergoes a staged reinforcement-learning curriculum spanning preference optimization, text, image, and omni-modal training to improve instruction following, reasoning, and safety alignment. Training combines modality-specific verifiers, filtered multimodal data, and safeguards against representational drift and speech-recognition degradation.

  • Curriculum: Five RL stages progress from Preference Optimization through Text-RL-stage-1, Image-RL, Omni-RL, and Text-RL-stage-2.The curriculum targets instruction following, reasoning, and safety alignment across text, image, and video modalities.
  • Preference Optimization: MPO combines DPO preference loss with BCO quality loss, using rejection-sampled vision responses labeled positive or negative by outcome correctness.This provides both preference-level and quality-level supervision during offline reinforcement learning.
  • Text-RL: Text-only RL trains only LM parameters with RLVR/RLHF while freezing LM input-token embeddings to mitigate representational drift across multimodal stages.The text-RL data and infrastructure are reused from Nemotron 3 Nano and Super post-training.
  • ImageRL: ImageRL applies outcome-based RL to visual reasoning, combining verifier-based outcome scores with format rewards for a single <think> block and boxed answer.The format reward preserves credit for correct answers despite surface-format errors while discouraging verbose multi-answer outputs.
  • OmniRL: Approximately 120K prompts across 113 sub-datasets form the OmniRL corpus, covering image, video, audio, and text-only reasoning with pass-rate filtering.Audio training uses an ASR verifier with reward 1 - WER, while prompts that are trivially solvable or entirely intractable are excluded.

4. Experiments … 4.3. Audio-Visual Evaluations

The experiments evaluate Nemotron 3 Nano Omni across vision, audio, and audio-visual reasoning, using established evaluation frameworks and benchmarks. The model improves over prior or competing models, including across audio-visual reasoning modes.

  • 4. Experiments: The evaluation spans reasoning over vision, audio, and text in Sections 4.1–4.4, with separate analyses of video sampling and quantization efficiency.Efficient Video Sampling is analyzed in Section 4.6, while quantization effects are examined in Sections 4.7 and 4.8.
  • 4. Experiments: Vision and audio evaluations use VLMEvalKit with a vLLM backend, while text evaluations use the NeMo-Skills framework.These frameworks support the evaluations in Sections 4.1–4.3.
  • 4.1. Visual Evaluations: Visual evaluations cover STEM reasoning, document understanding, OCR and charts, plus visual grounding and spatial reasoning.Benchmarks include MMMU, MathVista-Mini, MMLongBench-Doc, OCRBench, OCRBench-V2, ChartQA, AI2D, TextVQA, DocVQA, InfoVQA, OCR-Reasoning, CharXiv, and TreeBench.
  • 4.1. Visual Evaluations: Nemotron 3 Nano Omni shows significant improvements over Nemotron Nano V2 VL across visual benchmarks and outperforms Qwen3-Omni on several categories.The comparison is reported in Table 7.
  • 4.2. Audio Evaluations: Audio evaluations cover automatic speech recognition, long-form ASR, and audio understanding, including OpenASR, TED-LIUM Longform, and MMAU.OpenASR reports word error rate on English subsets, while TED-LIUM Longform tests transcription quality and long-context consistency.
  • 4.2. Audio Evaluations: Nemotron 3 Nano Omni outperforms Qwen family models on ASR and VoiceBench benchmarks.Table 8 also reports long-form ASR, MMAU, and VoiceBench, with word error rate lower-is-better and the other metrics higher-is-better.
  • 4.3. Audio-Visual Evaluations: Audio-visual evaluation uses DailyOmni and WorldSense to test cross-modal reasoning, temporal alignment, sound grounding, and long-range dependencies.DailyOmni contains 684 videos and 1,197 questions across six tasks; WorldSense contains 1,662 long-context videos and 3,172 questions across 26 tasks.
  • 4.3. Audio-Visual Evaluations: Nemotron 3 Nano Omni outperforms Qwen3-Omni in both reasoning-on and reasoning-off modes on Video+Audio benchmarks.Table 9 measures these comparisons using accuracy, where higher is better.

4.4. Text-only evaluations · 4.5. Reasoning budget control · 4.6. Conv3D and Efficient Video Sampling (EVS)

The evaluations show that Nemotron 3 Nano Omni preserves text-only capabilities while adding multimodal understanding, and that inference-time reasoning budgets and video token-reduction mechanisms improve efficiency without broadly harming accuracy.

  • 4.4. Text-only evaluations: Text-only evaluations use a 131,072-token maximum output, temperature 1.0, and top-p 1.0 across diverse academic, coding, instruction-following, and agentic benchmarks.AIME-2025 uses Pass@1 averaged over 8 runs; GPQA-Diamond uses 4 runs; the other listed benchmarks use 1 run.
  • 4.4. Text-only evaluations: Nemotron 3 Nano Omni aims to maintain the text-benchmark performance of its Nemotron 3 Nano 30B-A3B LLM backbone while adding vision and audio understanding.The comparison includes Nemotron 3 Nano Omni, Nemotron 3 Nano LLM, and Qwen3-Omni.
  • 4.5. Reasoning budget control: Reasoning-budget control compares a 16,384-token base configuration with reasoning enabled using a 13K budget, 1,024-token grace period, and the same maximum sequence length.The configurations are evaluated in Table 11 across several key benchmarks.
  • 4.5. Reasoning budget control: Reasoning-budget adjustment improves accuracy on selected benchmarks without degrading the remaining benchmarks.The gains may result from terminating malformed repetitive traces and truncating unnecessarily verbose reasoning chains.
  • 4.6. Conv3D and Efficient Video Sampling (EVS): Conv3D fuses every T=2 consecutive frames into one tubelet before the first ViT block, halving vision tokens through the ViT and LLM.This reduces ViT prefill cost, LLM-side prefill and attention compute, and KV-cache footprint during training and inference.
  • 4.6. Conv3D and Efficient Video Sampling (EVS): 5313 ms is the BF16 TTFT with Conv3D and EVS stacked, versus 7969 ms baseline, a 33% reduction costing about half a point of average accuracy.Conv3D alone reaches 5984 ms (−25%), while EVS alone reaches 6452 ms (−19%); the same ordering holds on NVFP4.
  • 4.6. Conv3D and Efficient Video Sampling (EVS): A 512-frame video produces ∼141k input tokens without either mechanism and ∼75k with Conv3D, a 47% reduction in LLM input tokens.The comparison uses a synthetic 512-frame, 512 × 512 video at 30 fps for concurrency-1 TTFT measurements.
  • 4.6. Conv3D and Efficient Video Sampling (EVS): EVS pruning keeps accuracy essentially flat through q=0.7 while improving TTFT by ∼14% versus no-EVS, with noticeable accuracy drops beyond q=0.8.LongVideoBench is the most sensitive benchmark to aggressive pruning, and TTFT improves monotonically through the tested range.

4.7. Quantization · 4.8. Inference efficiency

Nemotron 3 Nano Omni combines mixed-precision quantization with efficient inference, retaining near-BF16 accuracy while substantially improving throughput and latency across multimodal workloads. On NVIDIA B200, NVFP4 enables up to 7.5× output-token throughput, while the model surpasses Qwen3-Omni and Nemotron Nano V2 VL in serving efficiency.

  • 4.7. Quantization: Mixed-precision FP4 quantization uses NVFP4 for routed MoE experts, FP8 for selected projections and shared experts, and BF16 for remaining language-model layers.NVFP4 uses FP4 E2M1 values with per-block FP8 scales and a per-tensor FP32 global scale; selected FP8 layers use per-tensor E4M3 values with FP32 scaling.
  • 4.7. Quantization: Less than 1% median accuracy drop versus BF16 was observed for both FP8 and NVFP4 across 25 text, image, video, and audio benchmarks.The evaluation covered the multimodal benchmark suite described in the passage.
  • 4.8. Inference efficiency: 7.5× higher output-token throughput was achieved by NVFP4 than BF16 at iso-interactivity on NVIDIA B200 for single-image reasoning.Throughput was 18200 tok/s versus 2400 tok/s at 150 tok/s/user.
  • 4.8. Inference efficiency: More than 500 output tokens/s was sustained at concurrency 1, including longer sequences and larger multimodal inputs such as long videos and multi-document workloads.The passage attributes this sustained low-latency performance to the hybrid architecture.
  • 4.8. Inference efficiency: Approximately 1.3 s TTFT was achieved for multi-document workloads, compared with more than 2.5 s for Qwen3-Omni.This comparison uses time-to-first-token as the latency metric.
  • 4.8. Inference efficiency: 9× higher output throughput than Qwen3-Omni was provided on long-video workloads and 7.5× higher throughput on multi-document workloads at 50 output tokens/s per user.At the same interactivity target, throughput was also 3× higher than Nemotron Nano V2 VL.
  • 4.8. Inference efficiency: All measurements used one NVIDIA B200 GPU and vLLM nightly with EVS 50%, evaluating Nemotron 3 Nano Omni in NVFP4, Qwen3-Omni with dynamic FP8, and Nemotron Nano V2 VL in NVFP4.The multi-document workload contained 32 images at 1024×1536 resolution, while the long-video workload contained 512 frames at 512×512 resolution.

5. Conclusion

Nemotron 3 Nano Omni extends the Nemotron multimodal family with native audio support and stronger reasoning across text, images, video, and audio. Its architecture and training recipe support long heterogeneous inputs, while evaluations show consistent gains and leading or competitive results across multimodal tasks.

  • Conclusion: Nemotron 3 Nano Omni adds native audio support and stronger reasoning across text, images, video, and audio.It extends the Nemotron multimodal family as an efficient omni-modal model.
  • Conclusion: The model combines a Nemotron 3 Nano 30B-A3B MoE hybrid backbone, C-RADIOv4-H, Parakeet-TDT, dynamic image resolution, Conv3D video compression, and 256K context.These components enable processing of long, heterogeneous multimodal inputs with high accuracy.
  • Conclusion: Across evaluations, Nemotron 3 Nano Omni delivers consistent gains over Nemotron Nano V2 VL and leading or competitive results across document, GUI, audio-video, and voice tasks.The evaluated suites include OCRBench-V2, MMLongBench-Doc, ChartQA, CharXiv, ScreenSpot, ScreenSpot-Pro, OSWorld, WorldSense, DailyOmni, and VoiceBench.

6. Contributors

Section 6 credits a large contributor roster, presented across eight consecutive passages. The listed contributors include researchers, engineers, and collaborators whose names span the full section.

  • Opening contributor roster: The roster opens with Amala Sanjay Deshmukh, Kateryna Chumachenko, Tuomas Rintamaki, Matthieu Le, Tyler Poon, and Danial Mohseni Taheri.The passage continues with many additional contributors, including Ilia Karmanov, Guilin Liu, Jarno Seppanen, and Arushi Goel.
  • Continuing contributor roster: The next contributor group includes Yao Xu, Sarah Amiraslani, Nabin Mulepati, Eric Tramel, Jaehun Jung, Ximing Lu, and Brandon Cui.It also lists Jin Xu, Zhiqi Li, Shihao Wang, Yuanguo Kuang, Shaokun Zhang, and many others.
  • Continuing contributor roster: Further contributors include Yi-Fu Wu, Ali Roshan Ghias, Huiying Li, Brian Yu, Nima Tajbakhsh, Chen Cui, and Wenwen Gao.The roster continues through Roger Waleffe, Duncan Riach, Yunheng Zou, Jian Hu, Hao Zhang, Binfeng Xu, Yuhao Yang, and Zuhair Ahmed.
  • Additional contributors: Additional listed contributors include Alexandre Milesi, Carlo del Mundo, Chad Voegele, Zhiyu Cheng, Nave Assaf, Andrii Skliar, and Daniel Afrimi.The passage also names Natan Bagrov, Ran Zilberstein, Ofri Masad, Eugene Khvedchenia, Borys Tymchenko, and others.
  • Additional contributors: The roster continues with Michael Evans, Katherine Luna, Jie Lou, Pinky Xu, Guyue Huang, Negar Habibi, Michael Boone, and Pradeep Thalasta.It also includes Adeola Adesoba, Dina Yared, Christopher Parisien, Leon Derczynski, Shaona Ghosh, Wes Feely, and others.
  • Later contributor roster: Later entries include Aastha Jhunjhunwala, Sergei Kolchenko, Farzan Memarian, Haran Kumar, Shiv Kumar, Isabel Hulseman, and Anjali Shah.The final passage names Michael Lightstone, Jonathan Cohen, Jian Zhang, Oleksii Kuchaiev, Boris Ginsburg, Jan Kautz, Eileen Long, and others.
Loading 2604.24954v2…