Source-linked AI summary
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah, Abhanshu Gupta, Abhimanyu Bhatter, Abhinav Goyal, Abhinav Mathur, Abhinav Mohanty, Abhishek Kumar, Abhishek Sethi, Abi Komma, Abner Pena, Achin Jain, Adam Kunysz, Adam Opyrchal, Adarsh Singh, Aditya Rawal, Adok Achar Budihal Prasad, Adrià de Gispert, Agnika Kumar, Aishwarya Aryamane, Ajay Nair, Akilan M, Akshaya Iyengar, Akshaya Vishnu Kudlu Shanbhogue, Alan He, Alessandra Cervone, Alex Loeb, Alex Zhang, Alexander Fu, Alexander Lisnichenko, Alexander Zhipa, Alexandros Potamianos, Ali Kebarighotbi, Aliakbar Daronkolaei, Alok Parmesh, Amanjot Kaur Samra, Ameen Khan, Amer Rez, Amir Saffari, Amit Agarwalla, Amit Jhindal, Amith Mamidala, Ammar Asmro, Amulya Ballakur, Anand Mishra, Anand Sridharan, Anastasiia Dubinina, Andre Lenz, Andreas Doerr, Andrew Keating, Andrew Leaver, Andrew Smith, Andrew Wirth, Andy Davey, Andy Rosenbaum, Andy Sohn, Angela Chan, Aniket Chakrabarti, Anil Ramakrishna, Anirban Roy, Anita Iyer, Anjali Narayan-Chen, Ankith Yennu, Anna Dabrowska, Anna Gawlowska, Anna Rumshisky, Anna Turek, Anoop Deoras, Anton Bezruchkin, Anup Prasad, Anupam Dewan, Anwith Kiran, Apoorv Gupta, Aram Galstyan, Aravind Manoharan, Arijit Biswas, Arindam Mandal, Arpit Gupta, Arsamkhan Pathan, Arun Nagarajan, Arushan Rajasekaram, Arvind Sundararajan, Ashwin Ganesan, Ashwin Swaminathan, Athanasios Mouchtaris, Audrey Champeau, Avik Ray, Ayush Jaiswal, Ayush Sharma, Bailey Keefer, Balamurugan Muthiah, Beatriz Leon-Millan, Ben Koopman, Ben Li, Benjamin Biggs, Benjamin Ott, Bhanu Vinzamuri, Bharath Venkatesh, Bhavana Ganesh, Bhoomit Vasani, Bill Byrne, Bill Hsu, Bincheng Wang, Blake King, Blazej Gorny, Bo Feng, Bo Zheng, Bodhisattwa Paul, Bofan Sun, Bofeng Luo, Bowen Chen, Bowen Xie, Boya Yu, Brendan Jugan, Brett Panosh, Brian Collins, Brian Thompson, Can Karakus, Can Liu, Carl Lambrecht, Carly Lin, Carolyn Wang, Carrie Yuan, Casey Loyda, Cezary Walczak, Chalapathi Choppa, Chandana Satya Prakash, Chankrisna Richy Meas, Charith Peris, Charles Recaido, Charlie Xu, Charul Sharma, Chase Kernan, Chayut Thanapirom, Chengwei Su, Chenhao Xu, Chenhao Yin, Chentao Ye, Chenyang Tao, Chethan Parameshwara, Ching-Yun Chang, Chong Li, Chris Hench, Chris Tran, Christophe Dupuy, Christopher Davis, Christopher DiPersio, Christos Christodoulopoulos, Christy Li, Chun Chen, Claudio Delli Bovi, Clement Chung, Cole Hawkins, Connor Harris, Corey Ropell, Cynthia He, DK Joo, Dae Yon Hwang, Dan Rosen, Daniel Elkind, Daniel Pressel, Daniel Zhang, Danielle Kimball, Daniil Sorokin, Dave Goodell, Davide Modolo, Dawei Zhu, Deepikaa Suresh, Deepti Ragha, Denis Filimonov, Denis Foo Kune, Denis Romasanta Rodriguez, Devamanyu Hazarika, Dhananjay Ram, Dhawal Parkar, Dhawal Patel, Dhwanil Desai, Dinesh Singh Rajput, Disha Sule, Diwakar Singh, Dmitriy Genzel, Dolly Goldenberg, Dongyi He, Dumitru Hanciu, Dushan Tharmal, Dzmitry Siankovich, Edi Cikovic, Edwin Abraham, Ekraam Sabir, Elliott Olson, Emmett Steven, Emre Barut, Eric Jackson, Ethan Wu, Evelyn Chen, Ezhilan Mahalingam, Fabian Triefenbach, Fan Yang, Fangyu Liu, Fanzi Wu, Faraz Tavakoli, Farhad Khozeimeh, Feiyang Niu, Felix Hieber, Feng Li, Firat Elbey, Florian Krebs, Florian Saupe, Florian Sprünken, Frank Fan, Furqan Khan, Gabriela De Vincenzo, Gagandeep Kang, George Ding, George He, George Yeung, Ghada Qaddoumi, Giannis Karamanolakis, Goeric Huybrechts, Gokul Maddali, Gonzalo Iglesias, Gordon McShane, Gozde Sahin, Guangtai Huang, Gukyeong Kwon, Gunnar A. Sigurdsson, Gurpreet Chadha, Gururaj Kosuru, Hagen Fuerstenau, Hah Hah, Haja Maideen, Hajime Hosokawa, Han Liu, Han-Kai Hsu, Hann Wang, Hao Li, Hao Yang, Haofeng Zhu, Haozheng Fan, Harman Singh, Harshavardhan Kaluvala, Hashim Saeed, He Xie, Helian Feng, Hendrix Luo, Hengzhi Pei, Henrik Nielsen, Hesam Ilati, Himanshu Patel, Hongshan Li, Hongzhou Lin, Hussain Raza, Ian Cullinan, Imre Kiss, Inbarasan Thangamani, Indrayani Fadnavis, Ionut Teodor Sorodoc, Irem Ertuerk, Iryna Yemialyanava, Ishan Soni, Ismail Jelal, Ivan Tse, Jack FitzGerald, Jack Zhao, Jackson Rothgeb, Jacky Lee, Jake Jung, Jakub Debski, Jakub Tomczak, James Jeun, James Sanders, Jason Crowley, Jay Lee, Jayakrishna Anvesh Paidy, Jayant Tiwari, Jean Farmer, Jeff Solinsky, Jenna Lau, Jeremy Savareese, Jerzy Zagorski, Ji Dai, Jiacheng, Gu, Jiahui Li, Jian, Zheng, Jianhua Lu, Jianhua Wang, Jiawei Dai, Jiawei Mo, Jiaxi Xu, Jie Liang, Jie Yang, Jim Logan, Jimit Majmudar, Jing Liu, Jinghong Miao, Jingru Yi, Jingyang Jin, Jiun-Yu Kao, Jixuan Wang, Jiyang Wang, Joe Pemberton, Joel Carlson, Joey Blundell, John Chin-Jew, John He, Jonathan Ho, Jonathan Hueser, Jonathan Lunt, Jooyoung Lee, Joshua Tan, Joyjit Chatterjee, Judith Gaspers, Jue Wang, Jun Fang, Jun Tang, Jun Wan, Jun Wu, Junlei Wang, Junyi Shi, Justin Chiu, Justin Satriano, Justin Yee, Jwala Dhamala, Jyoti Bansal, Kai Zhen, Kai-Wei Chang, Kaixiang Lin, Kalyan Raman, Kanthashree Mysore Sathyendra, Karabo Moroe, Karan Bhandarkar, Karan Kothari, Karolina Owczarzak, Karthick Gopalswamy, Karthick Ravi, Karthik Ramakrishnan, Karthika Arumugam, Kartik Mehta, Katarzyna Konczalska, Kavya Ravikumar, Ke Tran, Kechen Qin, Kelin Li, Kelvin Li, Ketan Kulkarni, Kevin Angelo Rodrigues, Keyur Patel, Khadige Abboud, Kiana Hajebi, Klaus Reiter, Kris Schultz, Krishna Anisetty, Krishna Kotnana, Kristen Li, Kruthi Channamallikarjuna, Krzysztof Jakubczyk, Kuba Pierewoj, Kunal Pal, Kunwar Srivastav, Kyle Bannerman, Lahari Poddar, Lakshmi Prasad, Larry Tseng, Laxmikant Naik, Leena Chennuru Vankadara, Lenon Minorics, Leo Liu, Leonard Lausen, Leonardo F. R. Ribeiro, Li Zhang, Lili Gehorsam, Ling Qi, Lisa Bauer, Lori Knapp, Lu Zeng, Lucas Tong, Lulu Wong, Luoxin Chen, Maciej Rudnicki, Mahdi Namazifar, Mahesh Jaliminche, Maira Ladeira Tanke, Manasi Gupta, Mandeep Ahlawat, Mani Khanuja, Mani Sundaram, Marcin Leyk, Mariusz Momotko, Markus Boese, Markus Dreyer, Markus Mueller, Mason Fu, Mateusz Górski, Mateusz Mastalerczyk, Matias Mora, Matt Johnson, Matt Scott, Matthew Wen, Max Barysau, Maya Boumerdassi, Maya Krishnan, Mayank Gupta, Mayank Hirani, Mayank Kulkarni, Meganathan Narayanasamy, Melanie Bradford, Melanie Gens, Melissa Burke, Meng Jin, Miao Chen, Michael Denkowski, Michael Heymel, Michael Krestyaninov, Michal Obirek, Michalina Wichorowska, Michał Miotk, Milosz Watroba, Mingyi Hong, Mingzhi Yu, Miranda Liu, Mohamed Gouda, Mohammad El-Shabani, Mohammad Ghavamzadeh, Mohit Bansal, Morteza Ziyadi, Nan Xia, Nathan Susanj, Nav Bhasin, Neha Goswami, Nehal Belgamwar, Nicolas Anastassacos, Nicolas Bergeron, Nidhi Jain, Nihal Jain, Niharika Chopparapu, Nik Xu, Nikko Strom, Nikolaos Malandrakis, Nimisha Mishra, Ninad Parkhi, Ninareh Mehrabi, Nishita Sant, Nishtha Gupta, Nitesh Sekhar, Nithin Rajeev, Nithish Raja Chidambaram, Nitish Dhar, Noor Bhagwagar, Noy Konforty, Omar Babu, Omid Razavi, Orchid Majumder, Osama Dar, Oscar Hsu, Pablo Kvitca, Pallavi Pandey, Parker Seegmiller, Patrick Lange, Paul Ferraro, Payal Motwani, Pegah Kharazmi, Pei Wang, Pengfei Liu, Peter Bradtke, Peter Götz, Peter Zhou, Pichao Wang, Piotr Poskart, Pooja Sonawane, Pradeep Natarajan, Pradyun Ramadorai, Pralam Shah, Prasad Nirantar, Prasanthi Chavali, Prashan Wanigasekara, Prashant Saraf, Prashun Dey, Pratyush Pant, Prerak Pradhan, Preyaa Patel, Priyanka Dadlani, Prudhvee Narasimha Sadha, Qi Dong, Qian Hu, Qiaozi, Gao, Qing Liu, Quinn Lam, Quynh Do, R. Manmatha, Rachel Willis, Rafael Liu, Rafal Ellert, Rafal Kalinski, Rafi Al Attrach, Ragha Prasad, Ragini Prasad, Raguvir Kunani, Rahul Gupta, Rahul Sharma, Rahul Tewari, Rajaganesh Baskaran, Rajan Singh, Rajiv Gupta, Rajiv Reddy, Rajshekhar Das, Rakesh Chada, Rakesh Vaideeswaran Mahesh, Ram Chandrasekaran, Ramesh Nallapati, Ran Xue, Rashmi Gangadharaiah, Ravi Rachakonda, Renxian Zhang, Rexhina Blloshmi, Rishabh Agrawal, Robert Enyedi, Robert Lowe, Robik Shrestha, Robinson Piramuthu, Rohail Asad, Rohan Khanna, Rohan Mukherjee, Rohit Mittal, Rohit Prasad, Rohith Mysore Vijaya Kumar, Ron Diamant, Ruchita Gupta, Ruiwen Li, Ruoying Li, Rushabh Fegade, Ruxu Zhang, Ryan Arbow, Ryan Chen, Ryan Gabbard, Ryan Hoium, Ryan King, Sabarishkumar Iyer, Sachal Malick, Sahar Movaghati, Sai Balakavi, Sai Jakka, Sai Kashyap Paruvelli, Sai Muralidhar Jayanthi, Saicharan Shriram Mujumdar, Sainyam Kapoor, Sajjad Beygi, Saket Dingliwal, Saleh Soltan, Sam Ricklin, Sam Tucker, Sameer Sinha, Samridhi Choudhary, Samson Tan, Samuel Broscheit, Samuel Schulter, Sanchit Agarwal, Sandeep Atluri, Sander Valstar, Sanjana Shankar, Sanyukta Sanyukta, Sarthak Khanna, Sarvpriye Khetrapal, Satish Janakiraman, Saumil Shah, Saurabh Akolkar, Saurabh Giri, Saurabh Khandelwal, Saurabh Pawar, Saurabh Sahu, Sean Huang, Sejun Ra, Senthilkumar Gopal, Sergei Dobroshinsky, Shadi Saba, Shamik Roy, Shamit Lal, Shankar Ananthakrishnan, Sharon Li, Shashwat Srijan, Shekhar Bhide, Sheng Long Tang, Sheng Zha, Shereen Oraby, Sherif Mostafa, Shiqi Li, Shishir Bharathi, Shivam Prakash, Shiyuan Huang, Shreya Yembarwar, Shreyas Pansare, Shreyas Subramanian, Shrijeet Joshi, Shuai Liu, Shuai Tang, Shubham Chandak, Shubham Garg, Shubham Katiyar, Shubham Mehta, Shubham Srivastav, Shuo Yang, Siddalingesha D S, Siddharth Choudhary, Siddharth Singh Senger, Simon Babb, Sina Moeini, Siqi Deng, Siva Loganathan, Slawomir Domagala, Sneha Narkar, Sneha Wadhwa, Songyang Zhang, Songyao Jiang, Sony Trenous, Soumajyoti Sarkar, Soumya Saha, Sourabh Reddy, Sourav Dokania, Spurthideepika Sandiri, Spyros Matsoukas, Sravan Bodapati, Sri Harsha Reddy Wdaru, Sridevi Yagati Venkateshdatta, Srikanth Ronanki, Srinivasan R Veeravanallur, Sriram Venkatapathy, Sriramprabhu Sankaraguru, Sruthi Gorantla, Sruthi Karuturi, Stefan Schroedl, Subendhu Rongali, Subhasis Kundu, Suhaila Shakiah, Sukriti Tiwari, Sumit Bharti, Sumita Sami, Sumith Mathew, Sunny Yu, Sunwoo Kim, Suraj Bajirao Malode, Susana Cumplido Riel, Swapnil Palod, Swastik Roy, Syed Furqhan, Tagyoung Chung, Takuma Yoshitani, Taojiannan Yang, Tejaswi Chillakura, Tejwant Bajwa, Temi Lajumoke, Thanh Tran, Thomas Gueudre, Thomas Jung, Tianhui Li, Tim Seemman, Timothy Leffel, Tingting Xiang, Tirth Patel, Tobias Domhan, Tobias Falke, Toby Guo, Tom Li, Tomasz Horszczaruk, Tomasz Jedynak, Tushar Kulkarni, Tyst Marin, Tytus Metrycki, Tzu-Yen Wang, Umang Jain, Upendra Singh, Utkarsh Chirimar, Vaibhav Gupta, Vanshil Shah, Varad Deshpande, Varad Gunjal, Varsha Srikeshava, Varsha Vivek, Varun Bharadwaj, Varun Gangal, Varun Kumar, Venkatesh Elango, Vicente Ordonez, Victor Soto, Vignesh Radhakrishnan, Vihang Patel, Vikram Singh, Vinay Varma Kolanuvada, Vinayshekhar Bannihatti Kumar, Vincent Auvray, Vincent Cartillier, Vincent Ponzo, Violet Peng, Vishal Khandelwal, Vishal Naik, Vishvesh Sahasrabudhe, Vitaliy Korolev, Vivek Gokuladas, Vivek Madan, Vivek Subramanian, Volkan Cevher, Vrinda Gupta, Wael Hamza, Wei Zhang, Weitong Ruan, Weiwei Cheng, Wen Zhang, Wenbo Zhao, Wenyan Yao, Wenzhuo Ouyang, Wesley Dashner, William Campbell, William Lin, Willian Martin, Wyatt Pearson, Xiang Jiang, Xiangxing Lu, Xiangyang Shi, Xianwen Peng, Xiaofeng Gao, Xiaoge Jiang, Xiaohan Fei, Xiaohui Wang, Xiaozhou Joey Zhou, Xin Feng, Xinyan Zhao, Xinyao Wang, Xinyu Li, Xu Zhang, Xuan Wang, Xuandi Fu, Xueling Yuan, Xuning Wang, Yadunandana Rao, Yair Tavizon, Yan Rossiytsev, Yanbei Chen, Yang Liu, Yang Zou, Yangsook Park, Yannick Versley, Yanyan Zhang, Yash Patel, Yen-Cheng Lu, Yi Pan, Yi-Hsiang, Lai, Yichen Hu, Yida Wang, Yiheng Zhou, Yilin Xiang, Ying Shi, Ying Wang, Yishai Galatzer, Yongxin Wang, Yorick Shen, Yuchen Sun, Yudi Purwatama, Yue, Wu, Yue Gu, Yuechun Wang, Yujun Zeng, Yuncong Chen, Yunke Zhou, Yusheng Xie, Yvon Guy, Zbigniew Ambrozinski, Zhaowei Cai, Zhen Zhang, Zheng Wang, Zhenghui Jin, Zhewei Zhao, Zhiheng Li, Zhiheng Luo, Zhikang Zhang, Zhilin Fang, Zhiqi Bu, Zhiyuan Wang, Zhizhong Li, Zijian Wang, Zimeng, Qiu, Zishi Li
TL;DR
The paper introduces Amazon Nova, a foundation-model family evaluated across core capabilities, agentic performance, specialized domains, and runtime behavior. Its models demonstrate strong performance across text, multimodal, tool-use, and video evaluations, while benchmark currency limits interpretation of 3BFCL results.
Problem
Foundation models must be assessed across core capabilities, specialized domains, agentic tasks, and runtime performance.
Method
The paper benchmarks Amazon Nova models against selected publicly available models across core, multimodal, agentic, functional, runtime, and human-evaluation settings.
Results
Amazon Nova models demonstrate strong performance across core, multilingual, multimodal, and agentic benchmarks, while Nova Reel achieves video-consistency win rates of 67.0% versus Gen3 Alpha and 74.7% versus Luma 1.6.
Takeaways & Limitations
Amazon Nova provides models spanning multimodal, text-only, agentic, and video-generation capabilities with strong reported benchmark performance and runtime-focused evaluation.
Takeaways & Limitations
3BFCL results reflect the repository state and website leaderboard as of November 17, 2024, limiting their temporal comparability.
Abstract
from arXiv · showhide
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highly-capable multimodal model with the best combination of accuracy, speed, and cost for a wide range of tasks. Amazon Nova Lite is a low-cost multimodal model that is lightning fast for processing images, video, documents and text. Amazon Nova Micro is a text-only model that delivers our lowest-latency responses at very low cost. Amazon Nova Canvas is an image generation model that creates professional grade images with rich customization controls. Amazon Nova Reel is a video generation model offering high-quality outputs, customization, and motion control. Our models were built responsibly and with a commitment to customer trust, security, and reliability. We report benchmarking results for core capabilities, agentic performance, long context, functional adaptation, runtime performance, and human evaluation.
1 Introduction
Amazon Nova is a family of foundation and generative-media models designed to combine frontier intelligence, speed, customization, and price performance. The family spans text-only and multimodal understanding, agentic workflows, image generation and editing, and controllable video generation.
- Nova Pro, Lite, and Micro: Amazon Nova models target frontier intelligence, fast inference, agentic workflows, customizability, and low-cost price performance.Developers can fine-tune the models and distill larger models into smaller ones through Bedrock APIs.
- Nova Pro, Lite, and Micro: Amazon Nova Pro, Lite, and Micro are Transformer-based foundation models trained on large multilingual and multimodal data.Pro and Lite accept text, images, documents, and video; Micro is text-only.
- Nova Canvas and Reel: Amazon Nova Canvas generates images with adjustable resolution, aspect ratio, reference-image guidance, and image variations.It supports resolutions from 512 up to 2K horizontal resolution and aspect ratios between 1:4 and 4:1.
- Nova Canvas and Reel: Amazon Nova Canvas supports natural-language inpainting, outpainting, and background removal while preserving the image subject.Users specify the region to repaint through natural-language mask prompts.
- Nova Canvas and Reel: Amazon Nova Reel generates six-second 720p videos at 24 frames per second from text or reference images and supports more than 20 camera motions.Text prompts can control motions such as zoom and dolly forward.
- Nova Canvas and Reel: Canvas and Reel use latent diffusion, conditioning iterative denoising on tokenized text prompts and latent representations of images or video frames.A variational autoencoder maps visual inputs to latent variables before diffusion processing.
2 Amazon Nova Pro, Lite, and Micro Evaluations
Amazon Nova Pro, Lite, and Micro are evaluated across core capability, multimodal, agentic, long-context, financial reasoning, and runtime benchmarks. The reported results show strong performance across these evaluations, including high multimodal-agent scores, robust long-video question answering, and fast inference.
- Evaluation scope: The evaluation suite covers text-only and multimodal core capabilities, agentic workflows, long-context understanding, financial reasoning, and runtime performance.Text benchmarks span knowledge, reasoning, language understanding, multilinguality, and instruction following; multimodal benchmarks cover image and video understanding.
- Core capabilities: Nova Pro, Lite, and Micro demonstrate strong performance across core capability benchmarks, particularly on math, reasoning, and instruction following.Table 1 compares Nova models with publicly reported results for MMLU, ARC-C, DROP, GPQA, MATH, GSM8K, IFEval, and BBH.
- Multimodal capabilities: Nova Pro and Lite achieve high scores across image and video understanding benchmarks, ranking first or second on ChartQA and VATEX.The evaluated benchmarks include MMMU, ChartQA, DocVQA, TextVQA, VATEX, and EgoSchema.
- Agentic workflows: Nova Pro and Lite set a new state of the art on three visual-reasoning benchmarks requiring image understanding for correct function calling.These evaluations test image understanding capabilities in function-calling settings.
- Agentic and long-context performance: Nova Lite and Pro achieve high scores on all three multimodal agent benchmarks, while Nova models demonstrate robust LVBench long-video question answering.The multimodal agent benchmarks are VisualWebBench, MM-Mind2Web, and GroundUI-1K.
- Runtime performance: Nova Micro, Lite, and Pro are among the fastest models in their respective intelligence tiers under a 1,000-token input and 100-token output workload.Runtime evaluation measures TTFT, OTPS, and Total Response Time using results reported by Artificial Analysis.
3 Amazon Nova Canvas Evaluation
Amazon Nova Canvas is evaluated as a text-to-image diffusion model using automated metrics and single-blind human comparisons against public models. Human evaluation reports higher Canvas win rates than the compared text-to-image models.
- Model and evaluation strategy: Amazon Nova Canvas takes a text prompt and optional RGB image, generating an image conditioned on those inputs.The model is evaluated with both automated metrics and human evaluation.
- Automated metrics: ImageReward and TIFA measure generated-image quality and faithfulness to the input text.ImageReward uses a reward model aligned with human preference, while TIFA uses visual question answering on a reference-free benchmark.
- Automated metrics: Canvas is compared with DALL.E 3, Stable Diffusion 3 Medium, Stable Diffusion 3.5 Large, and Flux Schnell and Pro.The comparison results are reported in Table 8.
- Human evaluation: Human evaluation uses approximately 1,000 customer-oriented prompts spanning categories such as humans, landscapes, indoor environments, and creative themes.Images are generated at 1k x 1k resolution and evaluated in a single-blind pairwise setup.
- Human evaluation: Canvas achieves higher human-evaluation win rates than the other compared text-to-image models.Table 9 reports win, tie, and loss rates for comparisons with DALL.E 3 and Imagen 3.
4 Amazon Nova Reel Evaluation
Amazon Nova Reel is a text-and-image-conditioned video diffusion model evaluated through single-blind human pairwise comparisons. The evaluation separates video quality from temporal consistency and reports stronger consistency win rates against both comparison models.
- Model and evaluation: Amazon Nova Reel takes a text prompt and optional RGB image to generate a video conditioned on the inputs.The section evaluates the model’s generated videos.
- Model and evaluation: Human evaluators compare videos side by side on video quality and video consistency, choosing a preferred video or marking a tie.All videos are generated at 720p resolution.
- Evaluation axes: Video quality covers image quality, motion quality, image-text alignment, motion-text alignment, motion degree, entity size, composition, and likability.These components combine technical and perceptual aspects of generated videos.
- Evaluation axes: Video consistency measures temporal coherence of subjects and backgrounds, including stable entity appearance and spatial relationships.The axis focuses on avoiding unexpected morphing or changes throughout the video.
- Evaluation design: The prompt set spans six categories and includes motion-related instructions such as camera movements and dynamic attributes.This design targets broad real-world coverage and motion-text alignment.
- Results: 67.0% and 74.7% are Nova Reel’s video-consistency win rates against Gen3 Alpha and Luma 1.6, respectively.For video quality, Nova Reel’s win rates are 56.4% against Gen3 Alpha and 51.1% against Luma 1.6.
5 Responsible AI
Amazon Nova’s Responsible AI approach organizes development around eight dimensions, combines alignment and runtime safeguards, and uses benchmark-based and red-team evaluations. Feedback from repeated testing is used to guide subsequent model development.
- Responsible AI framework: Amazon Nova’s Responsible AI approach is organized around eight foundational dimensions spanning design objectives, adherence actions, and system evaluation.The final two components form a continuous model-development and verification loop.
- Responsible AI framework: Responsible AI objectives guide decisions from data collection and pretraining through post-deployment runtime mitigations.The objectives are informed by laws, regulations, voluntary frameworks, and customer commitments.
- Safeguards: Model behavior is aligned using curated pretraining data, Supervised Fine Tuning, and Reinforcement Learning from Human Feedback across core Responsible AI dimensions.The stated dimensions include Safety, Fairness, Veracity and Robustness, Controllability, and Privacy and Security.
- Safeguards: Runtime input and output moderation models detect malicious or insecure prompts and help ensure generated content adheres to requirements.Input moderation addresses prompt injection and jailbreaking attempts.
- Safeguards: Transparency measures include invisible watermarks and C2PA metadata for Canvas-generated content, with watermarks embedded in every video frame.The methods are designed to withstand alterations such as resizing, flipping, and H264 compression.
- Evaluation and iteration: Static benchmarks, proprietary dynamically updating benchmarks, and internal, external, and automated red teaming evaluate Responsible AI dimensions and masked user intent.Red-team feedback guides the next stage of model development.
- External red teaming: ActiveFence produced over 9,700 adversarial prompts across 20 categories for external safety and security testing.The testing covered harmful-content generation, malicious manipulation, and sensitive-information extraction.
6 Training Infrastructure
The Nova family was trained across Trainium1, NVIDIA A100, and H100 accelerators using parallel training infrastructure. System optimizations improved training goodput, checkpointing, restart recovery, and memory efficiency.
- Training platform: Nova models were trained on Trainium1 chips, NVIDIA A100 accelerators, and H100 accelerators with parallel trainings used to maintain performance parity.The clusters used petabit-scale non-blocking EFA networking.
- Runtime efficiency: Up to 97% weekly average goodput was achieved in pretraining runs through lower job failure rates, reduced checkpoint overhead, and shorter restart times.Checkpointing overhead reached approximately 1 second on H100 clusters and 0.1 seconds on Trainium1 clusters.
- Runtime efficiency: Checkpoint discovery fell from 3 minutes to 5 seconds through an asynchronous observer that maps checkpoint files to cluster nodes.Data-loading initialization was reduced to 205ms per restart.
- Memory efficiency: Super-Selective Activation Checkpointing reduces memory consumption by approximately 50% while adding approximately 2% recomputation overhead versus NVIDIA’s Selective Checkpointing.The scheme targets memory-constrained training environments.
A Amazon Nova Canvas Capabilities
Amazon Nova Canvas is an image content-generation model with multiple image-creation and editing capabilities. Figure 5 provides examples of these capabilities.
- Nova Canvas supports text-to-image generation from 512×512 up to 2K×2K resolution.
- It supports text- and mask-based editing, including inpainting, outpainting, and object removal.
- Image variation produces similar content with variations, while image conditioning follows a reference image’s layout and structure.
- Color-palette guidance uses supplied hex codes, and background removal automatically removes backgrounds from multi-object images.
- Figure 5 shows example capabilities of Amazon Nova Canvas for image content generation.
B Prompts and Scoring
The report provides prompt templates for Amazon Nova evaluations and selected public models. It also directs readers to additional report materials and evaluation results.
- Prompt templates are provided for Amazon Nova evaluations.
- The report includes templates used for select other public models where noted.
- Additional materials and evaluation results are available through a referenced location.
B.1 Text evaluation
The text-evaluation section specifies task prompts, demonstrations, and answer-format instructions across several benchmarks. It includes examples for arithmetic, sports, historical, demographic, and reasoning questions.
- Multiple-choice evaluation prompts ask models to select the best answer and often require a standardized concluding answer format.
- Examples cover question answering involving sports events, historical dates, demographic percentages, and numerical reasoning.
- The evaluation uses task-specific prompting, including preambles and few-shot examples for BBH and six shots for DROP.
- Subject-specific instructions define required answer forms for Boolean expressions, causal judgment, date understanding, and other tasks.
B.2 Multimodal evaluation
The multimodal evaluation section defines prompt formats for image, video, OCR, web, and navigation tasks. Prompts specify expected answer formats and, for navigation, require selecting one valid next action from the current webpage state.
- Image and video evaluations include multiple-choice and open-ended prompts with direct-answer or concise-summary requirements.
- Video prompts sample frames evenly for question answering and require concise captions without hallucinating objects.
- Web evaluation tasks cover OCR, action prediction, element grounding, action grounding, heading OCR, and web question answering.
- Navigation prompts require analyzing the current screenshot and previous actions before choosing the first next webpage operation.
- The navigation protocol restricts each step to one valid action, such as clicking, selecting, typing, or pressing Enter.
- Final navigation responses separate the element choice, action, and value into three lines, with NONE reserved for the none option.
B.3 Functional Capabilities
The functional-capabilities materials specify prompts and answer-processing procedures for finance question answering and quiz grading. They define extraction, numerical normalization, and correctness criteria for evaluated answers.
- Finance question answering: Finance questions require step-by-step analysis and a final answer in a specified format.The response must begin with “Lets think step-by-step” and end with a concise answer prefixed by “The answer is”.
- Finance question answering: Supporting facts for finance questions may include pre-text, tables, and post-text.
- Answer evaluation: Answers are extracted with the regex “The answer is (.*)” and converted to decimal numerical representations when needed.Examples include converting percentages and magnitude terms before evaluation.
- Answer evaluation: Numerical answers are judged correct when they match the ground truth after rounding to the same decimal places.
- Quiz grading: Quiz responses are graded as Correct or Incorrect based only on factual accuracy, ignoring punctuation and phrasing differences.The grading format includes the question, student answer, true answer, and a correctness label.
- Quiz grading: The grading output includes a concise one- or two-sentence justification and a correctness field.
C Qualitative examples of multimodal intelligence
The report provides qualitative examples of Nova-generated images and multimodal-agent behavior. The examples include outputs created with Nova Pro and Nova Lite, including images sourced from external datasets or team photography.
- Nova Pro examples: Figures 6, 8, and 9 show examples created with Nova Pro.Figure 6 uses a team member’s photograph, Figure 8 cites image source [88], and Figure 9 provides no source attribution in the passage.
- Multimodal agents: Figure 7 presents an example of a multimodal agent.
- Nova Lite examples: Figure 10 shows an image created with Nova Lite using an image from the ChartQA dataset.
D Correspondence and Contributors
The report identifies its correspondence information, organizational origin, and citation convention. It attributes the Nova family to Amazon AGI and partner teams and requests citation under “Amazon AGI” alone.
- Correspondence: The report provides a designated channel for correspondence.
- Organization: The Nova family was built by Amazon Artificial General Intelligence and partner teams.
- Citation: The report asks readers to cite “Amazon AGI” as the sole author.The requested author name is shown in the accompanying BibTeX entry.
D.1 Contributors
The contributor section states that listed individuals worked in the Nova program for at least one-fifth of its duration and measurably affected the reported models or services. It then provides an extensive alphabetical list of contributors.
- Contributor criteria: Contributors are defined as individuals who worked in the Nova program for at least one-fifth of its duration and measurably impacted a reported model or service.
- Contributor list: The section lists contributors alphabetically from Aaron Langford through Anoop Deoras.
- Contributor list: The alphabetical list continues through names beginning with Anton, Gonzalo, Daniel, and Bincheng.
- Contributor list: The list includes contributors beginning with Matthew, Katarzyna, and Jerzy.
- Contributor list: Additional listed contributors include Shamik, Robert, Parker, Yadunandana, Vignesh, and Sumit.