Source-linked AI summary

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, Danny Wyatt, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Francisco Guzmán, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Govind Thattai, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jack Zhang, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Karthik Prasad, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Kushal Lakhotia, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Maria Tsimpoukelli, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Ning Zhang, Olivier Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohan Maheswari, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vítor Albiero, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaofang Wang, Xiaoqing Ellen Tan, Xide Xia, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aayushi Srivastava, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Amos Teo, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Dong, Annie Franco, Anuj Goyal, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Ce Liu, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Cynthia Gao, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Eric-Tuan Le, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Filippos Kokkinos, Firat Ozgenel, Francesco Caggioni, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hakan Inan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Hongyuan Zhan, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Ilias Leontiadis, Irina-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Janice Lam, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kiran Jagadeesh, Kun Huang, Kunal Chawla, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Miao Liu, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikhil Mehta, Nikolay Pavlovich Laptev, Ning Dong, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Rangaprabhu Parthasarathy, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Russ Howes, Ruty Rinott, Sachin Mehta, Sachin Siby, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Mahajan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shishir Patil, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta, Summer Deng, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Koehler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaojian Wu, Xiaolan Wang, Xilun Wu, Xinbo Gao, Yaniv Kleinman, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yu Zhao, Yuchen Hao, Yundi Qian, Yunlu Li, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, Zhiwei Zhao, Zhiyu Ma

arXiv:2407.21783v3cs.AIcs.CLcs.CV

TL;DR

Foundation models power modern AI, but their capabilities across multilingual, coding, reasoning, tool-use, and broad language tasks require extensive evaluation. This paper develops and evaluates Llama 3, finding flagship performance on par with leading models such as GPT-4 across varied tasks while smaller models outperform similarly sized alternatives.

  • Problem

    The paper asks how a new foundation-model family performs across multilingual, coding, reasoning, tool-use, and broad language-understanding tasks.

  • Method

    It develops Llama 3, a multilingual herd of 8B, 70B, and 405B language models, and evaluates them on benchmarks and human comparisons.

  • Results

    Llama 3's flagship performs on par with leading models such as GPT-4 across varied tasks, while smaller models outperform similarly sized alternatives.

  • Takeaways & Limitations

    The authors publicly release Llama 3 language models to support research scrutiny and development of societally relevant AI systems.

  • Takeaways & Limitations

    Testing is not exhaustive: Llama 3 may still generate harmful content, especially beyond English or under skilled adversarial prompting.

Abstract

from arXiv · show

Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a context window of up to 128K tokens. This paper presents an extensive empirical evaluation of Llama 3. We find that Llama 3 delivers comparable quality to leading language models such as GPT-4 on a plethora of tasks. We publicly release Llama 3, including pre-trained and post-trained versions of the 405B parameter language model and our Llama Guard 3 model for input and output safety. The paper also presents the results of experiments in which we integrate image, video, and speech capabilities into Llama 3 via a compositional approach. We observe this approach performs competitively with the state-of-the-art on image, video, and speech recognition tasks. The resulting models are not yet being broadly released as they are still under development.

1 Introduction

The paper introduces Llama 3, a family of multilingual foundation models supporting coding, reasoning, and tool usage, including a 405B-parameter model with a 128K-token context window. It evaluates the models extensively, publicly releases the language and safety models, and reports initial multimodal experiments whose models remain under development.

  • Model family: Llama 3 is a herd of multilingual language models supporting multilinguality, coding, reasoning, and tool usage, with up to 405B parameters and a 128K-token context window.The herd includes 8B, 70B, and 405B parameter models.
  • Development approach: The models are developed around improving data, scale, and complexity management.The paper describes improved data quantity and quality, larger-scale training, and design choices intended to maximize scalable development.
  • Evaluation: The paper evaluates Llama 3 across numerous language-understanding benchmarks and through extensive human evaluations against competing models.An overview of flagship-model performance is provided in Table 2.
  • Release: All three Llama 3 models are publicly released under an updated Llama 3 Community License, including pre-trained and post-trained 405B models and Llama Guard for input and output safety.The release is intended to support research-community innovation and responsible AGI development.
  • Multimodal extensions: The paper reports initial multimodal experiments extending Llama 3 to image recognition, video recognition, and speech understanding, but these models are not yet ready for release.The multimodal extensions remain under active development.

2 General Overview

Llama 3 is developed through large-scale language-model pre-training followed by human-feedback post-training, yielding models with multilingual, coding, reasoning, and tool-use capabilities. Separate multimodal experiments add image, video, and speech capabilities through a compositional approach, but those models remain under development and unreleased.

  • Language model training: Llama 3 development comprises language-model pre-training and language-model post-training.Pre-training uses next-token prediction on multilingual text, while post-training aligns the model with human feedback.
  • Language model training: 405B parameters are pre-trained on 15.6T tokens to learn language structure and world knowledge.The pre-training corpus is large and multilingual, and training uses next-token prediction.
  • Language model training: Post-training uses repeated supervised finetuning and Direct Preference Optimization to align the model with human feedback.This stage also adds tool use, improves coding and reasoning, and incorporates safety mitigations.
  • Capabilities: The resulting models answer questions in at least eight languages, write high-quality code, solve complex reasoning problems, and use tools zero-shot.They can use tools out-of-the-box or in a zero-shot way.
  • Multimodal experiments: A compositional approach adds image, video, and speech capabilities through multimodal encoder and adapter training.The resulting models recognize image and video content and support speech interaction, but remain under development and unreleased.

3 Pre-Training

Llama 3 pre-training combines extensive data curation, a multilingual tokenizer, compute-optimal scaling, and large-scale training systems. The resulting pipeline supports a 405B-parameter flagship model with robust performance and over 90% effective training time.

  • Data curation: The dataset applies de-duplication, cleaning, and filters removing domains with substantial PII or known adult content.Filtering also targets unsafe websites and domains ranked harmful under Meta safety standards.
  • Data mix: The final data mix contains roughly 50% general-knowledge tokens, 25% mathematical and reasoning tokens, 17% code tokens, and 8% multilingual tokens.These proportions define the composition of the training corpus used for scaling-law analysis.
  • Pre-training recipe: Annealing improved Llama 3 8B validation performance by 24.0% on GSM8k and 6.4% on MATH, while improvements on the 405B model were negligible.The results suggest the flagship model did not require this specialized annealing for its in-context learning and reasoning capabilities.
  • Tokenizer: Adding 28K non-English tokenizer tokens improved English compression from 3.17 to 3.94 characters per token while improving multilingual compression and downstream performance without affecting English tokenization.The improved compression enables the model to read more text for the same training compute.
  • Scaling laws: Scaling-law extrapolation to 3.8 × 10^25 FLOPs suggested 402B parameters trained on 16.55T tokens, leading to the selected 405B-parameter flagship model.Flatter IsoFLOPs minima at higher compute budgets indicated robustness to small model-size versus token-count trade-offs.
  • Large-scale training: Training achieved 38-43% BF16 Model FLOPs Utilization across the reported configurations and higher than 90% effective training time despite 16K-GPU reliability challenges.Synchronous training made failures costly because a single GPU failure could require restarting the entire job.

4 Post-Training

Aligned Llama 3 models are produced through repeated post-training rounds that combine supervised finetuning with Direct Preference Optimization on human-annotated or synthetically generated examples.

  • Post-training procedure: Each post-training round applies supervised finetuning followed by Direct Preference Optimization to examples collected through human annotation or synthetic generation.The aligned models are trained on top of a pre-trained checkpoint through several rounds of alignment with human feedback.

4.1 Modeling

Llama 3’s post-training modeling pipeline defines a multi-message chat protocol and iteratively combines reward modeling, supervised finetuning, and Direct Preference Optimization. Across six rounds, each cycle gathers new preference and SFT data while sampling synthetic data from the latest models.

  • Post-training pipeline: The post-training pipeline trains a reward model, applies rejection sampling, finetunes with SFT, and further aligns models using DPO.SFT uses cross-entropy loss on target tokens while masking prompt tokens; DPO follows SFT for preference alignment.
  • Chat protocol: Llama 3 introduces a multi-message chat protocol with header and termination tokens to support tool use and message routing within a dialog turn.Header tokens identify message sources and destinations, while termination tokens mark turns between human and AI speakers.
  • Preference alignment: DPO primarily uses recent preference batches from prior alignment rounds, matching training data more closely to the policy model’s distribution.The authors found DPO more compute-efficient and better-performing than PPO for large-scale models, especially on instruction following.
  • Iterative training: The procedure repeats for six rounds, collecting new preference annotations and SFT data while sampling synthetic data from the latest models.Model averaging is also performed across experiments using different data versions or hyperparameters at each reward-model, SFT, or DPO stage.

4.2 Post-training Data

Llama 3 post-training data combines preference annotations with rejection-sampled, synthetic, and human-curated SFT data. The pipeline iteratively improves data quality through reward-based selection, rule-based cleaning, model-based pruning, and semantic deduplication.

  • Preference data: Preference training uses significantly better or better chosen responses, retaining all available preference data for reward modeling but only recent capability batches for DPO.Samples with similar chosen and rejected responses are discarded.
  • SFT data composition: SFT data combines human-annotation prompts with rejection-sampled responses, synthetic capability-targeted data, and small amounts of human-curated data.These sources form the internally collected alignment dataset described in Table 7.
  • Rejection sampling: Rejection sampling generates typically 10–30 outputs per prompt, then uses a reward model to select the best candidate from the latest chat-model policy.Later rounds add system prompts to control tone, style, or formatting across capabilities.
  • Data composition: The overall data mix is adjusted across topic, complexity, and quality to tune performance across a wide range of benchmarks.SFT and preference data can overlap in domain while retaining distinct curation and count statistics.
  • Data cleaning: Rule-based cleaning removes problematic patterns, including excessive emojis, exclamation points, and overused apologetic phrases.The process balances the proportion of samples containing phrases such as “I’m sorry” or “I apologize.”
  • Data pruning: Model-based pruning classifies topics, scores quality and difficulty, and semantically deduplicates dialogs using clustered quality × difficulty ranking and cosine-similarity filtering.Quality signals include reward models and Llama-based ratings, while difficulty uses Instag and Llama-based scoring.

4.3 Capabilities

Llama 3 improves capabilities across code, multilinguality, math and reasoning, long context, and tool use through specialized data, training strategies, and iterative feedback. The reported methods include multilingual specialization, translated math data, long-context data mixing, and multi-step tool use.

  • Code: Llama 3 targets code generation, documentation, debugging, and review across ten high-priority programming languages.The languages include Python, Java, Javascript, C/C++, Typescript, Rust, PHP, HTML/CSS, SQL, and bash/shell.
  • Multilinguality: Llama 3 improves multilinguality with a specialized expert, more multilingual data, multilingual instruction tuning, and improved language steering.Instruction-tuning data covers German, French, Italian, Portuguese, Hindi, Spanish, and Thai.
  • Math and reasoning: Translated math problems produce little to no quality issues and yield strong gains on MGSM.The passage attributes these gains to adding translated data.
  • Math and reasoning: Iteratively using feedback from incorrect attempts to correct them improves accurate reasoning and learning from mistakes.The iterative process uses feedback from incorrect attempts and corrections.
  • Long context: Mixing 0.1% synthetic long-context data with original short-context data optimizes performance across short-context and long-context benchmarks.Ablation studies support this data mixture.
  • Tool use: In chat, Llama 3 can solve queries with multiple sequential tool calls, step-by-step planning, and reasoning after each call.The resulting model supports these behaviors in multi-turn dialogs.

5 Results

Section 5 reports extensive evaluations of Llama 3 covering its pre-trained model, post-trained model, and safety characteristics. The results are organized into separate subsections for these evaluation areas.

  • The evaluation covers Llama 3’s pre-trained language model performance.
  • The evaluation covers Llama 3’s post-trained language model performance.
  • The evaluation investigates Llama 3’s safety characteristics, with results presented in separate subsections.

5.1 Pre-trained Language Model

Pre-trained Llama 3 models are evaluated across broad capability benchmarks, with strong performance for the 8B, 70B, and 405B models. The evaluation also examines statistical uncertainty, robustness to multiple-choice design choices, adversarial behavior, and training-data contamination.

  • Standard benchmarks: Llama 3 is evaluated across eight categories, including commonsense reasoning, knowledge, reading comprehension, math and problem solving, long context, code, adversarial evaluation, and aggregate tasks.The study compares models with other pre-trained models of comparable sizes.
  • Evaluation methodology: Benchmark scores are reported with 95% confidence intervals when applicable, but these intervals lower-bound actual capability-estimate variation because subsampling is not the only source of variation.The confidence intervals assume Gaussian-distributed benchmark scores, while the paper notes that this assumption is imperfect.
  • Standard benchmarks: Llama 3 8B outperforms competing models in virtually every evaluated category, while Llama 3 70B substantially improves over Llama 2 70B on most benchmarks.Figure 12 aggregates performance by averaging accuracies across benchmarks within each capability category.
  • Standard benchmarks: Llama 3 405B performs competitively with models in its class and substantially outperforms prior open-source models.Long-context results beyond the reported benchmark comparisons are presented separately in Section 5.2.
  • Robustness: Pre-trained Llama 3 models are very robust to changes in multiple-choice labels and few-shot label structure, especially the 405B model, and remain robust across answer orders and prompt formats.These experiments examine sensitivity to label variants, few-shot label bias, answer order, and prompt format in MMLU.

5.2 Post-trained Language Model

Llama 3 post-trained models are evaluated across general knowledge, instruction following, proficiency exams, coding, multilingual reasoning, long-context understanding, and tool use. Across these capabilities, the models generally outperform comparable systems, with especially strong results from the 405B and 70B variants.

  • General knowledge and instruction following: The 405B model outperforms GPT-4 and Nemotron 4 340B on general knowledge tasks, while Claude 3.5 Sonnet leads among larger models.The 8B and 70B variants also outperform similarly sized models on MMLU and MMLU-Pro.
  • General knowledge and instruction following: All Llama 3 variants outperform comparable models on IFEval instruction-following evaluations.IFEval measures prompt-level and instruction-level accuracy under strict and loose constraints.
  • Proficiency exams: The 70B model significantly outperforms GPT-3.5 Turbo and beats Nemotron 4 340B on many proficiency exams, while 405B performs similarly to Claude 3.5 Sonnet and GPT-4o.The evaluation covers GRE, LSAT, SAT, GMAT, and AP exams, reporting normalized GRE scores and accuracy for other exams.
  • Multilingual reasoning: For multilingual reasoning, Llama 3 405B averages 91.6% on MGSM but trails GPT-4o by 2% on MMLU, while 70B and 8B lead competitors on both tasks.The models demonstrate strong multilingual performance despite the 405B model’s MMLU gap.
  • Long-context understanding: Llama 3 models retrieve 100% of needles across document depths and context lengths, and the 405B model outperforms all others on InfiniteBench.On ZeroSCROLLS, the 405B and 70B models either match or surpass other models across various tasks.
  • Tool use: Llama 3 performs strongly on tool-use benchmarks: 8B leads its category on Nexus and BFCL, while 405B significantly beats GPT-4o on text-only code execution and plot generation but lags on file upload.On API-Bank, 8B and 70B outperform same-category models, while 405B trails Claude 3.5 Sonnet by 0.6%.

5.3 Human Evaluations

Human evaluations were designed to capture nuanced, user-facing qualities across diverse capabilities and difficulty levels. Llama 3 405B performed approximately on par with GPT-4, but had mixed results against GPT-4o and Claude 3.5 Sonnet, while human ratings remain susceptible to annotator subjectivity.

  • Motivation: Human evaluations measured subtle performance aspects, including tone, verbosity, nuance, and cultural-context understanding, to approximate real-world user experience.These evaluations complemented standard benchmark testing.
  • Prompt collection: About 7,000 prompts covered English, reasoning, coding, Hindi, Spanish, and Portuguese, plus English, reasoning, and coding multiturn capabilities.The prompt taxonomy spanned broad model capabilities and difficulties.
  • Prompt collection: 60% of prompts were hard, compared with 30% medium and 10% easy, and modeling teams lacked access to prompts to limit contamination and overfitting.All prompt sets underwent thorough quality assurance.
  • Results: Llama 3 405B performed approximately on par with GPT-4, while producing mixed wins and losses against GPT-4o and Claude 3.5 Sonnet.On nearly all capabilities, Llama 3 and GPT-4 win rates were within the margin described in the evaluation results.
  • Limitations: Human evaluation results can be inconsistent or unreliable because annotator biases, backgrounds, and preferences influence judgments of model responses.The paper notes that objective criteria for evaluating responses are difficult to define.

5.4 Safety

Llama 3’s safety development combines internal violation and false-refusal benchmarks with filtering, memorization analysis, and quality-focused safety-data practices. The models show competitive safety-helpfulness trade-offs across languages, tools, and systems, while remaining susceptible to several cyber-abuse behaviors and facing non-reproducible internal evaluations.

  • Safety evaluation: Llama 3 evaluates safety using internal violation and false-refusal benchmarks, whose combined size exceeds 4000 prompts.False refusal measures unhelpful refusals when a plausible, safe response is possible, using borderline prompts near the decision boundary.
  • Safety development: Safety controls span pre-training filters for likely personally identifiable information and discoverable-memorization analysis.The memorization analysis samples prompts and ground truths at different training-data frequencies using a rolling hash index of corpus n-grams.
  • Safety development: Safety-data quality matters more than quantity, motivating AI-assisted annotation tools and rigorous quality-assurance processes for human-generated data.Human-generated data can contain errors and inconsistencies, particularly for nuanced safety policies.
  • Overall safety: Smaller models require a larger proportion of safety data relative to helpfulness and make balancing violation and false-refusal rates more difficult.The model-size trade-off is examined through the relationship between FRR and VR.
  • Comparative safety: Llama 405B is Pareto-better than Comp. 2 across violation and false-refusal rates on both DocQA and Many-shot, while being significantly safer than Comp. 1 with a false-refusal trade-off.Against Comp. 1 in tool-usage search evaluations, Llama 405B is also significantly safer but has a slightly higher false-refusal rate; multilingual results find Llama 405B with Llama Guard at least as safe as comparable systems for supported languages.
  • Cybersecurity safety: Llama 3 405B complied with malicious code-interpreter prompts 10.4% of the time and succumbed to text-based prompt injection 21.7% of the time.Llama 3 70B complied with malicious prompts at 3.8%; overall, the paper reports no significant susceptibility to generating malicious code or exploiting vulnerabilities, while noting these internal benchmark results are not externally reproducible.

6 Inference

Llama 3 405B inference uses pipeline parallelism across two machines and FP8 quantization to address memory and efficiency constraints. Micro-batching improves throughput-latency trade-offs, while FP8 provides substantial pre-fill throughput gains but requires safeguards against corrupted responses.

  • Pipeline parallelism: Llama 3 405B does not fit in one machine’s 8 Nvidia H100 GPUs with BF16 parameters, so inference uses BF16 across 16 GPUs on two machines.Tensor parallelism is used within machines, while pipeline parallelism handles lower-bandwidth, higher-latency cross-node connectivity.
  • Pipeline parallelism: Micro-batching improves inference throughput at the same local batch size for 4,096 input tokens and 256 output tokens.It enables concurrent execution of micro-batches during both key-value cache pre-fill and decoding, while added synchronization points increase latency.
  • FP8 quantization: FP8 quantization is applied to most feedforward-network parameters and activations, covering roughly 50% of inference compute time, while self-attention parameters remain unquantized.The approach uses H100 native FP8 support and dynamic scaling factors.
  • FP8 quantization: Unbounded FP8 scaling factors can occasionally produce corrupted responses despite strong benchmark performance, so benchmark parity with BF16 is insufficient to assess quantization effects.The paper notes that standard benchmarks may show FP8 on par with BF16 even when distribution changes cause decoding errors.
  • FP8 quantization: FP8 inference improves pre-fill throughput by up to 50% compared with the two-machine BF16 approach for workloads with 4,096 input tokens and 256 output tokens.Figure 27 compares throughput-latency trade-offs for FP8 and BF16 inference during pre-filling and decoding.

7 Vision Experiments

Llama 3 adds visual-recognition capabilities compositionally by connecting pretrained vision and language components, then extending them for temporal video understanding. Across image and video benchmarks, the resulting models are competitive with leading multimodal systems, with Llama 3-V 405B outperforming GPT-4V on every reported image benchmark.

  • Compositional architecture: The vision system first trains cross-attention layers between a pretrained image encoder and language model, then adds temporal aggregators and video cross-attention layers.These stages use large image-text and video-text datasets to learn image recognition and temporal video processing.
  • Training data: Image-text data processing combines quality filtering, perceptual de-duplication, resampling, OCR enrichment, and safety mitigations.The pipeline removes low-quality pairs, reduces redundancy, improves low-frequency and fine-grained recognition, and adds text extracted from images.
  • Training data: 512-dimensional SSCD image representations support nearest-neighbor de-duplication at scale for efficiency and privacy.The representations are computed for all images before nearest-neighbor search across the dataset.
  • Image results: Llama 3-V 405B outperforms GPT-4V on all reported image-recognition benchmarks while trailing Gemini 1.5 Pro and Claude 3.5 Sonnet.The vision module remains competitive across benchmarks and capacities, with particularly strong performance on document understanding.
  • Video results: Zero-shot Llama 3 models with small video adapters are very competitive, sometimes outperforming models that may use native multimodal processing from pre-training.The reported comparison includes two Gemini and two GPT-4 models.

8 Speech Experiments

Llama 3 integrates speech capabilities compositionally, using an encoder and adapter for speech understanding and streaming TTS for speech generation. Evaluations show strong multilingual ASR, speech translation, spoken question answering, safety, and low-latency synthesis capabilities.

  • Speech interface: Llama 3 adds speech understanding through an encoder-adapter interface with text system prompts that enable different operating modes.Without a system prompt, it functions as a general-purpose spoken dialogue model.
  • Training data: Approximately 15M hours of multilingual unlabeled speech initialize the speech encoder, while supervised data targets recognition, translation, and spoken dialogue.The ASR corpus contains 230K hours across 34 languages, and the AST corpus contains 90K translation hours across 33-language directions to English and English-to-33-language directions.
  • Speech recognition: Llama 3 outperforms speech-specialized Whisper and SeamlessM4T on all reported ASR benchmarks and performs similarly to Gemini on MLS English.These results demonstrate strong speech-recognition performance for Llama 3 and multimodal foundation models more generally.
  • Speech translation: Speech translation evaluations use FLEURS and Covost 2 and measure translated English with BLEU scores, highlighting multimodal foundation models’ advantages.The task translates non-English speech into English text.
  • Spoken question answering: The speech interface handles code-switched speech without prior exposure and supports coherent extended multi-turn dialogue despite single-turn training.Examples in the paper illustrate multilingual and multi-turn capabilities.
  • Speech generation: The Llama 3 8B prosody model is preferred 63.6% of the time over a non-streaming baseline, while its streaming phone-only baseline receives 40.0%.Token-wise streaming reduces lookahead requirements and enables more responsive speech synthesis than non-streaming baselines.

9 Related Work

Llama 3 builds on prior work across foundation-model scaling, efficiency, architectures, open weights, post-training, and multimodal modeling. Its development combines increased compute and improved data, compute-for-inference tradeoffs, dense modeling, established alignment methods, and compositional approaches for images, videos, and speech.

  • Scale: The 405B model uses almost fifty times Llama 2 70B’s pre-training compute budget, reflecting scaling through increased compute and improved data.Despite its 405B parameters, it has fewer parameters than earlier, less performant models such as PALM, attributed to improved understanding of scaling laws.
  • Small models: Smaller Llama 3 models trade additional training compute for inference efficiency, while model distillation offers an alternative route to reducing parameter counts.Fewer-parameter models can reduce inference cost and simplify deployment.
  • Architectures: Llama 3 uses minimal architectural modifications relative to Llama 2, while mixture-of-experts models increase capacity efficiently through alternative designs.The passage states that Llama 3 outperforms these mixture-of-experts models, but the sentence is truncated before specifying the resulting architectural conclusion.
  • Open source: Llama3-405B is competitive with the current closed-weight state of the art, amid rapid improvement and proliferation of open-weights foundation-model families.The passage lists Mistral, Falcon, MPT, Pythia, Arctic, OpenELM, OLMo, StableLM, and OpenLLaMA among recent families.
  • Post-training: Llama 3 post-training follows instruction tuning and human-feedback alignment, using millions of human instructions and preference judgments across multiple rounds.The methods include rejection sampling, supervised finetuning, and Direct Preference Optimization, with earlier Llama 3 versions filtering, rewriting, or generating training examples.
  • Multimodal capabilities: Llama 3’s multimodal work combines prior compositional approaches for images, videos, and speech, including image-text integration, video-language adapters, and speech-language modeling.For images, the approach achieves results comparable with Gemini 1.0 Ultra and GPT-4 Vision; for speech, it avoids finetuning the language model itself for speech tasks.

10 Conclusion

The authors conclude that high-quality data, scale, and simplicity were consistently the most effective development principles for Llama 3, while substantial further improvements remain possible. They shared development details and preliminary multimodal results to inform research, and publicly released the language models following positive safety analyses.

  • Development lessons: High-quality data, scale, and simplicity consistently yielded the best results during Llama 3 development, while substantial further improvements remain possible.Preliminary experiments with more complex architectures and training recipes did not show benefits.
  • Sharing results: The authors shared their development process to clarify key factors in foundation-model development and support more informed public debate.They identify both research understanding and public discussion as motivations for sharing.
  • Sharing results: Preliminary multimodal integration experiments were shared to accelerate research, although the resulting models remain under development and are not ready for release.The multimodal capabilities discussed include image, video, and speech integration.
  • Public release: Following positive safety analyses, the authors publicly released Llama 3 language models to accelerate societally relevant AI development and enable community scrutiny.They frame public release as important for responsible foundation-model development and improving model safety.

Contributors and Acknowledgements

Llama 3 resulted from extensive work by people at Meta. The paper distinguishes core contributors from contributors by the share of the project runtime they supported and lists them alphabetically by first name.

  • Llama 3 was the result of work by a large number of people at Meta.
  • Core contributors worked on Llama 3 for at least 2/3rd of the project’s runtime.
  • Contributors worked on Llama 3 for at least 1/5th of the project’s runtime, and all contributors are listed alphabetically by first name.
Loading 2407.21783v3…