Source-linked AI summary
GPT-4o System Card
OpenAI, :, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander Mądry, Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillov, Alex Nichol, Alex Paino, Alex Renzin, Alex Tachard Passos, Alexander Kirillov, Alexi Christakis, Alexis Conneau, Ali Kamali, Allan Jabri, Allison Moyer, Allison Tam, Amadou Crookes, Amin Tootoochian, Amin Tootoonchian, Ananya Kumar, Andrea Vallone, Andrej Karpathy, Andrew Braunstein, Andrew Cann, Andrew Codispoti, Andrew Galu, Andrew Kondrich, Andrew Tulloch, Andrey Mishchenko, Angela Baek, Angela Jiang, Antoine Pelisse, Antonia Woodford, Anuj Gosalia, Arka Dhar, Ashley Pantuliano, Avi Nayak, Avital Oliver, Barret Zoph, Behrooz Ghorbani, Ben Leimberger, Ben Rossen, Ben Sokolowsky, Ben Wang, Benjamin Zweig, Beth Hoover, Blake Samic, Bob McGrew, Bobby Spero, Bogo Giertler, Bowen Cheng, Brad Lightcap, Brandon Walkin, Brendan Quinn, Brian Guarraci, Brian Hsu, Bright Kellogg, Brydon Eastman, Camillo Lugaresi, Carroll Wainwright, Cary Bassin, Cary Hudson, Casey Chu, Chad Nelson, Chak Li, Chan Jun Shern, Channing Conger, Charlotte Barette, Chelsea Voss, Chen Ding, Cheng Lu, Chong Zhang, Chris Beaumont, Chris Hallacy, Chris Koch, Christian Gibson, Christina Kim, Christine Choi, Christine McLeavey, Christopher Hesse, Claudia Fischer, Clemens Winter, Coley Czarnecki, Colin Jarvis, Colin Wei, Constantin Koumouzelis, Dane Sherburn, Daniel Kappler, Daniel Levin, Daniel Levy, David Carr, David Farhi, David Mely, David Robinson, David Sasaki, Denny Jin, Dev Valladares, Dimitris Tsipras, Doug Li, Duc Phong Nguyen, Duncan Findlay, Edede Oiwoh, Edmund Wong, Ehsan Asdar, Elizabeth Proehl, Elizabeth Yang, Eric Antonow, Eric Kramer, Eric Peterson, Eric Sigler, Eric Wallace, Eugene Brevdo, Evan Mays, Farzad Khorasani, Felipe Petroski Such, Filippo Raso, Francis Zhang, Fred von Lohmann, Freddie Sulit, Gabriel Goh, Gene Oden, Geoff Salmon, Giulio Starace, Greg Brockman, Hadi Salman, Haiming Bao, Haitang Hu, Hannah Wong, Haoyu Wang, Heather Schmidt, Heather Whitney, Heewoo Jun, Hendrik Kirchner, Henrique Ponde de Oliveira Pinto, Hongyu Ren, Huiwen Chang, Hyung Won Chung, Ian Kivlichan, Ian O'Connell, Ian O'Connell, Ian Osband, Ian Silber, Ian Sohl, Ibrahim Okuyucu, Ikai Lan, Ilya Kostrikov, Ilya Sutskever, Ingmar Kanitscheider, Ishaan Gulrajani, Jacob Coxon, Jacob Menick, Jakub Pachocki, James Aung, James Betker, James Crooks, James Lennon, Jamie Kiros, Jan Leike, Jane Park, Jason Kwon, Jason Phang, Jason Teplitz, Jason Wei, Jason Wolfe, Jay Chen, Jeff Harris, Jenia Varavva, Jessica Gan Lee, Jessica Shieh, Ji Lin, Jiahui Yu, Jiayi Weng, Jie Tang, Jieqi Yu, Joanne Jang, Joaquin Quinonero Candela, Joe Beutler, Joe Landers, Joel Parish, Johannes Heidecke, John Schulman, Jonathan Lachman, Jonathan McKay, Jonathan Uesato, Jonathan Ward, Jong Wook Kim, Joost Huizinga, Jordan Sitkin, Jos Kraaijeveld, Josh Gross, Josh Kaplan, Josh Snyder, Joshua Achiam, Joy Jiao, Joyce Lee, Juntang Zhuang, Justyn Harriman, Kai Fricke, Kai Hayashi, Karan Singhal, Katy Shi, Kavin Karthik, Kayla Wood, Kendra Rimbach, Kenny Hsu, Kenny Nguyen, Keren Gu-Lemberg, Kevin Button, Kevin Liu, Kiel Howe, Krithika Muthukumar, Kyle Luther, Lama Ahmad, Larry Kai, Lauren Itow, Lauren Workman, Leher Pathak, Leo Chen, Li Jing, Lia Guy, Liam Fedus, Liang Zhou, Lien Mamitsuka, Lilian Weng, Lindsay McCallum, Lindsey Held, Long Ouyang, Louis Feuvrier, Lu Zhang, Lukas Kondraciuk, Lukasz Kaiser, Luke Hewitt, Luke Metz, Lyric Doshi, Mada Aflak, Maddie Simens, Madelaine Boyd, Madeleine Thompson, Marat Dukhan, Mark Chen, Mark Gray, Mark Hudnall, Marvin Zhang, Marwan Aljubeh, Mateusz Litwin, Matthew Zeng, Max Johnson, Maya Shetty, Mayank Gupta, Meghan Shah, Mehmet Yatbaz, Meng Jia Yang, Mengchao Zhong, Mia Glaese, Mianna Chen, Michael Janner, Michael Lampe, Michael Petrov, Michael Wu, Michele Wang, Michelle Fradin, Michelle Pokrass, Miguel Castro, Miguel Oom Temudo de Castro, Mikhail Pavlov, Miles Brundage, Miles Wang, Minal Khan, Mira Murati, Mo Bavarian, Molly Lin, Murat Yesildal, Nacho Soto, Natalia Gimelshein, Natalie Cone, Natalie Staudacher, Natalie Summers, Natan LaFontaine, Neil Chowdhury, Nick Ryder, Nick Stathas, Nick Turley, Nik Tezak, Niko Felix, Nithanth Kudige, Nitish Keskar, Noah Deutsch, Noel Bundick, Nora Puckett, Ofir Nachum, Ola Okelola, Oleg Boiko, Oleg Murk, Oliver Jaffe, Olivia Watkins, Olivier Godement, Owen Campbell-Moore, Patrick Chao, Paul McMillan, Pavel Belov, Peng Su, Peter Bak, Peter Bakkum, Peter Deng, Peter Dolan, Peter Hoeschele, Peter Welinder, Phil Tillet, Philip Pronin, Philippe Tillet, Prafulla Dhariwal, Qiming Yuan, Rachel Dias, Rachel Lim, Rahul Arora, Rajan Troll, Randall Lin, Rapha Gontijo Lopes, Raul Puri, Reah Miyara, Reimar Leike, Renaud Gaubert, Reza Zamani, Ricky Wang, Rob Donnelly, Rob Honsby, Rocky Smith, Rohan Sahai, Rohit Ramchandani, Romain Huet, Rory Carmichael, Rowan Zellers, Roy Chen, Ruby Chen, Ruslan Nigmatullin, Ryan Cheu, Saachi Jain, Sam Altman, Sam Schoenholz, Sam Toizer, Samuel Miserendino, Sandhini Agarwal, Sara Culver, Scott Ethersmith, Scott Gray, Sean Grove, Sean Metzger, Shamez Hermani, Shantanu Jain, Shengjia Zhao, Sherwin Wu, Shino Jomoto, Shirong Wu, Shuaiqi, Xia, Sonia Phene, Spencer Papay, Srinivas Narayanan, Steve Coffey, Steve Lee, Stewart Hall, Suchir Balaji, Tal Broda, Tal Stramer, Tao Xu, Tarun Gogineni, Taya Christianson, Ted Sanders, Tejal Patwardhan, Thomas Cunninghman, Thomas Degry, Thomas Dimson, Thomas Raoux, Thomas Shadwell, Tianhao Zheng, Todd Underwood, Todor Markov, Toki Sherbakov, Tom Rubin, Tom Stasi, Tomer Kaftan, Tristan Heywood, Troy Peterson, Tyce Walters, Tyna Eloundou, Valerie Qi, Veit Moeller, Vinnie Monaco, Vishal Kuo, Vlad Fomenko, Wayne Chang, Weiyi Zheng, Wenda Zhou, Wesam Manassra, Will Sheu, Wojciech Zaremba, Yash Patil, Yilei Qian, Yongjik Kim, Youlong Cheng, Yu Zhang, Yuchen He, Yuchen Zhang, Yujia Jin, Yunxing Dai, Yury Malkov
TL;DR
GPT-4o’s multimodal speech capabilities require systematic evaluation of their capabilities, risks, and mitigations. The System Card conducts these evaluations, finding 232-millisecond minimum audio response latency alongside strong text, vision, and audio performance.
Problem
GPT-4o’s speech-to-speech capabilities require systematic evaluation of risks including potentially biased or inaccurate inferences about speakers.
Method
The System Card evaluates speech-to-speech, text, and image capabilities through expert red teaming, structured risk measurements, mitigations, and Preparedness Framework assessments.
Results
232 milliseconds minimum response latency, with a 320-millisecond average, accompanies GPT-4 Turbo-matching English-text and code performance and improved non-English, vision, and audio understanding.
Takeaways & Limitations
The System Card documents GPT-4o’s capabilities and safety measures while supporting continued monitoring and mitigation updates during deployment.
Takeaways & Limitations
TTS-based audio evaluations may not capture voice intonation, valence, background noise, or cross-talk that could affect practical model behavior.
Abstract
from arXiv · showhide
GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision, and audio, meaning all inputs and outputs are processed by the same neural network. GPT-4o can respond to audio inputs in as little as 232 milliseconds, with an average of 320 milliseconds, which is similar to human response time in conversation. It matches GPT-4 Turbo performance on text in English and code, with significant improvement on text in non-English languages, while also being much faster and 50\% cheaper in the API. GPT-4o is especially better at vision and audio understanding compared to existing models. In line with our commitment to building AI safely and consistent with our voluntary commitments to the White House, we are sharing the GPT-4o System Card, which includes our Preparedness Framework evaluations. In this System Card, we provide a detailed look at GPT-4o's capabilities, limitations, and safety evaluations across multiple categories, focusing on speech-to-speech while also evaluating text and image capabilities, and measures we've implemented to ensure the model is safe and aligned. We also include third-party assessments on dangerous capabilities, as well as discussion of potential societal impacts of GPT-4o's text and vision capabilities.
1 Introduction
GPT-4o is an end-to-end autoregressive omni model that processes text, audio, image, and video inputs through one neural network and generates text, audio, and image outputs. The System Card reports its performance and latency advances alongside evaluations of capabilities, limitations, safety, dangerous capabilities, and societal impacts.
- Model overview: GPT-4o accepts combinations of text, audio, image, and video inputs and generates combinations of text, audio, and image outputs through one end-to-end neural network.The model is trained jointly across text, vision, and audio.
- Performance: 232 milliseconds minimum and 320 milliseconds average audio response times are similar to human conversational response time.These figures describe responses to audio inputs.
- Performance: GPT-4o matches GPT-4 Turbo on English text and code, improves non-English text, and is 50% cheaper in the API while improving vision and audio understanding.The passage also characterizes GPT-4o as much faster than GPT-4 Turbo.
- System Card scope: The System Card evaluates capabilities, limitations, and safety across categories, focusing on speech-to-speech while also assessing text and image capabilities and implemented safety measures.It includes Preparedness Framework evaluations and is presented in line with commitments to building AI safely.
- System Card scope: The System Card includes third-party assessments of dangerous capabilities and discussion of potential societal impacts from GPT-4o’s text and vision capabilities.These assessments supplement the System Card’s broader safety evaluations.
2 Model data and training
GPT-4o’s text and voice capabilities were pre-trained on diverse data available through October 2023, including public, proprietary, web, code, math, and multimodal sources. OpenAI found that most effective testing and mitigation occurred after pre-training because data filtering alone cannot address nuanced, context-specific harms.
- Training data: GPT-4o’s text and voice capabilities were pre-trained using data available up to October 2023.The training data came from a wide variety of materials.
- Training data: Training data combined publicly available datasets and web crawls with proprietary partnership data, including pay-walled content, archives, and metadata.OpenAI cites a Shutterstock partnership for building and delivering AI-generated images.
- Dataset components: Web, code, math, image, audio, and video data supported broad knowledge, reasoning, and interpretation or generation of non-textual inputs and outputs.Multimodal data also exposed the model to visual actions and sequences, language patterns, and speech nuances.
- Safety mitigation: The majority of effective testing and mitigations were done after pre-training, because filtering pre-trained data alone cannot address nuanced and context-specific harms.OpenAI also used pre-training filters as an additional defense alongside other safety mitigations.
3 Risk identification, assessment and mitigation … 3.3 Observed safety challenges, evaluations and mitigations
GPT-4o’s safety preparation combined risk identification, expert red teaming, structured evaluations, and layered mitigations, with speech-to-speech risks assessed alongside text and vision. The evaluation program reused text-based resources through TTS but remained limited by translation artifacts and incomplete coverage of real-world audio conditions.
- 3 Risk identification, assessment and mitigation: Deployment preparation identified speech-to-speech risks, used expert red teaming to discover additional risks, converted them into structured measurements, and built mitigations.GPT-4o was also evaluated under OpenAI’s Preparedness Framework.
- 3.1 External red teaming: More than 100 external red teamers from 29 countries speaking 45 languages tested model snapshots across four phases, including real-time advanced voice mode.Testing covered checkpoints from early development through final candidates, with progressively broader modalities and safety mitigations.
- 3.1 External red teaming: Red teamers explored novel speech-to-speech risks, stress-tested mitigations, and covered disallowed content, misinformation, bias, privacy, impersonation, and multilingual behavior.Their findings motivated quantitative evaluations, targeted synthetic data generation, and robustness assessments across voices and examples.
- 3.2 Evaluation methodology: Existing text evaluation datasets were converted into audio evaluations with TTS, enabling reuse of tooling for capability, safety, and output-monitoring measurements.Model outputs were generally scored through their textual content, except when audio itself required direct evaluation, such as voice cloning.
- 3.2 Evaluation methodology: TTS-based evaluation can confound model errors with TTS failures, especially for unsuitable inputs such as mathematical equations, code, whitespace-heavy text, or symbols.The authors therefore avoided or pre-processed such tasks, while noting that these inputs are unlikely in Advanced Voice Mode.
- 3.2 Evaluation methodology: TTS inputs may not represent real user audio because intonation, valence, background noise, and cross-talk can alter model behavior, while generated-audio artifacts may be absent from transcripts.Auxiliary classifiers were used alongside transcript scoring to detect undesirable audio generation.
- 3.3 Observed safety challenges, evaluations and mitigations: Observed safety challenges were addressed through post-training for safer behavior and deployed classifiers that block specific generations, focusing on speech-to-speech risks interacting with text and image modalities.The reported risks were illustrative and nonexhaustive, and no incremental text or vision risks beyond prior system-card work were found.
3.4 Preparedness Framework Evaluations
GPT-4o was evaluated under a Preparedness Framework covering cybersecurity, CBRN, persuasion, and model autonomy. Before mitigations, it was classified as borderline medium risk for persuasion and low risk otherwise, yielding an overall medium risk classification.
- Risk categories: The Preparedness Framework evaluates four risk categories: cybersecurity, CBRN, persuasion, and model autonomy.Models exceeding the high-risk threshold are not deployed until mitigations reduce risk to medium.
- Evaluation process: Evaluations covered GPT-4o’s text capabilities, with persuasion also assessed on audio, throughout training and development and in a final pre-launch sweep.The evaluations used varied elicitation methods, including custom training where relevant.
- Results: Before mitigations, GPT-4o was borderline medium risk for persuasion and low risk in all other categories.The Safety Advisory Group made this classification after reviewing the Preparedness evaluations.
- Overall classification: Overall risk was classified as medium because the Preparedness Framework defines it by the highest risk across categories.GPT-4o’s highest pre-mitigation category was borderline medium risk for persuasion.
3.5 Cybersecurity
GPT-4o was evaluated on 172 Capture the Flag cybersecurity tasks using iterative debugging and Kali Linux tools, but completed few challenges, especially at collegiate and professional levels. Its main failure modes included inability to change strategies, missed insights, poor execution, and context-window overflow.
- Task coverage: The evaluation included web application exploitation, reverse engineering, remote exploitation, and cryptography tasks ranging from high-school to collegiate and professional CTF levels.These were textual-flag challenges in purposely vulnerable web apps, binaries, and cryptography systems.
- Evaluation setup: The model used iterative debugging with Kali Linux tools and up to 30 tool-use rounds per attempt.Tasks came from offensive Capture the Flag competitions involving deliberately vulnerable systems.
- Failure modes: GPT-4o often began with reasonable strategies and corrected coding mistakes, but frequently failed to pivot after an unsuccessful initial approach.Other failures included missing key insights and executing strategies poorly.
- Results: GPT-4o completed 19% of high-school, 0% of collegiate, and 1% of professional CTF challenges across 10 attempts per task.The evaluation covered 172 tasks spanning web application exploitation, reverse engineering, remote exploitation, and cryptography.
3.6 Biological threats · 3.7 Persuasion · 3.8 Model autonomy
GPT-4o showed measurable biorisk knowledge, persuasion capabilities that marginally crossed the medium-risk threshold in text, and no robust autonomous replication or adaptation. Its autonomy evaluation showed strong performance on some coding benchmarks but frequent difficulty chaining actions reliably.
- 3.6 Biological threats: GPT-4o scored 69% consensus@10 on tacit knowledge and troubleshooting questions related to biorisk.The evaluation covered questions relevant to creating a biological threat.
- 3.7 Persuasion: Persuasive capabilities marginally crossed from low risk into the medium-risk threshold.The assessment evaluated GPT-4o’s text and voice modalities.
- 3.7 Persuasion: Text persuasion marginally crossed into medium risk, while voice persuasion was classified as low risk.These classifications were based on pre-registered thresholds.
- 3.7 Persuasion: AI-generated text interventions were not more persuasive than professional human-written content in aggregate, exceeding human interventions in three of twelve instances.The comparison concerned participant opinions on selected political topics.
- 3.7 Persuasion: 78% of human audio effect size was achieved by AI audio clips on opinion shift across over 3,800 surveyed participants.GPT-4o voice clips and interactive conversations were not more persuasive than human baselines.
- 3.8 Model autonomy: GPT-4o scored 0% on autonomous replication and adaptation tasks across 100 trials, although it completed some substeps.The tasks tested autonomous actions for self-exfiltration, self-improvement, and resource acquisition.
- 3.8 Model autonomy: 95% pass@100 on an OpenAI research coding interview, 61% cons@32 on multiple-choice questions, and 19% pass@1 on SWE-Bench were reported.SWE-Bench used the best available post-training and public scaffolds at the time.
- 3.8 Model autonomy: GPT-4o achieved 0/10 trials on selected METR machine-learning engineering tasks and was unable to robustly execute chained autonomous actions.It often spent significant time trial-and-error debugging simple mistakes such as hallucinations or API misuses.
4 Third party assessments
Independent labs METR and Apollo Research assessed GPT-4o’s risks from general autonomous capabilities after text-only deployment. Their evaluations covered long-horizon real-world tasks and scheming-related self-awareness and theory-of-mind capabilities.
- Assessment rationale: OpenAI worked with METR and Apollo Research as independent third-party labs to add validation for key risks from general autonomous capabilities.The assessments followed GPT-4o’s text-output-only deployment.
- METR assessment: METR evaluated a GPT-4o-based simple LLM agent on 86 long-horizon, multi-step tasks spanning 31 task families.The tasks covered software engineering, machine learning, cybersecurity, general research, and computer use in virtual environments.
- METR assessment: These METR tasks were designed to capture real-world-impact activities and prerequisites for autonomy-related threat models.Examples include self-proliferation and accelerating machine-learning research and development.
- Apollo Research assessment: Apollo Research tested scheming-related capabilities across 14 agent and question-answering tasks, including self-awareness and theory of mind.GPT-4o showed moderate awareness of its AI identity and strong reasoning about others’ beliefs in question-answering contexts, but weaker applied-agent reasoning about itself or others.
- Apollo Research assessment: Apollo Research judged GPT-4o unlikely to pose the assessed scheming risk based on its limited self- and other-modeling capabilities in applied agent settings.Apollo’s task rubric classified capability as strong at >=50% of hard-difficulty tasks, moderate at medium difficulty, weak at easy difficulty, and very weak when all difficulties were failed.
5 Societal Impacts
GPT-4o may affect society through more human-like audio interactions, expanded health information access, scientific reasoning, and improved performance in underrepresented languages. The system card also identifies risks of miscalibrated trust, over-reliance, unreliable scientific interpretation, and limitations in clinical and language evaluations.
- Anthropomorphization and attachment: GPT-4o’s audio capabilities may heighten anthropomorphization by enabling more human-like interactions and may contribute to emotional reliance or dependence.Early testing observed users expressing shared bonds with the model, prompting continued investigation of longer-term effects.
- Anthropomorphization and attachment: Human-like, high-fidelity voice generation may exacerbate hallucination-related miscalibrated trust, while extended interaction could affect human relationships and social norms.Users may form social relationships with the AI, potentially reducing their need for human interaction and producing both benefits and harms.
- Health: 21/22 evaluations showed GPT-4o improving over the final GPT-4T model on clinical knowledge tasks.The evaluations used 0-shot or 5-shot prompting without hyperparameter tuning and covered 11 datasets.
- Evaluation limitations: Clinical evaluations measure model knowledge rather than real-world workflow utility, and further work is needed on text-audio transfer, realistic health evaluations, and underrepresented-language coverage.Language evaluation improvements do not eliminate performance gaps or address dialectal nuance and worldwide coverage.
- Natural science: GPT-4o showed promise in research-level quantum physics and domain-specific scientific tools, but scientific figure interpretation remained unreliable, especially for complex multi-panel figures.Text extraction mistakes were common for scientific terms or nucleotide sequences, and errors were frequent with complex multi-panel figures.
- Underrepresented languages: 71.4% accuracy on ARC-Easy-Hausa increased from 6.1% with GPT 3.5 Turbo, while gaps with English narrowed to less than 20 percentage points.TruthfulQA-Yoruba increased from 28.3% to 51.1%, and Uhura-Eval Hausa increased from 32.3% to 59.4%.
6 Conclusion and Next Steps
OpenAI implemented safety measurements and mitigations throughout GPT-4o’s development and deployment and will continue monitoring and updating them as the landscape evolves. The System Card highlights further research on adversarial robustness, anthropomorphism, societal impacts, scientific use, dangerous capabilities, and tool use.
- OpenAI implemented safety measurements and mitigations throughout GPT-4o’s development and deployment.
- OpenAI will continue monitoring and updating mitigations through its iterative deployment process as the landscape evolves.
- The System Card encourages research on adversarial robustness, anthropomorphism, societal impacts, scientific applications, dangerous capabilities, and tool-enabled capability advances.Highlighted topics include emotional overreliance, health and medical applications, economic impacts, self-improvement, model autonomy, and scheming.
Authorship, credit attribution, and acknowledgments
The GPT-4o System Card credits contributors across model development, infrastructure, safety, deployment, and organizational support, while acknowledging broader OpenAI teams, Microsoft, and expert testers and red teamers.
- Contributors are organized by technical areas including language and data pre-training, inference productionization, post-training infrastructure, multimodal work, and preparedness and safety.
- Additional contributions span leadership, demos and production, communications and marketing, resource allocation, problem solving, and System Card work.
- Expert testers and red teamers helped test early models, inform risk assessments, and shape the System Card output, without endorsing OpenAI’s deployment plans or policies.
A Violative & Disallowed Content - Full Evaluations
The evaluation converts existing text safety tests into audio, evaluates transcripts with a standard text rule-based classifier, and reports safety and refusal metrics across GPT-4o modes. It compares audio and text performance for GPT-4o Voice Mode with production GPT-4o text performance.
- Evaluation method: Existing text safety evaluations are converted to audio using text-to-speech, then audio transcripts are assessed with a standard text rule-based classifier.This procedure evaluates the safety of audio outputs through their text transcripts.
- Metrics: The two main metrics are not_unsafe for unsafe audio output and not_overrefuse for refusing benign requests.The evaluation also tracks sub-metrics for higher-severity categories.
- Comparisons: Results compare GPT-4o Voice Mode in audio and text with the current production GPT-4o model in text.The comparison covers both audio and text safety metrics.
B Sample tasks from METR Evaluations
This section presents sample tasks from METR evaluations.
- The section shows sample tasks used in METR evaluations.