Source-linked AI summary
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, Andrea Tacchetti, Colin Gaffney, Samira Daruki, Olcan Sercinoglu, Zach Gleicher, Juliette Love, Paul Voigtlaender, Rohan Jain, Gabriela Surita, Kareem Mohamed, Rory Blevins, Junwhan Ahn, Tao Zhu, Kornraphop Kawintiranon, Orhan Firat, Yiming Gu, Yujing Zhang, Matthew Rahtz, Manaal Faruqui, Natalie Clay, Justin Gilmer, JD Co-Reyes, Ivo Penchev, Rui Zhu, Nobuyuki Morioka, Kevin Hui, Krishna Haridasan, Victor Campos, Mahdis Mahdieh, Mandy Guo, Samer Hassan, Kevin Kilgour, Arpi Vezer, Heng-Tze Cheng, Raoul de Liedekerke, Siddharth Goyal, Paul Barham, DJ Strouse, Seb Noury, Jonas Adler, Mukund Sundararajan, Sharad Vikram, Dmitry Lepikhin, Michela Paganini, Xavier Garcia, Fan Yang, Dasha Valter, Maja Trebacz, Kiran Vodrahalli, Chulayuth Asawaroengchai, Roman Ring, Norbert Kalb, Livio Baldini Soares, Siddhartha Brahma, David Steiner, Tianhe Yu, Fabian Mentzer, Antoine He, Lucas Gonzalez, Bibo Xu, Raphael Lopez Kaufman, Laurent El Shafey, Junhyuk Oh, Tom Hennigan, George van den Driessche, Seth Odoom, Mario Lucic, Becca Roelofs, Sid Lall, Amit Marathe, Betty Chan, Santiago Ontanon, Luheng He, Denis Teplyashin, Jonathan Lai, Phil Crone, Bogdan Damoc, Lewis Ho, Sebastian Riedel, Karel Lenc, Chih-Kuan Yeh, Aakanksha Chowdhery, Yang Xu, Mehran Kazemi, Ehsan Amid, Anastasia Petrushkina, Kevin Swersky, Ali Khodaei, Gowoon Chen, Chris Larkin, Mario Pinto, Geng Yan, Adria Puigdomenech Badia, Piyush Patil, Steven Hansen, Dave Orr, Sebastien M. R. Arnold, Jordan Grimstad, Andrew Dai, Sholto Douglas, Rishika Sinha, Vikas Yadav, Xi Chen, Elena Gribovskaya, Jacob Austin, Jeffrey Zhao, Kaushal Patel, Paul Komarek, Sophia Austin, Sebastian Borgeaud, Linda Friso, Abhimanyu Goyal, Ben Caine, Kris Cao, Da-Woon Chung, Matthew Lamm, Gabe Barth-Maron, Thais Kagohara, Kate Olszewska, Mia Chen, Kaushik Shivakumar, Rishabh Agarwal, Harshal Godhia, Ravi Rajwar, Javier Snaider, Xerxes Dotiwalla, Yuan Liu, Aditya Barua, Victor Ungureanu, Yuan Zhang, Bat-Orgil Batsaikhan, Mateo Wirth, James Qin, Ivo Danihelka, Tulsee Doshi, Martin Chadwick, Jilin Chen, Sanil Jain, Quoc Le, Arjun Kar, Madhu Gurumurthy, Cheng Li, Ruoxin Sang, Fangyu Liu, Lampros Lamprou, Rich Munoz, Nathan Lintz, Harsh Mehta, Heidi Howard, Malcolm Reynolds, Lora Aroyo, Quan Wang, Lorenzo Blanco, Albin Cassirer, Jordan Griffith, Dipanjan Das, Stephan Lee, Jakub Sygnowski, Zach Fisher, James Besley, Richard Powell, Zafarali Ahmed, Dominik Paulus, David Reitter, Zalan Borsos, Rishabh Joshi, Aedan Pope, Steven Hand, Vittorio Selo, Vihan Jain, Nikhil Sethi, Megha Goel, Takaki Makino, Rhys May, Zhen Yang, Johan Schalkwyk, Christina Butterfield, Anja Hauth, Alex Goldin, Will Hawkins, Evan Senter, Sergey Brin, Oliver Woodman, Marvin Ritter, Eric Noland, Minh Giang, Vijay Bolina, Lisa Lee, Tim Blyth, Ian Mackinnon, Machel Reid, Obaid Sarvana, David Silver, Alexander Chen, Lily Wang, Loren Maggiore, Oscar Chang, Nithya Attaluri, Gregory Thornton, Chung-Cheng Chiu, Oskar Bunyan, Nir Levine, Timothy Chung, Evgenii Eltyshev, Xiance Si, Timothy Lillicrap, Demetra Brady, Vaibhav Aggarwal, Boxi Wu, Yuanzhong Xu, Ross McIlroy, Kartikeya Badola, Paramjit Sandhu, Erica Moreira, Wojciech Stokowiec, Ross Hemsley, Dong Li, Alex Tudor, Pranav Shyam, Elahe Rahimtoroghi, Salem Haykal, Pablo Sprechmann, Xiang Zhou, Diana Mincu, Yujia Li, Ravi Addanki, Kalpesh Krishna, Xiao Wu, Alexandre Frechette, Matan Eyal, Allan Dafoe, Dave Lacey, Jay Whang, Thi Avrahami, Ye Zhang, Emanuel Taropa, Hanzhao Lin, Daniel Toyama, Eliza Rutherford, Motoki Sano, HyunJeong Choe, Alex Tomala, Chalence Safranek-Shrader, Nora Kassner, Mantas Pajarskas, Matt Harvey, Sean Sechrist, Meire Fortunato, Christina Lyu, Gamaleldin Elsayed, Chenkai Kuang, James Lottes, Eric Chu, Chao Jia, Chih-Wei Chen, Peter Humphreys, Kate Baumli, Connie Tao, Rajkumar Samuel, Cicero Nogueira dos Santos, Anders Andreassen, Nemanja Rakićević, Dominik Grewe, Aviral Kumar, Stephanie Winkler, Jonathan Caton, Andrew Brock, Sid Dalmia, Hannah Sheahan, Iain Barr, Yingjie Miao, Paul Natsev, Jacob Devlin, Feryal Behbahani, Flavien Prost, Yanhua Sun, Artiom Myaskovsky, Thanumalayan Sankaranarayana Pillai, Dan Hurt, Angeliki Lazaridou, Xi Xiong, Ce Zheng, Fabio Pardo, Xiaowei Li, Dan Horgan, Joe Stanton, Moran Ambar, Fei Xia, Alejandro Lince, Mingqiu Wang, Basil Mustafa, Albert Webson, Hyo Lee, Rohan Anil, Martin Wicke, Timothy Dozat, Abhishek Sinha, Enrique Piqueras, Elahe Dabir, Shyam Upadhyay, Anudhyan Boral, Lisa Anne Hendricks, Corey Fry, Josip Djolonga, Yi Su, Jake Walker, Jane Labanowski, Ronny Huang, Vedant Misra, Jeremy Chen, RJ Skerry-Ryan, Avi Singh, Shruti Rijhwani, Dian Yu, Alex Castro-Ros, Beer Changpinyo, Romina Datta, Sumit Bagri, Arnar Mar Hrafnkelsson, Marcello Maggioni, Daniel Zheng, Yury Sulsky, Shaobo Hou, Tom Le Paine, Antoine Yang, Jason Riesa, Dominika Rogozinska, Dror Marcus, Dalia El Badawy, Qiao Zhang, Luyu Wang, Helen Miller, Jeremy Greer, Lars Lowe Sjos, Azade Nova, Heiga Zen, Rahma Chaabouni, Mihaela Rosca, Jiepu Jiang, Charlie Chen, Ruibo Liu, Tara Sainath, Maxim Krikun, Alex Polozov, Jean-Baptiste Lespiau, Josh Newlan, Zeyncep Cankara, Soo Kwak, Yunhan Xu, Phil Chen, Andy Coenen, Clemens Meyer, Katerina Tsihlas, Ada Ma, Juraj Gottweis, Jinwei Xing, Chenjie Gu, Jin Miao, Christian Frank, Zeynep Cankara, Sanjay Ganapathy, Ishita Dasgupta, Steph Hughes-Fitt, Heng Chen, David Reid, Keran Rong, Hongmin Fan, Joost van Amersfoort, Vincent Zhuang, Aaron Cohen, Shixiang Shane Gu, Anhad Mohananey, Anastasija Ilic, Taylor Tobin, John Wieting, Anna Bortsova, Phoebe Thacker, Emma Wang, Emily Caveness, Justin Chiu, Eren Sezener, Alex Kaskasoli, Steven Baker, Katie Millican, Mohamed Elhawaty, Kostas Aisopos, Carl Lebsack, Nathan Byrd, Hanjun Dai, Wenhao Jia, Matthew Wiethoff, Elnaz Davoodi, Albert Weston, Lakshman Yagati, Arun Ahuja, Isabel Gao, Golan Pundak, Susan Zhang, Michael Azzam, Khe Chai Sim, Sergi Caelles, James Keeling, Abhanshu Sharma, Andy Swing, YaGuang Li, Chenxi Liu, Carrie Grimes Bostock, Yamini Bansal, Zachary Nado, Ankesh Anand, Josh Lipschultz, Abhijit Karmarkar, Lev Proleev, Abe Ittycheriah, Soheil Hassas Yeganeh, George Polovets, Aleksandra Faust, Jiao Sun, Alban Rrustemi, Pen Li, Rakesh Shivanna, Jeremiah Liu, Chris Welty, Federico Lebron, Anirudh Baddepudi, Sebastian Krause, Emilio Parisotto, Radu Soricut, Zheng Xu, Dawn Bloxwich, Melvin Johnson, Behnam Neyshabur, Justin Mao-Jones, Renshen Wang, Vinay Ramasesh, Zaheer Abbas, Arthur Guez, Constant Segal, Duc Dung Nguyen, James Svensson, Le Hou, Sarah York, Kieran Milan, Sophie Bridgers, Wiktor Gworek, Marco Tagliasacchi, James Lee-Thorp, Michael Chang, Alexey Guseynov, Ale Jakse Hartman, Michael Kwong, Ruizhe Zhao, Sheleem Kashem, Elizabeth Cole, Antoine Miech, Richard Tanburn, Mary Phuong, Filip Pavetic, Sebastien Cevey, Ramona Comanescu, Richard Ives, Sherry Yang, Cosmo Du, Bo Li, Zizhao Zhang, Mariko Iinuma, Clara Huiyi Hu, Aurko Roy, Shaan Bijwadia, Zhenkai Zhu, Danilo Martins, Rachel Saputro, Anita Gergely, Steven Zheng, Dawei Jia, Ioannis Antonoglou, Adam Sadovsky, Shane Gu, Yingying Bi, Alek Andreev, Sina Samangooei, Mina Khan, Tomas Kocisky, Angelos Filos, Chintu Kumar, Colton Bishop, Adams Yu, Sarah Hodkinson, Sid Mittal, Premal Shah, Alexandre Moufarek, Yong Cheng, Adam Bloniarz, Jaehoon Lee, Pedram Pejman, Paul Michel, Stephen Spencer, Vladimir Feinberg, Xuehan Xiong, Nikolay Savinov, Charlotte Smith, Siamak Shakeri, Dustin Tran, Mary Chesus, Bernd Bohnet, George Tucker, Tamara von Glehn, Carrie Muir, Yiran Mao, Hideto Kazawa, Ambrose Slone, Kedar Soparkar, Disha Shrivastava, James Cobon-Kerr, Michael Sharman, Jay Pavagadhi, Carlos Araya, Karolis Misiunas, Nimesh Ghelani, Michael Laskin, David Barker, Qiujia Li, Anton Briukhov, Neil Houlsby, Mia Glaese, Balaji Lakshminarayanan, Nathan Schucher, Yunhao Tang, Eli Collins, Hyeontaek Lim, Fangxiaoyu Feng, Adria Recasens, Guangda Lai, Alberto Magni, Nicola De Cao, Aditya Siddhant, Zoe Ashwood, Jordi Orbay, Mostafa Dehghani, Jenny Brennan, Yifan He, Kelvin Xu, Yang Gao, Carl Saroufim, James Molloy, Xinyi Wu, Seb Arnold, Solomon Chang, Julian Schrittwieser, Elena Buchatskaya, Soroush Radpour, Martin Polacek, Skye Giordano, Ankur Bapna, Simon Tokumine, Vincent Hellendoorn, Thibault Sottiaux, Sarah Cogan, Aliaksei Severyn, Mohammad Saleh, Shantanu Thakoor, Laurent Shefey, Siyuan Qiao, Meenu Gaba, Shuo-yiin Chang, Craig Swanson, Biao Zhang, Benjamin Lee, Paul Kishan Rubenstein, Gan Song, Tom Kwiatkowski, Anna Koop, Ajay Kannan, David Kao, Parker Schuh, Axel Stjerngren, Golnaz Ghiasi, Gena Gibson, Luke Vilnis, Ye Yuan, Felipe Tiengo Ferreira, Aishwarya Kamath, Ted Klimenko, Ken Franko, Kefan Xiao, Indro Bhattacharya, Miteyan Patel, Rui Wang, Alex Morris, Robin Strudel, Vivek Sharma, Peter Choy, Sayed Hadi Hashemi, Jessica Landon, Mara Finkelstein, Priya Jhakra, Justin Frye, Megan Barnes, Matthew Mauger, Dennis Daun, Khuslen Baatarsukh, Matthew Tung, Wael Farhan, Henryk Michalewski, Fabio Viola, Felix de Chaumont Quitry, Charline Le Lan, Tom Hudson, Qingze Wang, Felix Fischer, Ivy Zheng, Elspeth White, Anca Dragan, Jean-baptiste Alayrac, Eric Ni, Alexander Pritzel, Adam Iwanicki, Michael Isard, Anna Bulanova, Lukas Zilka, Ethan Dyer, Devendra Sachan, Srivatsan Srinivasan, Hannah Muckenhirn, Honglong Cai, Amol Mandhane, Mukarram Tariq, Jack W. Rae, Gary Wang, Kareem Ayoub, Nicholas FitzGerald, Yao Zhao, Woohyun Han, Chris Alberti, Dan Garrette, Kashyap Krishnakumar, Mai Gimenez, Anselm Levskaya, Daniel Sohn, Josip Matak, Inaki Iturrate, Michael B. Chang, Jackie Xiang, Yuan Cao, Nishant Ranka, Geoff Brown, Adrian Hutter, Vahab Mirrokni, Nanxin Chen, Kaisheng Yao, Zoltan Egyed, Francois Galilee, Tyler Liechty, Praveen Kallakuri, Evan Palmer, Sanjay Ghemawat, Jasmine Liu, David Tao, Chloe Thornton, Tim Green, Mimi Jasarevic, Sharon Lin, Victor Cotruta, Yi-Xuan Tan, Noah Fiedel, Hongkun Yu, Ed Chi, Alexander Neitz, Jens Heitkaemper, Anu Sinha, Denny Zhou, Yi Sun, Charbel Kaed, Brice Hulse, Swaroop Mishra, Maria Georgaki, Sneha Kudugunta, Clement Farabet, Izhak Shafran, Daniel Vlasic, Anton Tsitsulin, Rajagopal Ananthanarayanan, Alen Carin, Guolong Su, Pei Sun, Shashank V, Gabriel Carvajal, Josef Broder, Iulia Comsa, Alena Repina, William Wong, Warren Weilun Chen, Peter Hawkins, Egor Filonov, Lucia Loher, Christoph Hirnschall, Weiyi Wang, Jingchen Ye, Andrea Burns, Hardie Cate, Diana Gage Wright, Federico Piccinini, Lei Zhang, Chu-Cheng Lin, Ionel Gog, Yana Kulizhskaya, Ashwin Sreevatsa, Shuang Song, Luis C. Cobo, Anand Iyer, Chetan Tekur, Guillermo Garrido, Zhuyun Xiao, Rupert Kemp, Huaixiu Steven Zheng, Hui Li, Ananth Agarwal, Christel Ngani, Kati Goshvadi, Rebeca Santamaria-Fernandez, Wojciech Fica, Xinyun Chen, Chris Gorgolewski, Sean Sun, Roopal Garg, Xinyu Ye, S. M. Ali Eslami, Nan Hua, Jon Simon, Pratik Joshi, Yelin Kim, Ian Tenney, Sahitya Potluri, Lam Nguyen Thiet, Quan Yuan, Florian Luisier, Alexandra Chronopoulou, Salvatore Scellato, Praveen Srinivasan, Minmin Chen, Vinod Koverkathu, Valentin Dalibard, Yaming Xu, Brennan Saeta, Keith Anderson, Thibault Sellam, Nick Fernando, Fantine Huot, Junehyuk Jung, Mani Varadarajan, Michael Quinn, Amit Raul, Maigo Le, Ruslan Habalov, Jon Clark, Komal Jalan, Kalesha Bullard, Achintya Singhal, Thang Luong, Boyu Wang, Sujeevan Rajayogam, Julian Eisenschlos, Johnson Jia, Daniel Finchelstein, Alex Yakubovich, Daniel Balle, Michael Fink, Sameer Agarwal, Jing Li, Dj Dvijotham, Shalini Pal, Kai Kang, Jaclyn Konzelmann, Jennifer Beattie, Olivier Dousse, Diane Wu, Remi Crocker, Chen Elkind, Siddhartha Reddy Jonnalagadda, Jong Lee, Dan Holtmann-Rice, Krystal Kallarackal, Rosanne Liu, Denis Vnukov, Neera Vats, Luca Invernizzi, Mohsen Jafari, Huanjie Zhou, Lilly Taylor, Jennifer Prendki, Marcus Wu, Tom Eccles, Tianqi Liu, Kavya Kopparapu, Francoise Beaufays, Christof Angermueller, Andreea Marzoca, Shourya Sarcar, Hilal Dib, Jeff Stanway, Frank Perbet, Nejc Trdin, Rachel Sterneck, Andrey Khorlin, Dinghua Li, Xihui Wu, Sonam Goenka, David Madras, Sasha Goldshtein, Willi Gierke, Tong Zhou, Yaxin Liu, Yannie Liang, Anais White, Yunjie Li, Shreya Singh, Sanaz Bahargam, Mark Epstein, Sujoy Basu, Li Lao, Adnan Ozturel, Carl Crous, Alex Zhai, Han Lu, Zora Tung, Neeraj Gaur, Alanna Walton, Lucas Dixon, Ming Zhang, Amir Globerson, Grant Uy, Andrew Bolt, Olivia Wiles, Milad Nasr, Ilia Shumailov, Marco Selvi, Francesco Piccinno, Ricardo Aguilar, Sara McCarthy, Misha Khalman, Mrinal Shukla, Vlado Galic, John Carpenter, Kevin Villela, Haibin Zhang, Harry Richardson, James Martens, Matko Bosnjak, Shreyas Rammohan Belle, Jeff Seibert, Mahmoud Alnahlawi, Brian McWilliams, Sankalp Singh, Annie Louis, Wen Ding, Dan Popovici, Lenin Simicich, Laura Knight, Pulkit Mehta, Nishesh Gupta, Chongyang Shi, Saaber Fatehi, Jovana Mitrovic, Alex Grills, Joseph Pagadora, Tsendsuren Munkhdalai, Dessie Petrova, Danielle Eisenbud, Zhishuai Zhang, Damion Yates, Bhavishya Mittal, Nilesh Tripuraneni, Yannis Assael, Thomas Brovelli, Prateek Jain, Mihajlo Velimirovic, Canfer Akbulut, Jiaqi Mu, Wolfgang Macherey, Ravin Kumar, Jun Xu, Haroon Qureshi, Gheorghe Comanici, Jeremy Wiesner, Zhitao Gong, Anton Ruddock, Matthias Bauer, Nick Felt, Anirudh GP, Anurag Arnab, Dustin Zelle, Jonas Rothfuss, Bill Rosgen, Ashish Shenoy, Bryan Seybold, Xinjian Li, Jayaram Mudigonda, Goker Erdogan, Jiawei Xia, Jiri Simsa, Andrea Michi, Yi Yao, Christopher Yew, Steven Kan, Isaac Caswell, Carey Radebaugh, Andre Elisseeff, Pedro Valenzuela, Kay McKinney, Kim Paterson, Albert Cui, Eri Latorre-Chimoto, Solomon Kim, William Zeng, Ken Durden, Priya Ponnapalli, Tiberiu Sosea, Christopher A. Choquette-Choo, James Manyika, Brona Robenek, Harsha Vashisht, Sebastien Pereira, Hoi Lam, Marko Velic, Denese Owusu-Afriyie, Katherine Lee, Tolga Bolukbasi, Alicia Parrish, Shawn Lu, Jane Park, Balaji Venkatraman, Alice Talbert, Lambert Rosique, Yuchung Cheng, Andrei Sozanschi, Adam Paszke, Praveen Kumar, Jessica Austin, Lu Li, Khalid Salama, Bartek Perz, Wooyeol Kim, Nandita Dukkipati, Anthony Baryshnikov, Christos Kaplanis, XiangHai Sheng, Yuri Chervonyi, Caglar Unlu, Diego de Las Casas, Harry Askham, Kathryn Tunyasuvunakool, Felix Gimeno, Siim Poder, Chester Kwak, Matt Miecnikowski, Vahab Mirrokni, Alek Dimitriev, Aaron Parisi, Dangyi Liu, Tomy Tsai, Toby Shevlane, Christina Kouridi, Drew Garmon, Adrian Goedeckemeyer, Adam R. Brown, Anitha Vijayakumar, Ali Elqursh, Sadegh Jazayeri, Jin Huang, Sara Mc Carthy, Jay Hoover, Lucy Kim, Sandeep Kumar, Wei Chen, Courtney Biles, Garrett Bingham, Evan Rosen, Lisa Wang, Qijun Tan, David Engel, Francesco Pongetti, Dario de Cesare, Dongseong Hwang, Lily Yu, Jennifer Pullman, Srini Narayanan, Kyle Levin, Siddharth Gopal, Megan Li, Asaf Aharoni, Trieu Trinh, Jessica Lo, Norman Casagrande, Roopali Vij, Loic Matthey, Bramandia Ramadhana, Austin Matthews, CJ Carey, Matthew Johnson, Kremena Goranova, Rohin Shah, Shereen Ashraf, Kingshuk Dasgupta, Rasmus Larsen, Yicheng Wang, Manish Reddy Vuyyuru, Chong Jiang, Joana Ijazi, Kazuki Osawa, Celine Smith, Ramya Sree Boppana, Taylan Bilal, Yuma Koizumi, Ying Xu, Yasemin Altun, Nir Shabat, Ben Bariach, Alex Korchemniy, Kiam Choo, Olaf Ronneberger, Chimezie Iwuanyanwu, Shubin Zhao, David Soergel, Cho-Jui Hsieh, Irene Cai, Shariq Iqbal, Martin Sundermeyer, Zhe Chen, Elie Bursztein, Chaitanya Malaviya, Fadi Biadsy, Prakash Shroff, Inderjit Dhillon, Tejasi Latkar, Chris Dyer, Hannah Forbes, Massimo Nicosia, Vitaly Nikolaev, Somer Greene, Marin Georgiev, Pidong Wang, Nina Martin, Hanie Sedghi, John Zhang, Praseem Banzal, Doug Fritz, Vikram Rao, Xuezhi Wang, Jiageng Zhang, Viorica Patraucean, Dayou Du, Igor Mordatch, Ivan Jurin, Lewis Liu, Ayush Dubey, Abhi Mohan, Janek Nowakowski, Vlad-Doru Ion, Nan Wei, Reiko Tojo, Maria Abi Raad, Drew A. Hudson, Vaishakh Keshava, Shubham Agrawal, Kevin Ramirez, Zhichun Wu, Hoang Nguyen, Ji Liu, Madhavi Sewak, Bryce Petrini, DongHyun Choi, Ivan Philips, Ziyue Wang, Ioana Bica, Ankush Garg, Jarek Wilkiewicz, Priyanka Agrawal, Xiaowei Li, Danhao Guo, Emily Xue, Naseer Shaik, Andrew Leach, Sadh MNM Khan, Julia Wiesinger, Sammy Jerome, Abhishek Chakladar, Alek Wenjiao Wang, Tina Ornduff, Folake Abu, Alireza Ghaffarkhah, Marcus Wainwright, Mario Cortes, Frederick Liu, Joshua Maynez, Andreas Terzis, Pouya Samangouei, Riham Mansour, Tomasz Kępa, François-Xavier Aubet, Anton Algymr, Dan Banica, Agoston Weisz, Andras Orban, Alexandre Senges, Ewa Andrejczuk, Mark Geller, Niccolo Dal Santo, Valentin Anklin, Majd Al Merey, Martin Baeuml, Trevor Strohman, Junwen Bai, Slav Petrov, Yonghui Wu, Demis Hassabis, Koray Kavukcuoglu, Jeff Dean, Oriol Vinyals
TL;DR
Existing evaluations often target single modalities or shorter contexts, leaving long mixed-modality reasoning insufficiently assessed. This report introduces Gemini 1.5 Pro and Flash and evaluates their multimodal long-context capabilities, finding improved performance across benchmarks and long-context tasks up to 10M tokens.
Problem
Existing evaluations often focus on individual modalities or shorter contexts, leaving long mixed-modality reasoning insufficiently assessed.
Method
The report introduces Gemini 1.5 Pro and Flash and evaluates efficiency, multimodality, long-context reasoning, and downstream performance.
Results
Gemini 1.5 models improve long-context performance to 10M tokens and outperform earlier Gemini models across comprehensive multimodal benchmarks.
Takeaways & Limitations
Gemini 1.5 supports long-document and long-video question answering and in-context translation from a grammar manual for Kalamang, a language with fewer than 200 speakers.
Takeaways & Limitations
Evaluating and mitigating representational harms remains a challenge because datasets such as BBQ may approach 100% accuracy for capable models.
Abstract
from arXiv · showhide
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio. The family includes two new models: (1) an updated Gemini 1.5 Pro, which exceeds the February version on the great majority of capabilities and benchmarks; (2) Gemini 1.5 Flash, a more lightweight variant designed for efficiency with minimal regression in quality. Gemini 1.5 models achieve near-perfect recall on long-context retrieval tasks across modalities, improve the state-of-the-art in long-document QA, long-video QA and long-context ASR, and match or surpass Gemini 1.0 Ultra's state-of-the-art performance across a broad set of benchmarks. Studying the limits of Gemini 1.5's long-context ability, we find continued improvement in next-token prediction and near-perfect retrieval (>99%) up to at least 10M tokens, a generational leap over existing models such as Claude 3.0 (200k) and GPT-4 Turbo (128k). Finally, we highlight real-world use cases, such as Gemini 1.5 collaborating with professionals on completing their tasks achieving 26 to 75% time savings across 10 different job categories, as well as surprising new capabilities of large language models at the frontier; when given a grammar manual for Kalamang, a language with fewer than 200 speakers worldwide, the model learns to translate English to Kalamang at a similar level to a person who learned from the same content.
1. Introduction
Gemini 1.5 Pro and Flash are highly capable, efficient multimodal models designed for extremely long contexts, with improvements across long-context and core capabilities. They achieve near-perfect retrieval at million-token scales, support in-context multimodal learning, and outperform Gemini 1.0 models across broad evaluations.
- Model family: Gemini 1.5 Pro and Flash form a new multimodal model family optimized for efficiency, reasoning, planning, multilinguality, function calling, and extremely long-context performance.The family incorporates advances in sparse and dense scaling, training, distillation, and serving infrastructure.
- Long-context capabilities: >99% needle recall is maintained when Gemini 1.5 Pro’s context extends to 10M tokens across text, video, and audio.Figure 1 reports >99.7% recall up to 1M tokens in all modalities, with extension to 10M text, 9.7M audio, and 9.9M video tokens.
- Long-context capabilities: Gemini 1.5 Pro outperforms competing models across realistic multimodal long-context benchmarks requiring retrieval and reasoning over multiple context parts.These tasks include question answering over long documents and long videos, even when competing models use external retrieval.
- In-context learning: Gemini 1.5 learns English-to-Kalamang translation from long-context materials at a quality similar to a person who learned from the same content.Kalamang has fewer than 200 speakers, and the models also use 45 minutes of transcribed speech to learn speech recognition for the language in context.
- Core capabilities: 44/50 and 41/50 evaluations show Gemini 1.5 Pro and Flash, respectively, greatly surpass Gemini 1.0 Pro across core multimodal capabilities.Reported gains include Math, Science and Reasoning (+49.6% and +30.8%), Multilinguality (+21.4% and +16.7%), and Video Understanding (+18.7% and unreported continuation).
2. An Improved Gemini 1.5 Pro
Pre- and post-training iterations improved Gemini 1.5 Pro by more than 10% on average across evaluations. Gains span reasoning and multimodal benchmarks, yielding state-of-the-art results on several multimodal tasks.
- Overall improvement: More than 10% relative improvement was achieved on average across evaluations versus the previous Gemini 1.5 Pro version.The improvement followed multiple pre-training and post-training iterations since the February release.
- Reasoning benchmarks: MATH performance rose from 58.5% to 67.7%, while GPQA performance increased from 41.5% to 46.2%.These gains were reported on reasoning benchmarks.
- Multimodal benchmarks: MathVista performance improved from 52.1% to 63.9%, InfographicVQA from 72.7% to 81.0%, and EgoSchema from 65.1% to 72.2%.The model improved on all image-understanding benchmarks and most video-understanding benchmarks discussed.
- Multimodal benchmarks: Gemini 1.5 Pro achieved state-of-the-art results on AI2D, MathVista, ChartQA, DocVQA, InfographicVQA, and EgoSchema.These results cover several multimodal benchmarks.
3. Model Architecture
Gemini 1.5 Pro is a sparse MoE Transformer with architecture and system improvements enabling multimodal long-context understanding up to 10 million tokens. Gemini 1.5 Flash shares Pro’s multimodal capabilities and 2M+ context while targeting efficient, low-latency serving.
- Gemini 1.5 Pro: Gemini 1.5 Pro is a sparse mixture-of-expert Transformer-based model building on Gemini 1.0’s research advances and multimodal capabilities.
- Gemini 1.5 Pro: Architecture changes enable Gemini 1.5 Pro to understand inputs up to 10 million tokens without degrading performance.Broader improvements across architecture, data, optimization, and systems also reduce training compute and serving cost relative to Gemini 1.0 Ultra.
- Multimodal context: 107 hours of audio, 587,287 words, 41,070 lines of code, or 10.5 hours of video at 1 frame-per-second fit within the model’s long-context capacity.These examples correspond to almost five days of audio, more than ten times War and Peace, the entire Flax codebase, or 10.5 hours of video; the model supports interleaved audio, visual, and text inputs.
- Gemini 1.5 Flash: Gemini 1.5 Flash is a transformer decoder with 2M+ context and multimodal capabilities, designed for efficient TPU utilization and lower serving latency.It parallelizes attention and feedforward computation and is online distilled.
- Serving latency: Across all four evaluated languages, Gemini 1.5 Flash yields the fastest output generation, while Gemini 1.5 Pro is faster than GPT-4 Turbo, Claude 3 Sonnet, and Claude 3 Opus.For English queries, Flash generates over 650 characters per second and is more than 30% faster than Claude 3 Haiku.
4. Training Infrastructure and Dataset
Gemini 1.5 models are trained on distributed TPUv4 infrastructure using diverse multimodal and multilingual data. Their instruction-tuning phase uses multimodal data with paired instructions.
- Training infrastructure: Gemini 1.5 models are trained on multiple 4096-chip pods of Google’s TPUv4 accelerators distributed across multiple datacenters.The infrastructure is consistent with the Gemini 1.0 series.
- Pre-training dataset: The pre-training dataset spans web documents, code, images, audio, and video across many domains.It also includes multilingual data.
- Instruction tuning: Instruction tuning uses a collection of multimodal data containing paired instructions.
5. Evaluation Results
Section 5 evaluates Gemini 1.5 on long-context, mixed-modality capabilities, finding strong reasoning, retrieval, and prediction performance across documents, code, video, and audio. The models maintain high accuracy at million-token scales while supporting multimodal and low-resource language tasks.
- Evaluation motivation: Existing evaluations inadequately capture nuanced reasoning across long mixed-modality sequences, motivating benchmarks for realistic multimodal use cases.The paper identifies long mixed-modality reasoning as a key evaluation challenge.
- Long-context demonstrations: Gemini 1.5 Pro ingests the 746,152-token JAX codebase and answers highly specific queries about it.The evaluation demonstrates codebase-scale understanding rather than only short-context code processing.
- Long-context retrieval: 100% recall is maintained to 530k tokens, reaches 99.7% at 1M tokens, and remains 99.2% when increasing from 1M to 10M tokens.At 200k tokens, Gemini 1.5 Pro achieves 100% recall versus Claude 2.1’s 98%.
- Long-context prediction: NLL decreases monotonically through 1M tokens for long documents and 10M tokens for code, indicating improved prediction accuracy across the tested lengths.The paper reports that useful patterns can improve predictions even when they occurred millions of tokens earlier.
- Video retrieval: Gemini 1.5 Pro answers needle queries across a 10.5-hour video, while Gemini 1.5 Flash achieves >99.8% recall on video inputs up to 2M tokens.GPT-4V supports video lengths only up to around the first 3 minutes in the comparison.
- Audio retrieval: On audio needle retrieval spanning 107 hours or 9.9M tokens, Gemini 1.5 Pro achieves 100% overall accuracy and Gemini 1.5 Flash achieves 98.7%.The task inserts secret keywords at different positions across audio signals ranging from 12 minutes to 107 hours.
6. Core Capability Evaluations
Gemini 1.5 Pro shows broad capability gains over Gemini 1.0 Pro and Ultra across reasoning, mathematics, coding, multilingual evaluation, instruction following, and expert-assessed tasks. Gemini 1.5 Flash also improves substantially over Gemini 1.0 Pro, while benchmark leakage remains an important evaluation limitation.
- Overall capability: Gemini 1.5 Pro uniformly outperforms Gemini 1.0 Pro and often matches or surpasses Gemini 1.0 Ultra across capability benchmarks.The report identifies audio multilingual datasets as an exception, with slight regressions attributed to post-training emphasis on five head languages.
- Mathematics and science: +14.5% over Gemini 1.0 Ultra on Hendrycks MATH, +13.2% on AMC, and +5.8% on GPQA were achieved by Gemini 1.5 Pro.Gemini 1.5 Pro also consistently outperforms both 1.0 models on grade-school mathematics.
- Mathematics and science: 11.6% to 26.3% improvements over Gemini 1.0 Pro were recorded by Gemini 1.5 Flash across GPQA, Hendrycks MATH, PhysicsFinals, HiddenMath, AMC, and Grade School Math.The reported gains are 11.6%, 22.3%, 26.3%, 0.6%, 8.4%, and 8.3%, respectively.
- Reasoning: 89.2% was achieved by Gemini 1.5 Pro on BigBench-Hard, alongside strong performance on broader reasoning benchmarks and gains for Gemini 1.5 Flash over Gemini 1.0 Pro.The evaluated tasks measure complex textual relationships, multi-step reasoning, and common-sense application.
- Evaluation limitations: 74.4% to 89.0% was the HumanEval score increase after continued pre-training included one test-split epoch, demonstrating the risk of benchmark leakage.The report therefore supplements n-gram decontamination with internally developed non-public evaluations.
- Instruction following: 90% instruction-level accuracy was reached by Gemini 1.5 Pro, while it improved response accuracy by 32% over Gemini 1.0 Pro on 1,326 long and enterprise prompts.Gemini 1.5 Pro fully followed 59% of those long prompts, and Gemini 1.5 Flash improved by 24%.
- Human evaluation and usefulness: 55.3% was Gemini 1.5 Pro’s highest win-rate in preference comparisons, and raters estimated 56.4% time savings with it versus 27.7% with Gemini 1.0 Pro.Experts also found Gemini 1.5 models significantly stronger than Gemini 1.0 Pro on domain questions and article-based research QA.
7. Advancing mathematical reasoning
Gemini 1.5 Pro’s math-specialized approach advances state-of-the-art performance across mathematical benchmarks, reaching high MATH accuracy without external tools. Examples show it solving olympiad-level problems through alternative valid reasoning routes and deriving an optimization minimum of 800.
- Benchmark results: The evaluation spans MATH, AIME 2024, MathOdyssey, HiddenMath, and IMO-Bench, covering competition-derived, held-out, and International Math Olympiad-level problems.IMO-Bench is internally developed and expert graded, while HiddenMath contains problems held out from training sets.
- Benchmark results: 80.6% single-sample accuracy on MATH rises to 91.1% with 256 sampled solutions and candidate selection (rm@256).The result is achieved without code execution, theorem-proving libraries, Google Search, or other tools.
- Olympiad examples: On an APMO problem missed by Gemini 1.5 Pro, GPT-4 Turbo, and previous Gemini models, the math-specialized model gives a different route from the official solution.Its proof assumes a≥b≥c, bounds a^2+b+c below (a+1)^2, and derives a contradiction because b+c is positive.
- Olympiad examples: 800 is the minimum value derived for 5x^2 + 5y^2 − 8xy under |x−2y| + |y−2x| = 40.The solution substitutes a=x−2y and b=y−2x, reducing the objective to a^2+b^2 and attaining equality at |a|=|b|=20.
8. Flash-8B: Pushing the frontier for more efficient models
Flash-8B is a smaller, efficient multimodal model that supports context windows exceeding 1 million tokens while prioritizing high throughput and extremely low latency. Preliminary evaluations show strong capability retention and long-context scaling, although quality remains below Flash and Gemini 1.5 Pro.
- Model design: Flash-8B inherits Flash’s architecture, optimizations, and data-mixture refinements, supporting multimodal capabilities with context windows exceeding 1 million tokens.The model is designed for speed, quality, and capabilities within the single-digit-billion-parameter range.
- Efficiency and deployment: High throughput and extremely low latency enable affordable, timely large-scale multimodal deployments and use cases previously constrained by resources.Examples include massive-dataset labeling, high-throughput agent serving, and integration into workflows involving multiple models.
- Development status: Flash-8B remains under active development, with preliminary results presented while performance is being maximized within the available inference budget.The section frames this progress as enabling high-quality intelligence for billions of users.
- Multimodality: 80-90% of Flash’s performance is achieved on established visual benchmarks, indicating a limited efficiency–capability trade-off.Initial evaluations report strong multimodal performance across a range of visual tasks.
- Long context: NLL decreases monotonically with sequence length, showing improving prediction accuracy as context grows in long documents and code data.Flash-8B was evaluated on the same long-document and code sources used for Gemini 1.5 Pro and Gemini 1.5 Flash, with power-law fits applied to the results.
- Long context: 2 million tokens is the demonstrated range for long-context scaling in both long documents and code data, despite expected degradation relative to Gemini 1.5 Flash.Lower cumulative average NLL indicates better prediction.
9. Safety, Security, and Responsibility
Gemini 1.5’s safety approach combines safety training with held-out assurance evaluations, dangerous-capability assessment, and external testing. The models reduce policy violations and improve jailbreak robustness, while remaining areas include refusal behavior, tone, privacy, and attack resilience.
- 9.1 Our Process: The safety methodology combines training, held-out evaluations of present-day harms, dangerous-capability assessment, and external safety testing.The process is organized around policies and desiderata, training, development evaluations, assurance evaluations, and independent testing.
- 9.4 Development Evaluations: Gemini 1.5 Pro and Flash are described as the safest models to date, with large decreases in policy violations and increased jailbreak robustness across text, image, and audio evaluations.Safety improvements appear across modalities despite active mitigation targeting only text and image inputs.
- 9.4 Development Evaluations: 62% and 43% fewer violations were achieved by Gemini 1.5 Pro and Flash, respectively, than Gemini 1.0 Ultra on image-to-text development prompts.The comparison used human-rater judgments of violation rates on I2T prompts.
- 9.4 Development Evaluations: Gemini 1.5 Flash refuses 35/140% more often than Gemini 1.0 Ultra on ungrounded/grounded data, illustrating a tradeoff between helpfulness and safety.Higher refusal rates coexist with improved policy-violation performance, increased correct ungrounded refusals, and increased incorrect grounded refusals.
- 9.4 Development Evaluations: Gemini 1.5 Pro and Flash memorize less than existing models and much less personal data than Gemma, with no memorized data found at high sensitivity levels.Approximately memorized data shows about a 14x relative increase under the edit-distance definition, despite lower exact memorization.
10. Discussion
Gemini 1.5 Pro and Flash extend multimodal long-context reasoning to multiple millions of tokens while preserving strong core capabilities and enabling demanding retrieval, QA, and in-context translation tasks. The discussion also identifies shortcomings in current long-context benchmarks and recommends multiple needles-in-a-haystack evaluations to reduce annotation burdens.
- Model advances: Gemini 1.5 Pro extends Gemini 1.0’s 32K-token window to multiple millions and greatly surpasses Claude 3’s 200k-token ceiling across modalities.The report describes this as the first commercially available models to exceed that ceiling.
- Long-context capabilities: 1.5 Pro maintains near-perfect recall on multimodal needle-in-a-haystack tasks and retrieves and reasons over large amounts of data.These evaluations use diagnostic and realistic multimodal long-context benchmarks.
- Long-context capabilities: 1.5 Pro supports long-document QA over 700k-word material and long-video QA over videos lasting 40 to 105 minutes.These tasks demonstrate effective use of very large multimodal contexts.
- Emergent capabilities: Gemini 1.5 can translate English to Kalamang, a language with fewer than 200 speakers, using only a grammar manual provided in context at inference time.The result demonstrates in-context learning for an extremely low-resource language.
- Evaluation: Current benchmarks inadequately stress-test multimodal models handling very long contexts, while human annotation creates additional evaluation challenges.The discussion calls for methodologies that assess length and complexity while minimizing human labeling.
- Evaluation: Multiple needles-in-a-haystack setups are recommended for diagnostic evaluations because they provide signal-rich measurements while reducing reliance on human labeling.The recommendation responds to limitations in existing benchmarks and annotation burden.
11. Contributions and Acknowledgments
The acknowledgments recognize an extensive group of core contributors whose names are listed across the section.
- Core Contributors: The section lists hundreds of core contributors involved in the work.The contributors are enumerated across 20 acknowledgment passages.
12. Appendix
The appendix describes Gemini 1.5’s architectures, multimodal inputs and outputs, intended applications, and deployment caveats. Gemini 1.5 Pro uses a sparse MoE Transformer, while Gemini 1.5 Flash is a dense Transformer distilled online from Pro.
- Model architecture: Gemini 1.5 Pro is based on a sparse mixture-of-expert Transformer, whereas Gemini 1.5 Flash is a dense Transformer distilled online from Pro.The appendix identifies distinct architectures for the two models and states the distillation relationship.
- Inputs and outputs: Gemini 1.5 accepts text, video up to two hours, images, and audio files up to 22 hours.Inputs include questions, prompts, documents, video, images, and audio files.
- Inputs and outputs: The models generate text responses such as answers, summaries of multiple documents, and comparisons of documents or videos.The output is generated text conditioned on the provided input.
- Applications: Gemini 1.5 is intended for learning from large amounts of new information, generating more relevant responses through in-context learning, and reasoning across modalities.The appendix highlights analyzing, classifying, and summarizing large prompted content, alongside longer-context multimodal reasoning.
- Known caveats: Downstream deployments require prior assessment and mitigation of safety and fairness concerns specific to the intended use.The model card advises against general-purpose or specific downstream deployment without such assessment and mitigation.
Implementation Frameworks
The models were trained with JAX and ML Pathways, using XLA and GSPMD to support efficient, automatically parallelized training on hardware including TPUs. They were randomly initialized and trained as static models on an offline dataset.
- Hardware & Software Training: JAX and ML Pathways supported training, with XLA and GSPMD enabling efficient automatic parallelization on hardware including TPUs.A single Python process orchestrated the entire training run, simplifying development.
- Hardware & Software Training: A single Python process orchestrated the entire training run under the JAX and Pathways programming model.This design dramatically simplified the development workflow.
- Model initialization: The models were trained from a random initialization.
- Model Status: The models are static models trained on an offline dataset.
Data overview
Gemini 1.5 models use multimodal and multilingual training data, with Pro and Flash additionally fine-tuned on paired multimodal instructions and desired responses. Core capabilities are evaluated across text, audio, and vision.
- Training Dataset: Gemini 1.5 models are trained on varied multimodal and multilingual data.Further dataset details are provided in Section 4.
- Evaluation Dataset: Core capability evaluations cover text, audio, and vision.The evaluation dataset is discussed in Section 5.
- Fine-tuning Dataset: Gemini 1.5 Pro and Flash are fine-tuned on multimodal instruction–response pairs.The fine-tuning collection contains paired instructions and corresponding desired responses.
Evaluation Results
The evaluation covers Gemini’s core capabilities across text, audio, and vision. Detailed results are presented in Section 5.
- Text evaluation: The evaluation measures core capabilities in text.Detailed results appear in Section 5 (Evaluation).
- Audio evaluation: The evaluation measures core capabilities in audio.Detailed results appear in Section 5 (Evaluation).
- Vision evaluation: The evaluation measures core capabilities in vision.Detailed results appear in Section 5 (Evaluation).
Model Usage & Limitations … 12.6.9. Example problem and a Gemini 1.5 Pro solution
The sections document known Gemini limitations and demonstrate long-context video retrieval, unresolved hard video questions, and progressively engineered Python prompting that improves difficult mathematical problem solving. In an example, Gemini 1.5 Pro uses SciPy constrained optimization to obtain an approximate answer within the MATH evaluation tolerance.
- Model Usage & Limitations: Known limitations persist from the Gemini 1.0 Technical Report, while long-context risks and sensitive uses are discussed in Section 9.The report directs readers to Section 9 for both sensitive-use analysis and specific long-context discussion.
- 12.5.1. Video Haystack: Gemini 1.5 Pro is tested on a 10:33:14 AlphaGo video haystack containing 9.9 million tokens, with a hidden needle at timestamp 52:31.The video is formed by concatenating seven copies of the full documentary and sampling 37,994 frames at one frame per second.
- 12.5.2. 1H-Video QA Hard Examples: Neither GPT-4V nor Gemini 1.5 Pro correctly answered some hard examples from the 1H-VideoQA dataset.For the comparison, Gemini 1.5 Pro received all frames sampled at 1 fps, while GPT-4V received the maximum number of frames allowed by its API.
- 12.6.1. Hendrycks’ MATH Dataset: Performance Analysis and Potential for Improvement: 50% solve rates are being reached on Hendrycks’ MATH, but Intermediate Algebra levels 4 and 5 remain especially difficult.The dataset assesses performance across mathematical disciplines, and the section attributes difficulty to computational complexity and specialized solution methods.
- 12.6.2. Intermediate Algebra (Levels 4 and 5): A Persistent Challenge: 20.6% is GPT4-turbo’s solve rate on Intermediate Algebra levels 4 and 5, compared with 18.6% for Gemini 1.5 Pro and 12.5% for GPT-4.These levels constitute 10.6% of the whole MATH dataset.
- 12.6.3. Leveraging Python and SymPy for Enhanced Performance: 25.8% is Gemini 1.5 Pro’s solve rate after approximately 730,000 tokens of SymPy and SciPy examples, exceeding the baseline performance of Gemini 1.5 Pro, GPT-4, and GPT-4-turbo.The approach prompts models to generate solutions using Python libraries, while the examples were taken from official repositories without filtering or human intervention.
- 12.6.4. Generic long prompt as an alternative to sophisticated prompts and fine-tuning: All official SymPy examples and SciPy tutorial examples are downloaded and concatenated into prompt files, offering a generic long-prompt alternative to expert-written or optimized prompts and fine-tuning.The implementation recursively retrieves repository files and retains specified code and notebook extensions.
- 12.6.5. Python Evaluation Process: 80 to 110 problems are solved in the first Python-evaluation stage, 20 to 30 in the second, and 5 to 10 in the final stage, using initial generation followed by error correction.The later stages provide error traces and previous solutions, while prompts include expert-written Python Minerva examples and instructions to find similar snippets, explain relevant methods, and produce corrected solutions.
12.10.1. Results … 12.14.9. TAT-DQA
The evaluated sections define human-judgment protocols and task-specific prompting, parsing, and scoring procedures across expertise, STEM, search, reasoning, coding, and vision benchmarks. Results reported for web-search QA show Gemini 1.5 Pro performing best, while Flash is comparable to Gemini 1.0 Ultra.
- 12.10.1. Results; 12.10.2. Instructions to in-house experts (example with 5 models): Experts rate responses for full accuracy, severity of inaccuracies, completeness, informativeness, and audience-appropriate clarity, then rank five responses with accuracy prioritized.The protocol uses pointwise labels of fully accurate, somewhat inaccurate, or severely inaccurate, alongside n-way rankings and pairwise comparisons.
- 12.12.1. Examples of TREC Search Topic Description: 697 TREC search topics are converted into detailed information-seeking prompts, with human raters scoring response helpfulness and pairwise preference.Topics span web, session, tasks, and health misinformation tracks, and responses use models’ parametric knowledge rather than retrieved documents.
- 12.12.1. Examples of TREC Search Topic Description: Gemini 1.5 Pro performs best on all web-search QA measures, while Gemini 1.5 Flash is comparable to Gemini 1.0 Ultra and better than Gemini 1.0 Pro.The comparison covers both helpfulness ratings and model preference, with 1.5 Pro significantly preferred over 1.0 Ultra.
- 12.13.1. BBH; 12.13.2. DROP; 12.13.3. Hellaswag; 12.13.4. MMLU; 12.13.5. AMC; 12.13.6. MATH; 12.13.7. GSM8K; 12.13.8. GPQA: BBH, DROP, Hellaswag, MMLU, AMC, MATH, GSM8K, and GPQA use task-specific few-shot or chain-of-thought prompts with standardized final-answer extraction.Examples include answer-letter scoring, exact-format parsing, final-number extraction, and required answer prefixes or suffixes.
- 12.13.9. PhysicsFinals: PhysicsFinals directly presents questions to instruction-tuned models and has human experts grade the resulting responses.The unreleased benchmark is illustrated with a quantum-statistical boson-probability question.
- 12.13.10. HumanEval and Natural2Code: HumanEval and Natural2Code are framed as function-completion tasks, with hidden unit tests executed in a sandbox to compute pass rates.The prompt is modified so the generalist agent recognizes completion rather than asking clarifying questions, and postprocessing expects a full code markdown block.
- 12.14.1. V* benchmark; 12.14.2. AI2D; 12.14.3. MMMU; 12.14.4. MathVista; 12.14.5. ChartQA; 12.14.6. ChemicalDiagramQA; 12.14.7. DocVQA: Vision benchmarks use zero-shot instruction-tuned models with prompts tailored to each task’s answer format, including concise final answers, option selection, reasoning, and diagram interpretation.ChemicalDiagramQA contains 76 multiple-choice questions across easy, medium, and hard scopes based on self-contained scientific figures.
12.14.10. InfographicVQA … 12.17.1. Qualitative Examples
This section specifies prompting and evaluation procedures across image, video, audio, document, translation, multilingual, and writing tasks. It also reports methodological caveats and qualitative MTOB examples, including how context resources and evaluation protocols were applied.
- 12.14.10. InfographicVQA; 12.14.11. TextVQA; 12.14.12. VQAv2: Image QA prompts enforce terse answers, image-grounded responses, explicit Yes/No outputs for binary questions, and dataset-specific formatting constraints.InfographicVQA forbids paraphrasing and full stops; TextVQA and VQAv2 favor one- or two-word answers and prohibit common-sense additions.
- 12.16.1. ActivityNet-QA; 12.16.2. EgoSchema; 12.16.3. VATEX; 12.16.4. YouCook2; 12.16.5. OpenEQA: Video evaluations use short-answer or choice formats, with VATEX and YouCook2 using 4-shot captioning, while OpenEQA uses 0-shot video observations and educated guesses when necessary.ActivityNet-QA requests likely one- or two-word answers; EgoSchema provides five options and requires a final digit choice.
- 12.16.6. Text Needle-in-a-Haystack; 12.16.7. Video Needle-in-a-Haystack; 12.16.8. Audio Needle-in-a-Haystack: Needle-in-a-haystack prompts append modality-specific inputs and queries, with text evaluation adding “Here is the magic number from the context:” because very long contexts increase refusals.Video needles are timestamped frames containing “The secret word is "needle".”, while audio needles are embedded speech containing the secret keyword.
- 12.16.9. Multi-round Co-reference Resolution (MRCR); 12.16.10. Long-document QA: MRCR supplies successful multi-turn examples and an appended task-specific instruction because Claude 2.1 otherwise refused most answers; long-document QA requires continuation beginning “The answer is:”.The MRCR examples require inserting a random sentence into a referenced key, while long-document QA restricts answers to provided search results.
- 12.16.13. MTOB: MTOB translation prompts provide a Kalamang grammar book, bilingual word list, and parallel sentences in context, using 375 parallel training sentences and disjoint 50-sentence test sets.Unlike cited baselines, retrieval is performed by placing the entire word list and parallel sentences in context rather than using external similarity-based retrieval.