Source-linked AI summary

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

BigScience Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, Dragomir Radev, Eduardo González Ponferrada, Efrat Levkovizh, Ethan Kim, Eyal Bar Natan, Francesco De Toni, Gérard Dupont, Germán Kruszewski, Giada Pistilli, Hady Elsahar, Hamza Benyamina, Hieu Tran, Ian Yu, Idris Abdulmumin, Isaac Johnson, Itziar Gonzalez-Dios, Javier de la Rosa, Jenny Chim, Jesse Dodge, Jian Zhu, Jonathan Chang, Jörg Frohberg, Joseph Tobing, Joydeep Bhattacharjee, Khalid Almubarak, Kimbo Chen, Kyle Lo, Leandro Von Werra, Leon Weber, Long Phan, Loubna Ben allal, Ludovic Tanguy, Manan Dey, Manuel Romero Muñoz, Maraim Masoud, María Grandury, Mario Šaško, Max Huang, Maximin Coavoux, Mayank Singh, Mike Tian-Jian Jiang, Minh Chien Vu, Mohammad A. Jauhar, Mustafa Ghaleb, Nishant Subramani, Nora Kassner, Nurulaqilla Khamis, Olivier Nguyen, Omar Espejel, Ona de Gibert, Paulo Villegas, Peter Henderson, Pierre Colombo, Priscilla Amuok, Quentin Lhoest, Rheza Harliman, Rishi Bommasani, Roberto Luis López, Rui Ribeiro, Salomey Osei, Sampo Pyysalo, Sebastian Nagel, Shamik Bose, Shamsuddeen Hassan Muhammad, Shanya Sharma, Shayne Longpre, Somaieh Nikpoor, Stanislav Silberberg, Suhas Pai, Sydney Zink, Tiago Timponi Torrent, Timo Schick, Tristan Thrush, Valentin Danchev, Vassilina Nikoulina, Veronika Laippala, Violette Lepercq, Vrinda Prabhu, Zaid Alyafeai, Zeerak Talat, Arun Raja, Benjamin Heinzerling, Chenglei Si, Davut Emre Taşar, Elizabeth Salesky, Sabrina J. Mielke, Wilson Y. Lee, Abheesht Sharma, Andrea Santilli, Antoine Chaffin, Arnaud Stiegler, Debajyoti Datta, Eliza Szczechla, Gunjan Chhablani, Han Wang, Harshit Pandey, Hendrik Strobelt, Jason Alan Fries, Jos Rozen, Leo Gao, Lintang Sutawika, M Saiful Bari, Maged S. Al-shaibani, Matteo Manica, Nihal Nayak, Ryan Teehan, Samuel Albanie, Sheng Shen, Srulik Ben-David, Stephen H. Bach, Taewoon Kim, Tali Bers, Thibault Fevry, Trishala Neeraj, Urmish Thakker, Vikas Raunak, Xiangru Tang, Zheng-Xin Yong, Zhiqing Sun, Shaked Brody, Yallow Uri, Hadar Tojarieh, Adam Roberts, Hyung Won Chung, Jaesung Tae, Jason Phang, Ofir Press, Conglong Li, Deepak Narayanan, Hatim Bourfoune, Jared Casper, Jeff Rasley, Max Ryabinin, Mayank Mishra, Minjia Zhang, Mohammad Shoeybi, Myriam Peyrounette, Nicolas Patry, Nouamane Tazi, Omar Sanseviero, Patrick von Platen, Pierre Cornette, Pierre François Lavallée, Rémi Lacroix, Samyam Rajbhandari, Sanchit Gandhi, Shaden Smith, Stéphane Requena, Suraj Patil, Tim Dettmers, Ahmed Baruwa, Amanpreet Singh, Anastasia Cheveleva, Anne-Laure Ligozat, Arjun Subramonian, Aurélie Névéol, Charles Lovering, Dan Garrette, Deepak Tunuguntla, Ehud Reiter, Ekaterina Taktasheva, Ekaterina Voloshina, Eli Bogdanov, Genta Indra Winata, Hailey Schoelkopf, Jan-Christoph Kalo, Jekaterina Novikova, Jessica Zosa Forde, Jordan Clive, Jungo Kasai, Ken Kawamura, Liam Hazan, Marine Carpuat, Miruna Clinciu, Najoung Kim, Newton Cheng, Oleg Serikov, Omer Antverg, Oskar van der Wal, Rui Zhang, Ruochen Zhang, Sebastian Gehrmann, Shachar Mirkin, Shani Pais, Tatiana Shavrina, Thomas Scialom, Tian Yun, Tomasz Limisiewicz, Verena Rieser, Vitaly Protasov, Vladislav Mikhailov, Yada Pruksachatkun, Yonatan Belinkov, Zachary Bamberger, Zdeněk Kasner, Alice Rueda, Amanda Pestana, Amir Feizpour, Ammar Khan, Amy Faranak, Ana Santos, Anthony Hevia, Antigona Unldreaj, Arash Aghagol, Arezoo Abdollahi, Aycha Tammour, Azadeh HajiHosseini, Bahareh Behroozi, Benjamin Ajibade, Bharat Saxena, Carlos Muñoz Ferrandis, Daniel McDuff, Danish Contractor, David Lansky, Davis David, Douwe Kiela, Duong A. Nguyen, Edward Tan, Emi Baylor, Ezinwanne Ozoani, Fatima Mirza, Frankline Ononiwu, Habib Rezanejad, Hessie Jones, Indrani Bhattacharya, Irene Solaiman, Irina Sedenko, Isar Nejadgholi, Jesse Passmore, Josh Seltzer, Julio Bonis Sanz, Livia Dutra, Mairon Samagaio, Maraim Elbadri, Margot Mieskes, Marissa Gerchick, Martha Akinlolu, Michael McKenna, Mike Qiu, Muhammed Ghauri, Mykola Burynok, Nafis Abrar, Nazneen Rajani, Nour Elkott, Nour Fahmy, Olanrewaju Samuel, Ran An, Rasmus Kromann, Ryan Hao, Samira Alizadeh, Sarmad Shubber, Silas Wang, Sourav Roy, Sylvain Viguier, Thanh Le, Tobi Oyebade, Trieu Le, Yoyo Yang, Zach Nguyen, Abhinav Ramesh Kashyap, Alfredo Palasciano, Alison Callahan, Anima Shukla, Antonio Miranda-Escalada, Ayush Singh, Benjamin Beilharz, Bo Wang, Caio Brito, Chenxi Zhou, Chirag Jain, Chuxin Xu, Clémentine Fourrier, Daniel León Periñán, Daniel Molano, Dian Yu, Enrique Manjavacas, Fabio Barth, Florian Fuhrimann, Gabriel Altay, Giyaseddin Bayrak, Gully Burns, Helena U. Vrabec, Imane Bello, Ishani Dash, Jihyun Kang, John Giorgi, Jonas Golde, Jose David Posada, Karthik Rangasai Sivaraman, Lokesh Bulchandani, Lu Liu, Luisa Shinzato, Madeleine Hahn de Bykhovetz, Maiko Takeuchi, Marc Pàmies, Maria A Castillo, Marianna Nezhurina, Mario Sänger, Matthias Samwald, Michael Cullan, Michael Weinberg, Michiel De Wolf, Mina Mihaljcic, Minna Liu, Moritz Freidank, Myungsun Kang, Natasha Seelam, Nathan Dahlberg, Nicholas Michio Broad, Nikolaus Muellner, Pascale Fung, Patrick Haller, Ramya Chandrasekhar, Renata Eisenberg, Robert Martin, Rodrigo Canalli, Rosaline Su, Ruisi Su, Samuel Cahyawijaya, Samuele Garda, Shlok S Deshmukh, Shubhanshu Mishra, Sid Kiblawi, Simon Ott, Sinee Sang-aroonsiri, Srishti Kumar, Stefan Schweter, Sushil Bharati, Tanmay Laud, Théo Gigant, Tomoya Kainuma, Wojciech Kusa, Yanis Labrak, Yash Shailesh Bajaj, Yash Venkatraman, Yifan Xu, Yingxin Xu, Yu Xu, Zhe Tan, Zhongli Xie, Zifan Ye, Mathilde Bras, Younes Belkada, Thomas Wolf

arXiv:2211.05100v4cs.CL

TL;DR

Large language model development is costly and often inaccessible, limiting broad research participation. BLOOM addresses this gap with a 176B-parameter multilingual model developed collaboratively and evaluated across varied benchmarks. The paper reports competitive performance that improves after multitask finetuning, while noting limits in its probing and summarization analyses.

  • Problem

    Training costs and restricted releases have excluded much of the research community from developing large language models.

  • Method

    BLOOM is a 176B-parameter multilingual language model developed collaboratively, with its dataset, architecture, training, and capabilities documented.

  • Results

    BLOOM achieves competitive performance across evaluations, with stronger results after multitask finetuning.

  • Takeaways & Limitations

    BLOOM provides a publicly released large-scale multilingual model and documents the collaborative process behind its development.

  • Takeaways & Limitations

    The analysis averages representations across layers at the end of training, leaving layer-specific and training-dynamics analyses for future work.

Abstract

from arXiv · show

Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed by resource-rich organizations and are frequently kept from the public. As a step towards democratizing this powerful technology, we present BLOOM, a 176B-parameter open-access language model designed and built thanks to a collaboration of hundreds of researchers. BLOOM is a decoder-only Transformer language model that was trained on the ROOTS corpus, a dataset comprising hundreds of sources in 46 natural and 13 programming languages (59 in total). We find that BLOOM achieves competitive performance on a wide variety of benchmarks, with stronger results after undergoing multitask prompted finetuning. To facilitate future research and applications using LLMs, we publicly release our models and code under the Responsible AI License.

Tokenization

The supplied passages contain author lists and a keywords line, but no substantive discussion of tokenization.

  • The passages primarily list contributors to the work.
  • Additional contributor lists continue across the supplied material.
  • The supplied keywords are “Language models” and “collaborative research.”

1. Introduction

BLOOM was introduced to address the concentration and restricted release of large language model development by providing a large, multilingual, openly released model. It was developed by hundreds of researchers using publicly funded French computing resources and documented its design and capabilities.

  • Large-model training costs and restricted model releases had excluded much of the research community from development.
  • 176 billion parameters and 46 natural plus 13 programming languages define BLOOM’s scale and multilingual scope.
  • BLOOM was developed and released by a collaboration of hundreds of researchers.
  • The model’s training compute came through a French public grant using IDRIS’ Jean Zay supercomputer.
  • The paper describes the training dataset, architecture, objective, distributed-learning engineering, and capability analysis.

2. Background

The background introduces autoregressive language modeling, traces its development from n-gram models through neural and Transformer approaches, and situates BLOOM within transfer learning and collaborative open research. It also describes social, environmental, and data-curation concerns surrounding large language models.

  • 2.1 Language Modeling: Autoregressive language modeling predicts each next token from the preceding tokens in a sequence.
  • 2.1 Language Modeling: N-gram models face exponential growth with sequence length and cannot directly assign probabilities to unseen sequences.
  • 2.1 Language Modeling: Neural language models estimate the next-token probability from prior tokens, progressing from fixed-window networks to recurrent and Transformer architectures.
  • 2.1 Language Modeling: Transfer learning pretrains model parameters on a data-rich task before finetuning them for a downstream task.
  • 2.1 Language Modeling: Pretrained models can perform tasks without subsequent training, motivating few- and zero-shot learning research.
  • 2.1 Language Modeling: Large-model development raises participation, carbon-footprint, and bias concerns, including risks from source material and filtering methods.

3. BLOOM

BLOOM’s design combines multilingual data curation, a decoder-only architecture, multilingual tokenization, distributed training, and responsible release practices. The model and its supporting infrastructure were developed to improve access to large-language-model research while documenting performance, environmental costs, and scope boundaries.

  • Training Dataset: BLOOM was trained on ROOTS, a 1.61-terabyte collection of 498 datasets spanning 46 natural and 13 programming languages.The corpus was accompanied by organizational and technical tools for data curation.
  • Data Governance: ROOTS curation combined multidisciplinary data-governance goals with explicit permissions, source separation, and filtering intended to retain text written by humans for humans.The project sought to account for developers, data subjects, and rights-holders while reducing non-natural-language content and bias from filtering decisions.
  • Prompted Datasets: T0 was trained on P3 prompts covering 170+ datasets and tasks including sentiment analysis, question answering, and natural language inference.The prompts were produced through hackathons involving BigScience collaborators and excluded harmful content and programming languages.
  • Architecture: BLOOM uses a decoder-only architecture selected after systematic evaluation found causal decoder-only models performed best immediately after pretraining for zero-shot generalization.The design also adopted ALiBi positional embeddings, which the authors found improved training smoothness and downstream performance relative to learned and rotary embeddings.
  • Tokenization: A 250k-token byte-level BPE tokenizer was selected to meet multilingual fertility objectives while avoiding unknown tokens and increasing vocabulary sharing across languages.Fertility was evaluated against monolingual tokenizers using Universal Dependencies and OSCAR subsets.
  • Environmental Impact: BLOOM’s estimated training emissions were approximately 81 tons of CO2eq, while its emissions were about two-thirds lower than OPT’s despite higher energy consumption.The authors attribute the difference to the lower carbon intensity of the French electricity grid used for BLOOM training.
  • Release: The model was released with behavioral-use clauses under a Responsible AI License to balance open access with restrictions on potentially harmful applications.The licensing approach explicitly addresses harmful-use concerns associated with releasing BLOOM.

4. Evaluation

The evaluation suite measures BLOOM in zero- and few-shot settings across classification, translation, and multilingual summarization, using varied human-developed prompts and relevant baselines.

  • Evaluations focus on zero-shot and few-shot prompting across SuperGLUE, machine translation, and summarization because these settings commonly reflect large-model use.
  • Prompts were generated through PromptSource, crowdsourced across tasks, and peer-reviewed for artifacts and consistency.
  • SuperGLUE: SuperGLUE evaluation covers seven English classification tasks, with five randomly selected prompts per task and maximum candidate-label likelihood used for predictions.
  • Machine Translation: Machine-translation evaluation includes WMT14, Flores-101, and DiaBLa, using sacrebleu BLEU scores and greedy decoding.
  • Summarization: Summarization is evaluated on WikiLingua across nine languages, testing abstractive generation in each source language rather than only English.

4.2 SuperGLUE

In zero- and one-shot SuperGLUE evaluation, BLOOM performs above chance on entailment tasks, while average performance on other tasks remains near chance; one-shot context reduces variability and improves BLOOM relative to OPT.

  • On BoolQ and CB, BLOOM, T0, OPT, and GPT-J perform well above random chance in both zero- and one-shot settings.
  • On other SuperGLUE tasks, average performance across prompts stays near chance, while individual best prompts can perform better.
  • T0 is an exception with strong performance, but its multitask finetuning makes it not directly comparable with the other evaluated models.
  • One-shot prompting reduces variability across prompts and models, with slight and inconsistent average performance increases.

4.3 Machine Translation

BLOOM’s translation quality varies with prompting, context, language resources, and language relatedness: few-shot examples improve key zero-shot problems, while many low-resource results are competitive but under-represented languages remain difficult.

  • Verbose prompts, especially “version-target,” perform best on WMT14, while “gpt3-target” and “xglm-source+target” perform poorly, particularly zero-shot.
  • Over-generation and incorrect output languages are major translation problems, and both improve as the number of few-shot examples increases.
  • Automatic DiaBLa results are inconclusive: previous-context examples yield higher BLEU but lower COMET than random examples.
  • Flores-101: Many low-resource translation results are comparable to or slightly better than supervised M2M, while Swahili–Yoruba quality is very poor.
  • Flores-101: Romance-language translation performs well across the board, including Galician despite its absence from BLOOM’s training data.
  • Flores-101: Under-represented training languages with fewer than 50k tokens each remain a concern for BLOOM’s translation quality.

4.4 Summarization

BLOOM outperforms OPT-175B on multilingual one-shot summarization, with performance increasing as model size grows. ROUGE-2 is used for comparability, but it often understates summary quality.

  • BLOOM achieves higher multilingual summarization performance than OPT-175B in one-shot evaluation.The authors suspect multilingual-focused training contributes to this result.
  • Performance increases as BLOOM’s parameter count increases.
  • ROUGE-2 is reported for comparability with prior work because generation evaluation lacks alternatives.The authors qualitatively observe that ROUGE-2 often understates generated-summary quality.

4.5 Code Generation

BLOOM’s code-generation performance is comparable to similarly sized GPT models trained on the Pile, while code-finetuned Codex models are substantially stronger. Multitask finetuning does not significantly improve BLOOM on HumanEval.

  • BLOOM pretrained models perform similarly to similarly sized GPT models trained on the Pile on HumanEval.ROOTS contains around 11% code, compared with around 13% code in the Pile.
  • Codex models are significantly stronger than other models because they were finetuned solely on code.
  • Multitask-finetuned BLOOMZ models do not improve significantly over BLOOM models on HumanEval.The authors hypothesize that xP3 lacks substantial pure code-completion data.

4.8 Embeddings

SGPT-BLOOM-7.1B-msmarco36 achieves state-of-the-art performance on several multilingual MTEB classification and semantic textual similarity splits. English-only models perform poorly on non-English languages.

  • SGPT-BLOOM-7.1B-msmarco36 provides state-of-the-art performance on several MTEB classification and semantic textual similarity splits.Its 7.1 billion parameters make it an order of magnitude larger than displayed multilingual MiniLM and MPNet models.
  • SGPT-BLOOM-1.7B-nli performs significantly worse, likely because it has fewer parameters and shorter finetuning.Its finetuning uses NLI, a much smaller dataset than MS-MARCO.
  • ST5-XL performs poorly on non-English languages because it is an English-only model.
  • The reported languages are part of the BLOOM pretraining corpus.

4.9 Multilingual Probing

BLOOM’s multilingual probing analysis evaluates morphosyntactic feature representations across languages, comparing model variants with count-based baselines and examining linguistic, data, and scale correlates. The results show strong but uneven multilingual representation, with performance patterns varying by language family, feature, and resource level.

  • Method: The probing framework evaluates BLOOM representations in 104 languages and 80 morphosyntactic features using language-specific Universal Dependencies datasets.The reported experiments focus on 17 languages from 7 language families present in BLOOM’s pretraining corpus.
  • Method: Binary logistic regression classifiers predict morphosyntactic-feature presence from layer representations, with weighted F1 used to address target-class imbalance.The study compares BLOOM-1B7 and BLOOM against random guessing and TF-IDF n-gram baselines.
  • Results: BLOOM-1B7 performs on par or better than BLOOM, and both language models outperform count-based baselines, with stronger results for Arabic, Basque, and Indo-European languages.The reported averages are computed over probing tasks and experiment runs within each language.
  • Results: Romance-language performance exceeds English, Indic results approach high-resource languages, while Bengali, Wolof, and Yoruba receive the lowest scores.The authors attribute these patterns to transfer from closely related languages with substantial pretraining data.
  • Results: Mood and Person are inferred well across languages, Number, NumType, and Voice moderately, while other categories are weaker; feature-value diversity may contribute to these differences.Mood and Person have similar values across languages, whereas Case values depend strongly on the language.
  • Results: BLOOM-1B7 is significantly better than BLOOM in probing results, but BLOOM has more stable performance across languages despite differing pretraining-data amounts.For BLOOM-1B7, probing results correlate strongly with language family, probing-dataset size, and pretraining-dataset size; the authors suggest larger models may generalize better.
  • Limitations and future work: Future work should probe under-resourced Indic and Niger-Congo languages, examine languages outside the pretraining corpus, and analyze representations across layers and training dynamics.The current analysis averages representations across all layers and evaluates only at the end of training.

4.10 Bias

BLOOM’s multilingual CrowS-Pairs evaluation adapts stereotype-pair testing for autoregressive models in English and French. Overall prompt accuracy is near chance, but several aggregate and category-level results differ significantly from 50%, and the evaluation has important scope limitations.

  • Method: CrowS-Pairs compares stereotyped and non-stereotyped minimal pairs to assess whether models systematically prefer stereotyped statements.The evaluation adapts a dataset originally designed for masked language models to autoregressive BLOOM using prompts.
  • Results: BLOOM’s overall prompt accuracy was close to .50 in both English and French, suggesting an overall absence of bias in this evaluation.The English and French scores were very close, indicating similar overall behavior across the two languages.
  • Results: CrowS-Pairs accuracy is homogeneous across bias categories, contrasting with earlier masked-language-model studies that found category-specific bias patterns.The table reports results averaged over eight runs for English and French.
  • Limitations: CrowS-Pairs results should be compared with other bias measures and assessed across all model languages because the evaluation covers limited situations, languages, and language variants.The authors also note that multilingual bias-assessment resources remain scarce.

5. Conclusion

The paper presents BLOOM as a 176B-parameter open-access multilingual model created through a large collaborative effort and documents its dataset, architecture, tokenizer, and evaluations. The authors report competitive performance that improves after multitask finetuning and hope the release supports further research and applications.

  • Conclusion: BLOOM is a 176B-parameter open-access multilingual language model created by BigScience through a collaboration of hundreds of researchers.It was trained for 3.5 months on the Jean Zay supercomputer.
  • Conclusion: The paper chronicles BLOOM’s development from the ROOTS training dataset through the model architecture and tokenizer, alongside evaluations of BLOOM and other large language models.The conclusion frames the work as documentation of the model’s development and evaluation.
  • Conclusion: BLOOM achieves competitive performance across evaluations, with stronger results after multitask finetuning.The conclusion reports this as the paper’s overall evaluation finding.
  • Implications: The authors hope releasing a powerful multilingual model will unlock new applications and research directions and support future large-scale collaborative projects.They also argue that collaboration enables results beyond the capacity of an individual research group and brings together researchers with different backgrounds.

6. Contributions

BLOOM’s contributions span model development, data, tokenization, prompting, evaluation, engineering, broader impacts, applications, and organization. The paper attributes these activities to distinct collaborative authorship categories.

  • Authorship categories: Major Contributors lists individuals considered essential to BLOOM or who spent more than 20% of their time on the BigScience effort overall.Authors may appear in multiple categories because they contributed in more than one way.
  • Technical contributions: The contributions are organized into categories covering model architecture and objective, dataset work, tokenization, prompt engineering, and evaluation and interpretability.The categories identify contributors to both technical components and model analysis.
  • Engineering: Engineering contributors supported the code and infrastructure required to train BLOOM on the Jean Zay supercomputer.The engineering category explicitly includes work on training infrastructure.
  • Broader impacts: Broader Impacts contributors worked on the ethical charter, license, model card, privacy, social impacts, and BLOOM’s carbon footprint.The category combines governance, documentation, and impact-related research.
  • Applications and organization: Applications contributors belonged to working groups focused on applications of BLOOM, while Organization contributors coordinated the BigScience effort.The applications category also includes authors of related application studies.

Appendix A. Prompts

Appendix A presents evaluation prompts, task-specific templates, bilingual translation prompts, and sample question-answer pairs used across the paper’s benchmarks.

  • Appendix A. Prompts: The appendix documents evaluation prompts available through PromptSource, showing both raw templates and examples with filled placeholders.Double curly brackets are replaced with sample content when prompts are used.
  • A.1.1 Data example: Coreference and word-sense tasks ask whether a pronoun refers to a specified noun or whether a word has the same meaning across two sentences.Examples include “they” referring to lemon trees and “check” used in two contexts.
  • A.3.1 Data example: Knowledge and reasoning prompts use short passages or premises followed by questions requiring answers such as true/false, entailment, or causal explanation.Examples cover phantom pain, tax-filing roles, and whether an object was fragile or small.
  • A.8.1 Data example: The appendix notes that XStoryCloze and Story Cloze are not publicly available and directs readers to their respective authors for samples.This availability constraint applies to the data examples in the appendix.
  • A.9 WMT: Translation prompts explicitly encode source and target languages, with examples for English–French, English–Hindi, and French–Catalan directions.Prompt names and content vary by language direction, and language codes or names are substituted into templates.
  • A.10 DiaBLa: DiaBLa prompts translate bilingual English–French dialogue and support few-shot examples in the same, opposite, or default language direction.The dialogue set contains native English and native French speakers, while additional tasks control the direction of the demonstration.
  • A.11 Flores-101 (MT): Flores-101 prompts are language-pair specific, illustrated by a French-to-Catalan template that places the source sentence before the target field.The template uses sentence_fra and sentence_cat placeholders for the two languages.
  • A.6.1 Data example: Several tasks present two answer choices and request the most plausible interpretation, including stereotype recognition, pronoun resolution, and cause-effect selection.The demonstrated answers identify a stereotype, select the city councilmen as the referent, and choose fragility as the cause of bubble-wrap packaging.
Loading 2211.05100v4…