paper-with-me

홈 › Papers

Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

2024-02-09 · Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzemiński, Hakimeh Fadaei, Irem Ergün, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Vu Minh Chien, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, Sara Hooker

Datasets are foundational to many breakthroughs in modern artificial intelligence. Many recent achievements in the space of natural language processing (NLP) can be attributed to the finetuning of pre-trained models on a diverse set of tasks that enables a large language model (LLM) to respond to instructions. Instruction fine-tuning (IFT) requires specifically constructed and annotated datasets. However, existing datasets are almost all in the English language. In this work, our primary goal is to bridge the language gap by building a human-curated instruction-following dataset spanning 65 languages. We worked with fluent speakers of languages from around the world to collect natural instances of instructions and completions. Furthermore, we create the most extensive multilingual collection to date, comprising 513 million instances through templating and translating existing datasets across 114 languages. In total, we contribute four key resources: we develop and open-source the Aya Annotation Platform, the Aya Dataset, the Aya Collection, and the Aya Evaluation Suite. The Aya initiative also serves as a valuable case study in participatory research, involving collaborators from 119 countries. We see this as a valuable framework for future research collaborations that aim to bridge gaps in resources.

📄 PDF Abstract BibTeX arXiv:2402.06619

Code (1)

for-ai/language-confusion

Tasks

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages

2024-12-12 · Chester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä 외

We present OpenNER 1.0, a standardized collection of openly available named entity recognition (NER) datasets. OpenNER contains 34 datasets spanning 51 languages, annotated in varying named entity ontologies. We correct …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

2024-02-12 · Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko 외

Recent breakthroughs in large language models (LLMs) have centered around a handful of data-rich languages. What does it take to broaden access to breakthroughs beyond first-class citizen languages? Our work introduces A…

Language ModelingLanguage Modellingmodel

Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

2024-10-21 · Xiang Yue, Yueqi Song, Akari Asai, Seungone Kim 외

Despite recent advances in multimodal large language models (MLLMs), their development has predominantly focused on English- and western-centric datasets and tasks, leaving most of the world's languages and diverse cultu…

M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

2023-06-07 · Lei LI, Yuwei Yin, Shicheng Li, Liang Chen 외

Instruction tuning has significantly advanced large language models (LLMs) such as ChatGPT, enabling them to align with human instructions across diverse tasks. However, progress in open vision-language models (VLMs) has…

World Knowledge

Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback

2023-07-29 · Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen 외

A key technology for the development of large language models (LLMs) involves instruction tuning that helps align the models' responses with human expectations to realize impressive learning abilities. Two major approach…