paper-with-me

홈 › Papers

Zyda-2: a 5 Trillion Token High-Quality Dataset

2024-11-09 · Yury Tokpanov, Paolo Glorioso, Quentin Anthony, Beren Millidge

In this technical report, we present Zyda-2: a five trillion token dataset for language model pretraining. Zyda-2 was used to train our Zamba2 series of models which are state-of-the-art for their weight class. We build Zyda-2 by collating high-quality open-source tokens such as FineWeb and DCLM, then distilling them to the highest-quality subset via cross-deduplication and model-based quality filtering. Zyda-2 is released under a permissive open license, and is available at https://huggingface.co/datasets/Zyphra/Zyda-2

📄 PDF Abstract BibTeX arXiv:2411.06068

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Zyda: A 1.3T Dataset for Open Language Modeling

2024-06-04 · Yury Tokpanov, Beren Millidge, Paolo Glorioso, Jonathan Pilault 외

The size of large language models (LLMs) has scaled dramatically in recent years and their computational and data requirements have surged correspondingly. State-of-the-art language models, even at relatively smaller siz…

Language ModelingLanguage Modelling

The Zamba2 Suite: Technical Report

2024-11-22 · Paolo Glorioso, Quentin Anthony, Yury Tokpanov, Anna Golubeva 외

In this technical report, we present the Zamba2 series -- a suite of 1.2B, 2.7B, and 7.4B parameter hybrid Mamba2-transformer models that achieve state of the art performance against the leading open-weights models of th…

LazyDAgger: Reducing Context Switching in Interactive Imitation Learning

2021-03-31 · Ryan Hoque, Ashwin Balakrishna, Carl Putterman, Michael Luo 외

Corrective interventions while a robot is learning to automate a task provide an intuitive method for a human supervisor to assist the robot and convey information about desired behavior. However, these interventions can…

continuous-controlContinuous ControlImitation Learning

GneissWeb: Preparing High Quality Data for LLMs at Scale

2025-02-19 · Hajar Emami Gohari, Swanand Ravindra Kadhe, Syed Yousaf Shah. Constantin Adam, Abdulhamid Adebayo 외

Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's ability to generalize on a wide range of dow…

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only

2023-09-26 · NeurIPS 2023 11

Large language models are commonly trained on a mixture of filtered web data and curated ``high-quality'' corpora, such as social media conversations, books, or technical papers. This curation process is believed to be n…