paper-with-me

Papers

Spokes: Optimizing for Diverse Pretraining Data Selection

2026-06-13 · Clarence Lee, Yejin Choi, Luke Zettlemoyer, Pang Wei Koh, Hai Leong Chieu arxiv

Diversity plays a critical role in data selection, improving performance under fixed data budgets by reducing redundancy and repetition. However, optimizing for diversity is inherently challenging, as it is a set-level property that depends on interactions between data points rather than individual examples. As a result, existing approaches typically rely on proxies or approximations, which often fail to ensure sufficiently diverse subsets. In this work, we directly optimize diversity by introducing a probabilistic diversification framework based on the G-Vendi score, optimized via exponentiated gradient descent. Our method produces subsets that are substantially more diverse than those obtained via random sampling, achieving a +489 increase in G-Vendi score on a 500k-sample subset. We evaluate our approach on FineWeb and DCLM, where it consistently outperforms existing methods. Notably, SPOKES (diversity-only) improves average downstream performance by +0.4 and +0.5 points over random sampling on DCLM and FineWeb, respectively. More importantly, jointly optimizing for both quality and diversity yields the strongest results: SPOKES achieves gains of +1.5 and +1.4 points on DCLM and FineWeb, outperforming all baselines, including semantic deduplication and quality filtering.

📄 PDF Abstract BibTeX arXiv:2606.15216

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SpokesBiz -- an Open Corpus of Conversational Polish

2023-12-19 · Piotr Pęzik, Sylwia Karasińska, Anna Cichosz, Łukasz Jałowiecki 외

This paper announces the early release of SpokesBiz, a freely available corpus of conversational Polish developed within the CLARIN-BIZ project and comprising over 650 hours of recordings. The transcribed recordings have…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Increasing the Accessibility of Time-Aligned Speech Corpora with Spokes Mix

2018-05-01 · LREC 2018 5 · Piotr P{\k{e}}zik
Speech Recognition

Hubs and Spokes Learning: Efficient and Scalable Collaborative Machine Learning

2025-04-29 · Atul Sharma, Kavindu Herath, Saurabh Bagchi, Chaoyue Liu 외

We introduce the Hubs and Spokes Learning (HSL) framework, a novel paradigm for collaborative machine learning that combines the strengths of Federated Learning (FL) and Decentralized Learning (P2PL). HSL employs a two-t…

Federated Learning

DAVE: A Policy-Enforcing LLM Spokesperson for Secure Multi-Document Data Sharing

2026-02-19 · René Brinkhege, Prahlad Menon arxiv

In current inter-organizational data spaces, usage policies are enforced mainly at the asset level: a whole document or dataset is either shared or withheld. When only parts of a document are sensitive, providers who wan…

The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities

2024-11-07 · Zhaofeng Wu, Xinyan Velocity Yu, Dani Yogatama, Jiasen Lu 외

Modern language models can process inputs across diverse languages and modalities. We hypothesize that models acquire this capability through learning a shared representation space across heterogeneous data types (e.g., …