paper-with-me

홈 › Papers

SEED: Targeted Data Selection by Weighted Independent Set

2026-05-15 · Yuan Zhang, Lifeng Guo, Junwen Pan, Wenzhao Zheng, Wen Zhou, Kuan Cheng, Kurt Keutzer, Shanghang Zhang arxiv

Data selection seeks to identify a compact yet informative subset from large-scale training corpora, balancing sample quality against collection diversity. We formulate this problem as a Weighted Independent Set (WIS) on a similarity graph, where nodes represent data samples weighted by influence, and edges connect semantically redundant pairs. This formulation naturally yields subsets that are simultaneously high-quality and diverse. However, two challenges arise in practice: naive node weights fail to distinguish informative signals from gradient noise, and edge construction under heterogeneous domain distributions produces structurally imbalanced graphs that bias selection toward sparse regions. To address these issues, we introduce two principled refinements from a unified graph perspective: (1) \textit{node value calibration} that restricts influence estimation to the bilateral salient subspace to ground node importance in task-relevant signals rather than surface-level statistics; (2) \textit{local scale normalization} that adapts edge thresholds to local neighborhood density, mitigating graph imbalance induced by cross-domain distribution shifts. Together, these components yield a robust and scalable data selection pipeline dubbed SEED. We further construct \texttt{Honeybee-Remake-SEED-200K}, a compact multimodal dataset curated by SEED. Extensive experiments show that SEED consistently outperforms state-of-the-art methods on instruction tuning, visual instruction tuning, and semantic segmentation across diverse model families.

📄 PDF Abstract BibTeX arXiv:2605.15691

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Segmentation

Similar Papers 제목 키워드 기반

The effects of ecological selection on species diversity and trait distribution: predictions and an empirical test

2019-08-21 · DeMalach Niv, Po-Ju Ke, Tadashi Fukami

Ecological selection is a major driver of community assembly. Selection is classified as stabilizing when species with intermediate trait values gain the highest reproductive success, whereas selection is considered dire…

AttributeDiversity

Pack and Measure: An Effective Approach for Influence Propagation in Social Networks

2023-12-31 · Faisal N. Abu-Khzam, Ghinwa Bou Matar, Sergio Thoumi

The Influence Maximization problem under the Independent Cascade model (IC) is considered. The problem asks for a minimal set of vertices to serve as "seed set" from which a maximum influence propagation is expected. New…

Community Quality and Influence Maximization: An Empirical Study

2025-12-01 · Motaz Ben Hassine arxiv

Influence maximization in social networks plays a vital role in applications such as viral marketing, epidemiology, product recommendation, opinion mining, and counter-terrorism. A common approach identifies seed nodes b…

Product RecommendationCommunity DetectionOpinion Mining

Effects of population- and seed bank noise on neutral evolution and efficacy of natural selection

2017-12-11

Population genetics models typically consider a fixed population size and a unique selection coefficient. However, population dynamics inherently generate noise in numbers of individuals and selection acts on various com…

Dimensionality Reduction

Fisher-Wright model with deterministic seed bank and selection

2016-07-28

Seed banks are a common characteristics to many plant species, which allow storage of genetic diversity in the soil as dormant seeds for various periods of time. We investigate an above-ground population following a Fish…

Diversity