paper-with-me

홈 › Papers

SCARV: Structure-Constrained Aggregation for Stable Sample Ranking in Redundant NLP Datasets

2026-05-01 · Xu Zheng, Feiyu Wu, Linhong Wu, Zhuocheng Wang, Hui Li arxiv

Sample-level rankings are increasingly used in data-centric NLP for analysis, filtering, debugging, and curation, yet existing pipelines typically score training examples pointwise and rank them as if they were independent. This assumption is fragile in the presence of exact duplicates, near-duplicates, paraphrases, and other redundant structure common in NLP corpora, where stochastic training can make highly similar examples receive unstable relative orderings across random seeds. We study stable sample-level ranking under redundancy and propose \textsc{SCARV}, a modular aggregation framework that operates on top of an existing scoring proxy. \textsc{SCARV} combines robust multi-seed aggregation with a structure-aware aggregation/allocation step over redundancy clusters. Across synthetic redundancy, naturally mined QQP redundancy, multiple proxy families, several NLP tasks, and end-to-end DistilBERT fine-tuning, \textsc{SCARV} substantially improves over bare proxy rankings in global and local stability and yields more reproducible ranking-based decisions such as subset selection and suspicious-example retrieval. Our decomposition and compute-aware frontier sharpen the mechanism: robust multi-seed aggregation is the dominant generic stabilizer, while the structure-aware component adds value mainly under low aggregation budgets or when redundancy clusters are informative, naturally occurring, or sufficiently covered. These results position \textsc{SCARV} not as a universal data selector or a universally dominant replacement for seed-only aggregation, but as a stability-oriented aggregation layer for proxy-induced rankings in redundant NLP datasets.

📄 PDF Abstract BibTeX arXiv:2605.00944

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stable Causal Discovery via Directed Acyclic Graph Aggregation

2026-05-18 · Yunan Wu, Yue Wang, Chunlin Li, Chenglong Ye arxiv

Directed Acyclic Graphs (DAGs) are central to uncovering causal structure in complex systems, yet learning a single DAG from data is often challenging: model uncertainty, finite samples, and a combinatorially large searc…

Large-scale Datasets: Faces with Partial Occlusions and Pose Variations in the Wild

2017-06-27 · Tarik Alafif, Zeyad Hailat, Melih Aslan, Xue-wen Chen

Face detection methods have relied on face datasets for training. However, existing face datasets tend to be in small scales for face learning in both constrained and unconstrained environments. In this paper, we first i…

Face DetectionFace Recognition

Data Aggregation for Hierarchical Clustering

2023-09-05 · Erich Schubert, Andreas Lang

Hierarchical Agglomerative Clustering (HAC) is likely the earliest and most flexible clustering method, because it can be used with many distances, similarities, and various linkage strategies. It is often used when the …

ClusteringVector Quantization (k-means problem)

Snowveil: A Framework for Decentralised Preference Discovery

2025-12-20 · Grammateia Kotsialou arxiv

Aggregating subjective preferences in social choice traditionally assumes a trusted central authority. In contrast, this paper formalises Decentralised Preference Discovery (DPD): the reliable identification of a social …

DINOv3 Visual Representations for Blueberry Perception Toward Robotic Harvesting

2026-03-02 · Rui-Feng Wang, Daniel Petti, Yue Chen, Changying Li arxiv

Vision Foundation Models trained via large-scale self-supervised learning have demonstrated strong generalization in visual perception; however, their practical role and performance limits in agricultural settings remain…

Self-Supervised Learning