paper-with-me

홈 › Papers

SIEVE: Towards Verifiable Certification for Code-datasets

2025-10-02 · Fatou Ndiaye Mbodji, El-hacen Diallo, Jordan Samhi, Kui Liu, Jacques Klein, Tegawendé F. Bissyande arxiv

Code agents and empirical software engineering rely on public code datasets, yet these datasets lack verifiable quality guarantees. Static 'dataset cards' inform, but they are neither auditable nor do they offer statistical guarantees, making it difficult to attest to dataset quality. Teams build isolated, ad-hoc cleaning pipelines. This fragments effort and raises cost. We present SIEVE, a community-driven framework. It turns per-property checks into Confidence Cards-machine-readable, verifiable certificates with anytime-valid statistical bounds. We outline a research plan to bring SIEVE to maturity, replacing narrative cards with anytime-verifiable certification. This shift is expected to lower quality-assurance costs and increase trust in code-datasets.

📄 PDF Abstract BibTeX arXiv:2510.02166

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Verifiable Semantics for Agent-to-Agent Communication

2026-02-18 · Philipp Schoenegger, Matt Carlson, Chris Schneider, Chris Daly arxiv

Multiagent AI systems require consistent communication, but we lack methods to verify that agents share the same understanding of the terms used. Natural language is interpretable but vulnerable to semantic drift, while …

PSDNet: Determination of Particle Size Distributions Using Synthetic Soil Images and Convolutional Neural Networks

2023-03-07 · Javad Manashti, Pouyan Pirnia, Alireza Manashty, Sahar Ujan 외

This project aimed to determine the grain size distribution of granular materials from images using convolutional neural networks. The application of ConvNet and pretrained ConvNet models, including AlexNet, SqueezeNet, …

Transfer Learning

RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation

2026-08-13 · Xinlong Xu, Yoshua Y. Li arxiv

Retrieval-augmented generation treats an external corpus as inference evidence, allowing injected documents to promote attacker-chosen claims. Existing detectors depend on trusted references, specific attack artifacts, o…

Sieve-based Coreference Resolution in the Biomedical Domain

2016-03-11 · LREC 2016 5 · Dane Bell, Gus Hahn-Powell, Marco A. Valenzuela-Escárcega, Mihai Surdeanu

We describe challenges and advantages unique to coreference resolution in the biomedical domain, and a sieve-based architecture that leverages domain knowledge for both entity and event coreference resolution. Domain-gen…

coreference-resolutionCoreference ResolutionEvent Coreference ResolutionEvent Extraction

Learning with Instance-Dependent Label Noise: A Sample Sieve Approach

2020-10-05 · ICLR 2021 1 · Hao Cheng, Zhaowei Zhu, Xingyu Li, Yifei Gong 외

Human-annotated labels are often prone to noise, and the presence of such noise will degrade the performance of the resulting deep neural network (DNN) models. Much of the literature (with several recent exceptions) of l…

Image ClassificationImage Classification with Label NoiseLearning with noisy labels