paper-with-me

홈 › Papers

LARP: Learner-Agnostic Robust Data Prefiltering

2025-06-25 · Kristian Minchev, Dimitar Iliev Dimitrov, Nikola Konstantinov

The widespread availability of large public datasets is a key factor behind the recent successes of statistical inference and machine learning methods. However, these datasets often contain some low-quality or contaminated data, to which many learning procedures are sensitive. Therefore, the question of whether and how public datasets should be prefiltered to facilitate accurate downstream learning arises. On a technical level this requires the construction of principled data prefiltering methods which are learner-agnostic robust, in the sense of provably protecting a set of pre-specified downstream learners from corrupted data. In this work, we formalize the problem of Learner-Agnostic Robust data Prefiltering (LARP), which aims at finding prefiltering procedures that minimize a worst-case loss over a pre-specified set of learners. We first instantiate our framework in the context of scalar mean estimation with Huber estimators under the Huber data contamination model. We provide a hardness result on a specific problem instance and analyze several natural prefiltering procedures. Our theoretical results indicate that performing LARP on a heterogeneous set of learners leads to some loss in model performance compared to the alternative of prefiltering data for each learner/use-case individually. We explore the resulting utility loss and its dependence on the problem parameters via extensive experiments on real-world image and tabular data, observing statistically significant reduction in utility. Finally, we model the trade-off between the utility drop and the cost of repeated (learner-specific) prefiltering within a game-theoretic framework and showcase benefits of LARP for large datasets.

📄 PDF Abstract BibTeX arXiv:2506.20573

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Applications of Artificial Intelligence in Live Action Role-Playing Games (LARP)

2020-08-25 · Christoph Salge, Emily Short, Mike Preuss, Spyridion Samothrakis 외

Live Action Role-Playing (LARP) games and similar experiences are becoming a popular game genre. Here, we discuss how artificial intelligence techniques, particularly those commonly used in AI for Games, could be applied…

LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

2024-10-28 · Hanyu Wang, Saksham Suri, Yixuan Ren, Hao Chen 외

We present LARP, a novel video tokenizer designed to overcome limitations in current video tokenization methods for autoregressive (AR) generative models. Unlike traditional patchwise tokenizers that directly encode loca…

Video GenerationVideo Reconstruction

ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review

2026-01-30 · Palash Goyal, Mihir Parmar, Yiwen Song, Hamid Palangi 외 arxiv

The exponential growth of machine learning submissions has strained the traditional peer review process, resulting in slow feedback loops for authors and an immense burden on reviewers to rigorously audit technical sound…

Large Language Models as Generalizable Policies for Embodied Tasks

2023-10-26 · Andrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure 외

We show that large language models (LLMs) can be adapted to be generalizable policies for embodied visual tasks. Our approach, called Large LAnguage model Reinforcement Learning Policy (LLaRP), adapts a pre-trained froze…

Language ModelingLanguage ModellingLarge Language Modelreinforcement-learning+1

LARP: Language-Agent Role Play for Open-World Games

2023-12-24 · Ming Yan, Ruihao Li, Hao Zhang, Hao Wang 외

Language agents have shown impressive problem-solving skills within defined settings and brief timelines. Yet, with the ever-evolving complexities of open-world simulations, there's a pressing need for agents that can fl…

Decision Making