paper-with-me

홈 › Papers

Is Memorization Helpful or Harmful? Prior Information Sets the Threshold

2026-02-10 · Chen Cheng, Rina Foygel Barber arxiv

We examine the connection between training error and generalization error for arbitrary estimating procedures, working in an overparameterized linear model under general priors in a Bayesian setup. We find determining factors inherent to the prior distribution $π$, giving explicit conditions under which optimal generalization necessitates that the training error be (i) near interpolating relative to the noise size (i.e., memorization is necessary), or (ii) close to the noise level (i.e., overfitting is harmful). Remarkably, these phenomena occur when the noise reaches thresholds determined by the Fisher information and the variance parameters of the prior $π$.

📄 PDF Abstract BibTeX arXiv:2602.09405

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?

2026-06-18 · Sihui Dai, Mann Patel arxiv

Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models interpret different types of compliance demonstrations. We study this by mixing benign compliance demon…

MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learning

2026-06-20 · Wenhao Wang, Franziska Boenisch, Michael Backes, Adam Dziedzic arxiv

Memorization in machine learning models enables high performance on rare in-distribution samples by capturing their atypical patterns. However, it also causes harmful retention of noise and outliers, degrading generaliza…

Self-Supervised LearningContrastive Learning

How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation

2025-02-20 · Zhuohang Long, Siyuan Wang, Shujun Liu, Yuhang Lai 외

Jailbreak attacks, where harmful prompts bypass generative models' built-in safety, raise serious concerns about model vulnerability. While many defense methods have been proposed, the trade-offs between safety and helpf…

Binary Classification

An Investigation of Memorization Risk in Healthcare Foundation Models

2025-10-14 · Sana Tonekaboni, Lena Stempfle, Adibvafa Fallahpour, Walter Gerych 외 arxiv

Foundation models trained on large-scale de-identified electronic health records (EHRs) hold promise for clinical applications. However, their capacity to memorize patient information raises important privacy concerns. I…

Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs

2025-09-22 · Alexander Panfilov, Evgenii Kortukov, Kristina Nikolić, Matthias Bethge 외 arxiv

Large language model (LLM) developers aim for their models to be honest, helpful, and harmless. However, when faced with malicious requests, models are trained to refuse, sacrificing helpfulness. We show that frontier LL…