paper-with-me

홈 › Papers

Slower Generalization, Faster Memorization: A Sweet Spot in Algorithmic Learning

2026-05-14 · Shin So, Kyelim Lee, Albert No arxiv

Critical-data-size accounts of grokking suggest a natural post-threshold intuition: once training data is sufficient to identify the underlying rule, additional data should accelerate validation convergence. We show that this intuition can fail in a controlled structured-output task. In Needleman--Wunsch (NW) matrix generation, small Transformers reach high validation exact-match accuracy fastest at an intermediate dataset size, not at the largest one. Past this dataset-size sweet spot, generalization remains achievable but requires more gradient updates. Conversely, in the regime where partial validation competence first appears, larger datasets can require fewer updates to reach high training accuracy, suggesting that emerging rule structure can accelerate fitting beyond example-wise memorization. A multiplication baseline does not show the same post-threshold slowdown. These results separate the critical data size for the onset of generalization from the dataset size that optimizes update-based convergence, and identify structured-output tasks where learning the rule and completing exact-fitting can diverge.

📄 PDF Abstract BibTeX arXiv:2605.14659

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Maximizing a Perceptual Sweet Spot

2022-01-05 · Pedro Izquierdo Lehmann, Rodrigo F. Cadiz, Carlos A. Sing Long

The sweet spot can be interpreted as the region where acoustic sources create a spatial auditory illusion. We study the problem of maximizing this sweet spot when reproducing a desired sound wave using an array of loudsp…

Morphology Informed Selections for Subword Vocabulary Size

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Currently, guidance around selection of an optimal or appropriate subword vocabulary size is incomplete and confusing at best. Using a measure of subword-morpheme overlap, our analysis shows that one can find a "sweet sp…

Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models

2025-08-27 · Laksh Patel, Neel Shanbhag arxiv

Modern generative models risk overfitting and unintentionally memorizing rare training examples, which can be extracted by adversaries or inflate benchmark performance. We propose Generative Data Cartography (GenDataCart…

Managing Data Lineage of O&G Machine Learning Models: The Sweet Spot for Shale Use Case

2020-03-10 · Raphael Thiago, Renan Souza, L. Azevedo, E. Soares 외

Machine Learning (ML) has increased its role, becoming essential in several industries. However, questions around training data lineage, such as "where has the dataset used to train this model come from?"; the introducti…

BIG-bench Machine Learning

Architecture Is All You Need: Diversity-Enabled Sweet Spots for Robust Humanoid Locomotion

2025-10-16 · Blake Werner, Lizhi Yang, Aaron D. Ames arxiv

Robust humanoid locomotion in unstructured environments requires architectures that balance fast low-level stabilization with slower perceptual decision-making. We show that a simple layered control architecture (LCA), a…