paper-with-me

홈 › Papers

Optimal Regularization for a Data Source

2022-12-27 · Oscar Leong, Eliza O'Reilly, Yong Sheng Soh, Venkat Chandrasekaran

In optimization-based approaches to inverse problems and to statistical estimation, it is common to augment criteria that enforce data fidelity with a regularizer that promotes desired structural properties in the solution. The choice of a suitable regularizer is typically driven by a combination of prior domain information and computational considerations. Convex regularizers are attractive computationally but they are limited in the types of structure they can promote. On the other hand, nonconvex regularizers are more flexible in the forms of structure they can promote and they have showcased strong empirical performance in some applications, but they come with the computational challenge of solving the associated optimization problems. In this paper, we seek a systematic understanding of the power and the limitations of convex regularization by investigating the following questions: Given a distribution, what is the optimal regularizer for data drawn from the distribution? What properties of a data source govern whether the optimal regularizer is convex? We address these questions for the class of regularizers specified by functionals that are continuous, positively homogeneous, and positive away from the origin. We say that a regularizer is optimal for a data distribution if the Gibbs density with energy given by the regularizer maximizes the population likelihood (or equivalently, minimizes cross-entropy loss) over all regularizer-induced Gibbs densities. As the regularizers we consider are in one-to-one correspondence with star bodies, we leverage dual Brunn-Minkowski theory to show that a radial function derived from a data distribution is akin to a ``computational sufficient statistic'' as it is the key quantity for identifying optimal regularizers and for assessing the amenability of a data source to convex regularization.

📄 PDF Abstract BibTeX arXiv:2212.13597

Code (0)

등록된 구현이 없습니다.

Tasks

Dictionary Learning

Similar Papers 제목 키워드 기반

Source-Optimal Training is Transfer-Suboptimal

2025-11-11 · C. Evans Hedges arxiv

We prove that training a source model optimally for its own task is generically suboptimal when the objective is downstream transfer. We study the source-side optimization problem in L2-SP ridge regression and show a fun…

Transfer Learning

On the Interaction of Regularization Factors in Low-resource Neural Machine Translation

2022-06-01 · EAMT 2022 6 · None Àlex R. Atrio, Andrei Popescu-Belis

We explore the roles and interactions of the hyper-parameters governing regularization, and propose a range of values applicable to low-resource neural machine translation. We demonstrate that default or recommended valu…

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationTranslation

Optimal Multi-Task Learning at Regularization Horizon for Speech Translation Task

2025-09-04 · JungHo Jung, Junhyun Lee arxiv

End-to-end speech-to-text translation typically suffers from the scarcity of paired speech-text data. One way to overcome this shortcoming is to utilize the bitext data from the Machine Translation (MT) task and perform …

Speech-to-Text TranslationMachine TranslationMulti-Task Learning

Decreasing Entropic Regularization Averaged Gradient for Semi-Discrete Optimal Transport

2025-10-31 · Ferdinand Genans, Antoine Godichon-Baggioni, François-Xavier Vialard, Olivier Wintenberger arxiv

Adding entropic regularization to Optimal Transport (OT) problems has become a standard approach for designing efficient and scalable solvers. However, regularization introduces a bias from the true solution. To mitigate…

Optimal Rates For Regularization Of Statistical Inverse Learning Problems

2016-04-14 · Gilles Blanchard, Nicole Mücke

We consider a statistical inverse learning problem, where we observe the image of a function $f$ through a linear operator $A$ at i.i.d. random design points $X_i$, superposed with an additive noise. The distribution of …