paper-with-me

홈 › Papers

Learning the Hypotheses Space from data Part II: Convergence and Feasibility

2020-01-30 · Diego Marcondes, Adilson Simonis, Junior Barrera

In part \textit{I} we proposed a structure for a general Hypotheses Space $\mathcal{H}$, the Learning Space $\mathbb{L}(\mathcal{H})$, which can be employed to avoid \textit{overfitting} when estimating in a complex space with relative shortage of examples. Also, we presented the U-curve property, which can be taken advantage of in order to select a Hypotheses Space without exhaustively searching $\mathbb{L}(\mathcal{H})$. In this paper, we carry further our agenda, by showing the consistency of a model selection framework based on Learning Spaces, in which one selects from data the Hypotheses Space on which to learn. The method developed in this paper adds to the state-of-the-art in model selection, by extending Vapnik-Chervonenkis Theory to \textit{random} Hypotheses Spaces, i.e., Hypotheses Spaces learned from data. In this framework, one estimates a random subspace $\hat{\mathcal{M}} \in \mathbb{L}(\mathcal{H})$ which converges with probability one to a target Hypotheses Space $\mathcal{M}^{\star} \in \mathbb{L}(\mathcal{H})$ with desired properties. As the convergence implies asymptotic unbiased estimators, we have a consistent framework for model selection, showing that it is feasible to learn the Hypotheses Space from data. Furthermore, we show that the generalization errors of learning on $\hat{\mathcal{M}}$ are lesser than those we commit when learning on $\mathcal{H}$, so it is more efficient to learn on a subspace learned from data.

📄 PDF Abstract BibTeX arXiv:2001.11578

Code (0)

등록된 구현이 없습니다.

Tasks

Model Selection

Similar Papers 제목 키워드 기반

Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Science

2025-06-04 · Peter Jansen, Samiah Hassan, Ruoyao Wang

Contemporary approaches to assisted scientific discovery use language models to automatically generate large numbers of potential hypothesis to test, while also automatically generating code-based experiments to test tho…

ArticlesCode GenerationRetrieval-augmented Generationscientific discovery

HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation

2025-10-01 · Rosni Vasu, Peter Jansen, Pao Siangliulue, Cristina Sarasua 외 arxiv

While there has been a surge of interest in automated scientific discovery (ASD), especially with the emergence of LLMs, it remains challenging for tools to generate hypotheses that are both testable and grounded in the …

Bayes-Entropy Collaborative Driven Agents for Research Hypotheses Generation and Optimization

2025-08-03 · Shiyang Duan, Yuan Tian, Qi Bing, Xiaowei Shao arxiv

The exponential growth of scientific knowledge has made the automated generation of scientific hypotheses that combine novelty, feasibility, and research value a core challenge. Existing methods based on large language m…

Action Mapping for Reinforcement Learning in Continuous Environments with Constraints

2024-12-05 · Mirco Theile, Lukas Dirnberger, Raphael Trumpp, Marco Caccamo 외

Deep reinforcement learning (DRL) has had success across various domains, but applying it to environments with constraints remains challenging due to poor sample efficiency and slow convergence. Recent literature explore…

Deep Reinforcement Learning

A Fine-Grained Understanding of Uniform Convergence for Halfspaces

2026-05-07 · Aryeh Kontorovich, Kasper Green Larsen arxiv

We study the fine-grained uniform convergence behavior of halfspaces beyond worst-case VC bounds. For inhomogeneous halfspaces in $\mathbb{R}^d$ with $d\ge 2$, we show that standard first-order VC bounds are essentially …