paper-with-me

홈 › Papers

Sublinearly Structured Deep Neural Networks Achieve Feature Learning Consistency for Compositional Functions

2026-06-22 · Sehwan Kim, Yan Sun, Faming Liang arxiv

Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine-learning tasks, yet the theoretical foundations of their performance remain incomplete. From a statistical viewpoint, a natural question is: can DNNs attain feature-learning and prediction consistency comparable to that of classical models? While a full characterization is open, we provide positive results for a broad subclass. We establish feature-learning consistency guarantees for sublinearly structured DNNs-architectures whose input/output dimensions and number of hidden neurons grow sublinearly with the sample size-when learning hierarchically compositional target functions. Importantly, this consistency still holds even in the conventional "over-parameterized" regime where the total number of parameters exceeds the number of training samples. Empirically, sublinearly structured DNNs match or surpass wide DNNs in prediction. A structural audit further indicates that widely used convolutional neural networks (CNNs), including AlexNet, VGGNet, ResNet, GoogLeNet, are sublinearly structured on their image classification benchmarks. We further prove that the sublinearly structured DNNs achieve universal approximation for hierarchically compositional functions in the large-sample limit. Moreover, images exhibit an inherent hierarchical, compositional structure. Taken together, these results explain, through a statistical lens, why many large-scale deep learning models succeed after adequate training on massive image datasets.

📄 PDF Abstract BibTeX arXiv:2606.23477

Code (0)

등록된 구현이 없습니다.

Tasks

Image Classification

Similar Papers 제목 키워드 기반

Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension

2026-07-06 · Lian Xu, Mohammed Bennamoun, Farid Boussaid, Hamid Laga 외 arxiv

Referring expression comprehension (REC) aims to localize the object in an image described by natural language. In Weakly supervised REC (WREC), existing approaches primarily operate on anchor-level visual representation…

Referring Expression

CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information Alignment

2025-02-16 · Nura Aljaafari, Danilo S. Carvalho, André Freitas

Large language models (LLMs) struggle with compositional generalisation, limiting their ability to systematically combine learned components to interpret novel inputs. While architectural modifications, fine-tuning, and …

Data AugmentationSentiment AnalysisSentiment Classification

Consistency of Compositional Generalization across Multiple Levels

2024-12-18 · Chuanhao Li, Zhen Li, Chenchen Jing, Xiaomeng Fan 외

Compositional generalization is the capability of a model to understand novel compositions composed of seen concepts. There are multiple levels of novel compositions including phrase-phrase level, phrase-word level, and …

Meta-LearningQuestion AnsweringVideo GroundingVisual Question Answering

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching

2026-05-20 · Xuefei Sun, Xujia Zhang, Brendan Crowe, Doncey Albin 외 arxiv

Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VLM) approaches achieve promising results but rely on view-dependent r…

Spatial ReasoningVisual GroundingGraph Matching

3D-SSGAN: Lifting 2D Semantics for 3D-Aware Compositional Portrait Synthesis

2024-01-08 · Ruiqi Liu, Peng Zheng, Ye Wang, Rui Ma

Existing 3D-aware portrait synthesis methods can generate impressive high-quality images while preserving strong 3D consistency. However, most of them cannot support the fine-grained part-level control over synthesized i…

DisentanglementImage Generation