paper-with-me

Papers

Understanding Learning Dynamics Through Structured Representations

2025-08-04 · Saleh Nikooroo, Thomas Engel arxiv

While modern deep networks have demonstrated remarkable versatility, their training dynamics remain poorly understood--often driven more by empirical tweaks than architectural insight. This paper investigates how internal structural choices shape the behavior of learning systems. Building on prior efforts that introduced simple architectural constraints, we explore the broader implications of structure for convergence, generalization, and adaptation. Our approach centers on a family of enriched transformation layers that incorporate constrained pathways and adaptive corrections. We analyze how these structures influence gradient flow, spectral sensitivity, and fixed-point behavior--uncovering mechanisms that contribute to training stability and representational regularity. Theoretical analysis is paired with empirical studies on synthetic and structured tasks, demonstrating improved robustness, smoother optimization, and scalable depth behavior. Rather than prescribing fixed templates, we emphasize principles of tractable design that can steer learning behavior in interpretable ways. Our findings support a growing view that architectural design is not merely a matter of performance tuning, but a critical axis for shaping learning dynamics in scalable and trustworthy neural systems.

📄 PDF Abstract BibTeX arXiv:2508.02126

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Self-Supervised Learning of Structured Dynamics from Videos

2026-07-23 · Lukas Knobel, Andrew Zisserman, Yuki M. Asano arxiv

Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in …

Self-Supervised LearningRepresentation Learning

Unsupervised Learning of Object Structure and Dynamics from Videos

2019-06-19 · NeurIPS 2019 12 · Matthias Minderer, Chen Sun, Ruben Villegas, Forrester Cole 외

Extracting and predicting object structure and dynamics from videos without supervision is a major challenge in machine learning. To address this challenge, we adopt a keypoint-based image representation and learn a stoc…

Action Recognitioncontinuous-controlContinuous ControlObject+2

What's In Your Field? Mapping Scientific Research with Knowledge Graphs and Large Language Models

2025-03-12 · Abhipsha Das, Nicholas Lourie, Siavash Golkar, Mariel Pettee

The scientific literature's exponential growth makes it increasingly challenging to navigate and synthesize knowledge across disciplines. Large language models (LLMs) are powerful tools for understanding scientific text,…

Knowledge GraphsNavigateRetrieval-augmented Generation

GeoSem-WAM: Geometry- and Semantic-Aware World Action Models

2026-06-02 · Fulong Ma, Daojie Peng, Wenjun Yue, Jiahang Cao 외 arxiv

Recent World Action Models (WAMs) have demonstrated impressive capabilities in embodied decision-making. However, whether their effectiveness stems from explicit future imagination during inference or representation lear…

Representation LearningScene UnderstandingVideo Generation

Multi-turn Physics-informed Vision-language Model for Physics-grounded Anomaly Detection

2026-03-16 · Yao Gu, Xiaohao Xu, Yingna Wu arxiv

Vision-Language Models (VLMs) demonstrate strong general-purpose reasoning but remain limited in physics-grounded anomaly detection, where causal understanding of dynamics is essential. Existing VLMs, trained predominant…

Anomaly Detection