paper-with-me

홈 › Papers

Sparsity in long-time control of neural ODEs

2021-02-26 · Carlos Esteve-Yagüe, Borjan Geshkovski

We consider the neural ODE and optimal control perspective of supervised learning, with $\ell^1$-control penalties, where rather than only minimizing a final cost (the \emph{empirical risk}) for the state, we integrate this cost over the entire time horizon. We prove that any optimal control (for this cost) vanishes beyond some positive stopping time. When seen in the discrete-time context, this result entails an \emph{ordered} sparsity pattern for the parameters of the associated residual neural network: ordered in the sense that these parameters are all $0$ beyond a certain layer. Furthermore, we provide a polynomial stability estimate for the empirical risk with respect to the time horizon. This can be seen as a \emph{turnpike property}, for nonsmooth dynamics and functionals with $\ell^1$-penalties, and without any smallness assumptions on the data, both of which are new in the literature.

📄 PDF Abstract BibTeX arXiv:2102.13566

Code (1)

borjanG/dynamical.systems 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference

2026-02-09 · Yifei Gao, Lei Wang, Rong-Cheng Tu, Qixin Zhang 외 arxiv

A core bottleneck in large language model (LLM) inference is the cost of attending over the ever-growing key-value (KV) cache. Although near-oracle top-k KV selection can preserve the quality of dense attention while sha…

HashAttention: Semantic Sparsity for Faster Inference

2024-12-19 · Aditya Desai, Shuo Yang, Alejandro Cuadron, Ana Klimovic 외

Utilizing longer contexts is increasingly essential to power better AI systems. However, the cost of attending to long contexts is high due to the involved softmax computation. While the scaled dot-product attention (SDP…

GPUSemantic SimilaritySemantic Textual Similarity

Fast and Controllable Post-training Sparsity: Learning Optimal Sparsity Allocation with Global Constraint in Minutes

2024-05-09 · Ruihao Gong, Yang Yong, Zining Wang, Jinyang Guo 외

Neural network sparsity has attracted many research interests due to its similarity to biological schemes and high energy efficiency. However, existing methods depend on long-time training or fine-tuning, which prevents …

VideoSAGE: Video Summarization with Graph Representation Learning

2024-04-14 · Jose M. Rojas Chaves, Subarna Tripathi

We propose a graph-based representation learning framework for video summarization. First, we convert an input video to a graph where nodes correspond to each of the video frames. Then, we impose sparsity on the graph by…

Graph Representation LearningNode ClassificationRepresentation LearningVideo Summarization

Scaling Attention via Feature Sparsity

2026-03-17 · Yan Xie, Tiansheng Wen, Tangda Huang, Bo Chen 외 arxiv

Scaling Transformers to ultra-long contexts is bottlenecked by the $O(n^2 d)$ cost of self-attention. Existing methods reduce this cost along the sequence axis through local windows, kernel approximations, or token-level…