paper-with-me

홈 › Papers

Training Memory in Deep Neural Networks: Mechanisms, Evidence, and Measurement Gaps

2026-01-29 · Vasileios Sevetlidis, George Pavlidis arxiv

Modern deep-learning training is not memoryless. Updates depend on optimizer moments and averaging, data-order policies (random reshuffling vs with-replacement, staged augmentations and replay), the nonconvex path, and auxiliary state (teacher EMA/SWA, contrastive queues, BatchNorm statistics). This survey organizes mechanisms by source, lifetime, and visibility. It introduces seed-paired, function-space causal estimands; portable perturbation primitives (carry/reset of momentum/Adam/EMA/BN, order-window swaps, queue/teacher tweaks); and a reporting checklist with audit artifacts (order hashes, buffer/BN checksums, RNG contracts). The conclusion is a protocol for portable, causal, uncertainty-aware measurement that attributes how much training history matters across models, data, and regimes.

📄 PDF Abstract BibTeX arXiv:2601.21624

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

2026-05-25 · Tianshuo Xu, Yichen Xie, Depu Meng, Chensheng Peng 외 arxiv

Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity problem: pretrained video diffusion trans…

PolicyBank: Evolving Policy Understanding for LLM Agents

2026-04-16 · Jihye Choi, Jinsung Yoon, Long T. Le, Somesh Jha 외 arxiv

LLM agents operating under organizational policies must comply with authorization constraints typically specified in natural language. In practice, such specifications inevitably contain ambiguities and logical or semant…

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering

2026-05-25 · Ming Xie, Zizheng Huang, Xudong Tan, Chao Wang 외 arxiv

While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal methods are inherently designed for offlin…

Question AnsweringVisual Reasoning

Deep Learning to Attend to Risk in ICU

2017-07-17 · Phuoc Nguyen, Truyen Tran, Svetha Venkatesh

Modeling physiological time-series in ICU is of high clinical importance. However, data collected within ICU are irregular in time and often contain missing measurements. Since absence of a measure would signify its lack…

Decision MakingDeep LearningICU MortalityTime Series+1

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

2026-09-09 · Killian Steunou, Yannis Tevissen, Mounîm A. El Yacoubi arxiv

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their …

Question Answering