paper-with-me

Papers

Dissociating model architectures from inference computations

2025-07-21 · Noor Sajid, Johan Medrano arxiv

Parr et al., 2025 examines how auto-regressive and deep temporal models differ in their treatment of non-Markovian sequence modelling. Building on this, we highlight the need for dissociating model architectures, i.e., how the predictive distribution factorises, from the computations invoked at inference. We demonstrate that deep temporal computations are mimicked by autoregressive models by structuring context access during iterative inference. Using a transformer trained on next-token prediction, we show that inducing hierarchical temporal factorisation during iterative inference maintains predictive capacity while instantiating fewer computations. This emphasises that processes for constructing and refining predictions are not necessarily bound to their underlying model architectures.

📄 PDF Abstract BibTeX arXiv:2507.15776

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Training Recurrent Neural Networks against Noisy Computations during Inference

2018-07-17 · Minghai Qin, Dejan Vucinic

We explore the robustness of recurrent neural networks when the computations within the network are noisy. One of the motivations for looking into this problem is to reduce the high power cost of conventional computing o…

CPUGPUspeech-recognitionSpeech Recognition

Recurrent computations for visual pattern completion

2017-06-07 · Hanlin Tang, Martin Schrimpf, Bill Lotter, Charlotte Moerman 외

Making inferences from partial information constitutes a critical aspect of cognition. During visual perception, pattern completion enables recognition of poorly visible or occluded objects. We combined psychophysics, ph…

Image Classification

Serving Deep Learning Model in Relational Databases

2023-10-07 · Lixi Zhou, Qi Lin, Kanchan Chowdhury, Saif Masood 외

Serving deep learning (DL) models on relational data has become a critical requirement across diverse commercial and scientific domains, sparking growing interest recently. In this visionary paper, we embark on a compreh…

Deep LearningManagementmodel

Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation

2025-11-16 · Yushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou 외 arxiv

While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, …

Dual-stream spatiotemporal networks with feature sharing for monitoring animals in the home cage

2022-06-01 · Ezechukwu I. Nwokedi, Rasneer S. Bains, Luc Bidaut, Xujiong Ye 외

This paper presents a spatiotemporal deep learning approach for mouse behavioural classification in the home-cage. Using a series of dual-stream architectures with assorted modifications to increase performance, we intro…

Anomaly DetectionUnsupervised Anomaly Detection