paper-with-me

Papers

MambaEye: A Size-Agnostic Visual Encoder with Causal Sequential Processing

2025-11-25 · Changho Choi, Minho Kim, Jinkyu Kim arxiv

Despite decades of progress, a truly input-size agnostic visual encoder-a fundamental characteristic of human vision-has remained elusive. We address this limitation by proposing \textbf{MambaEye}, a novel, causal sequential encoder that leverages the low complexity and causal-process based pure Mamba2 backbone. Unlike previous Mamba-based vision encoders that often employ bidirectional processing, our strictly unidirectional approach preserves the inherent causality of State Space Models, enabling the model to generate a prediction at any point in its input sequence. A core innovation is our use of relative move embedding, which encodes the spatial shift between consecutive patches, providing a strong inductive bias for translation invariance and making the model inherently adaptable to arbitrary image resolutions and scanning patterns. To achieve this, we introduce a novel diffusion-inspired loss function that provides dense, step-wise supervision, training the model to build confidence as it gathers more visual evidence. We demonstrate that MambaEye exhibits robust performance across a wide range of image resolutions, especially at higher resolutions such as $1536^2$ on the ImageNet-1K classification task. This feat is achieved while maintaining linear time and memory complexity relative to the number of patches.

📄 PDF Abstract BibTeX arXiv:2511.19963

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unsupervised Causal Prototypical Networks for De-biased Interpretable Dermoscopy Diagnosis

2026-02-27 · Junhao Jia, Yueyi Wu, Huangwei Chen, Haodong Jing 외 arxiv

Despite the success of deep learning in dermoscopy image analysis, its inherent black-box nature hinders clinical trust, motivating the use of prototypical networks for case-based visual transparency. However, inevitable…

Vision encoders should be image size agnostic and task driven

2025-08-22 · Nedyalko Prisadnikov, Danda Pani Paudel, Yuqian Fu, Luc Van Gool arxiv

This position paper argues that the next generation of vision encoders should be image size agnostic and task driven. The source of our inspiration is biological. Not a structural aspect of biological vision, but a behav…

Image Classification

Estimating Visual Attribute Effects in Advertising from Observational Data: A Deepfake-Informed Double Machine Learning Approach

2026-03-02 · Yizhi Liu, Balaji Padmanabhan, Siva Viswanathan arxiv

Digital advertising increasingly relies on visual content, yet marketers lack rigorous methods for understanding how specific visual attributes causally affect consumer engagement. This paper addresses a fundamental meth…

Causal Inference

DeepSeek-OCR 2: Visual Causal Flow

2026-01-28 · Haoran Wei, Yaofeng Sun, Yukun Li arxiv

We present DeepSeek-OCR 2 to investigate the feasibility of a novel encoder-DeepEncoder V2-capable of dynamically reordering visual tokens upon image semantics. Conventional vision-language models (VLMs) invariably proce…

A Unified Cascaded Encoder ASR Model for Dynamic Model Sizes

2022-04-13 · Shaojin Ding, Weiran Wang, Ding Zhao, Tara N. Sainath 외

In this paper, we propose a dynamic cascaded encoder Automatic Speech Recognition (ASR) model, which unifies models for different deployment scenarios. Moreover, the model can significantly reduce model size and power co…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)modelspeech-recognition+1