paper-with-me

홈 › Papers

Structure over Pixels: Learning Variable-Length Visual Programs

2026-05-26 · Piotr Wyrwiński, Kacper Dobek, Krzysztof Krawiec arxiv

Discrete visual tokenizers translate images into ordered sequences of codes, providing a natural representation for structural description of scenes. Yet existing adaptive tokenizers either require post-hoc search or select among a discrete set of pre-trained rates, rather than learning a continuous per-image sequence length coupled to the model and scene, and they typically train against pixel reconstruction, emphasizing texture rather than structure. We propose STROP, a discrete visual tokenizer architecture that forms structural scene representations and simultaneously learns how long an image's visual program should be. Using a four-phase curriculum supervised by local rate--distortion probes against frozen DINOv3 features, STROP optimizes a dedicated length head that estimates the active prefix length in a single forward pass. By bypassing pixel-level reconstruction gradients, the codebook is shaped entirely by the quality of higher-level latent representations. Program length grows with scene complexity, and signs of compositional structure emerge both in downstream dense-prediction transfer and in direct inspection of the learned code vocabulary.

📄 PDF Abstract BibTeX arXiv:2605.27696

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Systematic Evaluation of Causal Discovery in Visual Model Based Reinforcement Learning

2021-07-02 · Nan Rosemary Ke, Aniket Didolkar, Sarthak Mittal, Anirudh Goyal 외

Inducing causal relationships from observations is a classic problem in machine learning. Most work in causality starts from the premise that the causal variables themselves are observed. However, for AI agents such as r…

BenchmarkingCausal DiscoveryModel-based Reinforcement Learningreinforcement-learning+2

Adaptive Digital Scan Variable Pixels

2015-06-22 · Sherin Sugathan, Reshma Scaria, Alex Pappachen James

The square and rectangular shape of the pixels in the digital images for sensing and display purposes introduces several inaccuracies in the representation of digital images. The major disadvantage of square pixel shapes…

Visualizing Semantic Structures of Sequential Data by Learning Temporal Dependencies

2019-01-20 · Kyoung-Woon On, Eun-Sol Kim, Yu-Jung Heo, Byoung-Tak Zhang

While conventional methods for sequential learning focus on interaction between consecutive inputs, we suggest a new method which captures composite semantic flows with variable-length dependencies. In addition, the sema…

VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization

2026-04-14 · Andrei Atanov, Jesse Allardice, Roman Bachmann, Oğuzhan Fatih Kar 외 arxiv

Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what information is preserved and how it is organized. A de facto standar…

Video Generation

Constrained Manifold Learning for Hyperspectral Imagery Visualization

2017-11-24 · Danping Liao, Yuntao Qian, Yuan Yan Tang

Displaying the large number of bands in a hyper- spectral image (HSI) on a trichromatic monitor is important for HSI processing and analysis system. The visualized image shall convey as much information as possible from …