paper-with-me

Papers

Subobject-level Image Tokenization

2024-02-22 · Delong Chen, Samuel Cahyawijaya, Jianfeng Liu, Baoyuan Wang, Pascale Fung

Transformer-based vision models typically tokenize images into fixed-size square patches as input units, which lacks the adaptability to image content and overlooks the inherent pixel grouping structure. Inspired by the subword tokenization widely adopted in language models, we propose an image tokenizer at a subobject level, where the subobjects are represented by semantically meaningful image segments obtained by segmentation models (e.g., segment anything models). To implement a learning system based on subobject tokenization, we first introduced a Direct Segment Anything Model (DirectSAM) that efficiently produces comprehensive segmentation of subobjects, then embed subobjects into compact latent vectors and fed them into a large language model for vision language learning. Empirical results demonstrated that our subobject-level tokenization significantly facilitates efficient learning of translating images into object and attribute descriptions compared to the traditional patch-level tokenization. Codes and models are open-sourced at https://github.com/ChenDelong1999/subobjects.

📄 PDF Abstract BibTeX arXiv:2402.14327

Code (1)

chendelong1999/subobjects 공식 구현 pytorch

Tasks

AttributeLanguage ModelingLanguage ModellingLarge Language ModelSegmentation

Similar Papers 제목 키워드 기반

Topos Theory for Generative AI and LLMs

2025-08-05 · Sridhar Mahadevan arxiv

We propose the design of novel categorical generative AI architectures (GAIAs) using topos theory, a type of category that is ``set-like": a topos has all (co)limits, is Cartesian closed, and has a subobject classifier. …

Topos Causal Models

2025-08-05 · Sridhar Mahadevan arxiv

We propose topos causal models (TCMs), a novel class of causal models that exploit the key properties of a topos category: they are (co)complete, meaning all (co)limits exist, they admit a subobject classifier, and allow…

Causal Inference

Language-Guided Image Tokenization for Generation

2024-12-08 · CVPR 2025 1 · Kaiwen Zha, Lijun Yu, Alireza Fathi, David A. Ross 외

Image tokenization, the process of transforming raw image pixels into a compact low-dimensional latent representation, has proven crucial for scalable and efficient image generation. However, mainstream image tokenizatio…

DescriptiveImage GenerationText to Image GenerationText-to-Image Generation

A Spitting Image: Modular Superpixel Tokenization in Vision Transformers

2024-08-14 · Marius Aasan, Odd Kolbjørnsen, Anne Schistad Solberg, Adín Ramirez Rivera

Vision Transformer (ViT) architectures traditionally employ a grid-based approach to tokenization independent of the semantic content of an image. We propose a modular superpixel tokenization strategy which decouples tok…

Intuitionistic $j$-Do-Calculus in Topos Causal Models

2025-10-20 · Sridhar Mahadevan arxiv

In this paper, we generalize Pearl's do-calculus to an Intuitionistic setting called $j$-stable causal inference inside a topos of sheaves. Our framework is an elaboration of the recently proposed framework of Topos Caus…

Causal Inference