paper-with-me

홈 › Papers

Teaching Metric Distance to Autoregressive Multimodal Foundational Models

2025-03-04 · Jiwan Chung, Saejin Kim, Yongrae Jo, Jaewoo Park, Dongjun Min, Youngjae Yu

As large language models expand beyond natural language to domains such as mathematics, multimodal understanding, and embodied agents, tokens increasingly reflect metric relationships rather than purely linguistic meaning. We introduce DIST2Loss, a distance-aware framework designed to train autoregressive discrete models by leveraging predefined distance relationships among output tokens. At its core, DIST2Loss transforms continuous exponential family distributions derived from inherent distance metrics into discrete, categorical optimization targets compatible with the models' architectures. This approach enables the models to learn and preserve meaningful distance relationships during token generation while maintaining compatibility with existing architectures. Empirical evaluations show consistent performance gains in diverse multimodal applications, including visual grounding, robotic manipulation, generative reward modeling, and image generation using vector-quantized features. These improvements are most notable in low-data regimes, demonstrating DIST2Loss's strength under resource constraints.

📄 PDF Abstract BibTeX arXiv:2503.02379

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationVisual Grounding

Similar Papers 제목 키워드 기반

Structural Analysis of Vector Autoregressive Models

2023-12-11 · Christis Katsouris

This set of lecture notes discuss key concepts for the Structural Analysis of Vector Autoregressive models for the teaching of a course on Applied Macroeconometrics with Advanced Topics.

PROSE-FD: A Multimodal PDE Foundation Model for Learning Multiple Operators for Forecasting Fluid Dynamics

2024-09-15 · Yuxuan Liu, Jingmin Sun, Xinjie He, Griffin Pinney 외

We propose PROSE-FD, a zero-shot multimodal PDE foundational model for simultaneous prediction of heterogeneous two-dimensional physical systems related to distinct fluid dynamics settings. These systems include shallow …

Operator learningPrediction

A Study of Autoregressive Decoders for Multi-Tasking in Computer Vision

2023-03-30 · Lucas Beyer, Bo Wan, Gagan Madan, Filip Pavetic 외

There has been a recent explosion of computer vision models which perform many tasks and are composed of an image encoder (usually a ViT) and an autoregressive decoder (usually a Transformer). However, most of this work …

DecoderMulti-Task LearningOptical Character RecognitionQuestion Answering+2

Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities

2025-05-05 · Xinjie Zhang, Jintao Guo, Shanshan Zhao, Minghao Fu 외

Recent years have seen remarkable progress in both multimodal understanding models and image generation models. Despite their respective successes, these two domains have evolved independently, leading to distinct archit…

Image GenerationSurveyText to Image GenerationText-to-Image Generation

Focus Plus: Detect Learner's Distraction by Web Camera in Distance Teaching

2022-10-10 · Eason Chen, Yuen Hsien Tseng, Kuo-Ping Lo

Distance teaching has become popular these years because of the COVID-19 epidemic. However, both students and teachers face several challenges in distance teaching, like being easy to distract. We proposed Focus+, a syst…