paper-with-me

홈 › Papers

Visual Bridge: Universal Visual Perception Representations Generating

2025-11-11 · Yilin Gao, Shuguang Dou, Junzhou Li, Zhiheng Yu, Yin Li, Dongsheng Jiang, Shugong Xu arxiv

Recent advances in diffusion models have achieved remarkable success in isolated computer vision tasks such as text-to-image generation, depth estimation, and optical flow. However, these models are often restricted by a ``single-task-single-model'' paradigm, severely limiting their generalizability and scalability in multi-task scenarios. Motivated by the cross-domain generalization ability of large language models, we propose a universal visual perception framework based on flow matching that can generate diverse visual representations across multiple tasks. Our approach formulates the process as a universal flow-matching problem from image patch tokens to task-specific representations rather than an independent generation or regression problem. By leveraging a strong self-supervised foundation model as the anchor and introducing a multi-scale, circular task embedding mechanism, our method learns a universal velocity field to bridge the gap between heterogeneous tasks, supporting efficient and flexible representation transfer. Extensive experiments on classification, detection, segmentation, depth estimation, and image-text retrieval demonstrate that our model achieves competitive performance in both zero-shot and fine-tuned settings, outperforming prior generalist and several specialist models. Ablation studies further validate the robustness, scalability, and generalization of our framework. Our work marks a significant step towards general-purpose visual perception, providing a solid foundation for future research in universal vision modeling.

📄 PDF Abstract BibTeX arXiv:2511.07877

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image GenerationDomain GeneralizationDepth EstimationText Retrieval

Similar Papers 제목 키워드 기반

MetaUAS: Universal Anomaly Segmentation with One-Prompt Meta-Learning

2025-05-14 · Bin-Bin Gao

Zero- and few-shot visual anomaly segmentation relies on powerful vision-language models that detect unseen anomalies using manually designed textual prompts. However, visual representations are inherently independent of…

Anomaly DetectionAnomaly SegmentationMeta-LearningSegmentation+1

Visually Descriptive Language Model for Vector Graphics Reasoning

2024-04-09 · Zhenhailong Wang, Joy Hsu, Xingyao Wang, Kuan-Hao Huang 외

Despite significant advancements, large multimodal models (LMMs) still struggle to bridge the gap between low-level visual perception -- focusing on shapes, sizes, and layouts -- and high-level language reasoning, such a…

DescriptiveLanguage ModelingLanguage ModellingQuestion Answering+3

Unifying Visual Perception by Dispersible Points Learning

2022-08-18 · Jianming Liang, Guanglu Song, Biao Leng, Yu Liu

We present a conceptually simple, flexible, and universal visual perception head for variant visual tasks, e.g., classification, object detection, instance segmentation and pose estimation, and different frameworks, such…

Instance SegmentationObjectobject-detectionObject Detection+3

AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models

2023-09-19 · Yuan Tseng, Layne Berry, Yi-Ting Chen, I-Hsiang Chiu 외

Audio-visual representation learning aims to develop systems with human-like perception by utilizing correlation between auditory and visual information. However, current models often focus on a limited set of tasks, and…

audio-visual learningRepresentation Learning

Aligning and Prompting Everything All at Once for Universal Visual Perception

2023-12-04 · CVPR 2024 1 · Yunhang Shen, Chaoyou Fu, Peixian Chen, Mengdan Zhang 외

Vision foundation models have been explored recently to build general-purpose vision systems. However, predominant paradigms, driven by casting instance-level tasks as an object-word alignment, bring heavy cross-modality…

AllObjectobject-detectionObject Detection+5