paper-with-me

Papers

Delving into Multi-modal Multi-task Foundation Models for Road Scene Understanding: From Learning Paradigm Perspectives

2024-02-05 · Sheng Luo, Wei Chen, Wanxin Tian, Rui Liu, Luanxuan Hou, Xiubao Zhang, Haifeng Shen, Ruiqi Wu, Shuyi Geng, Yi Zhou, Ling Shao, Yi Yang, Bojun Gao, Qun Li, Guobin Wu

Foundation models have indeed made a profound impact on various fields, emerging as pivotal components that significantly shape the capabilities of intelligent systems. In the context of intelligent vehicles, leveraging the power of foundation models has proven to be transformative, offering notable advancements in visual understanding. Equipped with multi-modal and multi-task learning capabilities, multi-modal multi-task visual understanding foundation models (MM-VUFMs) effectively process and fuse data from diverse modalities and simultaneously handle various driving-related tasks with powerful adaptability, contributing to a more holistic understanding of the surrounding scene. In this survey, we present a systematic analysis of MM-VUFMs specifically designed for road scenes. Our objective is not only to provide a comprehensive overview of common practices, referring to task-specific models, unified multi-modal models, unified multi-task models, and foundation model prompting techniques, but also to highlight their advanced capabilities in diverse learning paradigms. These paradigms include open-world understanding, efficient transfer for road scenes, continual learning, interactive and generative capability. Moreover, we provide insights into key challenges and future trends, such as closed-loop driving systems, interpretability, embodied driving agents, and world models. To facilitate researchers in staying abreast of the latest developments in MM-VUFMs for road scenes, we have established a continuously updated repository at https://github.com/rolsheng/MM-VUFM4DS

📄 PDF Abstract BibTeX arXiv:2402.02968

Code (1)

rolsheng/mm-vufm4ds 공식 구현

Tasks

Continual LearningMulti-Task Learningroad scene understandingScene Understanding

Similar Papers 제목 키워드 기반

Knowledge Graphs Meet Multi-Modal Learning: A Comprehensive Survey

2024-02-08 · Zhuo Chen, Yichi Zhang, Yin Fang, Yuxia Geng 외

Knowledge Graphs (KGs) play a pivotal role in advancing various AI applications, with the semantic web community's exploration into multi-modal dimensions unlocking new avenues for innovation. In this survey, we carefull…

ArticlesEntity Alignmentimage-classificationImage Classification+7

SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models

2025-02-18 · Xianfu Cheng, Wei zhang, Shiwei Zhang, Jian Yang 외

The increasing application of multi-modal large language models (MLLMs) across various sectors have spotlighted the essence of their output reliability and accuracy, particularly their ability to produce content grounded…

Image ComprehensionQuestion AnsweringText GenerationVisual Question Answering

DynRefer: Delving into Region-level Multi-modality Tasks via Dynamic Resolution

2024-05-25 · Yuzhong Zhao, Feng Liu, Yue Liu, Mingxiang Liao 외

Region-level multi-modality methods can translate referred image regions to human preferred language descriptions. Unfortunately, most of existing methods using fixed visual inputs remain lacking the resolution adaptabil…

Attribute

DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution

2025-01-01 · CVPR 2025 1 · Yuzhong Zhao, Feng Liu, Yue Liu, Mingxiang Liao 외

One important task of multimodal models is to translate referred image regions to human preferred language descriptions. Existing methods, however, ignore the resolution adaptability needs of different tasks, which h…

Attribute

Delving into Out-of-Distribution Detection with Vision-Language Representations

2022-11-24 · Yifei Ming, Ziyang Cai, Jiuxiang Gu, Yiyou Sun 외

Recognizing out-of-distribution (OOD) samples is critical for machine learning systems deployed in the open world. The vast majority of OOD detection methods are driven by a single modality (e.g., either vision or langua…

Out-of-Distribution Detection