paper-with-me

Papers

Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent

2024-11-08 · Linfeng He, Yiming Sun, Sihao Wu, Jiaxu Liu, Xiaowei Huang

In this paper, we propose a novel framework for enhancing visual comprehension in autonomous driving systems by integrating visual language models (VLMs) with additional visual perception module specialised in object detection. We extend the Llama-Adapter architecture by incorporating a YOLOS-based detection network alongside the CLIP perception network, addressing limitations in object detection and localisation. Our approach introduces camera ID-separators to improve multi-view processing, crucial for comprehensive environmental awareness. Experiments on the DriveLM visual question answering challenge demonstrate significant improvements over baseline models, with enhanced performance in ChatGPT scores, BLEU scores, and CIDEr metrics, indicating closeness of model answer to ground truth. Our method represents a promising step towards more capable and interpretable autonomous driving systems. Possible safety enhancement enabled by detection modality is also discussed.

📄 PDF Abstract BibTeX arXiv:2411.05898

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingLanguage ModelingLanguage Modellingobject-detectionObject DetectionQuestion AnsweringVisual Question Answering

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Integrating Audio-Visual Features for Multimodal Deepfake Detection

2023-10-05 · Sneha Muppalla, Shan Jia, Siwei Lyu

Deepfakes are AI-generated media in which an image or video has been digitally modified. The advancements made in deepfake technology have led to privacy and security issues. Most deepfake detection techniques rely on th…

Binary ClassificationDeepFake DetectionFace Swapping

TSJNet: A Multi-modality Target and Semantic Awareness Joint-driven Image Fusion Network

2024-02-02 · Yuchan Jie, Yushen Xu, Xiaosong Li, Haishu Tan

Multi-modality image fusion involves integrating complementary information from different modalities into a single image. Current methods primarily focus on enhancing image fusion with a single advanced task such as inco…

Objectobject-detectionObject DetectionSegmentation

Formula-Supervised Visual-Geometric Pre-training

2024-09-20 · Ryosuke Yamada, Kensho Hara, Hirokatsu Kataoka, Koshi Makihara 외

Throughout the history of computer vision, while research has explored the integration of images (visual) and point clouds (geometric), many advancements in image and 3D object recognition have tended to process these mo…

3D Object Classification3D Object RecognitionObjectObject Recognition+1

ModalPatch: A Plug-and-Play Module for Robust Multi-Modal 3D Object Detection under Modality Drop

2026-03-03 · Shuangzhi Li, Lei Ma, Xingyu Li arxiv

Multi-modal 3D object detection is pivotal for autonomous driving, integrating complementary sensors like LiDAR and cameras. However, its real-world reliability is challenged by transient data interruptions and missing, …

3D Object DetectionAutonomous Driving

VirPro: Visual-referred Probabilistic Prompt Learning for Weakly-Supervised Monocular 3D Detection

2026-03-18 · Chupeng Liu, Jiyong Rao, Shangquan Sun, Runkai Zhao 외 arxiv

Monocular 3D object detection typically relies on pseudo-labeling techniques to reduce dependency on real-world annotations. Recent advances demonstrate that deterministic linguistic cues can serve as effective auxiliary…

Monocular 3D Object Detection