paper-with-me

홈 › Papers

NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave Radar

2024-08-30 · Runwei Guan, Jianan Liu, Liye Jia, Haocheng Zhao, Shanliang Yao, Xiaohui Zhu, Ka Lok Man, Eng Gee Lim, Jeremy Smith, Yutao Yue

Recently, visual grounding and multi-sensors setting have been incorporated into perception system for terrestrial autonomous driving systems and Unmanned Surface Vehicles (USVs), yet the high complexity of modern learning-based visual grounding model using multi-sensors prevents such model to be deployed on USVs in the real-life. To this end, we design a low-power multi-task model named NanoMVG for waterway embodied perception, guiding both camera and 4D millimeter-wave radar to locate specific object(s) through natural language. NanoMVG can perform both box-level and mask-level visual grounding tasks simultaneously. Compared to other visual grounding models, NanoMVG achieves highly competitive performance on the WaterVG dataset, particularly in harsh environments and boasts ultra-low power consumption for long endurance.

📄 PDF Abstract BibTeX arXiv:2408.17207

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingVisual Grounding

Similar Papers 제목 키워드 기반

CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks

2026-01-19 · Mingshuang Luo, Ruibing Hou, Bo Chao, Hong Chang 외 arxiv

Human-centric visual analysis plays a pivotal role in diverse applications, including surveillance, healthcare, and human-computer interaction. With the emergence of large-scale unlabeled human image datasets, there is a…

Unsupervised Pre-trainingSelf-Supervised LearningRepresentation Learning

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

2025-08-18 · Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to…

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

2026-07-16 · Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu, Xu Zheng 외 arxiv

Egocentric Visual Question Answering (VQA) has attracted widespread attention as an important task for enabling Multimodal Large Language Models (MLLMs) to interact with the real world. However, existing MLLMs struggle t…

Visual Question AnsweringSpatial Reasoning

Visual Jigsaw Post-Training Improves MLLMs

2025-09-29 · Penghao Wu, Yushan Zhang, Haiwen Diao, Bo Li 외 arxiv

Reinforcement learning based post-training has recently emerged as a powerful paradigm for enhancing the alignment and reasoning capabilities of multimodal large language models (MLLMs). While vision-centric post-trainin…

Reinforcement Learning

MM-Ego: Towards Building Egocentric Multimodal LLMs

2024-10-09 · Hanrong Ye, Haotian Zhang, Erik Daxberger, Lin Chen 외

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric …

Video Understanding