paper-with-me

Papers

LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMs

2025-03-10 · Hanyu Zhou, Gim Hee Lee

Large multimodal models (LMMs) excel in scene understanding but struggle with fine-grained spatiotemporal reasoning due to weak alignment between linguistic and visual representations. Existing methods map textual positions and durations into the visual space encoded from frame-based videos, but suffer from temporal sparsity that limits language-vision temporal coordination. To address this issue, we introduce LLaFEA (Large Language and Frame-Event Assistant) to leverage event cameras for temporally dense perception and frame-event fusion. Our approach employs a cross-attention mechanism to integrate complementary spatial and temporal features, followed by self-attention matching for global spatio-temporal associations. We further embed textual position and duration tokens into the fused visual space to enhance fine-grained alignment. This unified framework ensures robust spatio-temporal coordinate alignment, enabling LMMs to interpret scenes at any position and any time. In addition, we construct a dataset of real-world frames-events with coordinate instructions and conduct extensive experiments to validate the effectiveness of the proposed method.

📄 PDF Abstract BibTeX arXiv:2503.06934

Code (0)

등록된 구현이 없습니다.

Tasks

PositionScene Understanding

Similar Papers 제목 키워드 기반

VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection

2026-03-18 · Bo-Cheng Qiu, Yu-Fan Lin, Yu-Zhe Pien, Chia-Ming Lee 외 arxiv

Capsule endoscopy event detection is challenging because diagnostically relevant findings are sparse, visually heterogeneous, and embedded in long, noisy video streams, while evaluation is performed at the event level ra…

Spatially-guided Temporal Aggregation for Robust Event-RGB Optical Flow Estimation

2025-01-01 · Qianang Zhou, Junhui Hou, Meiyi Yang, Yongjian Deng 외

Current optical flow methods exploit the stable appearance of frame (or RGB) data to establish robust correspondences across time. Event cameras, on the other hand, provide high-temporal-resolution motion cues and excel …

Optical Flow Estimation

Efficient Event Stream Super-Resolution with Recursive Multi-Branch Fusion

2024-06-28 · Quanmin Liang, Zhilin Huang, Xiawu Zheng, Feidiao Yang 외

Current Event Stream Super-Resolution (ESR) methods overlook the redundant and complementary information present in positive and negative events within the event stream, employing a direct mixing approach for super-resol…

Object RecognitionSuper-ResolutionVideo Reconstruction

Bring Event into RGB and LiDAR: Hierarchical Visual-Motion Fusion for Scene Flow

2024-03-12 · CVPR 2024 1 · Hanyu Zhou, Yi Chang, Zhiwei Shi, Luxin Yan

Single RGB or LiDAR is the mainstream sensor for the challenging scene flow, which relies heavily on visual features to match motion features. Compared with single modality, existing methods adopt a fusion strategy to di…

MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic Segmentation

2025-12-30 · Fuqiang Gu, Yuanke Li, Xianlei Long, Kangping Ji 외 arxiv

Semantic segmentation is a fundamental task in computer vision with wide-ranging applications, including autonomous driving and robotics. While RGB-based methods have achieved strong performance with CNNs and Transformer…

Semantic SegmentationAutonomous Driving