paper-with-me

홈 › Papers

EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning

2024-04-22 · Mingjie Ma, zhihuan yu, Yichao Ma, GuoHui Li

Visual Commonsense Reasoning (VCR) is a cognitive task, challenging models to answer visual questions requiring human commonsense, and to provide rationales explaining why the answers are correct. With emergence of Large Language Models (LLMs), it is natural and imperative to explore their applicability to VCR. However, VCR task demands more external knowledge to tackle its challenging questions, necessitating special designs to activate LLMs' commonsense reasoning abilities. Also, most existing Multimodal LLMs adopted an abstraction of entire input image, which makes it difficult to comprehend VCR's unique co-reference tags between image regions and text, posing challenges for fine-grained alignment. To address these issues, we propose EventLens that leverages Event-Aware Pretraining and Cross-modal Linking and EnhanceS VCR. First, by emulating the cognitive process of human reasoning, an Event-Aware Pretraining auxiliary task is introduced to better activate LLM's global comprehension of intricate scenarios. Second, during fine-tuning, we further utilize reference tags to bridge RoI features with texts, while preserving both modality semantics. Finally, we use instruct-style prompts to narrow the gap between pretraining and fine-tuning, and task-specific adapters to better integrate LLM's inherent knowledge with new commonsense. Experimental results show the effectiveness of our proposed auxiliary task and fine-grained linking strategy.

📄 PDF Abstract BibTeX arXiv:2404.13847

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Commonsense Reasoning

Similar Papers 제목 키워드 기반

Generative Event Pretraining with Foundation Model Alignment

2026-03-24 · Jianwen Cao, Jiaxu Xing, Nico Messikommer, Davide Scaramuzza arxiv

Event cameras provide robust visual signals under fast motion and challenging illumination conditions thanks to their microsecond latency and high dynamic range. However, their unique sensing characteristics and limited …

Object RecognitionDepth Estimation

Scaling Recurrence-aware Foundation Models for Clinical Records via Next-Visit Prediction

2026-03-25 · Haresh Rengaraj Rajamohan, Xiang Gao, Weicheng Zhu, Shih-Lun Huang 외 arxiv

While large-scale pretraining has revolutionized language modeling, its potential remains underexplored in healthcare with structured electronic health records (EHRs). We present RAVEN, a novel generative pretraining str…

Semantic-Aware Pretraining for Dense Video Captioning

2022-04-13 · Teng Wang, Zhu Liu, Feng Zheng, Zhichao Lu 외

This report describes the details of our approach for the event dense-captioning task in ActivityNet Challenge 2021. We present a semantic-aware pretraining method for dense video captioning, which empowers the learned f…

Dense CaptioningDense Video CaptioningVideo Captioning

Scaling Dense Event-Stream Pretraining from Visual Foundation Models

2026-03-04 · Zhiwen Chen, Junhui Hou, Zhiyu Zhu, Jinjian Wu 외 arxiv

Learning versatile, fine-grained representations from irregular event streams is pivotal yet nontrivial, primarily due to the heavy annotation that hinders scalability in dataset size, semantic richness, and application …

Cross-Modal learning for Audio-Visual Video Parsing

2021-04-03 · Jatin Lamba, abhishek, Jayaprakash Akula, Rishabh Dabral 외

In this paper, we present a novel approach to the audio-visual video parsing (AVVP) task that demarcates events from a video separately for audio and visual modalities. The proposed parsing approach simultaneously detect…

Event DetectionMultiple Instance LearningVideo Grounding