paper-with-me

홈 › Papers

EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions

2025-08-31 · Dinh-Khoi Vo, Van-Loc Nguyen, Minh-Triet Tran, Trung-Nghia Le arxiv

Event-based image retrieval from free-form captions presents a significant challenge: models must understand not only visual features but also latent event semantics, context, and real-world knowledge. Conventional vision-language retrieval approaches often fall short when captions describe abstract events, implicit causality, temporal context, or contain long, complex narratives. To tackle these issues, we introduce a multi-stage retrieval framework combining dense article retrieval, event-aware language model reranking, and efficient image collection, followed by caption-guided semantic matching and rank-aware selection. We leverage Qwen3 for article search, Qwen3-Reranker for contextual alignment, and Qwen2-VL for precise image scoring. To further enhance performance and robustness, we fuse outputs from multiple configurations using Reciprocal Rank Fusion (RRF). Our system achieves the top-1 score on the private test set of Track 2 in the EVENTA 2025 Grand Challenge, demonstrating the effectiveness of combining language-based reasoning and multimodal retrieval for complex, real-world image understanding. The code is available at https://github.com/vdkhoi20/EVENT-Retriever.

📄 PDF Abstract BibTeX arXiv:2509.00751

Code (0)

등록된 구현이 없습니다.

Tasks

Image Retrieval

Similar Papers 제목 키워드 기반

Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation

2024-07-21 · Liwen Sun, James Zhao, Megan Han, Chenyan Xiong

Multimodal foundation models hold significant potential for automating radiology report generation, thereby assisting clinicians in diagnosing cardiac diseases. However, generated reports often suffer from serious factua…

DiagnosticRAGRetrievalText Generation

Generating Event-oriented Attribution for Movies via Two-Stage Prefix-Enhanced Multimodal LLM

2024-09-14 · Yuanjie Lyu, Tong Xu, Zihan Niu, Bo Peng 외

The prosperity of social media platforms has raised the urgent demand for semantic-rich services, e.g., event and storyline attribution. However, most existing research focuses on clip-level event understanding, primaril…

MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

2024-11-04 · Sheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin 외

State-of-the-art retrieval models typically address a straightforward search scenario, in which retrieval tasks are fixed (e.g., finding a passage to answer a specific question) and only a single modality is supported fo…

Cross-Modal RetrievalInformation RetrievalRerankingRetrieval+1

ChangeBridge: Spatiotemporal Image Generation with Multimodal Controls for Remote Sensing

2025-07-07 · Zhenghui Zhao, Chen Wu, Xiangyong Cao, Di Wang 외 arxiv

Spatiotemporal image generation is a highly meaningful task, which can generate future scenes conditioned on given observations. However, existing change generation methods can only handle event-driven changes (e.g., new…

Change DetectionImage Generation

Event-Enriched Image Analysis Grand Challenge at ACM Multimedia 2025

2025-08-26 · Thien-Phuc Tran, Minh-Quang Nguyen, Minh-Triet Tran, Tam V. Nguyen 외 arxiv

The Event-Enriched Image Analysis (EVENTA) Grand Challenge, hosted at ACM Multimedia 2025, introduces the first large-scale benchmark for event-level multimodal understanding. Traditional captioning and retrieval tasks l…

Image Retrieval