paper-with-me

Papers

To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection

2026-06-04 · Erfan Loweimi, Mengjie Qian, Kate Knill, Guanfeng Wu, Chi-Ho Chan, Abbas Haider, Muhammad Awan, Josef Kittler, Hui Wang, Mark Gales arxiv

When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curated benchmarks, a target may be heard but unseen, seen but unheard, or both. Fusing scores from an absent modality injects noise, degrading precision below the best unimodal system. We propose a query-adaptive framework that detects active modalities via cross-modal score consistency: when both modalities are active, files retrieved by one also score highly on the other; this agreement breaks down when a modality is absent. Classifiers driven by these cross-modal features achieve 89% detection accuracy. On the BBC Rewind corpus (with over 12,000 broadcast videos) the adaptive system attains 94.2% P@1, outperforming speaker-only (82.9%), face-only (93.4%), and fixed fusion (90.0%), recovering 64% of the gap to an oracle with ground-truth modality labels (96.6%).

📄 PDF Abstract BibTeX arXiv:2606.05931

Code (0)

등록된 구현이 없습니다.

Tasks

Person Retrieval

Similar Papers 제목 키워드 기반

LightMem-Ego: Your AI Memory for Everyday Life

2026-07-13 · Yijun Chen, Boyi Xiao, Yixian Zhao, Haoting Xia 외 hf

Personal AI assistants on mobile and wearable devices continuously perceive users' daily lives through visual and audio streams. However, answering queries about past experiences requires lightweight multimodal memory th…

CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing

2024-01-22 · Xianghu Yue, Xiaohai Tian, Lu Lu, Malu Zhang 외

There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represe…

AudioCapsAudio-Visual SynchronizationLanguage ModelingLanguage Modelling+3

Deep Impression: Audiovisual Deep Residual Networks for Multimodal Apparent Personality Trait Recognition

2016-09-16 · Yağmur Güçlütürk, Umut Güçlü, Marcel A. J. van Gerven, Rob Van Lier

Here, we develop an audiovisual deep residual network for multimodal apparent personality trait recognition. The network is trained end-to-end for predicting the Big Five personality traits of people from their videos. T…

Facial Expression RecognitionFeature EngineeringPersonality Trait Recognition

Personal Visual Context Learning in Large Multimodal Models

2026-05-11 · Zihui Xue, Ami Baid, Sangho Kim, Mi Luo 외 arxiv

As wearable devices like smart glasses integrate Large Multimodal Models (LMMs) into the continuous first-person visual streams of individual users, the evolution of these models into true personal assistants hinges on v…

Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models

2026-07-26 · Yiming Zhong, Chang Nie, Caifeng Shan arxiv

Omnimodal large language models (OmniLLMs) are rapidly extending multimodal reasoning to cover synchronized audio and video. However, the resulting audio-video token sequences are long, leading to high prefill latency an…

Multimodal Reasoning