paper-with-me

Papers

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata

2026-02-13 · Prasanna Sridhar, Horace Lee, David M. S. Pinto, Andrew Zisserman, Abhishek Dutta arxiv

In this paper, we present WISE, an open-source audiovisual search engine which integrates a range of multimodal retrieval capabilities into a single, practical tool accessible to users without machine learning expertise. WISE supports natural-language and reverse-image queries at both the scene level (e.g. empty street) and object level (e.g. horse) across images and videos; face-based search for specific individuals; audio retrieval of acoustic events using text (e.g. wood creak) or an audio file; search over automatically transcribed speech; and filtering by user-provided metadata. Rich insights can be obtained by combining queries across modalities -- for example, retrieving German trains from a historical archive by applying the object query "train" and the metadata query "Germany", or searching for a face in a place. By employing vector search techniques, WISE can scale to support efficient retrieval over millions of images or thousands of hours of video. Its modular architecture facilitates the integration of new models. WISE can be deployed locally for private or sensitive collections, and has been applied to various real-world use cases. Our code is open-source and available at https://gitlab.com/vgg/wise/wise.

📄 PDF Abstract BibTeX arXiv:2602.12819

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models

2026-01-29 · Wenxuan Huang, Yu Zeng, Qiuchen Wang, Zhen Fang 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable success across a broad range of vision tasks. However, constrained by the capacity of their internal world knowledge, prior work has proposed augmenting M…

Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes

2024-07-15 · Yaoting Wang, Peiwen Sun, Dongzhan Zhou, Guangyao Li 외

Traditional reference segmentation tasks have predominantly focused on silent visual scenes, neglecting the integral role of multimodal perception and interaction in human experiences. In this work, we introduce a novel …

Segmentation

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes

2025-06-02 · CVPR 2025 1 · Yuji Wang, Haoran Xu, Yong liu, Jiaze Li 외

Reference Audio-Visual Segmentation (Ref-AVS) aims to provide a pixel-wise scene understanding in Language-aided Audio-Visual Scenes (LAVS). This task requires the model to continuously segment objects referred to by tex…

Scene Understanding

4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models

2025-03-13 · CVPR 2025 1 · Wanhua Li, Renping Zhou, Jiawei Zhou, Yingwei Song 외

Learning 4D language fields to enable time-sensitive, open-ended language queries in dynamic scenes is essential for many real-world applications. While LangSplat successfully grounds CLIP features into 3D Gaussian repre…

Large Language ModelObjectSentence Embeddings

What Looks Good with my Sofa: Multimodal Search Engine for Interior Design

2017-07-21 · Ivona Tautkute, Aleksandra Możejko, Wojciech Stokowiec, Tomasz Trzciński 외

In this paper, we propose a multi-modal search engine for interior design that combines visual and textual queries. The goal of our engine is to retrieve interior objects, e.g. furniture or wall clocks, that share visual…