paper-with-me

Papers

The VISIONE Video Search System: Exploiting Off-the-Shelf Text Search Engines for Large-Scale Video Retrieval

2020-08-06 · Giuseppe Amato, Paolo Bolettieri, Fabio Carrara, Franca Debole, Fabrizio Falchi, Claudio Gennaro, Lucia Vadicamo, Claudio Vairo

In this paper, we describe in details VISIONE, a video search system that allows users to search for videos using textual keywords, occurrence of objects and their spatial relationships, occurrence of colors and their spatial relationships, and image similarity. These modalities can be combined together to express complex queries and satisfy user needs. The peculiarity of our approach is that we encode all the information extracted from the keyframes, such as visual deep features, tags, color and object locations, using a convenient textual encoding indexed in a single text retrieval engine. This offers great flexibility when results corresponding to various parts of the query (visual, text and locations) have to be merged. In addition, we report an extensive analysis of the system retrieval performance, using the query logs generated during the Video Browser Showdown (VBS) 2019 competition. This allowed us to fine-tune the system by choosing the optimal parameters and strategies among the ones that we tested.

📄 PDF Abstract BibTeX arXiv:2008.02749

Code (0)

등록된 구현이 없습니다.

Tasks

RetrievalText RetrievalVideo Retrieval

Similar Papers 제목 키워드 기반

DeepCache: Principled Cache for Mobile Deep Vision

2017-12-01 · Mengwei Xu, Mengze Zhu, Yunxin Liu, Felix Xiaozhu Lin 외

We present DeepCache, a principled cache design for deep learning inference in continuous mobile vision. DeepCache benefits model execution efficiency by exploiting temporal locality in input video streams. It addresses …

Video Compression

Lightweight Attentional Feature Fusion: A New Baseline for Text-to-Video Retrieval

2021-12-03 · Fan Hu, Aozhu Chen, Ziyue Wang, Fangming Zhou 외

In this paper we revisit feature fusion, an old-fashioned topic, in the new context of text-to-video retrieval. Different from previous research that considers feature fusion only at one end, let it be video or text, we …

Ad-hoc video searchfeature selectionRetrievalText to Video Retrieval+1

Cognitive Semantic Communication Systems Driven by Knowledge Graph

2022-02-24 · Fuhui Zhou, Yihao Li, Xinyuan Zhang, Qihui Wu 외

Semantic communication is envisioned as a promising technique to break through the Shannon limit. However, the existing semantic communication frameworks do not involve inference and error correction, which limits the ac…

Data CompressionSemantic Communication

New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark

2024-03-28 · Nadège Alavoine, Gaëlle Laperriere, Christophe Servan, Sahar Ghannay 외

Intent classification and slot-filling are essential tasks of Spoken Language Understanding (SLU). In most SLUsystems, those tasks are realized by independent modules. For about fifteen years, models achieving both of th…

intent-classificationIntent ClassificationIntent Classification and Slot Fillingslot-filling+2

Zero-Shot Transfer of Haptics-Based Object Insertion Policies

2023-01-29 · Samarth Brahmbhatt, Ankur Deka, Andrew Spielberg, Matthias Müller

Humans naturally exploit haptic feedback during contact-rich tasks like loading a dishwasher or stocking a bookshelf. Current robotic systems focus on avoiding unexpected contact, often relying on strategically placed en…