paper-with-me

Papers

An Active Learning Based Approach For Effective Video Annotation And Retrieval

2015-04-27 · Moitreya Chatterjee, Anton Leuski

Conventional multimedia annotation/retrieval systems such as Normalized Continuous Relevance Model (NormCRM) [16] require a fully labeled training data for a good performance. Active Learning, by determining an order for labeling the training data, allows for a good performance even before the training data is fully annotated. In this work we propose an active learning algorithm, which combines a novel measure of sample uncertainty with a novel clustering-based approach for determining sample density and diversity and integrate it with NormCRM. The clusters are also iteratively refined to ensure both feature and label-level agreement among samples. We show that our approach outperforms multiple baselines both on a recent, open character animation dataset and on the popular TRECVID corpus at both the tasks of annotation and text-based retrieval of videos.

📄 PDF Abstract BibTeX arXiv:1504.07004

Code (0)

등록된 구현이 없습니다.

Tasks

Active LearningClusteringDiversityRetrieval

Similar Papers 제목 키워드 기반

Are Binary Annotations Sufficient? Video Moment Retrieval via Hierarchical Uncertainty-Based Active Learning

2023-01-01 · CVPR 2023 1 · Wei Ji, Renjie Liang, Zhedong Zheng, Wenqiao Zhang 외

Recent research on video moment retrieval has mostly focused on enhancing the performance of accuracy, efficiency, and robustness, all of which largely rely on the abundance of high-quality annotations. While the pre…

Active LearningMoment RetrievalRetrieval

Interactive Multi-Turn Retrieval for Health Videos

2026-05-02 · Chengzheng Wu, Ke Qiu, Baoming Zhang, Ruiyu Mao 외 arxiv

The growing availability of health-related instructional videos creates new opportunities for clinical training, patient rehabilitation, and health education, yet existing retrieval systems remain largely single-turn: a …

Semantic RetrievalVideo Retrieval

Simple Baselines for Interactive Video Retrieval with Questions and Answers

2023-08-21 · ICCV 2023 1 · Kaiqu Liang, Samuel Albanie

To date, the majority of video retrieval systems have been optimized for a "single-shot" scenario in which the user submits a query in isolation, ignoring previous interactions with the system. Recently, there has been r…

Question AnsweringRetrievalVideo Retrieval

LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding

2026-01-21 · Xiaodong Wang, Langling Huang, Zhirong Wu, Xu Zhao 외 arxiv

The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interactive videos, such as movies and recordings…

VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool

2024-07-07 · Yan Wang, Yawen Zeng, Jingsheng Zheng, Xiaofen Xing 외

Multimodal large language models (MLLMs) are flourishing, but mainly focus on images with less attention than videos, especially in sub-fields such as prompt engineering, video chain-of-thought (CoT), and instruction tun…

Active LearningHallucinationPrompt Engineering