paper-with-me

홈 › Papers

A Data-Driven Approach for Tag Refinement and Localization in Web Videos

2014-07-02 · Lamberto Ballan, Marco Bertini, Giuseppe Serra, Alberto del Bimbo

Tagging of visual content is becoming more and more widespread as web-based services and social networks have popularized tagging functionalities among their users. These user-generated tags are used to ease browsing and exploration of media collections, e.g. using tag clouds, or to retrieve multimedia content. However, not all media are equally tagged by users. Using the current systems is easy to tag a single photo, and even tagging a part of a photo, like a face, has become common in sites like Flickr and Facebook. On the other hand, tagging a video sequence is more complicated and time consuming, so that users just tag the overall content of a video. In this paper we present a method for automatic video annotation that increases the number of tags originally provided by users, and localizes them temporally, associating tags to keyframes. Our approach exploits collective knowledge embedded in user-generated tags and web sources, and visual similarity of keyframes and images uploaded to social sites like YouTube and Flickr, as well as web sources like Google and Bing. Given a keyframe, our method is able to select on the fly from these visual sources the training exemplars that should be the most relevant for this test sample, and proceeds to transfer labels across similar images. Compared to existing video tagging approaches that require training classifiers for each tag, our system has few parameters, is easy to implement and can deal with an open vocabulary scenario. We demonstrate the approach on tag refinement and localization on DUT-WEBV, a large dataset of web videos, and show state-of-the-art results.

📄 PDF Abstract BibTeX arXiv:1407.0623

Code (0)

등록된 구현이 없습니다.

Tasks

TAG

Similar Papers 제목 키워드 기반

Explainable Forensics of Manipulated Segments in Untrimmed Long Videos

2026-06-01 · Yue Feng, Jingjing Li, Qijia Lu, Wei Ji 외 arxiv

The rapid advancement of AI-driven video generation has transformed content creation, while simultaneously increasing the risk of misinformation through localized manipulations in long-form videos. Existing video forensi…

Video Generation

PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization

2025-02-11 · Bing Fan, Yunhe Feng, Yapeng Tian, Yuewei Lin 외

Egocentric visual query localization (EgoVQL) focuses on localizing the target of interest in space and time from first-person videos, given a visual query. Despite recent progressive, existing methods often struggle to …

FMI-TAL: Few-shot Multiple Instances Temporal Action Localization by Probability Distribution Learning and Interval Cluster Refinement

2024-08-25 · Fengshun Wang, Qiurui Wang, Yuting Wang

The present few-shot temporal action localization model can't handle the situation where videos contain multiple action instances. So the purpose of this paper is to achieve manifold action instances localization in a le…

Action LocalizationFew Shot Temporal Action LocalizationTemporal Action Localization

Precise Temporal Action Localization by Evolving Temporal Proposals

2018-04-13 · Haonan Qiu, Yingbin Zheng, Hao Ye, Yao Lu 외

Locating actions in long untrimmed videos has been a challenging problem in video content analysis. The performances of existing action localization approaches remain unsatisfactory in precisely determining the beginning…

Action LocalizationTemporal Action Localization

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

2026-06-30 · Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado López 외 arxiv

Vision foundation models such as SAM 3 can provide transferable object-level structure across diverse surgical video conditions, but segmentation outputs do not explicitly encode the action-conditioned semantics that def…