paper-with-me

Papers

Automating Video Thumbnails Selection and Generation with Multimodal and Multistage Analysis

2024-10-18 · Elia Fantini

This thesis presents an innovative approach to automate video thumbnail selection for traditional broadcast content. Our methodology establishes stringent criteria for diverse, representative, and aesthetically pleasing thumbnails, considering factors like logo placement space, incorporation of vertical aspect ratios, and accurate recognition of facial identities and emotions. We introduce a sophisticated multistage pipeline that can select candidate frames or generate novel images by blending video elements or using diffusion models. The pipeline incorporates state-of-the-art models for various tasks, including downsampling, redundancy reduction, automated cropping, face recognition, closed-eye and emotion detection, shot scale and aesthetic prediction, segmentation, matting, and harmonization. It also leverages large language models and visual transformers for semantic consistency. A GUI tool facilitates rapid navigation of the pipeline's output. To evaluate our method, we conducted comprehensive experiments. In a study of 69 videos, 53.6% of our proposed sets included thumbnails chosen by professional designers, with 73.9% containing similar images. A survey of 82 participants showed a 45.77% preference for our method, compared to 37.99% for manually chosen thumbnails and 16.36% for an alternative method. Professional designers reported a 3.57-fold increase in valid candidates compared to the alternative method, confirming that our approach meets established criteria. In conclusion, our findings affirm that the proposed method accelerates thumbnail creation while maintaining high-quality standards and fostering greater user engagement.

📄 PDF Abstract BibTeX arXiv:2410.19825

Code (0)

등록된 구현이 없습니다.

Tasks

Face RecognitionImage Matting

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

From Thumbnails to Summaries - A single Deep Neural Network to Rule Them All

2018-08-01 · Hongxiang Gu, Viswanathan Swaminathan

Video summaries come in many forms, from traditional single-image thumbnails, animated thumbnails, storyboards, to trailer-like video summaries. Content creators use the summaries to display the most attractive portion o…

AllDecoderManagement

Sentence Specified Dynamic Video Thumbnail Generation

2019-08-12 · Yitian Yuan, Lin Ma, Wenwu Zhu

With the tremendous growth of videos over the Internet, video thumbnails, providing video content previews, are becoming increasingly crucial to influencing users' online searching experiences. Conventional video thumbna…

Sentence

A Benchmark Dataset for Micro-video Thumbnail Selection

2021-12-30 · Liu Bo

The thumbnail, as the first sight of a micro-video, plays a pivotal role in attracting users to click and watch. Although several pioneer efforts have been dedicated to jointly considering the quality and representativen…

FRC-GIF: Frame Ranking-based Personalized Artistic Media Generation Method for Resource Constrained Devices

2023-11-30 · journal 2023 11 · Ghulam Mujtaba, Sunder Ali Khowaja, Muhammad Aslam Jarwar, Jaehyuk Choi 외

Generating video highlights in the form of animated graphics interchange formats (GIFs) has significantly simplified the process of video browsing. Animated GIFs have paved the way for applications concerning streaming p…

Animated GIF GenerationClassification

Detecting Cultural Differences in News Video Thumbnails via Computational Aesthetics

2025-05-28 · Marvin Limpijankit, John Kender

We propose a two-step approach for detecting differences in the style of images across sources of differing cultural affinity, where images are first clustered into finer visual themes based on content before their aesth…