paper-with-me

홈 › Papers

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting

2025-07-22 · Jeongeun Lee, Youngjae Yu, Dongha Lee arxiv

The exponential growth of video content has made personalized video highlighting an essential task, as user preferences are highly variable and complex. Existing video datasets, however, often lack personalization, relying on isolated videos or simple text queries that fail to capture the intricacies of user behavior. In this work, we introduce HIPPO-Video, a novel dataset for personalized video highlighting, created using an LLM-based user simulator to generate realistic watch histories reflecting diverse user preferences. The dataset includes 2,040 (watch history, saliency score) pairs, covering 20,400 videos across 170 semantic categories. To validate our dataset, we propose HiPHer, a method that leverages these personalized watch histories to predict preference-conditioned segment-wise saliency scores. Through extensive experiments, we demonstrate that our method outperforms existing generic and query-based approaches, showcasing its potential for highly user-centric video highlighting in real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2507.16873

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

2026-06-24 · Baiqi Li, Ce Zhang, Yu Fang, Yue Yang 외 arxiv

A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet toda…

Robot Manipulation

A Large Language Model Enhanced Sequential Recommender for Joint Video and Comment Recommendation

2024-03-20 · Bowen Zheng, Zihan Lin, Enze Liu, Chen Yang 외

In online video platforms, reading or writing comments on interesting videos has become an essential part of the video watching experience. However, existing video recommender systems mainly model users' interaction beha…

Language ModelingLanguage ModellingLarge Language ModelRecommendation Systems+1

A Transformer-Based Model for the Prediction of Human Gaze Behavior on Videos

2024-04-10 · Suleyman Ozdel, Yao Rong, Berat Mert Albaba, Yen-Ling Kuo 외

Eye-tracking applications that utilize the human gaze in video understanding tasks have become increasingly important. To effectively automate the process of video analysis based on eye-tracking data, it is important to …

Activity RecognitionGaze PredictionVideo Understanding

ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis

2025-09-28 · Congzhi Zhang, Zhibin Wang, Yinchao Ma, Jiawei Peng 외 arxiv

While Reinforcement Learning with Verifiable Reward (RLVR) significantly advances image reasoning in Large Vision-Language Models (LVLMs), its application to complex video reasoning remains underdeveloped. This gap stems…

Reinforcement LearningInformation Retrieval

SWaT: Statistical Modeling of Video Watch Time through User Behavior Analysis

2024-08-14 · Shentao Yang, Haichuan Yang, Linna Du, Adithya Ganesh 외

The significance of estimating video watch time has been highlighted by the rising importance of (short) video recommendation, which has become a core product of mainstream social media platforms. Modeling video watch ti…

Binary Classification