paper-with-me

홈 › Papers

Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features

2025-10-19 · Shihao Ji, Zihui Song arxiv

The remarkable zero-shot reasoning capabilities of large-scale Visual Language Models (VLMs) on static images have yet to be fully translated to the video domain. Conventional video understanding models often rely on extensive, task-specific training on annotated datasets, a process that is both costly and limited in scalability. This paper introduces a novel, training-free framework for video understanding that circumvents end-to-end training by synergistically combining the rich semantic priors of pre-trained VLMs with classic machine learning algorithms for pattern discovery. Our core idea is to reframe video understanding as a self-supervised spatio-temporal clustering problem within a high-dimensional semantic feature space. The proposed pipeline first transforms a video stream into a semantic feature trajectory using the frozen visual encoder of a pre-trained VLM. Subsequently, we employ Kernel Temporal Segmentation (KTS), a robust machine learning technique, to partition the continuous feature stream into discrete, semantically coherent event segments. These segments are then subjected to unsupervised density-based clustering to identify recurring macroscopic scenes and themes throughout the video. By selecting representative keyframes from each discovered cluster and leveraging the VLM's generative capabilities for textual description, our framework automatically produces a structured, multi-modal summary of the video content. This approach provides an effective, interpretable, and model-agnostic pathway for zero-shot, automated structural analysis of video content.

📄 PDF Abstract BibTeX arXiv:2510.16781

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Xiaoicesing 2: A High-Fidelity Singing Voice Synthesizer Based on Generative Adversarial Network

2022-10-26 · Interspeech 2023 8 · Chunhui Wang, Chang Zeng, Xing He

XiaoiceSing is a singing voice synthesis (SVS) system that aims at generating 48kHz singing voices. However, the mel-spectrogram generated by it is over-smoothing in middle- and high-frequency areas due to no special des…

Generative Adversarial NetworkSinging Voice Synthesis

The Design and Implementation of XiaoIce, an Empathetic Social Chatbot

2018-12-21 · CL 2020 3 · Li Zhou, Jianfeng Gao, Di Li, Heung-Yeung Shum

This paper describes the development of Microsoft XiaoIce, the most popular social chatbot in the world. XiaoIce is uniquely designed as an AI companion with an emotional connection to satisfy the human need for communic…

ChatbotDecision Making

XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System

2020-06-11 · Peiling Lu, Jie Wu, Jian Luan, Xu Tan 외

This paper presents XiaoiceSing, a high-quality singing voice synthesis system which employs an integrated network for spectrum, F0 and duration modeling. We follow the main architecture of FastSpeech while proposing som…

RhythmSinging Voice SynthesisVocal Bursts Intensity Prediction

From Eliza to XiaoIce: Challenges and Opportunities with Social Chatbots

2018-01-06 · Heung-Yeung Shum, Xiaodong He, Di Li

Conversational systems have come a long way since their inception in the 1960s. After decades of research and development, we've seen progress from Eliza and Parry in the 60's and 70's, to task-completion systems as in t…

Chatbot

SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding

2025-12-04 · Chang-Hsun Wu, Kai-Po Chang, Yu-Yang Sheng, Hung-Kai Chung 외 arxiv

Video Large Language Models (VideoLLMs) have shown remarkable progress in video understanding. However, these models still struggle to effectively perceive and exploit rich temporal information in videos when responding …