paper-with-me

홈 › Papers

STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding

2026-03-29 · Junho Kim, Hosu Lee, James M. Rehg, Minsu Kim, Yong Man Ro arxiv

Recent progress in video large language models (Video-LLMs) has enabled strong offline reasoning over long and complex videos. However, real-world deployments increasingly require streaming perception and proactive interaction, where video frames arrive online and the system must decide not only what to respond, but also when to respond. In this work, we revisit proactive activation in streaming video as a structured sequence modeling problem, motivated by the observation that temporal transitions in streaming video naturally form span-structured activation patterns. To capture this span-level structure, we model activation signals jointly over a sliding temporal window and update them iteratively as new frames arrive. We propose STRIDE (Structured Temporal Refinement with Iterative DEnoising), which employs a lightweight masked diffusion module at the activation interface to jointly predict and progressively refine activation signals across the window. Extensive experiments on diverse streaming benchmarks and downstream models demonstrate that STRIDE shows more reliable and temporally coherent proactive responses, significantly improving when-to-speak decision quality in online streaming scenarios.

📄 PDF Abstract BibTeX arXiv:2603.27593

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Efficient Coarse-to-Fine Diffusion Models with Time Step Sequence Redistribution

2026-03-22 · Yu-Shan Tai, An-Yeu, Wu arxiv

Recently, diffusion models (DMs) have made significant strides in high-quality image generation. However, the multi-step denoising process often results in considerable computational overhead, impeding deployment on reso…

Image Generation

Golden Gemini is All You Need: Finding the Sweet Spots for Speaker Verification

2023-12-06 · Tianchi Liu, Kong Aik Lee, Qiongqiong Wang, Haizhou Li

Previous studies demonstrate the impressive performance of residual neural networks (ResNet) in speaker verification. The ResNet models treat the time and frequency dimensions equally. They follow the default stride conf…

AllSpeaker Verification

Gait Recognition under Speed Transition

2014-06-01 · CVPR 2014 6 · Al Mansur, Yasushi Makihara, Rasyid Aqmar, Yasushi Yagi

This paper describes a method of gait recognition from image sequences wherein a subject is accelerating or decelerating. As a speed change occurs due to a change of pitch (the first-order derivative of a phase, namely, …

Gait Recognition

When Image Denoising Meets High-Level Vision Tasks: A Deep Learning Approach

2017-06-14 · Ding Liu, Bihan Wen, Xianming Liu, Zhangyang Wang 외

Conventionally, image denoising and high-level vision tasks are handled separately in computer vision. In this paper, we cope with the two jointly and explore the mutual influence between them. First we propose a convolu…

DenoisingImage Denoising

HumanOmni-Speaker: Identifying Who said What and When

2026-03-23 · Detao Bai, Zhiheng Ma, Xihan Wei arxiv

While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, multi-person conversational dynamics to accu…

Natural Language QueriesSpeaker Diarization