paper-with-me

홈 › Papers

WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs

2026-02-25 · Yulin Zhang, Cheng Shi, Sibei Yang arxiv

Recent advances in Multimodal Large Language Models have greatly improved visual understanding and reasoning, yet their quadratic attention and offline training protocols make them ill-suited for streaming settings where frames arrive sequentially and future observations are inaccessible. We diagnose a core limitation of current Video-LLMs, namely Time-Agnosticism, in which videos are treated as an unordered bag of evidence rather than a causally ordered sequence, yielding two failures in streams: temporal order ambiguity, in which the model cannot follow or reason over the correct chronological order, and past-current focus blindness where it fails to distinguish present observations from accumulated history. We present WeaveTime, a simple, efficient, and model agnostic framework that first teaches order and then uses order. We introduce a lightweight Temporal Reconstruction objective-our Streaming Order Perception enhancement-that instills order aware representations with minimal finetuning and no specialized streaming data. At inference, a Past-Current Dynamic Focus Cache performs uncertainty triggered, coarse-to-fine retrieval, expanding history only when needed. Plugged into exsiting Video-LLM without architectural changes, WeaveTime delivers consistent gains on representative streaming benchmarks, improving accuracy while reducing latency. These results establish WeaveTime as a practical path toward time aware stream Video-LLMs under strict online, time causal constraints. Code and weights will be made publicly available. Project Page: https://zhangyl4.github.io/publications/weavetime/

📄 PDF Abstract BibTeX arXiv:2602.22142

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns

2026-06-23 · Vatsal Baherwani, Zixi Chen, Shikai Qiu, Andrew Gordon Wilson 외 arxiv

Neural scaling laws for transformer language models predict smooth improvements in pretraining loss with increasing parameters, but downstream capabilities such as in-context learning are known to emerge abruptly past a …

Monitoring Emergent Reward Hacking During Generation via Internal Activations

2026-03-04 · Patrick Wilhelm, Thorsten Wittkopp, Odej Kao arxiv

Fine-tuned large language models can exhibit reward-hacking behavior arising from emergent misalignment, which is difficult to detect from final outputs alone. While prior work has studied reward hacking at the level of …

InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter

2026-08-21 · Yunze Tong, Mushui Liu, Canyu Zhao, Shiyi Zhang 외 arxiv

With large pretrained models, existing methods have effectively improved instruction-based video editing. However, most of them rely on an in-place editing assumption. They align the edited video with the given source cl…

Detecting and Summarizing Emergent Events in Microblogs and Social Media Streams by Dynamic Centralities

2016-10-20 · Avudaiappan Neela, Herzog Alexander, Kadam Sneha, Du Yuheng 외

Methods for detecting and summarizing emergent keywords have been extensively studied since social media and microblogging activities have started to play an important role in data analysis and decision making. We presen…

Decision Making

XferBench: a Data-Driven Benchmark for Emergent Language

2024-07-03 · Brendon Boldt, David Mortensen

In this paper, we introduce a benchmark for evaluating the overall quality of emergent languages using data-driven methods. Specifically, we interpret the notion of the "quality" of an emergent language as its similarity…