paper-with-me

홈 › Papers

Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition

2026-02-09 · Yuhao Dong, Shulin Tian, Shuai Liu, Shuangrui Ding, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Jiaqi Wang, Ziwei Liu arxiv

Despite the growing video understanding capabilities of recent Multimodal Large Language Models (MLLMs), existing video benchmarks primarily assess understanding based on models' static, internal knowledge, rather than their ability to learn and adapt from dynamic, novel contexts from few examples. To bridge this gap, we present Demo-driven Video In-Context Learning, a novel task focused on learning from in-context demonstrations to answer questions about the target videos. Alongside this, we propose Demo-ICL-Bench, a challenging benchmark designed to evaluate demo-driven video in-context learning capabilities. Demo-ICL-Bench is constructed from 1200 instructional YouTube videos with associated questions, from which two types of demonstrations are derived: (i) summarizing video subtitles for text demonstration; and (ii) corresponding instructional videos as video demonstrations. To effectively tackle this new challenge, we develop Demo-ICL, an MLLM with a two-stage training strategy: video-supervised fine-tuning and information-assisted direct preference optimization, jointly enhancing the model's ability to learn from in-context examples. Extensive experiments with state-of-the-art MLLMs confirm the difficulty of Demo-ICL-Bench, demonstrate the effectiveness of Demo-ICL, and thereby unveil future research directions.

📄 PDF Abstract BibTeX arXiv:2602.08439

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction

2026-03-12 · Shuzhen Bi, Mengsong Wu, Hao Hao, Keqian Li 외 arxiv

The transition from monolithic large language models (LLMs) to modular, skill-equipped agents represents a fundamental architectural shift in artificial intelligence deployment. While general-purpose models demonstrate r…

Procedure-Aware Pretraining for Instructional Video Understanding

2023-03-31 · CVPR 2023 1 · Honglu Zhou, Roberto Martín-Martín, Mubbasir Kapadia, Silvio Savarese 외

Our goal is to learn a video representation that is useful for downstream procedure understanding tasks in instructional videos. Due to the small amount of available annotations, a key challenge in procedure understandin…

Video Understanding

ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding

2026-04-13 · Xucheng Wang, Xiaoman Zhang, Sung Eun Kim, Ankit Pal 외 arxiv

Ultrasound acquisition requires skilled probe manipulation and real-time adjustments. Vision-language models (VLMs) could enable autonomous ultrasound systems, but existing benchmarks evaluate only static images, not dyn…

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

2026-05-22 · Yingtie Lei, Zhongwei Wan, Jiankun Zhang, Samiul Alam 외 arxiv

Large language model (LLM) agents accumulate rich episodic trajectories while solving real-world tasks, but it remains unclear whether such experience can be distilled into reusable procedural skills. We introduce SkillE…

Procedural Pretraining: Warming Up Language Models with Abstract Data

2026-01-29 · Liangze Jiang, Zachary Shinnick, Anton van den Hengel, Hemanth Saratchandran 외 arxiv

Pretraining language models directly on web-scale corpora is the de facto paradigm. We study an alternative where the model is initially exposed to abstract structured data to ease the subsequent acquisition of rich sema…