paper-with-me

Papers Video Understanding

“Video Understanding” 태그가 달린 논문 1,149편 · 필터 해제

VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding

2025-07-17 · Shihao Wang, Guo Chen, De-An Huang, Zhiqi Li 외

Recent studies have revealed that selecting informative and relevant video frames can significantly improve the performance of Video Large Language Models (Video-LLMs). Current methods, such as reducing inter-frame redun…

Video GroundingVideo Understanding

UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks

2025-07-15 · Peiran Wu, Yunze Liu, Zhengdong Zhu, Enmin Zhou 외

Real-world user-generated videos, especially on platforms like TikTok, often feature rich and intertwined audio visual content. However, existing video captioning benchmarks and models remain predominantly visual centric…

Video CaptioningVideo Understanding

EmbRACE-3K: Embodied Reasoning and Action in Complex Environments

2025-07-14 · Mingxian Lin, Wei Huang, Yitang Li, Chengjie Jiang 외

Recent advanced vision-language models(VLMs) have demonstrated strong performance on passive, offline image and video understanding tasks. However, their effectiveness in embodied settings, which require online interacti…

Scene UnderstandingSpatial ReasoningVideo Understanding

Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI

2025-07-14 · Jiangkai Wu, Zhiyuan Ren, LiMing Liu, Xinggong Zhang

AI Video Chat emerges as a new paradigm for Real-time Communication (RTC), where one peer is not a human, but a Multimodal Large Language Model (MLLM). This makes interaction between humans and AI more intuitive, as if c…

Large Language ModelMultimodal Large Language ModelVideo Understanding

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation

2025-07-08 · Quanzhu Niu, Yikang Zhou, Shihao Chen, Tao Zhang 외

Video Instance Segmentation (VIS) fundamentally struggles with pervasive challenges including object occlusions, motion blur, and appearance variations during temporal association. To overcome these limitations, this wor…

Depth EstimationDepth PredictionInstance SegmentationMonocular Depth Estimation+4

Omni-Video: Democratizing Unified Video Understanding and Generation

2025-07-08 · Zhiyu Tan, Hao Yang, Luozheng Qin, Jia Gong 외

Notable breakthroughs in unified understanding and generation modeling have led to remarkable advancements in image understanding, reasoning, production and editing, yet current foundational models predominantly focus on…

Video GenerationVideo Understanding

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding

2025-07-08 · Tongtong Cheng, Rongzhen Li, Yixin Xiong, Tao Zhang 외

Accurate driving behavior recognition and reasoning are critical for autonomous driving video understanding. However, existing methods often tend to dig out the shallow causal, fail to address spurious correlations acros…

Autonomous DrivingVideo Understanding

Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models

2025-07-08 · L'ea Dubois, Klaus Schmidt, Chengyu Wang, Ji-Hoon Park 외

Current video understanding models excel at recognizing "what" is happening but fall short in high-level cognitive tasks like causal reasoning and future prediction, a limitation rooted in their lack of commonsense world…

Future predictionLarge Language ModelVideo UnderstandingWorld Knowledge+1

Kwai Keye-VL Technical Report

2025-07-02 · Kwai Keye Team, Biao Yang, Bin Wen, Changyi Liu 외

While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form videos, a dominant medium in today's digit…

Instruction FollowingReinforcement Learning (RL)Video Understanding

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges

2025-07-02 · Sanjeda Akter, Ibne Farabi Shihab, Anuj Sharma

Crash detection from video feeds is a critical problem in intelligent transportation systems. Recent developments in large language models (LLMs) and vision-language models (VLMs) have transformed how we process, reason …

Video Understanding

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs

2025-07-01 · Jiaming Zhang, Rui Hu, Qing Guo, Wei Yang Bryan Lim

Video Multimodal Large Language Models (V-MLLMs) have shown impressive capabilities in temporal reasoning and cross-modal understanding, yet their vulnerability to adversarial attacks remains underexplored due to unique …

Text GenerationVideo Understanding

GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

2025-07-01 · GLM-V Team, :, Wenyi Hong, Wenmeng Yu 외

We present GLM-4.1V-Thinking, a vision-language model (VLM) designed to advance general-purpose multimodal understanding and reasoning. In this report, we share our key findings in the development of the reasoning-centri…

document understandingMultimodal ReasoningVideo Understanding

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

2025-06-30 · Haoji Zhang, Yiqin Wang, Yansong Tang, Yong liu 외

Benefiting from the advances in large language models and cross-modal alignment, existing multimodal large language models have achieved prominent performance in image and short video understanding. However, the understa…

cross-modal alignmentEgoSchemaMMEMVBench+2

ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment

2025-06-28 · Amir Aghdam, Vincent Tao Hu

We address the task of zero-shot fine-grained video classification, where no video examples or temporal annotations are available for unseen action classes. While contrastive vision-language models such as SigLIP demonst…

Dynamic Time WarpingLarge Language ModelOpen Set Learningtext similarity+2

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

2025-06-27 · Boyuan Sun, Jiaxing Zhao, Xihan Wei, Qibin Hou

In this paper, we present LLaVA-Scissor, a training-free token compression strategy designed for video multimodal large language models. Previous methods mostly attempt to compress tokens based on attention scores, but f…

Question AnsweringVideo Question AnsweringVideo Understanding

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

2025-06-27 · Shaojie Zhang, Jiahui Yang, Jianqin Yin, Zhenbo Luo 외

Multimodal Large Language Models (MLLMs) have demonstrated significant success in visual understanding tasks. However, challenges persist in adapting these models for video comprehension due to the large volume of data a…

MMEVideo MMEVideo Understanding

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes

2025-06-26 · Yujia Liang, Jile Jiao, Zhicheng Wang, Xuetao Feng 외

Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with varying camera angles or scene changes. This…

AttributeQuestion AnsweringVideo Understanding

Task-Aware KV Compression For Cost-Effective Long Video Understanding

2025-06-26 · Minghao Qin, Yan Shu, Peitian Zhang, Kun Lun 외

Long-video understanding (LVU) remains a severe challenge for existing multimodal large language models (MLLMs), primarily due to the prohibitive computational cost. Recent approaches have explored KV compression to miti…

Video Understanding

PEVLM: Parallel Encoding for Vision-Language Models

2025-06-24 · Letian Kang, Shixian Luo, Yiqiang Li, Xiaoyang Yu 외

Vision-Language Models (VLMs) have demonstrated strong performance in video-language tasks, yet their application to long video understanding remains constrained by the quadratic complexity of standard attention mechanis…

Autonomous DrivingVideo Understanding

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning

2025-06-19 · Yi Chen, Yuying Ge, Rui Wang, Yixiao Ge 외

Recent reinforcement learning approaches, such as outcome-supervised GRPO, have advanced Chain-of-Thought reasoning in large language models (LLMs), yet their adaptation to multimodal LLMs (MLLMs) is unexplored. To addre…

Multimodal Reasoningreinforcement-learningReinforcement LearningVideo Understanding
1–20 / 1,149 다음 →