paper-with-me

홈 › Papers

PiTe: Pixel-Temporal Alignment for Large Video-Language Model

2024-09-11 · Yang Liu, Pengxiang Ding, Siteng Huang, Min Zhang, Han Zhao, Donglin Wang

Fueled by the Large Language Models (LLMs) wave, Large Visual-Language Models (LVLMs) have emerged as a pivotal advancement, bridging the gap between image and text. However, video making it challenging for LVLMs to perform adequately due to the complexity of the relationship between language and spatial-temporal data structure. Recent Large Video-Language Models (LVidLMs) align feature of static visual data like image into latent space of language feature, by general multi-modal tasks to leverage abilities of LLMs sufficiently. In this paper, we explore fine-grained alignment approach via object trajectory for different modalities across both spatial and temporal dimensions simultaneously. Thus, we propose a novel LVidLM by trajectory-guided Pixel-Temporal Alignment, dubbed PiTe, that exhibits promising applicable model property. To achieve fine-grained video-language alignment, we curate a multi-modal pre-training dataset PiTe-143k, the dataset provision of moving trajectories in pixel level for all individual objects, that appear and mention in the video and caption both, by our automatic annotation pipeline. Meanwhile, PiTe demonstrates astounding capabilities on myriad video-related multi-modal tasks through beat the state-of-the-art methods by a large margin.

📄 PDF Abstract BibTeX arXiv:2409.07239

Code (1)

yliu-cs/pite 공식 구현 pytorch

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Alignment is All You Need: A Training-free Augmentation Strategy for Pose-guided Video Generation

2024-08-29 · Xiaoyu Jin, Zunnan Xu, Mingwen Ou, Wenming Yang

Character animation is a transformative field in computer graphics and vision, enabling dynamic and realistic video animations from static images. Despite advancements, maintaining appearance consistency in animations re…

AllVideo Generation

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

2024-11-07 · Shehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao 외

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with preci…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+3

VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

2025-01-01 · CVPR 2025 1 · Shehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao 외

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with p…

Large Language ModelVideo SegmentationVideo Semantic SegmentationVisual Grounding

Behavior Discovery and Alignment of Articulated Object Classes from Unstructured Video

2015-11-30 · Luca Del Pero, Susanna Ricco, Rahul Sukthankar, Vittorio Ferrari

We propose an automatic system for organizing the content of a collection of unstructured videos of an articulated object class (e.g. tiger, horse). By exploiting the recurring motion patterns of the class across videos,…

Retrieval

Learning Spatial and Spatio-Temporal Pixel Aggregations for Image and Video Denoising

2021-01-26 · Xiangyu Xu, Muchen Li, Wenxiu Sun, Ming-Hsuan Yang

Existing denoising methods typically restore clear results by aggregating pixels from the noisy input. Instead of relying on hand-crafted aggregation schemes, we propose to explicitly learn this process with deep neural …

DenoisingImage DenoisingVideo Denoising