OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog
We present the Object Language Video Transformer (OLViT) - a novel model for video dialog operating over a multi-modal attention-based dialog state tracker. Existing video dialog models struggle with questions requiring both spatial and temporal localization within videos, long-term temporal reasoning, and accurate object tracking across multiple dialog turns. OLViT addresses these challenges by maintaining a global dialog state based on the output of an Object State Tracker (OST) and a Language State Tracker (LST): while the OST attends to the most important objects within the video, the LST keeps track of the most important linguistic co-references to previous dialog turns. In stark contrast to previous works, our approach is generic by nature and is therefore capable of learning continuous multi-modal dialog state representations of the most relevant objects and rounds. As a result, they can be seamlessly integrated into Large Language Models (LLMs) and offer high flexibility in dealing with different datasets and tasks. Evaluations on the challenging DVD (response classification) and SIMMC 2.1 (response generation) datasets show that OLViT achieves new state-of-the-art performance across both datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectObject TrackingResponse GenerationTemporal LocalizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution
Large language models (LLMs) still struggle with the rigorous reasoning demands of hard competitive programming. While recent multi-agent frameworks attempt to bridge this reliability gap, they remain fundamentally state…
Reinforcement LearningProgram SynthesisLightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms
Currently, lightweight hybrid backbone networks have partially alleviated the issue of computational saturation, but the imbalance in computational efficiencys between convolutional neural networks (CNNs) and attention m…
UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking
Multi-modal object tracking has attracted considerable attention by integrating multiple complementary inputs (e.g., thermal, depth, and event data) to achieve outstanding performance. Although current general-purpose mu…
Object TrackingCross-modulated Attention Transformer for RGBT Tracking
Existing Transformer-based RGBT trackers achieve remarkable performance benefits by leveraging self-attention to extract uni-modal features and cross-attention to enhance multi-modal feature interaction and template-sear…
Rgb-T Tracking3D Multi-Object Tracking Using Graph Neural Networks with Cross-Edge Modality Attention
Online 3D multi-object tracking (MOT) has witnessed significant research interest in recent years, largely driven by demand from the autonomous systems community. However, 3D offline MOT is relatively less explored. Labe…
3D Multi-Object TrackingGraph Neural NetworkMulti-Object TrackingObject Tracking