paper-with-me

Papers

DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM

2024-10-03 · Xuchen Li, Shiyu Hu, Xiaokun Feng, Dailing Zhang, Meiqi Wu, Jing Zhang, Kaiqi Huang

Visual language tracking (VLT) has emerged as a cutting-edge research area, harnessing linguistic data to enhance algorithms with multi-modal inputs and broadening the scope of traditional single object tracking (SOT) to encompass video understanding applications. Despite this, most VLT benchmarks still depend on succinct, human-annotated text descriptions for each video. These descriptions often fall short in capturing the nuances of video content dynamics and lack stylistic variety in language, constrained by their uniform level of detail and a fixed annotation frequency. As a result, algorithms tend to default to a "memorize the answer" strategy, diverging from the core objective of achieving a deeper understanding of video content. Fortunately, the emergence of large language models (LLMs) has enabled the generation of diverse text. This work utilizes LLMs to generate varied semantic annotations (in terms of text lengths and granularities) for representative SOT benchmarks, thereby establishing a novel multi-modal benchmark. Specifically, we (1) propose a new visual language tracking benchmark with diverse texts, named DTVLT, based on five prominent VLT and SOT benchmarks, including three sub-tasks: short-term tracking, long-term tracking, and global instance tracking. (2) We offer four granularity texts in our benchmark, considering the extent and density of semantic information. We expect this multi-granular generation strategy to foster a favorable environment for VLT and video understanding research. (3) We conduct comprehensive experimental analyses on DTVLT, evaluating the impact of diverse text on tracking performance and hope the identified performance bottlenecks of existing algorithms can support further research in VLT and video understanding. The proposed benchmark, experimental results and toolkit will be released gradually on http://videocube.aitestunion.com/.

📄 PDF Abstract BibTeX arXiv:2410.02492

Code (0)

등록된 구현이 없습니다.

Tasks

Object TrackingVideo Understanding

Similar Papers 제목 키워드 기반

MCiteBench: A Multimodal Benchmark for Generating Text with Citations

2025-03-04 · Caiyu Hu, Yikai Zhang, Tinghui Zhu, Yiwei Ye 외

Multimodal Large Language Models (MLLMs) have advanced in integrating diverse modalities but frequently suffer from hallucination. A promising solution to mitigate this issue is to generate text with citations, providing…

HallucinationText Generation

Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data

2024-12-19 · Zhiqiang Tang, Zihan Zhong, Tong He, Gerald Friedland

This paper studies the best practices for automatic machine learning (AutoML). While previous AutoML efforts have predominantly focused on unimodal data, the multimodal aspect remains under-explored. Our study delves int…

AutoMLcross-modal alignmentData Augmentation

Benchmarking Diverse-Modal Entity Linking with Generative Models

2023-05-27 · Sijia Wang, Alexander Hanbo Li, Henry Zhu, Sheng Zhang 외

Entities can be expressed in diverse formats, such as texts, images, or column names and cell values in tables. While existing entity linking (EL) models work well on per modality configuration, such as text-only EL, vis…

BenchmarkingDecoderEntity LinkingVisual Grounding

VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark

2025-08-19 · Vy Tuong Dang, An Vo, Emilio Villa-Cueva, Quang Tau 외 arxiv

We introduce VMMU, a Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark designed to evaluate how vision-language models (VLMs) interpret and reason over visual and textual information beyond English. V…

Visual Reasoning

M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks

2024-07-04 · Florian Schneider, Sunayana Sitaram

Since the release of ChatGPT, the field of Natural Language Processing has experienced rapid advancements, particularly in Large Language Models (LLMs) and their multimodal counterparts, Large Multimodal Models (LMMs). D…

Outlier Detection