paper-with-me

홈 › Papers

MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance

2025-09-10 · Kaikai Zhao, Zhaoxiang Liu, Peng Wang, Xin Wang, Zhicheng Ma, Yajun Xu, Wenjing Zhang, Yibing Nan, Kai Wang, Shiguo Lian arxiv

General-domain large multimodal models (LMMs) have achieved significant advances in various image-text tasks. However, their performance in the Intelligent Traffic Surveillance (ITS) domain remains limited due to the absence of dedicated multimodal datasets. To address this gap, we introduce MITS (Multimodal Intelligent Traffic Surveillance), the first large-scale multimodal benchmark dataset specifically designed for ITS. MITS includes 170,400 independently collected real-world ITS images sourced from traffic surveillance cameras, annotated with eight main categories and 24 subcategories of ITS-specific objects and events under diverse environmental conditions. Additionally, through a systematic data generation pipeline, we generate high-quality image captions and 5 million instruction-following visual question-answer pairs, addressing five critical ITS tasks: object and event recognition, object counting, object localization, background analysis, and event reasoning. To demonstrate MITS's effectiveness, we fine-tune mainstream LMMs on this dataset, enabling the development of ITS-specific applications. Experimental results show that MITS significantly improves LMM performance in ITS applications, increasing LLaVA-1.5's performance from 0.494 to 0.905 (+83.2%), LLaVA-1.6's from 0.678 to 0.921 (+35.8%), Qwen2-VL's from 0.584 to 0.926 (+58.6%), and Qwen2.5-VL's from 0.732 to 0.930 (+27.0%). We release the dataset, code, and models as open-source, providing high-value resources to advance both ITS and LMM research.

📄 PDF Abstract BibTeX arXiv:2509.09730

Code (0)

등록된 구현이 없습니다.

Tasks

Object LocalizationObject Counting

Similar Papers 제목 키워드 기반

SVIT: Scaling up Visual Instruction Tuning

2023-07-09 · Bo Zhao, Boya Wu, Muyang He, Tiejun Huang

Thanks to the emerging of foundation models, the large language and vision models are integrated to acquire the multimodal ability of visual captioning, question answering, etc. Although existing multimodal models presen…

DiversityImage CaptioningQuestion Answering

Multimodal Banking Dataset: Understanding Client Needs through Event Sequences

2024-09-26 · Mollaev Dzhambulat, Alexander Kostin, Postnova Maria, Ivan Karpukhin 외

Financial organizations collect a huge amount of data about clients that typically has a temporal (sequential) structure and is collected from various sources (modalities). Due to privacy issues, there are no large-scale…

Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning

2025-06-16 · Can Polat, Hasan Kurban, Erchin Serpedin, Mustafa Kurban

Evaluating foundation models for crystallographic reasoning requires benchmarks that isolate generalization behavior while enforcing physical constraints. This work introduces a multiscale multicrystal dataset with two p…

HallucinationSpatial Interpolation

mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus

2024-06-13 · Matthieu Futeral, Armel Zebaze, Pedro Ortiz Suarez, Julien Abadji 외

Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. [2022] showed that additionally training them on interleav…

Few-Shot LearningIn-Context Learning

DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning

2026-02-18 · Haoxiang Sun, Lizhen Xu, Bing Zhao, Wotao Yin 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has been shown effective in enhancing the visual reflection and reasoning capabilities of Large Multimodal Models (LMMs). However, existing datasets are predominantly…

Reinforcement LearningMultimodal Reasoning