paper-with-me

홈 › Papers

Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenarios

2025-09-04 · Jingen Qu, Lijun Li, Bo Zhang, Yichen Yan, Jing Shao arxiv

Multimodal large language models (MLLMs) are rapidly evolving, presenting increasingly complex safety challenges. However, current dataset construction methods, which are risk-oriented, fail to cover the growing complexity of real-world multimodal safety scenarios (RMS). And due to the lack of a unified evaluation metric, their overall effectiveness remains unproven. This paper introduces a novel image-oriented self-adaptive dataset construction method for RMS, which starts with images and end constructing paired text and guidance responses. Using the image-oriented method, we automatically generate an RMS dataset comprising 35k image-text pairs with guidance responses. Additionally, we introduce a standardized safety dataset evaluation metric: fine-tuning a safety judge model and evaluating its capabilities on other safety datasets.Extensive experiments on various tasks demonstrate the effectiveness of the proposed image-oriented pipeline. The results confirm the scalability and effectiveness of the image-oriented approach, offering a new perspective for the construction of real-world multimodal safety datasets. The dataset is presented at https://huggingface.co/datasets/NewCityLetter/RMS2/tree/main.

📄 PDF Abstract BibTeX arXiv:2509.04403

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Not All Data are Good Labels: On the Self-supervised Labeling for Time Series Forecasting

2025-02-20 · Yuxuan Yang, Dalin Zhang, Yuxuan Liang, Hua Lu 외

Time Series Forecasting (TSF) is a crucial task in various domains, yet existing TSF models rely heavily on high-quality data and insufficiently exploit all available data. This paper explores a novel self-supervised app…

AllSelf-Supervised LearningTime SeriesTime Series Forecasting

Improving Neural Indoor Surface Reconstruction with Mask-Guided Adaptive Consistency Constraints

2023-09-18 · Xinyi Yu, Liqin Lu, Jintao Rong, Guangkai Xu 외

3D scene reconstruction from 2D images has been a long-standing task. Instead of estimating per-frame depth maps and fusing them in 3D, recent research leverages the neural implicit surface as a unified representation fo…

3D Reconstruction3D Scene ReconstructionSurface Reconstruction

Towards Real-World HDR Video Reconstruction: A Large-Scale Benchmark Dataset and A Two-Stage Alignment Network

2024-04-30 · CVPR 2024 1 · Yong Shu, Liquan Shen, Xiangyu Hu, Mengyao Li 외

As an important and practical way to obtain high dynamic range (HDR) video, HDR video reconstruction from sequences with alternating exposures is still less explored, mainly due to the lack of large-scale real-world data…

Video Reconstruction

Self-supervised Cloth Reconstruction via Action-conditioned Cloth Tracking

2023-02-19 · Zixuan Huang, Xingyu Lin, David Held

State estimation is one of the greatest challenges for cloth manipulation due to cloth's high dimensionality and self-occlusion. Prior works propose to identify the full state of crumpled clothes by training a mesh recon…

Self-Supervised LearningState Estimation

EVDI++: Event-based Video Deblurring and Interpolation via Self-Supervised Learning

2025-09-10 · Chi Zhang, Xiang Zhang, Chenxu Jiang, Gui-Song Xia 외 arxiv

Frame-based cameras with extended exposure times often produce perceptible visual blurring and information loss between frames, significantly degrading video quality. To address this challenge, we introduce EVDI++, a uni…

Self-Supervised Learning