paper-with-me

Papers

Data Factory with Minimal Human Effort Using VLMs

2025-10-07 · Jiaojiao Ye, Jiaxing Zhong, Qian Xie, Yuzhou Zhou, Niki Trigoni, Andrew Markham arxiv

Generating enough and diverse data through augmentation offers an efficient solution to the time-consuming and labour-intensive process of collecting and annotating pixel-wise images. Traditional data augmentation techniques often face challenges in manipulating high-level semantic attributes, such as materials and textures. In contrast, diffusion models offer a robust alternative, by effectively utilizing text-to-image or image-to-image transformation. However, existing diffusion-based methods are either computationally expensive or compromise on performance. To address this issue, we introduce a novel training-free pipeline that integrates pretrained ControlNet and Vision-Language Models (VLMs) to generate synthetic images paired with pixel-level labels. This approach eliminates the need for manual annotations and significantly improves downstream tasks. To improve the fidelity and diversity, we add a Multi-way Prompt Generator, Mask Generator and High-quality Image Selection module. Our results on PASCAL-5i and COCO-20i present promising performance and outperform concurrent work for one-shot semantic segmentation.

📄 PDF Abstract BibTeX arXiv:2510.05722

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationData Augmentation

Similar Papers 제목 키워드 기반

DatasetGAN: Efficient Labeled Data Factory with Minimal Human Effort

2021-04-13 · CVPR 2021 1 · Yuxuan Zhang, Huan Ling, Jun Gao, Kangxue Yin 외

We introduce DatasetGAN: an automatic procedure to generate massive datasets of high-quality semantically segmented images requiring minimal human effort. Current deep networks are extremely data-hungry, benefiting from …

DecoderImage SegmentationSemantic Segmentation

Network Space Search for Pareto-Efficient Spaces

2021-04-22 · Min-Fong Hong, Hao-Yun Chen, Min-Hung Chen, Yu-Syuan Xu 외

Network spaces have been known as a critical factor in both handcrafted network designs or defining search spaces for Neural Architecture Search (NAS). However, an effective space involves tremendous prior knowledge and/…

Neural Architecture Search

Diffusion Graph Neural Networks for Robustness in Olfaction Sensors and Datasets

2025-05-31 · Kordel K. France, Ovidiu Daescu

Robotic odour source localization (OSL) is a critical capability for autonomous systems operating in complex environments. However, current OSL methods often suffer from ambiguities, particularly when robots misattribute…

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning

2025-04-28 · Run Luo, Renke Shan, Longze Chen, Ziqiang Liu 외

Large Vision-Language Models (LVLMs) are pivotal for real-world AI tasks like embodied intelligence due to their strong vision-language reasoning abilities. However, current LVLMs process entire images at the token level…

Contrastive Learning

Texture or Semantics? Vision-Language Models Get Lost in Font Recognition

2025-03-31 · Zhecheng Li, Guoxian Song, Yujun Cai, Zhen Xiong 외

Modern Vision-Language Models (VLMs) exhibit remarkable visual and linguistic capabilities, achieving impressive performance in various tasks such as image recognition and object localization. However, their effectivenes…

Few-Shot LearningFont RecognitionObject Localization