paper-with-me

홈 › Papers

Interactive Data Synthesis for Systematic Vision Adaptation via LLMs-AIGCs Collaboration

2023-05-22 · Qifan Yu, Juncheng Li, Wentao Ye, Siliang Tang, Yueting Zhuang

Recent text-to-image generation models have shown promising results in generating high-fidelity photo-realistic images. In parallel, the problem of data scarcity has brought a growing interest in employing AIGC technology for high-quality data expansion. However, this paradigm requires well-designed prompt engineering that cost-less data expansion and labeling remain under-explored. Inspired by LLM's powerful capability in task guidance, we propose a new paradigm of annotated data expansion named as ChatGenImage. The core idea behind it is to leverage the complementary strengths of diverse models to establish a highly effective and user-friendly pipeline for interactive data augmentation. In this work, we extensively study how LLMs communicate with AIGC model to achieve more controllable image generation and make the first attempt to collaborate them for automatic data augmentation for a variety of downstream tasks. Finally, we present fascinating results obtained from our ChatGenImage framework and demonstrate the powerful potential of our synthetic data for systematic vision adaptation. Our codes are available at https://github.com/Yuqifan1117/Labal-Anything-Pipeline.

📄 PDF Abstract BibTeX arXiv:2305.12799

Code (1)

yuqifan1117/labal-anything-pipeline 공식 구현 pytorch

Tasks

Data AugmentationImage GenerationPrompt EngineeringText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Measuring How (Not Just Whether) VLMs Build Common Ground

2025-09-04 · Saki Imai, Mert İnan, Anthony Sicilia, Malihe Alikhani arxiv

Large vision language models (VLMs) increasingly claim reasoning skills, yet current benchmarks evaluate them in single-turn or question answering settings. However, grounding is an interactive process in which people gr…

Question Answering

Psychological and behavioural responses in human-agent vs. human-human interactions: a systematic review and meta-analysis

2025-09-25 · Jianan Zhou, Fleur Corbett, Joori Byun, Talya Porat 외 arxiv

Interactive intelligent agents are being integrated across society. Despite achieving human-like capabilities, humans' responses to these agents remain poorly understood, with research fragmented across disciplines. We c…

Evidence-Based Text-Conditioned 3D CT Synthesis for Ovarian Cancer

2026-06-27 · Francesca Pia Panaccione, Eugenio Lomurno, Francesca Fati, Carlotta Pecchiari 외 arxiv

Ovarian cancer is frequently diagnosed at an advanced stage, making preoperative contrast-enhanced computed tomography (CT) central to staging and surgical planning; yet the scarcity of annotated imaging data, compounded…

Domain Adaptation

AsgardBench -- Evaluating Visually Grounded Interactive Planning Under Minimal Feedback

2026-03-16 · Andrea Tupini, Lars Liden, Reuben Tan, Yu Wang 외 arxiv

With AsgardBench we aim to evaluate visually grounded, high-level action sequence generation and interactive planning, focusing specifically on plan adaptation during execution based on visual observations rather than na…

Visual Grounding

What's the next frontier for Data-centric AI? Data Savvy Agents

2025-11-02 · Nabeel Seedat, Jiashuo Liu, Mihaela van der Schaar arxiv

The recent surge in AI agents that autonomously communicate, collaborate with humans and use diverse tools has unlocked promising opportunities in various real-world settings. However, a vital aspect remains underexplore…