paper-with-me

홈 › Papers

TMI: Text-to-Image Meets Image-to-Image for Complementary Data Synthesis to Boost Long-Tailed Instance Segmentation

2026-07-09 · Hyeonseop Song, Seokhun Choi, Hoseok Do arxiv

Large-vocabulary instance segmentation is constrained by long-tailed category distributions and fine-grained inter-class ambiguity. While data synthesis offers a promising alternative, current paradigms have complementary limitations: text-to-image (T2I) methods inherit noisy pseudo-labels and struggle on rare classes, whereas copy-paste methods compromise contextual realism. To address these issues, we propose a hybrid pipeline coupling T2I generation with context-aware image-to-image (I2I) editing. The T2I branch provides broad category and scene diversity, while a teacher-student scheme ensures label reliability by selectively retaining only prompt-specified categories. To strengthen supervision for rare classes, we introduce VRAIN (Verified Rare-class Augmentation via INstructed editing), a novel I2I editor. VRAIN inserts high-confidence instances at semantically appropriate locations within in-the-wild scenes, yielding semantically coherent and visually natural edits that reduce domain gaps and enable targeted augmentation. On the LVIS benchmark, our method surpasses existing baselines, improving overall AP by up to +4.0 points and rare-class AP by up to +9.5 points, while scaling effectively with backbone capacity. Our project page is available at https://seokhunchoi.github.io/TMI

📄 PDF Abstract BibTeX arXiv:2607.08201

Code (2)

InsomaniacElf/sg-tamil-tts-resources-
Tavish9/awesome-daily-AI-arxiv ★ 111

Tasks

Instance Segmentation

Similar Papers 제목 키워드 기반

JPEG Meets PDE-based Image Compression

2021-02-01 · Sarah Andris, Joachim Weickert, Tobias Alt, Pascal Peter

Inpainting-based image compression is emerging as a promising competitor to transform-based compression techniques. Its key idea is to reconstruct image information from only few known regions through inpainting. Specifi…

Image Compression

LViT: Language meets Vision Transformer in Medical Image Segmentation

2022-06-29 · Zihan Li, Yunxiang Li, Qingde Li, Puyang Wang 외

Deep learning has been widely used in medical image segmentation and other aspects. However, the performance of existing medical image segmentation models has been limited by the challenge of obtaining sufficient high-qu…

Image SegmentationMedical Image SegmentationPseudo LabelSegmentation+2

Diffusion Meets Few-shot Class Incremental Learning

2025-03-30 · Junsu Kim, Yunhoe Ku, Dongyoon Han, Seungryul Baek

Few-shot class-incremental learning (FSCIL) is challenging due to extremely limited training data; while aiming to reduce catastrophic forgetting and learn new information. We propose Diffusion-FSCIL, a novel approach th…

class-incremental learningClass Incremental LearningFew-Shot Class-Incremental LearningIncremental Learning

Enlighten Anything: When Segment Anything Model Meets Low-Light Image Enhancement

2023-06-17 · Qihan Zhao, Xiaofeng Zhang, Hao Tang, Chaochen Gu 외

Image restoration is a low-level visual task, and most CNN methods are designed as black boxes, lacking transparency and intrinsic aesthetics. Many unsupervised approaches ignore the degradation of visible information in…

Image EnhancementImage RestorationLow-Light Image EnhancementSSIM+1

Enlighten-Your-Voice: When Multimodal Meets Zero-shot Low-light Image Enhancement

2023-12-15 · Xiaofeng Zhang, Zishan Xu, Hao Tang, Chaochen Gu 외

Low-light image enhancement is a crucial visual task, and many unsupervised methods tend to overlook the degradation of visible information in low-light scenes, which adversely affects the fusion of complementary informa…

Image EnhancementLow-Light Image Enhancement