paper-with-me

Papers

Enrich the content of the image Using Context-Aware Copy Paste

2024-07-11 · Qiushi Guo

Data augmentation remains a widely utilized technique in deep learning, particularly in tasks such as image classification, semantic segmentation, and object detection. Among them, Copy-Paste is a simple yet effective method and gain great attention recently. However, existing Copy-Paste often overlook contextual relevance between source and target images, resulting in inconsistencies in generated outputs. To address this challenge, we propose a context-aware approach that integrates Bidirectional Latent Information Propagation (BLIP) for content extraction from source images. By matching extracted content information with category information, our method ensures cohesive integration of target objects using Segment Anything Model (SAM) and You Only Look Once (YOLO). This approach eliminates the need for manual annotation, offering an automated and user-friendly solution. Experimental evaluations across diverse datasets demonstrate the effectiveness of our method in enhancing data diversity and generating high-quality pseudo-images across various computer vision tasks.

📄 PDF Abstract BibTeX arXiv:2407.08151

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDiversityimage-classificationImage Classificationobject-detectionObject DetectionSemantic Segmentation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Copy-Paste 설명 없음

Similar Papers 제목 키워드 기반

CIAN: Multi-Stage Framework for Event-Enriched Image Captioning via Retrieval-Augmented Generation

2026-06-16 · Trinh Thi Thu Hien, Trung-Nghia Le arxiv

Event-enriched image captioning describes not only visible content but also the broader context of events, including timing, location, and participants, capabilities missing in most pixel-bound models. We propose the Con…

Image Captioning

Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content?

2025-12-26 · Naen Xu, Jinghuai Zhang, Changjiang Li, Hengyu An 외 arxiv

Large vision-language models (LVLMs) have achieved remarkable advancements in multimodal reasoning tasks. However, their widespread accessibility raises critical concerns about potential copyright infringement. Will LVLM…

Multimodal Reasoning

Polyp-SES: Automatic Polyp Segmentation with Self-Enriched Semantic Model

2024-10-02 · Quang Vinh Nguyen, Thanh Hoang Son Vo, Sae-Ryung Kang, Soo-Hyung Kim

Automatic polyp segmentation is crucial for effective diagnosis and treatment in colonoscopy images. Traditional methods encounter significant challenges in accurately delineating polyps due to limitations in feature rep…

Instance SegmentationMedical Image SegmentationSegmentationVideo Polyp Segmentation

Shakespearizing Modern Language Using Copy-Enriched Sequence-to-Sequence Models

2017-07-04 · Harsh Jhamtani, Varun Gangal, Eduard Hovy, Eric Nyberg

Variations in writing styles are commonly used to adapt the content to a specific context, audience, or purpose. However, applying stylistic variations is still by and large a manual process, and there have been little e…

Shakespearizing Modern Language Using Copy-Enriched Sequence to Sequence Models

2017-09-01 · WS 2017 9 · Harsh Jhamtani, Varun Gangal, Eduard Hovy, Eric Nyberg

Variations in writing styles are commonly used to adapt the content to a specific context, audience, or purpose. However, applying stylistic variations is still by and large a manual process, and there have been little e…