paper-with-me

홈 › Papers

Weakly Supervised Annotations for Multi-modal Greeting Cards Dataset

2022-12-01 · Sidra Hanif, Longin Jan Latecki

In recent years, there is a growing number of pre-trained models trained on a large corpus of data and yielding good performance on various tasks such as classifying multimodal datasets. These models have shown good performance on natural images but are not fully explored for scarce abstract concepts in images. In this work, we introduce an image/text-based dataset called Greeting Cards. Dataset (GCD) that has abstract visual concepts. In our work, we propose to aggregate features from pretrained images and text embeddings to learn abstract visual concepts from GCD. This allows us to learn the text-modified image features, which combine complementary and redundant information from the multi-modal data streams into a single, meaningful feature. Secondly, the captions for the GCD dataset are computed with the pretrained CLIP-based image captioning model. Finally, we also demonstrate that the proposed the dataset is also useful for generating greeting card images using pre-trained text-to-image generation model.

📄 PDF Abstract BibTeX arXiv:2212.00847

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningImage GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

MWSIS: Multimodal Weakly Supervised Instance Segmentation with 2D Box Annotations for Autonomous Driving

2023-12-12 · Guangfeng Jiang, Jun Liu, Yuzhi Wu, Wenlong Liao 외

Instance segmentation is a fundamental research in computer vision, especially in autonomous driving. However, manual mask annotation for instance segmentation is quite time-consuming and costly. To address this problem,…

3D Instance SegmentationAutonomous DrivingInstance SegmentationSegmentation+2

CIEC: Coupling Implicit and Explicit Cues for Multimodal Weakly Supervised Manipulation Localization

2026-02-02 · Xinquan Yu, Wei Lu, Xiangyang Luo, Rui Yang arxiv

To mitigate the threat of misinformation, multimodal manipulation localization has garnered growing attention. Consider that current methods rely on costly and time-consuming fine-grained annotations, such as patch/token…

Prompt learning with bounding box constraints for medical image segmentation

2025-07-03 · Mélanie Gaillochet, Mehrdad Noori, Sahar Dastani, Christian Desrosiers 외 arxiv

Pixel-wise annotations are notoriously labourious and costly to obtain in the medical domain. To mitigate this burden, weakly supervised approaches based on bounding box annotations-much easier to acquire-offer a practic…

Medical Image Segmentation

Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection

2024-05-08 · Shengyang Sun, Xiaojin Gong

Weakly supervised multimodal violence detection aims to learn a violence detection model by leveraging multiple modalities such as RGB, optical flow, and audio, while only video-level annotations are available. In the pu…

Anomaly Detection In Surveillance VideosOptical Flow Estimation

VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment

2022-10-09 · Shraman Pramanick, Li Jing, Sayan Nag, Jiachen Zhu 외

Vision-language pre-training (VLP) has recently proven highly effective for various uni- and multi-modal downstream applications. However, most existing end-to-end VLP methods use high-resolution image-text box data to p…

object-detectionObject DetectionReferring ExpressionReferring Expression Comprehension