paper-with-me

홈 › Papers

Sketch Me if You Can: Towards Generating Detailed Descriptions of Object Shape by Grounding in Images and Drawings

2019-10-01 · WS 2019 10 · Ting Han, Sina Zarrie{\ss}

A lot of recent work in Language {\&} Vision has looked at generating descriptions or referring expressions for objects in scenes of real-world images, though focusing mostly on relatively simple language like object names, color and location attributes (e.g., brown chair on the left). This paper presents work on Draw-and-Tell, a dataset of detailed descriptions for common objects in images where annotators have produced fine-grained attribute-centric expressions distinguishing a target object from a range of similar objects. Additionally, the dataset comes with hand-drawn sketches for each object. As Draw-and-Tell is medium-sized and contains a rich vocabulary, it constitutes an interesting challenge for CNN-LSTM architectures used in state-of-the-art image captioning models. We explore whether the additional modality given through sketches can help such a model to learn to accurately ground detailed language referring expressions to object shapes. Our results are encouraging.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeImage CaptioningObject

Similar Papers 제목 키워드 기반

Sketch and Text Guided Diffusion Model for Colored Point Cloud Generation

2023-08-05 · ICCV 2023 1 · Zijie Wu, Yaonan Wang, Mingtao Feng, He Xie 외

Diffusion probabilistic models have achieved remarkable success in text guided image generation. However, generating 3D shapes is still challenging due to the lack of sufficient data containing 3D models along with their…

DenoisingImage GenerationPoint Cloud Generation

StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks

2016-12-10 · ICCV 2017 10 · Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang 외

Synthesizing high-quality images from text descriptions is a challenging problem in computer vision and has many practical applications. Samples generated by existing text-to-image approaches can roughly reflect the mean…

Image GenerationText-to-Image Generation

CustomSketching: Sketch Concept Extraction for Sketch-based Image Synthesis and Editing

2024-02-27 · Chufeng Xiao, Hongbo Fu

Personalization techniques for large text-to-image (T2I) models allow users to incorporate new concepts from reference images. However, existing methods primarily rely on textual descriptions, leading to limited control …

Image Generation

Magic3DSketch: Create Colorful 3D Models From Sketch-Based 3D Modeling Guided by Text and Language-Image Pre-Training

2024-07-27 · Ying Zang, Yidong Han, Chaotao Ding, Jianqi Zhang 외

The requirement for 3D content is growing as AR/VR application emerges. At the same time, 3D modelling is only available for skillful experts, because traditional methods like Computer-Aided Design (CAD) are often too la…

Text to 3D

SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches

2025-02-11 · Haichuan Lin, Yilin Ye, Jiazhi Xia, Wei Zeng

Text-to-image models can generate visually appealing images from text descriptions. Efforts have been devoted to improving model controls with prompt tuning and spatial conditioning. However, our formative study highligh…

Image GenerationText to Image GenerationText-to-Image Generation