SemStyle: Learning to Generate Stylised Image Captions using Unaligned Text
Linguistic style is an essential part of written communication, with the power to affect both clarity and attractiveness. With recent advances in vision and language, we can start to tackle the problem of generating image captions that are both visually grounded and appropriately styled. Existing approaches either require styled training captions aligned to images or generate captions with low relevance. We develop a model that learns to generate visually relevant styled captions from a large corpus of styled text without aligned images. The core idea of this model, called SemStyle, is to separate semantics and style. One key component is a novel and concise semantic term representation generated using natural language processing techniques and frame semantics. In addition, we develop a unified language model that decodes sentences with diverse word choices and syntax for different styles. Evaluations, both automatic and manual, show captions from SemStyle preserve image semantics, are descriptive, and are style shifted. More broadly, this work provides possibilities to learn richer image descriptions from the plethora of linguistic data available on the web.
Code (1)
Tasks
DescriptiveImage CaptioningLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems
Natural language generation (NLG) is a critical component of spoken dialogue and it has a significant impact both on usability and perceived quality. Most NLG systems in common use employ rules and heuristics and tend to…
InformativenessSentenceSpoken Dialogue SystemsText GenerationAn Agent-Based Model With Realistic Financial Time Series: A Method for Agent-Based Models Validation
This paper proposes a methodology to empirically validate an agent-based model (ABM) that generates artificial financial time series data comparable with real-world financial data. The approach is based on comparing the …
Time SeriesTime Series AnalysisTikZero: Zero-Shot Text-Guided Graphics Program Synthesis
With the rise of generative AI, synthesizing figures from text captions becomes a compelling application. However, achieving high geometric precision and editability requires representing figures as graphics programs in …
Program SynthesisUnsupervised Audio-Caption Aligning Learns Correspondences between Individual Sound Events and Textual Phrases
We investigate unsupervised learning of correspondences between sound events and textual phrases through aligning audio clips with textual captions describing the content of a whole audio clip. We align originally unalig…
Event DetectionRetrievalSound Event DetectionThird Time's the Charm? Image and Video Editing with StyleGAN3
StyleGAN is arguably one of the most intriguing and well-studied generative models, demonstrating impressive performance in image generation, inversion, and manipulation. In this work, we explore the recent StyleGAN3 arc…
DisentanglementImage GenerationVideo Editing