Generating Description for Sequential Images with Local-Object Attention Conditioned on Global Semantic Context
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningText GenerationSimilar Papers 제목 키워드 기반
Generating Descriptions for Sequential Images with Local-Object Attention and Global Semantic Context Modelling
In this paper, we propose an end-to-end CNN-LSTM model for generating descriptions for sequential images with a local-object attention mechanism. To generate coherent descriptions, we capture global semantic context usin…
Text2Scene: Generating Compositional Scenes from Textual Descriptions
In this paper, we propose Text2Scene, a model that generates various forms of compositional scene representations from natural language descriptions. Unlike recent works, our method does NOT use Generative Adversarial Ne…
Text-to-Image Generation Grounded by Fine-Grained User Attention
Localized Narratives is a dataset with detailed natural language descriptions of images paired with mouse traces that provide a sparse, fine-grained visual grounding for phrases. We propose TReCS, a sequential model that…
Image GenerationPositionRetrievalSegmentation+3Image Captioning with Object Detection and Localization
Automatically generating a natural language description of an image is a task close to the heart of image understanding. In this paper, we present a multi-model neural network method closely related to the human visual s…
Image CaptioningObjectobject-detectionObject DetectionGenerating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze
When speakers describe an image, they tend to look at objects before mentioning them. In this paper, we investigate such sequential cross-modal alignment by modelling the image description generation process computationa…
cross-modal alignmentImage CaptioningImage Description