paper-with-me

홈 › Papers

Generating Description for Sequential Images with Local-Object Attention Conditioned on Global Semantic Context

2018-11-01 · WS 2018 11 · Jing Su, Chenghua Lin, Mian Zhou, QingYun Dai, Haoyu Lv
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningText Generation

Similar Papers 제목 키워드 기반

Generating Descriptions for Sequential Images with Local-Object Attention and Global Semantic Context Modelling

2020-12-02 · Jing Su, Chenghua Lin, Mian Zhou, Qingyun Dai 외

In this paper, we propose an end-to-end CNN-LSTM model for generating descriptions for sequential images with a local-object attention mechanism. To generate coherent descriptions, we capture global semantic context usin…

Text2Scene: Generating Compositional Scenes from Textual Descriptions

2018-09-04 · CVPR 2019 6 · Fuwen Tan, Song Feng, Vicente Ordonez

In this paper, we propose Text2Scene, a model that generates various forms of compositional scene representations from natural language descriptions. Unlike recent works, our method does NOT use Generative Adversarial Ne…

Text-to-Image Generation Grounded by Fine-Grained User Attention

2020-11-07 · Jing Yu Koh, Jason Baldridge, Honglak Lee, Yinfei Yang

Localized Narratives is a dataset with detailed natural language descriptions of images paired with mouse traces that provide a sparse, fine-grained visual grounding for phrases. We propose TReCS, a sequential model that…

Image GenerationPositionRetrievalSegmentation+3

Image Captioning with Object Detection and Localization

2017-06-08 · Zhongliang Yang, Yu-Jin Zhang, Sadaqat ur Rehman, Yongfeng Huang

Automatically generating a natural language description of an image is a task close to the heart of image understanding. In this paper, we present a multi-model neural network method closely related to the human visual s…

Image CaptioningObjectobject-detectionObject Detection

Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze

2020-11-09 · EMNLP 2020 11 · Ece Takmaz, Sandro Pezzelle, Lisa Beinborn, Raquel Fernández

When speakers describe an image, they tend to look at objects before mentioning them. In this paper, we investigate such sequential cross-modal alignment by modelling the image description generation process computationa…

cross-modal alignmentImage CaptioningImage Description