paper-with-me

홈 › Papers

Generating Multi-Sentence Lingual Descriptions of Indoor Scenes

2015-02-28 · Dahua Lin, Chen Kong, Sanja Fidler, Raquel Urtasun

This paper proposes a novel framework for generating lingual descriptions of indoor scenes. Whereas substantial efforts have been made to tackle this problem, previous approaches focusing primarily on generating a single sentence for each image, which is not sufficient for describing complex scenes. We attempt to go beyond this, by generating coherent descriptions with multiple sentences. Our approach is distinguished from conventional ones in several aspects: (1) a 3D visual parsing system that jointly infers objects, attributes, and relations; (2) a generative grammar learned automatically from training text; and (3) a text generation algorithm that takes into account the coherence among sentences. Experiments on the augmented NYU-v2 dataset show that our framework can generate natural descriptions with substantially higher ROGUE scores compared to those produced by the baseline.

📄 PDF Abstract BibTeX arXiv:1503.00064

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceText Generation

Similar Papers 제목 키워드 기반

Less is More: Generating Grounded Navigation Instructions from Landmarks

2021-11-25 · CVPR 2022 1 · Su Wang, Ceslee Montgomery, Jordi Orbay, Vighnesh Birodkar 외

We study the automatic generation of navigation instructions from 360-degree images captured on indoor routes. Existing generators suffer from poor visual grounding, causing them to rely on language priors and hallucinat…

DecoderInstruction FollowingVisual Grounding

Generating Image Descriptions using Multilingual Data

2017-09-01 · WS 2017 9 · Alan Jaffe
Image CaptioningLanguage ModelingLanguage ModellingMachine Translation+1

On generating coherent multilingual descriptions of museum objects from Semantic Web ontologies

2012-05-01 · WS 2012 5 · Dana Dann{\'e}lls
Text Generation

Image Pivoting for Learning Multilingual Multimodal Representations

2017-07-24 · EMNLP 2017 9 · Spandana Gella, Rico Sennrich, Frank Keller, Mirella Lapata

In this paper we propose a model to learn multimodal multilingual representations for matching images and sentences in different languages, with the aim of advancing multilingual versions of image search and image unders…

Image DescriptionImage RetrievalSemantic Textual Similarity

Sentence-Level Multilingual Multi-modal Embedding for Natural Language Processing

2017-09-01 · RANLP 2017 9 · Iacer Calixto, Qun Liu

We propose a novel discriminative ranking model that learns embeddings from multilingual and multi-modal data, meaning that our model can take advantage of images and descriptions in multiple languages to improve embeddi…

Machine TranslationNMTRe-RankingSemantic Textual Similarity+3