Im2Text: Describing Images Using 1 Million Captioned Photographs
We develop and demonstrate automatic image description methods using a large captioned photo collection. One contribution is our technique for the automatic collection of this new dataset -- performing a huge number of Flickr queries and then filtering the noisy results down to 1 million images with associated visually relevant captions. Such a collection allows us to approach the extremely challenging problem of description generation using relatively simple non-parametric methods and produces surprisingly effective results. We also develop methods incorporating many state of the art, but fairly noisy, estimates of image content to produce even more pleasing results. Finally we introduce a new objective performance measure for image captioning.
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningImage DescriptionSimilar Papers 제목 키워드 기반
1 Million Captioned Dutch Newspaper Images
Images naturally appear alongside text in a wide variety of media, such as books, magazines, newspapers, and in online articles. This type of multi-modal data offers an interesting basis for vision and language research …
ArticlesData-to-Text GenerationImage CaptioningObject Recognition+3Text-to-Image Synthesis Based on Machine Generated Captions
Text to Image Synthesis refers to the process of automatic generation of a photo-realistic image starting from a given text and is revolutionizing many real-world applications. In order to perform such process it is nece…
Image CaptioningImage GenerationCapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to real-world applications remains challen…
Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisA robot-assisted pipeline to rapidly scan 1.7 million historical aerial photographs
During the 20th Century, aerial surveys captured hundreds of millions of high-resolution photographs of the earth's surface. These images, the precursors to modern satellite imagery, represent an extraordinary visual rec…
RetrievalSynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
We present SynthCLIP, a CLIP model trained on entirely synthetic text-image pairs. Leveraging recent text-to-image (TTI) networks and large language models (LLM), we generate synthetic datasets of images and correspondin…