paper-with-me

홈 › Papers

Im2Text: Describing Images Using 1 Million Captioned Photographs

2011-12-01 · NeurIPS 2011 12 · Vicente Ordonez, Girish Kulkarni, Tamara L. Berg

We develop and demonstrate automatic image description methods using a large captioned photo collection. One contribution is our technique for the automatic collection of this new dataset -- performing a huge number of Flickr queries and then filtering the noisy results down to 1 million images with associated visually relevant captions. Such a collection allows us to approach the extremely challenging problem of description generation using relatively simple non-parametric methods and produces surprisingly effective results. We also develop methods incorporating many state of the art, but fairly noisy, estimates of image content to produce even more pleasing results. Finally we introduce a new objective performance measure for image captioning.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningImage Description

Similar Papers 제목 키워드 기반

1 Million Captioned Dutch Newspaper Images

2016-05-01 · LREC 2016 5 · Desmond Elliott, Martijn Kleppe

Images naturally appear alongside text in a wide variety of media, such as books, magazines, newspapers, and in online articles. This type of multi-modal data offers an interesting basis for vision and language research …

ArticlesData-to-Text GenerationImage CaptioningObject Recognition+3

Text-to-Image Synthesis Based on Machine Generated Captions

2019-10-09 · Marco Menardi, Alex Falcon, Saida S. Mohamed, Lorenzo Seidenari 외

Text to Image Synthesis refers to the process of automatic generation of a photo-realistic image starting from a given text and is revolutionizing many real-world applications. In order to perform such process it is nece…

Image CaptioningImage Generation

CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech

2025-06-03 · Helin Wang, Jiarui Hai, Dading Chong, Karan Thakkar 외

Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to real-world applications remains challen…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

A robot-assisted pipeline to rapidly scan 1.7 million historical aerial photographs

2025-03-31 · Sheila Masson, Alan Potts, Allan Williams, Steve Berggreen 외

During the 20th Century, aerial surveys captured hundreds of millions of high-resolution photographs of the earth's surface. These images, the precursors to modern satellite imagery, represent an extraordinary visual rec…

Retrieval

SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?

2024-02-02 · Hasan Abed Al Kader Hammoud, Hani Itani, Fabio Pizzati, Philip Torr 외

We present SynthCLIP, a CLIP model trained on entirely synthetic text-image pairs. Leveraging recent text-to-image (TTI) networks and large language models (LLM), we generate synthetic datasets of images and correspondin…