paper-with-me

홈 › Papers

1 Million Captioned Dutch Newspaper Images

2016-05-01 · LREC 2016 5 · Desmond Elliott, Martijn Kleppe

Images naturally appear alongside text in a wide variety of media, such as books, magazines, newspapers, and in online articles. This type of multi-modal data offers an interesting basis for vision and language research but most existing datasets use crowdsourced text, which removes the images from their original context. In this paper, we introduce the KBK-1M dataset of 1.6 million images in their original context, with co-occurring texts found in Dutch newspapers from 1922 - 1994. The images are digitally scanned photographs, cartoons, sketches, and weather forecasts; the text is generated from OCR scanned blocks. The dataset is suitable for experiments in automatic image captioning, image―article matching, object recognition, and data-to-text generation for weather forecasting. It can also be used by humanities scholars to analyse photographic style changes, the representation of people and societal issues, and new tools for exploring photograph reuse via image-similarity-based search.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesData-to-Text GenerationImage CaptioningObject RecognitionOptical Character Recognition (OCR)Text GenerationWeather Forecasting

Similar Papers 제목 키워드 기반

Using Word Embeddings to Examine Gender Bias in Dutch Newspapers, 1950-1990

2019-07-21 · WS 2019 8 · Melvin Wevers

Contemporary debates on filter bubbles and polarization in public and social media raise the question to what extent news media of the past exhibited biases. This paper specifically examines bias related to gender in six…

Word Embeddings

Im2Text: Describing Images Using 1 Million Captioned Photographs

2011-12-01 · NeurIPS 2011 12 · Vicente Ordonez, Girish Kulkarni, Tamara L. Berg

We develop and demonstrate automatic image description methods using a large captioned photo collection. One contribution is our technique for the automatic collection of this new dataset -- performing a huge number of …

Image CaptioningImage Description

Collection of a corpus of Dutch SMS

2012-05-01 · LREC 2012 5 · Maaske Treurniet, Orph{\'e}e De Clercq, Henk van den Heuvel, Nelleke Oostdijk

In this paper we present the first freely available corpus of Dutch text messages containing data originating from the Netherlands and Flanders. This corpus has been collected in the framework of the SoNaR project and co…

The Newspaper Navigator Dataset: Extracting And Analyzing Visual Content from 16 Million Historic Newspaper Pages in Chronicling America

2020-05-04 · Benjamin Charles Germain Lee, Jaime Mears, Eileen Jakeway, Meghan Ferriter 외

Chronicling America is a product of the National Digital Newspaper Program, a partnership between the Library of Congress and the National Endowment for the Humanities to digitize historic newspapers. Over 16 million pag…

Optical Character Recognition (OCR)

CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech

2025-06-03 · Helin Wang, Jiarui Hai, Dading Chong, Karan Thakkar 외

Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to real-world applications remains challen…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis