paper-with-me

홈 › Papers

Learning to generate one-sentence biographies from Wikidata

2017-02-21 · EACL 2017 4 · Andrew Chisholm, Will Radford, Ben Hachey

We investigate the generation of one-sentence Wikipedia biographies from facts derived from Wikidata slot-value pairs. We train a recurrent neural network sequence-to-sequence model with attention to select facts and generate textual summaries. Our model incorporates a novel secondary objective that helps ensure it generates sentences that contain the input facts. The model achieves a BLEU score of 41, improving significantly upon the vanilla sequence-to-sequence model and scoring roughly twice that of a simple template baseline. Human preference evaluation suggests the model is nearly as good as the Wikipedia reference. Manual analysis explores content selection, suggesting the model can trade the ability to infer knowledge against the risk of hallucinating incorrect information.

📄 PDF Abstract BibTeX arXiv:1702.06235

Code (1)

andychisholm/mimo 공식 구현 pytorch

Tasks

Sentence

Similar Papers 제목 키워드 기반

Towards a Brazilian History Knowledge Graph

2024-03-28 · Valeria de Paiva, Alexandre Rademaker

This short paper describes the first steps in a project to construct a knowledge graph for Brazilian history based on the Brazilian Dictionary of Historical Biographies (DHBB) and Wikipedia/Wikidata. We contend that larg…

Classical Chinese Sentence Segmentation for Tomb Biographies of Tang Dynasty

2019-08-28 · Chao-Lin Liu, Yi Chang

Tomb biographies of the Tang dynasty provide invaluable information about Chinese history. The original biographies are classical Chinese texts which contain neither word boundaries nor sentence boundaries. Relying on th…

BIG-bench Machine LearningSentenceSentence segmentation

Explicit vs. Implicit Biographies: Evaluating and Adapting LLM Information Extraction on Wikidata-Derived Texts

2025-09-18 · Alessandra Stramiglio, Andrea Schimmenti, Valentina Pasqual, Marieke van Erp 외 arxiv

Text Implicitness has always been challenging in Natural Language Processing (NLP), with traditional methods relying on explicit statements to identify entities and their relationships. From the sentence "Zuhdi attends c…

Information Extraction

WDV: A Broad Data Verbalisation Dataset Built from Wikidata

2022-05-05 · Gabriel Amaral, Odinaldo Rodrigues, Elena Simperl

Data verbalisation is a task of great importance in the current field of natural language processing, as there is great benefit in the transformation of our abundant structured and semi-structured data into human-readabl…

Learning to Generate Wikipedia Summaries for Underserved Languages from Wikidata

2018-03-19 · NAACL 2018 6 · Lucie-Aimée Kaffee, Hady Elsahar, Pavlos Vougiouklis, Christophe Gravier 외

While Wikipedia exists in 287 languages, its content is unevenly distributed among them. In this work, we investigate the generation of open domain Wikipedia summaries in underserved languages using structured data from …

Sentence