paper-with-me

홈 › Papers

WebBrain: Learning to Generate Factually Correct Articles for Queries by Grounding on Large Web Corpus

2023-04-10 · Hongjing Qian, Yutao Zhu, Zhicheng Dou, Haoqi Gu, Xinyu Zhang, Zheng Liu, Ruofei Lai, Zhao Cao, Jian-Yun Nie, Ji-Rong Wen

In this paper, we introduce a new NLP task -- generating short factual articles with references for queries by mining supporting evidence from the Web. In this task, called WebBrain, the ultimate goal is to generate a fluent, informative, and factually-correct short article (e.g., a Wikipedia article) for a factual query unseen in Wikipedia. To enable experiments on WebBrain, we construct a large-scale dataset WebBrain-Raw by extracting English Wikipedia articles and their crawlable Wikipedia references. WebBrain-Raw is ten times larger than the previous biggest peer dataset, which can greatly benefit the research community. From WebBrain-Raw, we construct two task-specific datasets: WebBrain-R and WebBrain-G, which are used to train in-domain retriever and generator, respectively. Besides, we empirically analyze the performances of the current state-of-the-art NLP techniques on WebBrain and introduce a new framework ReGen, which enhances the generation factualness by improved evidence retrieval and task-specific pre-training for generation. Experiment results show that ReGen outperforms all baselines in both automatic and human evaluations.

📄 PDF Abstract BibTeX arXiv:2304.04358

Code (1)

qhjqhj00/webbrain 공식 구현 pytorch

Tasks

ArticlesRetrievalText Generation

Similar Papers 제목 키워드 기반

Automatic Detection of Entity-Manipulated Text using Factual Knowledge

2022-03-19 · ACL 2022 5 · Ganesh Jawahar, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan

In this work, we focus on the problem of distinguishing a human written news article from a news article that is created by manipulating entities in a human written news article (e.g., replacing entities with factually i…

Articles

Assessing Factual Music Comprehension in Large Audio Language Models

2025-11-02 · Daniel Chenyu Lin, Michael Freeman, John Thickstun arxiv

Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio. In this paper, we (1) provide empirical evidence that assessment of LALMs us…

Natural Language QueriesInformation Retrieval

Knowledge-Augmented Language Model Verification

2023-10-19 · Jinheon Baek, Soyeong Jeong, Minki Kang, Jong C. Park 외

Recent Language Models (LMs) have shown impressive capabilities in generating texts with the knowledge internalized in parameters. Yet, LMs often generate the factually incorrect responses to the given queries, since the…

Language ModelingLanguage ModellingmodelQuestion Answering+1

Rome was built in 1776: A Case Study on Factual Correctness in Knowledge-Grounded Response Generation

2021-10-11 · Sashank Santhanam, Behnam Hedayatnia, Spandana Gella, Aishwarya Padmakumar 외

Recently neural response generation models have leveraged large pre-trained transformer models and knowledge snippets to generate relevant and informative responses. However, this does not guarantee that generated respon…

Response Generation

CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization

2021-09-19 · EMNLP 2021 11 · Shuyang Cao, Lu Wang

We study generating abstractive summaries that are faithful and factually consistent with the given articles. A novel contrastive learning formulation is presented, which leverages both reference summaries, as positive t…

Abstractive Text SummarizationArticlesContrastive LearningReranking