Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
Large language models (LLMs) have been shown to be proficient in correctly answering questions in the context of online discourse. However, the study of using LLMs to model human-like answers to fact-driven social media questions is still under-explored. In this work, we investigate how LLMs model the wide variety of human answers to fact-driven questions posed on several topic-specific Reddit communities, or subreddits. We collect and release a dataset of 409 fact-driven questions and 7,534 diverse, human-rated answers from 15 r/Ask{Topic} communities across 3 categories: profession, social identity, and geographic location. We find that LLMs are considerably better at modeling highly-rated human answers to such questions, as opposed to poorly-rated human answers. We present several directions for future research based on our initial findings.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Human-Aligned Enhancement of Programming Answers with LLMs Guided by User Feedback
Large Language Models (LLMs) are widely used to support software developers in tasks such as code generation, optimization, and documentation. However, their ability to improve existing programming answers in a human-lik…
Code GenerationLLM-Driven Personalized Answer Generation and Evaluation
Online learning has experienced rapid growth due to its flexibility and accessibility. Personalization, adapted to the needs of individual learners, is crucial for enhancing the learning experience, particularly in onlin…
Answer GenerationEvaluation of Attribution Bias in Retrieval-Augmented Large Language Models
Attributing answers to source documents is an approach used to enhance the verifiability of a model's output in retrieval augmented generation (RAG). Prior work has mainly focused on improving and evaluating the attribut…
AttributecounterfactualRAGRetrieval+2Potemkin Understanding in Large Language Models
Large language models (LLMs) are regularly evaluated using benchmark datasets. But what justifies making inferences about an LLM's capabilities based on its answers to a curated set of questions? This paper first introdu…
validLRQ-Fact: LLM-Generated Relevant Questions for Multimodal Fact-Checking
Human fact-checkers have specialized domain knowledge that allows them to formulate precise questions to verify information accuracy. However, this expert-driven approach is labor-intensive and is not scalable, especiall…
Fact CheckingMisinformation