paper-with-me

Papers

Assessing Web Search Credibility and Response Groundedness in Chat Assistants

2025-10-15 · Ivan Vykopal, Matúš Pikuliak, Simon Ostermann, Marián Šimko arxiv

Chat assistants increasingly integrate web search functionality, enabling them to retrieve and cite external sources. While this promises more reliable answers, it also raises the risk of amplifying misinformation from low-credibility sources. In this paper, we introduce a novel methodology for evaluating assistants' web search behavior, focusing on source credibility and the groundedness of responses with respect to cited sources. Using 100 claims across five misinformation-prone topics, we assess GPT-4o, GPT-5, Perplexity, and Qwen Chat. Our findings reveal differences between the assistants, with Perplexity achieving the highest source credibility, whereas GPT-4o exhibits elevated citation of non-credibility sources on sensitive topics. This work provides the first systematic comparison of commonly used chat assistants for fact-checking behavior, offering a foundation for evaluating AI systems in high-stakes information environments.

📄 PDF Abstract BibTeX arXiv:2510.13749

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Search Arena: Analyzing Search-Augmented LLMs

2025-06-05 · Mihran Miroyan, Tsung-Han Wu, Logan King, Tianle Li 외

Search-augmented language models combine web search with Large Language Models (LLMs) to improve response groundedness and freshness. However, analyzing these systems remains challenging: existing datasets are limited in…

Fact Checking

A Revisit of Fake News Dataset with Augmented Fact-checking by ChatGPT

2023-12-19 · Zizhong Li, Haopeng Zhang, Jiawei Zhang

The proliferation of fake news has emerged as a critical issue in recent years, requiring significant efforts to detect it. However, the existing fake news detection datasets are sourced from human journalists, which are…

Fact CheckingFake News Detection

Building Multimodal AI Chatbots

2023-04-21 · Min Young Lee

This work aims to create a multimodal AI system that chats with humans and shares relevant photos. While earlier works were limited to dialogues about specific objects or scenes within images, recent works have incorpora…

ChatbotMultimodal Deep Learning

Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility

2023-05-15 · Wentao Ye, Mingfeng Ou, Tianyi Li, Yipeng chen 외

The recent popularity of large language models (LLMs) has brought a significant impact to boundless fields, particularly through their open-ended ecosystem such as the APIs, open-sourced models, and plugins. However, wit…

Memorization

Decoding AI Judgment: How LLMs Assess News Credibility and Bias

2025-02-06 · Edoardo Loru, Jacopo Nudo, Niccolò Di Marco, Matteo Cinelli 외

Large Language Models (LLMs) are increasingly used to assess news credibility, yet little is known about how they make these judgments. While prior research has examined political bias in LLM outputs or their potential f…

Fact Checking