paper-with-me

홈 › Papers

CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking

2025-09-04 · Ruiling Guo, Xinwei Yang, Chen Huang, Tong Zhang, Yong Hu arxiv

The effectiveness of large language models (LLMs) to fact-check misinformation remains uncertain, despite their growing use. To this end, we present CANDY, a benchmark designed to systematically evaluate the capabilities and limitations of LLMs in fact-checking Chinese misinformation. Specifically, we curate a carefully annotated dataset of ~20k instances. Our analysis shows that current LLMs exhibit limitations in generating accurate fact-checking conclusions, even when enhanced with chain-of-thought reasoning and few-shot prompting. To understand these limitations, we develop a taxonomy to categorize flawed LLM-generated explanations for their conclusions and identify factual fabrication as the most common failure mode. Although LLMs alone are unreliable for fact-checking, our findings indicate their considerable potential to augment human performance when deployed as assistive tools in scenarios. Our dataset and code can be accessed at https://github.com/SCUNLP/CANDY

📄 PDF Abstract BibTeX arXiv:2509.03957

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AIxcellent Vibes at GermEval 2025 Shared Task on Candy Speech Detection: Improving Model Performance by Span-Level Training

2025-09-09 · Christian Rene Thelen, Patrick Gustav Blaneck, Tobias Bornheim, Niklas Grieger 외 arxiv

Positive, supportive online communication in social media (candy speech) has the potential to foster civility, yet automated detection of such language remains underexplored, limiting systematic analysis of its impact. W…

A Multimodal Data Collection Framework for Dialogue-Driven Assistive Robotics to Clarify Ambiguities: A Wizard-of-Oz Pilot Study

2026-01-23 · Guangping Liu, Nicholas Hawkins, Billy Madden, Tipu Sultan 외 arxiv

Integrated control of wheelchairs and wheelchair-mounted robotic arms (WMRAs) has strong potential to increase independence for users with severe motor limitations, yet existing interfaces often lack the flexibility need…

MIP Candy: A Modular PyTorch Framework for Medical Image Processing

2026-02-24 · Tianhao Fu, Yucheng Chen arxiv

Medical image processing demands specialized software that handles high-dimensional volumetric data, heterogeneous file formats, and domain-specific training procedures. Existing frameworks either provide low-level compo…

Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications

2026-06-23 · Shayon Dasgupta, Avijit Dasgupta, C. V. Jawahar arxiv

Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified architecture enables strong performance on tasks like image caption…

Visual Question AnsweringImage Captioning

Saccharina latissima, candy-factory waste, and digestate from full-scale biogas plant as alternative carbohydrate and nutrient sources for lactic acid production

2023-08-07 · Eleftheria Papadopoulou, Charlene Vance, Paloma S. Rozene Vallespin, Panagiotis Tsapekos 외

To substitute petroleum-based materials with bio-based alternatives, microbial fermentation combined with inexpensive biomass is suggested. In this study Saccharina latissima hydrolysate, candy-factory waste, and digesta…