paper-with-me

Papers

Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies

2026-06-30 · Xudong Shen, Li Yuan, Ye Chen, Xin Wu, Yi Cai, Zhiyong Wu arxiv

Large Language Models (LLMs) exhibit strong semantic capabilities, yet their resilience to manipulative linguistic patterns such as logical fallacies remains underexplored. Prior work has primarily examined whether LLMs can identify or classify fallacies, leaving their robustness against fallacious persuasion insufficiently studied. To address this gap, we introduce LoFa (Logical Fallacy), a comprehensive benchmark for evaluating LLM robustness against fallacies. LoFa is constructed through a multi-agent pipeline that pairs factual questions with fallacious arguments, and is accompanied by a multi-round debate framework for assessing model resilience under sustained adversarial persuasion. To disentangle fallacy robustness from a model's inherent knowledge limitations, we further propose Logical Fallacy Resistance at k (LFR@k), a metric that quantifies resistance to fallacious attacks. Experiments show that LLMs exhibit varying levels of robustness across different fallacy types, revealing distinct vulnerability profiles among models.

📄 PDF Abstract BibTeX arXiv:2606.31039

Code (0)

등록된 구현이 없습니다.

Tasks

Logical Fallacies

Similar Papers 제목 키워드 기반

Language Models Learn to Mislead Humans via RLHF

2024-09-19 · Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez 외

Language models (LMs) can produce errors that are hard to detect for humans, especially when the task is complex. RLHF, the most popular post-training method, may exacerbate this problem: to achieve higher rewards, LMs m…

Question Answering

Underwater Target Recognition based on Multi-Decision LOFAR Spectrum Enhancement: A Deep Learning Approach

2021-04-26 · Jie Chen, Jie Liu, Chang Liu, Jian Zhang 외

The Low frequency analysis and recording (LOFAR) spectrum is one of the key features of the under water target, which can be used for underwater target recognition. However, the underwater environment noise is complicate…

Segmenting Maxillofacial Structures in CBCT Volumes

2025-01-01 · CVPR 2025 1 · Federico Bolelli, Kevin Marchesini, Niels van Nistelrooij, Luca Lumetti 외

Cone-beam computed tomography (CBCT) is a standard imaging modality in orofacial and dental practices, providing essential 3D volumetric imaging of anatomical structures, including jawbones, teeth, sinuses, and neuro…

AnatomyBenchmarkingMambaSegmentation

HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO Debiasing

2026-05-23 · Ruyi Chen, Lu Zhou, Xiaogang Xu, Chiyu Zhang 외 arxiv

Text-to-Image (T2I) models have made significant strides in visual realism and semantic consistency, yet they often perpetuate and amplify societal biases. Existing evaluation methods typically address only single-dimens…

Physics-guided Emulators Reveal Resilience and Fragility under Operational Latencies and Outages

2025-10-21 · Sarth Dubey, Subimal Ghosh, Udit Bhatia arxiv

Reliable hydrologic and flood forecasting requires models that remain stable when input data are delayed, missing, or inconsistent. However, most advances in rainfall-runoff prediction have been evaluated under ideal dat…