paper-with-me

Papers

DetectBench: Can Large Language Model Detect and Piece Together Implicit Evidence?

2024-06-18 · Zhouhong Gu, Lin Zhang, Xiaoxuan Zhu, Jiangjie Chen, Wenhao Huang, Yikai Zhang, Shusen Wang, Zheyu Ye, Yan Gao, Hongwei Feng, Yanghua Xiao

Detecting evidence within the context is a key step in the process of reasoning task. Evaluating and enhancing the capabilities of LLMs in evidence detection will strengthen context-based reasoning performance. This paper proposes a benchmark called DetectBench for verifying the ability to detect and piece together implicit evidence within a long context. DetectBench contains 3,928 multiple-choice questions, with an average of 994 tokens per question. Each question contains an average of 4.55 pieces of implicit evidence, and solving the problem typically requires 7.62 logical jumps to find the correct answer. To enhance the performance of LLMs in evidence detection, this paper proposes Detective Reasoning Prompt and Finetune. Experiments demonstrate that the existing LLMs' abilities to detect evidence in long contexts are far inferior to humans. However, the Detective Reasoning Prompt effectively enhances the capability of powerful LLMs in evidence detection, while the Finetuning method shows significant effects in enhancing the performance of weaker LLMs. Moreover, when the abilities of LLMs in evidence detection are improved, their final reasoning performance is also enhanced accordingly.

📄 PDF Abstract BibTeX arXiv:2406.12641

Code (1)

MikeGu721/DetectBench 공식 구현

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultiple-choice

Similar Papers 제목 키워드 기반

Piecing Together Clues: A Benchmark for Evaluating the Detective Skills of Large Language Models

2023-07-11 · Zhouhong Gu, Lin Zhang, Jiangjie Chen, Haoning Ye 외

Detectives frequently engage in information detection and reasoning simultaneously when making decisions across various cases, especially when confronted with a vast amount of information. With the rapid development of l…

Common Sense ReasoningDecision MakingPrompt EngineeringReading Comprehension

VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models

2024-06-11 · Yu Liu, Lang Gao, Mingxin Yang, Yu Xie 외

Large Language Models (LLMs) have training corpora containing large amounts of program code, greatly improving the model's code comprehension and generation capabilities. However, sound comprehensive research on detectin…

Vulnerability Detection

AI-generated Image Detection: Passive or Watermark?

2024-11-20 · Moyang Guo, Yuepeng Hu, Zhengyuan Jiang, Zeyu Li 외

While text-to-image models offer numerous benefits, they also pose significant societal risks. Detecting AI-generated images is crucial for mitigating these risks. Detection methods can be broadly categorized into passiv…

PuzzleNet: Scene Text Detection by Segment Context Graph Learning

2020-02-26 · Hao Liu, Antai Guo, Deqiang Jiang, Yiqing Hu 외

Recently, a series of decomposition-based scene text detection methods has achieved impressive progress by decomposing challenging text regions into pieces and linking them in a bottom-up manner. However, most of them me…

Graph LearningScene Text DetectionText Detection

S^2F-Net:A Robust Spatial-Spectral Fusion Framework for Cross-Model AIGC Detection

2026-01-18 · Xiangyu Hu, Yicheng Hong, Hongchuang Zheng, Wenjun Zeng 외 arxiv

The rapid development of generative models has imposed an urgent demand for detection schemes with strong generalization capabilities. However, existing detection methods generally suffer from overfitting to specific sou…