paper-with-me

홈 › Papers

Benchmarking of LLM Detection: Comparing Two Competing Approaches

2024-06-17 · Thorsten Pröhl, Erik Putzier, Rüdiger Zarnekow

This article gives an overview of the field of LLM text recognition. Different approaches and implemented detectors for the recognition of LLM-generated text are presented. In addition to discussing the implementations, the article focuses on benchmarking the detectors. Although there are numerous software products for the recognition of LLM-generated text, with a focus on ChatGPT-like LLMs, the quality of the recognition (recognition rate) is not clear. Furthermore, while it can be seen that scientific contributions presenting their novel approaches strive for some kind of comparison with other approaches, the construction and independence of the evaluation dataset is often not comprehensible. As a result, discrepancies in the performance evaluation of LLM detectors are often visible due to the different benchmarking datasets. This article describes the creation of an evaluation dataset and uses this dataset to investigate the different detectors. The selected detectors are benchmarked against each other.

📄 PDF Abstract BibTeX arXiv:2406.11670

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Varco Arena: A Tournament Approach to Reference-Free Benchmarking Large Language Models

2024-11-02 · Seonil Son, Ju-Min Oh, Heegon Jin, Cheolhun Jang 외

The rapid advancement of Large Language Models (LLMs) necessitates robust evaluation methodologies. Current benchmarking approaches often rely on comparing model outputs against predefined prompts and reference outputs. …

Benchmarking

Towards Benchmarking and Evaluating Deepfake Detection

2022-03-04 · Chenhao Lin, Jingyi Deng, Pengbin Hu, Chao Shen 외

Deepfake detection automatically recognizes the manipulated medias through the analysis of the difference between manipulated and non-altered videos. It is natural to ask which are the top performers among the existing d…

BenchmarkingDeepFake DetectionFace Swapping

Efficient Benchmarking of NLP APIs using Multi-armed Bandits

2017-04-01 · EACL 2017 4 · Gholamreza Haffari, Tuan Dung Tran, Mark Carman

Comparing NLP systems to select the best one for a task of interest, such as named entity recognition, is critical for practitioners and researchers. A rigorous approach involves setting up a hypothesis testing scenario …

BenchmarkingMulti-Armed Banditsnamed-entity-recognitionNamed Entity Recognition+3

OrionBench: Benchmarking Time Series Generative Models in the Service of the End-User

2023-10-26 · Sarah Alnegheimish, Laure Berti-Equille, Kalyan Veeramachaneni

Time series anomaly detection is a vital task in many domains, including patient monitoring in healthcare, forecasting in finance, and predictive maintenance in energy industries. This has led to a proliferation of anoma…

Anomaly DetectionBenchmarkingTime SeriesTime Series Anomaly Detection

Benchmarking Individual Tree Mapping with Sub-meter Imagery

2023-11-14 · Dimitri Gominski, Ankit Kariryaa, Martin Brandt, Christian Igel 외

There is a rising interest in mapping trees using satellite or aerial imagery, but there is no standardized evaluation protocol for comparing and enhancing methods. In dense canopy areas, the high variability of tree siz…

BenchmarkingSegmentation