paper-with-me

홈 › Papers

HLB: Benchmarking LLMs' Humanlikeness in Language Use

2024-09-24 · Xufeng Duan, Bei Xiao, Xuemei Tang, Zhenguang G. Cai

As synthetic data becomes increasingly prevalent in training language models, particularly through generated dialogue, concerns have emerged that these models may deviate from authentic human language patterns, potentially losing the richness and creativity inherent in human communication. This highlights the critical need to assess the humanlikeness of language models in real-world language use. In this paper, we present a comprehensive humanlikeness benchmark (HLB) evaluating 20 large language models (LLMs) using 10 psycholinguistic experiments designed to probe core linguistic aspects, including sound, word, syntax, semantics, and discourse (see https://huggingface.co/spaces/XufengDuan/HumanLikeness). To anchor these comparisons, we collected responses from over 2,000 human participants and compared them to outputs from the LLMs in these experiments. For rigorous evaluation, we developed a coding algorithm that accurately identified language use patterns, enabling the extraction of response distributions for each task. By comparing the response distributions between human participants and LLMs, we quantified humanlikeness through distributional similarity. Our results reveal fine-grained differences in how well LLMs replicate human responses across various linguistic levels. Importantly, we found that improvements in other performance metrics did not necessarily lead to greater humanlikeness, and in some cases, even resulted in a decline. By introducing psycholinguistic methods to model evaluation, this benchmark offers the first framework for systematically assessing the humanlikeness of LLMs in language use.

📄 PDF Abstract BibTeX arXiv:2409.15890

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Does GPT-4 pass the Turing test?

2023-10-31 · Cameron R. Jones, Benjamin K. Bergen

We evaluated GPT-4 in a public online Turing test. The best-performing GPT-4 prompt passed in 49.7% of games, outperforming ELIZA (22%) and GPT-3.5 (20%), but falling short of the baseline set by human participants (66%)…

TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking

2025-02-16 · Shahriar Kabir Nahin, Rabindra Nath Nandi, Sagor Sarker, Quazi Sarwar Muhtaseem 외

In this paper, we present TituLLMs, the first large pretrained Bangla LLMs, available in 1B and 3B parameter sizes. Due to computational constraints during both training and inference, we focused on smaller models. To tr…

Benchmarking

Varco Arena: A Tournament Approach to Reference-Free Benchmarking Large Language Models

2024-11-02 · Seonil Son, Ju-Min Oh, Heegon Jin, Cheolhun Jang 외

The rapid advancement of Large Language Models (LLMs) necessitates robust evaluation methodologies. Current benchmarking approaches often rely on comparing model outputs against predefined prompts and reference outputs. …

Benchmarking

HypoTermQA: Hypothetical Terms Dataset for Benchmarking Hallucination Tendency of LLMs

2024-02-25 · Cem Uluoglakci, Tugba Taskaya Temizel

Hallucinations pose a significant challenge to the reliability and alignment of Large Language Models (LLMs), limiting their widespread acceptance beyond chatbot applications. Despite ongoing efforts, hallucinations rema…

BenchmarkingChatbotHallucinationLanguage Modeling+1

Beyond Benchmarking: A New Paradigm for Evaluation and Assessment of Large Language Models

2024-07-10 · Jin Liu, Qingquan Li, Wenlong Du

In current benchmarks for evaluating large language models (LLMs), there are issues such as evaluation content restriction, untimely updates, and lack of optimization guidance. In this paper, we propose a new paradigm fo…

Benchmarking