paper-with-me

Papers

FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs

2024-11-21 · Shinbok Lee, Gaeun Seo, Daniel Lee, Byeongil Ko, Sunghee Jung, Myeongcheol Shin

This study investigates language models' generative capabilities in tool-use dialogs. We categorize the models' outputs in tool-use dialogs into four distinct types: Tool Call, Answer Completion, Slot Question, and Relevance Detection, which serve as aspects for evaluation. We introduce FunctionChat-Bench, comprising 700 evaluation items and automated assessment programs. Using this benchmark, we evaluate several language models that support function calling. Our findings indicate that while language models may exhibit high accuracy in single-turn Tool Call scenarios, this does not necessarily translate to superior generative performance in multi-turn environments. We argue that the capabilities required for function calling extend beyond generating tool call messages; they must also effectively generate conversational messages that engage the user.

📄 PDF Abstract BibTeX arXiv:2411.14054

Code (1)

kakao/functionchat-bench 공식 구현

Tasks

Relevance Detection

Similar Papers 제목 키워드 기반

TurkBench: A Benchmark for Evaluating Turkish Large Language Models

2026-01-11 · Çağrı Toraman, Ahmet Kaan Sever, Ayse Aysu Cengiz, Elif Ecem Arslan 외 arxiv

With the recent surge in the development of large language models, the need for comprehensive and language-specific evaluation benchmarks has become critical. While significant progress has been made in evaluating Englis…

Instruction Following

Environmental large language model Evaluation (ELLE) dataset: A Benchmark for Evaluating Generative AI applications in Eco-environment Domain

2025-01-10 · Jing Guo, Nan Li, Ming Xu

Generative AI holds significant potential for ecological and environmental applications such as monitoring, data analysis, education, and policy support. However, its effectiveness is limited by the lack of a unified eva…

Language Model EvaluationLanguage ModelingLanguage ModellingLarge Language Model

RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

2024-10-13 · Enyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi 외

Reward models (RMs) guide the alignment of large language models (LLMs), steering them toward behaviors preferred by humans. Evaluating RMs is the key to better aligning LLMs. However, the current evaluation of RMs may n…

Benchmarking

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

2023-07-30 · Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge 외

Based on powerful Large Language Models (LLMs), recent generative Multimodal Large Language Models (MLLMs) have gained prominence as a pivotal research area, exhibiting remarkable capability for both comprehension and ge…

BenchmarkingMultiple-choice

WritingBench: A Comprehensive Benchmark for Generative Writing

2025-03-07 · Yuning Wu, Jiahao Mei, Ming Yan, Chenliang Li 외

Recent advancements in large language models (LLMs) have significantly enhanced text generation capabilities, yet evaluating their performance in generative writing remains a challenge. Existing benchmarks primarily focu…

Text Generation