paper-with-me

Papers

clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations

2025-05-08 · Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen

The emergence of instruction-tuned large language models (LLMs) has advanced the field of dialogue systems, enabling both realistic user simulations and robust multi-turn conversational agents. However, existing research often evaluates these components in isolation-either focusing on a single user simulator or a specific system design-limiting the generalisability of insights across architectures and configurations. In this work, we propose clem todd (chat-optimized LLMs for task-oriented dialogue systems development), a flexible framework for systematically evaluating dialogue systems under consistent conditions. clem todd enables detailed benchmarking across combinations of user simulators and dialogue systems, whether existing models from literature or newly developed ones. It supports plug-and-play integration and ensures uniform datasets, evaluation metrics, and computational constraints. We showcase clem todd's flexibility by re-evaluating existing task-oriented dialogue systems within this unified setup and integrating three newly proposed dialogue systems into the same evaluation pipeline. Our results provide actionable insights into how architecture, scale, and prompting strategies affect dialogue performance, offering practical guidance for building efficient and effective conversational AI systems.

📄 PDF Abstract BibTeX arXiv:2505.05445

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingTask-Oriented Dialogue Systems

Similar Papers 제목 키워드 기반

Towards Embodied AI with MuscleMimic: Unlocking full-body musculoskeletal motor learning at scale

2026-03-26 · Chengkun Li, Cheryl Wang, Bianca Ziliotto, Merkourios Simos 외 arxiv

Learning motor control for muscle-driven musculoskeletal models is hindered by the computational cost of biomechanically accurate simulation and the scarcity of validated, open full-body models. Here we present MuscleMim…

TERL: Large-Scale Multi-Target Encirclement Using Transformer-Enhanced Reinforcement Learning

2025-03-16 · Heng Zhang, Guoxiang Zhao, Xiaoqiang Ren

Pursuit-evasion (PE) problem is a critical challenge in multi-robot systems (MRS). While reinforcement learning (RL) has shown its promise in addressing PE tasks, research has primarily focused on single-target pursuit, …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Multiple Toddler Tracking in Indoor Videos

2023-11-29 · Somaieh Amraee, Bishoy Galoaa, Matthew Goodwin, Elaheh Hatamimajoumerd 외

Multiple toddler tracking (MTT) involves identifying and differentiating toddlers in video footage. While conventional multi-object tracking (MOT) algorithms are adept at tracking diverse objects, toddlers pose unique ch…

Multi-Object TrackingMultiple Object TrackingObject Tracking

CycleMix: A Holistic Strategy for Medical Image Segmentation from Scribble Supervision

2022-03-03 · CVPR 2022 1 · Ke Zhang, Xiahai Zhuang

Curating a large set of fully annotated training data can be costly, especially for the tasks of medical image segmentation. Scribble, a weaker form of annotation, is more obtainable in practice, but training segmentatio…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

CycleMLP: A MLP-like Architecture for Dense Prediction

2021-07-21 · ICLR 2022 4 · Shoufa Chen, Enze Xie, Chongjian Ge, Runjian Chen 외

This paper presents a simple MLP-like architecture, CycleMLP, which is a versatile backbone for visual recognition and dense predictions. As compared to modern MLP architectures, e.g., MLP-Mixer, ResMLP, and gMLP, whose …

Image ClassificationInstance Segmentationobject-detectionObject Detection+3