paper-with-me

Papers

Walking a Tightrope -- Evaluating Large Language Models in High-Risk Domains

2023-11-25 · Chia-Chien Hung, Wiem Ben Rim, Lindsay Frost, Lars Bruckner, Carolin Lawrence

High-risk domains pose unique challenges that require language models to provide accurate and safe responses. Despite the great success of large language models (LLMs), such as ChatGPT and its variants, their performance in high-risk domains remains unclear. Our study delves into an in-depth analysis of the performance of instruction-tuned LLMs, focusing on factual accuracy and safety adherence. To comprehensively assess the capabilities of LLMs, we conduct experiments on six NLP datasets including question answering and summarization tasks within two high-risk domains: legal and medical. Further qualitative analysis highlights the existing limitations inherent in current LLMs when evaluating in high-risk domains. This underscores the essential nature of not only improving LLM capabilities but also prioritizing the refinement of domain-specific metrics, and embracing a more human-centric approach to enhance safety and factual reliability. Our findings advance the field toward the concerns of properly evaluating LLMs in high-risk domains, aiming to steer the adaptability of LLMs in fulfilling societal obligations and aligning with forthcoming regulations, such as the EU AI Act.

📄 PDF Abstract BibTeX arXiv:2311.14966

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Walking the Tightrope: Disentangling Beneficial and Detrimental Drifts in Non-Stationary Custom-Tuning

2025-05-19 · Xiaoyu Yang, Jie Lu, En Yu

This paper uncovers a critical yet overlooked phenomenon in multi-modal large language models (MLLMs): detrimental concept drift within chain-of-thought (CoT) reasoning during non-stationary reinforcement fine-tuning (RF…

counterfactualCounterfactual Reasoning

Walking the Tightrope of LLMs for Software Development: A Practitioners' Perspective

2025-11-09 · Samuel Ferino, Rashina Hoda, John Grundy, Christoph Treude arxiv

Background: Large Language Models emerged with the potential of provoking a revolution in software development (e.g., automating processes, workforce transformation). Although studies have started to investigate the perc…

Walking the Tightrope: An Investigation of the Convolutional Autoencoder Bottleneck

2019-11-18 · Ilja Manakov, Markus Rohm, Volker Tresp

In this paper, we present an in-depth investigation of the convolutional autoencoder (CAE) bottleneck. Autoencoders (AE), and especially their convolutional variants, play a vital role in the current deep learning toolbo…

Outlier DetectionRepresentation LearningTransfer Learning

HCVR Scene Generation: High Compatibility Virtual Reality Environment Generation for Extended Redirected Walking

2026-01-21 · Yiran Zhang, Xingpeng Sun, Aniket Bera arxiv

Natural walking enhances immersion in virtual environments (VEs), but physical space limitations and obstacles hinder exploration, especially in large virtual scenes. Redirected Walking (RDW) techniques mitigate this by …

Scene Generation

Understanding the Stability of Deep Control Policies for Biped Locomotion

2020-07-30 · Hwangpil Park, Ri Yu, Yoonsang Lee, Kyungho Lee 외

Achieving stability and robustness is the primary goal of biped locomotion control. Recently, deep reinforce learning (DRL) has attracted great attention as a general methodology for constructing biped control policies a…