paper-with-me

Papers

Assessing and Verifying Task Utility in LLM-Powered Applications

2024-05-03 · Negar Arabzadeh, Siqing Huo, Nikhil Mehta, Qinqyun Wu, Chi Wang, Ahmed Awadallah, Charles L. A. Clarke, Julia Kiseleva

The rapid development of Large Language Models (LLMs) has led to a surge in applications that facilitate collaboration among multiple agents, assisting humans in their daily tasks. However, a significant gap remains in assessing to what extent LLM-powered applications genuinely enhance user experience and task execution efficiency. This highlights the need to verify utility of LLM-powered applications, particularly by ensuring alignment between the application's functionality and end-user needs. We introduce AgentEval, a novel framework designed to simplify the utility verification process by automatically proposing a set of criteria tailored to the unique purpose of any given application. This allows for a comprehensive assessment, quantifying the utility of an application against the suggested criteria. We present a comprehensive analysis of the effectiveness and robustness of AgentEval for two open source datasets including Math Problem solving and ALFWorld House-hold related tasks. For reproducibility purposes, we make the data, code and all the logs publicly available at https://bit.ly/3w3yKcS .

📄 PDF Abstract BibTeX arXiv:2405.02178

Code (0)

등록된 구현이 없습니다.

Tasks

Math

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications

2024-02-14 · Negar Arabzadeh, Julia Kiseleva, Qingyun Wu, Chi Wang 외

The rapid development in the field of Large Language Models (LLMs) has led to a surge in applications that facilitate collaboration among multiple agents to assist humans in their daily tasks. However, a significant gap …

Math

Stress Testing Concept Erasure with Large Language Model Agents

2026-07-20 · Yuyang Xue, Feng Chen, Zhihua Liu, Edward Moroshko 외 arxiv

Concept erasure aims to remove semantic concepts from a trained generative model and is increasingly important for responsible AI deployment. However, verifying whether a model has robustly removed targeted concepts rema…

Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment

2026-06-13 · Lipeng He, Yihan Wang, Jiawen Zhang, N. Asokan arxiv

Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses report near-zero attack success rate on …

Reinforcement Learning

Cascaded Composite Turbulence and Misalignment: Statistical Characterization and Applications to Reconfigurable Intelligent Surface-Empowered Wireless Systems

2021-06-29 · Alexandros-Apostolos A. Boulogeorgos, Nestor Chatzidiamantis, Harilaos G. Sandalidis, Angeliki Alexiou 외

Reconfigurable intelligent surfaces (RISs) empowered high-frequency (HF) wireless systems are expected to become the supporting pillar for several reliability and data rate hungry applications. Such systems are, however,…

Enhancing nonnative speech perception and production through an AI-powered application

2025-03-18 · Georgios P. Georgiou

While research on using Artificial Intelligence (AI) through various applications to enhance foreign language pronunciation is expanding, it has primarily focused on aspects such as comprehensibility and intelligibility,…

Sentence