paper-with-me

Papers

EconEvals: Benchmarks and Litmus Tests for LLM Agents in Unknown Environments

2025-03-24 · Sara Fish, Julia Shephard, Minkai Li, Ran I. Shorrer, Yannai A. Gonczarowski

We develop benchmarks for LLM agents that act in, learn from, and strategize in unknown environments, the specifications of which the LLM agent must learn over time from deliberate exploration. Our benchmarks consist of decision-making tasks derived from key problems in economics. To forestall saturation, the benchmark tasks are synthetically generated with scalable difficulty levels. Additionally, we propose litmus tests, a new kind of quantitative measure for LLMs and LLM agents. Unlike benchmarks, litmus tests quantify differences in character, values, and tendencies of LLMs and LLM agents, by considering their behavior when faced with tradeoffs (e.g., efficiency versus equality) where there is no objectively right or wrong behavior. Overall, our benchmarks and litmus tests assess the abilities and tendencies of LLM agents in tackling complex economic problems in diverse settings spanning procurement, scheduling, task allocation, and pricing -- applications that should grow in importance as such agents are further integrated into the economy.

📄 PDF Abstract BibTeX arXiv:2503.18825

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingScheduling

Similar Papers 제목 키워드 기반

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

2026-05-11 · Chiyu Zhang, Huiqin Yang, Bendong Jiang, Xiaolei Zhang 외 arxiv

The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content safety: behavior jailbreak, where an adversary induces an agent to exec…

Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems

2026-06-22 · Prajjwal Gupta, Prasang Gupta, Vishal Bhutani, Apoorva Sharma 외 arxiv

As agentic LLM systems move from prototypes to deployment across increasingly diverse domains, evaluating them has become both more important and more difficult. The challenge is not only that individual metrics may be u…

Universal Litmus Patterns: Revealing Backdoor Attacks in CNNs

2019-06-26 · CVPR 2020 6 · Soheil Kolouri, Aniruddha Saha, Hamed Pirsiavash, Heiko Hoffmann

The unprecedented success of deep neural networks in many applications has made these networks a prime target for adversarial exploitation. In this paper, we introduce a benchmark technique for detecting backdoor attacks…

Traffic Sign Recognition

Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas

2025-05-20 · Yu Ying Chiu, Zhilin Wang, Sharan Maiya, Yejin Choi 외

Detecting AI risks becomes more challenging as stronger models emerge and find novel methods such as Alignment Faking to circumvent these detection attempts. Inspired by how risky behaviors in humans (i.e., illegal activ…

Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations

2025-06-16 · Abhilekh Borah, Chhavi Sharma, Danush Khanna, Utkarsh Bhatt 외

Alignment is no longer a luxury, it is a necessity. As large language models (LLMs) enter high-stakes domains like education, healthcare, governance, and law, their behavior must reliably reflect human-aligned values and…

Diagnostic