paper-with-me

Papers AI Agent

“AI Agent” 태그가 달린 논문 391편 · 필터 해제

Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI

2025-07-13 · Phat Nguyen, Ngai-Man Cheung

Token compression techniques have recently emerged as powerful tools for accelerating Vision Transformer (ViT) inference in computer vision. Due to the quadratic computational complexity with respect to the token sequenc…

AI Agent

OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety

2025-07-08 · Sanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang 외

Recent advances in AI agents capable of solving complex, everyday tasks, from scheduling to customer service, have enabled deployment in real-world settings, but their possibilities for unsafe behavior demands rigorous e…

AI AgentScheduling

STELLA: Self-Evolving LLM Agent for Biomedical Research

2025-07-01 · Ruofan Jin, Zaixi Zhang, Mengdi Wang, Le Cong

The rapid growth of biomedical data, tools, and literature has created a fragmented research landscape that outpaces human expertise. While AI agents offer a solution, they typically rely on static, manually curated tool…

AI AgentHumanity's Last ExamSelf-Evolving AI

Prover Agent: An Agent-based Framework for Formal Mathematical Proofs

2025-06-24 · Kaito Baba, Chaoran Liu, Shuhei Kurita, Akiyoshi Sannai

We present Prover Agent, a novel AI agent for automated theorem proving that integrates large language models (LLMs) with a formal proof assistant, Lean. Prover Agent coordinates an informal reasoning LLM, a formal prove…

AI AgentAutomated Theorem ProvingMathematical Proofs

AI Agents-as-Judge: Automated Assessment of Accuracy, Consistency, Completeness and Clarity for Enterprise Documents

2025-06-23 · Sudip Dasgupta, Himanshu Shankar

This study presents a modular, multi-agent system for the automated review of highly structured enterprise business documents using AI agents. Unlike prior solutions focused on unstructured texts or limited compliance ch…

AI Agent

Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues

2025-06-19 · Myke C. Cohen, Zhe Su, Hsien-Te Kao, Daniel Nguyen 외

This paper presents an evaluation framework for agentic AI systems in mission-critical negotiation contexts, addressing the need for AI agents that can adapt to diverse human operators and stakeholders. Using Sotopia as …

AI AgentCausal Discovery

xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

2025-06-16 · Kaiyuan Chen, Yixin Ren, Yang Liu, Xiaobo Hu 외

We introduce xbench, a dynamic, profession-aligned evaluation suite designed to bridge the gap between AI agent capabilities and real-world productivity. While existing benchmarks often focus on isolated technical skills…

AI AgentInformation RetrievalMarketing

IndoorWorld: Integrating Physical Task Solving and Social Simulation in A Heterogeneous Multi-Agent Environment

2025-06-14 · Dekun Wu, Frederik Brudy, Bang Liu, Yi Wang

Virtual environments are essential to AI agent research. Existing environments for LLM agent research typically focus on either physical task solving or social simulation, with the former oversimplifying agent individual…

AI Agent

ADAgent: LLM Agent for Alzheimer's Disease Analysis with Collaborative Coordinator

2025-06-11 · Wenlong Hou, Guangqian Yang, Ye Du, Yeung Lau 외

Alzheimer's disease (AD) is a progressive and irreversible neurodegenerative disease. Early and precise diagnosis of AD is crucial for timely intervention and treatment planning to alleviate the progressive neurodegenera…

AI AgentLarge Language ModelPrognosis

$τ^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

2025-06-09 · Victor Barres, Honghua Dong, Soham Ray, Xujie Si 외

Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs…

AI Agent

Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce

2025-06-06 · Yijia Shao, Humishka Zope, Yucheng Jiang, Jiaxin Pei 외

The rapid rise of compound AI systems (a.k.a., AI agents) is reshaping the labor market, raising concerns about job displacement, diminished human agency, and overreliance on automation. Yet, we lack a systematic underst…

AI Agent

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

2025-06-06 · Jiachen Zhu, Menghui Zhu, Renting Rui, Rong Shan 외

The advent of large language models (LLMs), such as GPT, Gemini, and DeepSeek, has significantly advanced natural language processing, giving rise to sophisticated chatbots capable of diverse language-related tasks. The …

AI Agent

The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective

2025-06-04 · Jiin Kim, Byeongjun Shin, Jinha Chung, Minsoo Rhu

Large-language-model (LLM)-based AI agents have recently showcased impressive versatility by employing dynamic reasoning, an adaptive, multi-step process that coordinates with external tools. This shift from static, sing…

AI AgentLarge Language Model

AI Agent Behavioral Science

2025-06-04 · Lin Chen, Yunke Zhang, Jie Feng, Haoye Chai 외

Recent advances in large language models (LLMs) have enabled the development of AI agents that exhibit increasingly human-like behaviors, including planning, adaptation, and social dynamics across diverse, interactive, a…

AI AgentFairness

AI Agents for Conversational Patient Triage: Preliminary Simulation-Based Evaluation with Real-World EHR Data

2025-06-04 · Sina Rashidian, Nan Li, Jonathan Amar, Jong Ha Lee 외

Background: We present a Patient Simulator that leverages real world patient encounters which cover a broad range of conditions and symptoms to provide synthetic test subjects for development and testing of healthcare ag…

AI Agent

Comparative Analysis of AI Agent Architectures for Entity Relationship Classification

2025-06-03 · Maryam Berijanian, Kuldeep Singh, Amin Sehati

Entity relationship classification remains a challenging task in information extraction, especially in scenarios with limited labeled data and complex relational structures. In this study, we conduct a comparative analys…

AI AgentRelationRelation ClassificationRelation Extraction

ATAG: AI-Agent Application Threat Assessment with Attack Graphs

2025-06-03 · Parth Atulbhai Gandhi, Akansha Shukla, David Tayouri, Beni Ifland 외

Evaluating the security of multi-agent systems (MASs) powered by large language models (LLMs) is challenging, primarily because of the systems' complex internal dynamics and the evolving nature of LLM vulnerabilities. Tr…

AI Agent

Designing Algorithmic Delegates: The Role of Indistinguishability in Human-AI Handoff

2025-06-03 · Sophie Greenwood, Karen Levy, Solon Barocas, Hoda Heidari 외

As AI technologies improve, people are increasingly willing to delegate tasks to AI agents. In many cases, the human decision-maker chooses whether to delegate to an AI agent based on properties of the specific instance …

AI AgentDecision Making

ThinkTank: A Framework for Generalizing Domain-Specific AI Agent Systems into Universal Collaborative Intelligence Platforms

2025-06-03 · Praneet Sai Madhu Surabhi, Dheeraj Reddy Mudireddy, Jian Tao

This paper presents ThinkTank, a comprehensive and scalable framework designed to transform specialized AI agent systems into versatile collaborative intelligence platforms capable of supporting complex problem-solving a…

AI AgentRetrieval-augmented Generation

AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning

2025-06-02 · Zhong Zhang, Yaxi Lu, Yikun Fu, Yupeng Huo 외

The recent progress of large language model agents has opened new possibilities for automating tasks through graphical user interfaces (GUIs), especially in mobile environments where intelligent interaction can greatly e…

AI AgentDiversityGUI Element DetectionLarge Language Model
1–20 / 391 다음 →