paper-with-me

Papers

Evaluating Agent Interactions Through Episodic Knowledge Graphs

2022-09-22 · CCGPK (COLING) 2022 10 · Selene Báez Santamaría, Piek Vossen, Thomas Baier

We present a new method based on episodic Knowledge Graphs (eKGs) for evaluating (multimodal) conversational agents in open domains. This graph is generated by interpreting raw signals during conversation and is able to capture the accumulation of knowledge over time. We apply structural and semantic analysis of the resulting graphs and translate the properties into qualitative measures. We compare these measures with existing automatic and manual evaluation metrics commonly used for conversational agents. Our results show that our Knowledge-Graph-based evaluation provides more qualitative insights into interaction and the agent's behavior.

📄 PDF Abstract BibTeX arXiv:2209.11746

Code (1)

selbaez/evaluating-conversations-as-ekg 공식 구현

Tasks

Knowledge Graphs

Similar Papers 제목 키워드 기반

Evaluating Long-Term Memory for Long-Context Question Answering

2025-10-27 · Alessandra Terranova, Björn Ross, Alexandra Birch arxiv

In order for large language models to achieve true conversational continuity and benefit from experiential learning, they need memory. While research has focused on the development of complex memory systems, it remains u…

Question Answering

SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment

2026-04-10 · Sihang Jiang, Lipeng Ma, Zhonghua Hong, Keyi Wang 외 arxiv

Current LLM-based agents demonstrate strong performance in episodic task execution but remain constrained by static toolsets and episodic amnesia, failing to accumulate experience across task boundaries. This paper forma…

If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs

2025-03-30 · Siqi Fan, Xiusheng Huang, Yiqun Yao, Xuezhi Fang 외

Large language models (LLMs) can carry out human-like dialogue, but unlike humans, they are stateless due to the superposition property. However, during multi-turn, multi-agent interactions, LLMs begin to exhibit consist…

Fact CheckingLifelong learning

Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions

2026-05-25 · Jeongeun Lee, Chanyoung Park, Dongha Lee arxiv

Multimodal large language model (MLLM)-based embodied agents have shown strong potential for solving complex tasks in physical environments. However, personalized assistance requires more than following generic instructi…

EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents

2026-01-23 · Xinze Li, Ziyue Zhu, Siyuan Liu, Yubo Ma 외 arxiv

We introduce EMemBench, a programmatic benchmark generator for evaluating long-term episodic memory of agents through interactive games. Rather than using a fixed set of questions, EMemBench generates questions from envi…

Spatial Reasoning