paper-with-me

홈 › Papers

Synthesis and Evaluation of Long-term History-aware Medical Dialogue

2026-05-19 · Hebin Hu, Renke Dai, Ah-Hwee Tan, Yilin Kang arxiv

An effective healthcare agent must be able to recall and reason over a patient's longitudinal medical history. However, the absence of datasets with realistic long-term dialogue timelines limits systematic evaluation. Real clinical text is constrained by privacy and ethics, while existing benchmarks focus on isolated interactions, failing to capture cross-session reasoning. We introduce a framework for synthesizing high-quality, long-term medical dialogues with LLMs. Our approach entails a knowledge-guided decomposition into three stages: constructing synthetic patient profiles with diverse disease and complication trajectories, generating multi-turn dialogues per encounter, and integrating them into a coherent longitudinal history dataset, MediLongChat. We establish three benchmark tasks-In-dialogue Reasoning, Cross-dialogue Reasoning, and Synthesis Reasoning-to evaluate the memory capabilities of healthcare agents. To assess data quality, we introduce a multi-dimensional evaluation framework combining vector-based metrics with LLM-as-a-judge assessments. Specifically, we define automatic measures-Faithfulness, Coherence, and Diversity-together with two LLM-based evaluations: Correctness and Realism. Benchmark experiments show that even state-of-the-art LLMs struggle with MediLongChat. These findings highlight the benchmark's applicability and underscore the need for tailored methods to advance healthcare agents.

📄 PDF Abstract BibTeX arXiv:2605.19766

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Robust Dancer: Long-term 3D Dance Synthesis Using Unpaired Data

2023-03-29 · Bin Feng, Tenglong Ao, Zequn Liu, Wei Ju 외

How to automatically synthesize natural-looking dance movements based on a piece of music is an incrementally popular yet challenging task. Most existing data-driven approaches require hard-to-get paired training data an…

Disentanglement

History-Aware Hierarchical Transformer for Multi-session Open-domain Dialogue System

2023-02-02 · Tong Zhang, Yong liu, Boyang Li, Zhiwei Zeng 외

With the evolution of pre-trained language models, current open-domain dialogue systems have achieved great progress in conducting one-session conversations. In contrast, Multi-Session Conversation (MSC), which consists …

History-Aware Reasoning for GUI Agents

2025-11-12 · Ziwei Wang, Leyang Yang, Xiaoxuan Tang, Sheng Zhou 외 arxiv

Advances in Multimodal Large Language Models have significantly enhanced Graphical User Interface (GUI) automation. Equipping GUI agents with reliable episodic reasoning capabilities is essential for bridging the gap bet…

Reinforcement Learning

History-Aware Visuomotor Policy Learning via Point Tracking

2025-09-21 · Jingjing Chen, Hongjie Fang, Chenxi Wang, Shiquan Wang 외 arxiv

Many manipulation tasks require memory beyond the current observation, yet most visuomotor policies rely on the Markov assumption and thus struggle with repeated states or long-horizon dependencies. Existing methods atte…

Computational EfficiencyPoint Tracking

Long History Short-Term Memory for Long-Term Video Prediction

2019-09-25 · Wonmin Byeon, Jan Kautz

While video prediction approaches have advanced considerably in recent years, learning to predict long-term future is challenging — ambiguous future or error propagation over time yield blurry predictions. To address thi…

Video Prediction