paper-with-me

홈 › Papers

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

2026-05-24 · Xiang Cheng, Yulan Hu, Lulu Zheng, Zheng Pan, Xin Li, Yong Liu arxiv

Travel planning in the real world is overwhelmingly a \textit{group} activity, yet existing LLM travel-planning benchmarks reduce it to a single user, where the field is approaching saturation. This single-user assumption sidesteps what makes group planning hard for an agent: discovering private preferences across multiple users, surfacing conflicts, and balancing utility against fairness. To bring the task back to its multi-user reality, we introduce \textbf{\textit{GroupTravelBench}}, the first benchmark for \textbf{multi-user, multi-turn} travel planning. Built from real user profiles, POI data, and ticket prices, it comprises 650 tasks across three difficulty levels, each running in a synchronous group-chat sandbox with cached tool data for reproducible offline evaluation. Beyond the multi-step reasoning and tool use that single-user benchmarks already test, GroupTravelBench probes three group-specific capabilities: \textit{(i) elicitation} of private preferences through multi-turn dialogue; \textit{(ii) coordination} of inter-user conflicts via compromise or subgrouping; and \textit{(iii) planning} that balances group utility against fairness. We pair this with a complementary evaluation framework combining rule-based outcome metrics and LLM-judge process metrics. Across a wide range of frontier models, even the strongest agents fall short on all four rule-based outcome metrics, with plan validity below 12\%, suggesting that group-level outcome quality is a key open challenge for LLM travel-planning agents.

📄 PDF Abstract BibTeX arXiv:2605.25200

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Trip+: Benchmarking Agents in Personalized Interactive Travel Planning

2026-06-19 · Junle Chen, Wei Chen, Yehong Xu, Zhengjun Huang 외 arxiv

Interactive travel planning has become a popular use case for language models. Agents are deployed to manage evolving preferences and unexpected disruptions over multiple turns. Such settings require models to make compl…

Towards Full Delegation: Designing Ideal Agentic Behaviors for Travel Planning

2024-11-21 · Song Jiang, Da Ju, Andrew Cohen, Sasha Mitts 외

How are LLM-based agents used in the future? While many of the existing work on agents has focused on improving the performance of a specific family of objective and challenging tasks, in this work, we take a different p…

AI Tour Meeting: Group Travel Planning by LLM Agents

2026-07-21 · Daisuke Kikuta hf

This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)-based agents. The agents are instantiated with distinct personas and collaboratively seek an itinerary…

Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory

2026-08-27 · Zihao Cheng, Yingyu Shan, Hongru Wang, Zeming Liu 외 arxiv

Travel planning agents assist users in generating personalized travel plans by modeling their individual preferences. Existing agents either rely on explicit user instructions or engage in multi-turn clarification to eli…

Reinforcement Learning

TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning

2025-02-27 · Soumyabrata Chaudhuri, Pranav Purkar, Ritwik Raghav, Shubhojit Mallick 외

Recent advancements in probing Large Language Models (LLMs) have explored their latent potential as personalized travel planning agents, yet existing benchmarks remain limited in real world applicability. Existing datase…

Scheduling