paper-with-me

홈 › Papers

Trip+: Benchmarking Agents in Personalized Interactive Travel Planning

2026-06-19 · Junle Chen, Wei Chen, Yehong Xu, Zhengjun Huang, Yuqian Wu, Zhoujin Tian, Kai Wang, Lei Wang, Xiaofang Zhou arxiv

Interactive travel planning has become a popular use case for language models. Agents are deployed to manage evolving preferences and unexpected disruptions over multiple turns. Such settings require models to make complex, profile-conditioned planning decisions. However, existing benchmarks often evaluate feasibility, personalization, or interaction in relatively isolated settings. We therefore introduce Trip+ to measure the ability of agents to plan travel holistically. In Trip+, given traveler profiles and dynamic interactions, agents must generate and revise minute-level itineraries. End-to-end traveler experiences are evaluated via an LLM-based simulator, enabling the assessment of subjective metrics like fatigue. Our scenarios range from simple request resolutions to complex environment-driven replanning. We evaluate 18 LMs and find a consistent gap in experiential quality. Models favor technically feasible but exhausting itineraries that diverge sharply from profiled traveler preferences.

📄 PDF Abstract BibTeX arXiv:2606.21169

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TripTailor: A Real-World Benchmark for Personalized Travel Planning

2025-08-02 · Yuanzhe Shen, Kaimin Wang, Changze Lv, Xiaoqing Zheng 외 arxiv

The continuous evolution and enhanced reasoning capabilities of large language models (LLMs) have elevated their role in complex tasks, notably in travel planning, where demand for personalized, high-quality itineraries …

Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory

2026-08-27 · Zihao Cheng, Yingyu Shan, Hongru Wang, Zeming Liu 외 arxiv

Travel planning agents assist users in generating personalized travel plans by modeling their individual preferences. Existing agents either rely on explicit user instructions or engage in multi-turn clarification to eli…

Reinforcement Learning

TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning

2025-02-27 · Soumyabrata Chaudhuri, Pranav Purkar, Ritwik Raghav, Shubhojit Mallick 외

Recent advancements in probing Large Language Models (LLMs) have explored their latent potential as personalized travel planning agents, yet existing benchmarks remain limited in real world applicability. Existing datase…

Scheduling

TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios

2026-02-02 · Yuanzhe Shen, Zisu Huang, Zhengyuan Wang, Muzhao Tian 외 arxiv

As LLM-based agents are deployed in increasingly complex real-world settings, existing benchmarks underrepresent key challenges such as enforcing global constraints, coordinating multi-tool reasoning, and adapting to evo…

Reinforcement Learning

SynthTRIPs: A Knowledge-Grounded Framework for Benchmark Query Generation for Personalized Tourism Recommenders

2025-04-12 · Ashmi Banerjee, Adithi Satish, Fitri Nur Aisyah, Wolfgang Wörndl 외

Tourism Recommender Systems (TRS) are crucial in personalizing travel experiences by tailoring recommendations to users' preferences, constraints, and contextual factors. However, publicly available travel datasets often…

HallucinationRecommendation Systems