paper-with-me

홈 › Papers

Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search

2026-03-25 · Hadar Peer, Carlos Hernandez, Sven Koenig, Ariel Felner, Oren Salzman arxiv

Empirical evaluation in multi-objective search (MOS) has historically suffered from fragmentation, relying on heterogeneous problem instances with incompatible objective definitions that make cross-study comparisons difficult. This standardization gap is further exacerbated by the realization that DIMACS road networks, a historical default benchmark for the field, exhibit highly correlated objectives that fail to capture diverse Pareto-front structures. To address this, we introduce the first comprehensive, standardized benchmark suite for exact and approximate MOS. Our suite spans four structurally diverse domains: real-world road networks, structured synthetic graphs, game-based grid environments, and high-dimensional robotic motion-planning roadmaps. By providing fixed graph instances, standardized start-goal queries, and both exact and approximate reference Pareto-optimal solution sets, this suite captures a full spectrum of objective interactions: from strongly correlated to strictly independent. Ultimately, this benchmark provides a common foundation to ensure future MOS evaluations are robust, reproducible, and structurally comprehensive.

📄 PDF Abstract BibTeX arXiv:2603.24084

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bridging the Knowledge-Action Gap by Evaluating LLMs in Dynamic Dental Clinical Scenarios

2026-01-19 · Hongyang Ma, Tiantian Gu, Huaiyuan Sun, Huilin Zhu 외 arxiv

The transition of Large Language Models (LLMs) from passive knowledge retrievers to autonomous clinical agents demands a shift in evaluation-from static accuracy to dynamic behavioral reliability. To explore this boundar…

Tracking the Trackers: An Analysis of the State of the Art in Multiple Object Tracking

2017-04-10 · Laura Leal-Taixé, Anton Milan, Konrad Schindler, Daniel Cremers 외

Standardized benchmarks are crucial for the majority of computer vision applications. Although leaderboards and ranking tables should not be over-claimed, benchmarks often provide the most objective measure of performanc…

Multiple Object TrackingMultiple People TrackingObjectObject Tracking

Bridging Resolution: A Survey of the State of the Art

2020-12-01 · COLING 2020 8 · Hideo Kobayashi, Vincent Ng

Bridging reference resolution is an anaphora resolution task that is arguably more challenging and less studied than entity coreference resolution. Given that significant progress has been made on coreference resolution …

coreference-resolutionCoreference ResolutionSurvey

Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks

2025-09-24 · Hailay Kidu Teklehaymanot, Gebrearegawi Gidey, Wolfgang Nejdl arxiv

Despite advances in Neural Machine Translation (NMT), low-resource languages like Tigrinya remain underserved due to persistent challenges, including limited corpora, inadequate tokenization strategies, and the lack of s…

Machine TranslationTransfer Learning

MOT16: A Benchmark for Multi-Object Tracking

2016-03-02 · Anton Milan, Laura Leal-Taixe, Ian Reid, Stefan Roth 외

Standardized benchmarks are crucial for the majority of computer vision applications. Although leaderboards and ranking tables should not be over-claimed, benchmarks often provide the most objective measure of performanc…

Multi-Object TrackingMultiple Object TrackingMultiple People TrackingObject+1