paper-with-me

홈 › Papers

An Open-World, Diverse, Cross-Spatial-Temporal Benchmark for Dynamic Wild Person Re-Identification

2024-03-22 · Lei Zhang, Xiaowei Fu, Fuxiang Huang, Yi Yang, Xinbo Gao

Person re-identification (ReID) has made great strides thanks to the data-driven deep learning techniques. However, the existing benchmark datasets lack diversity, and models trained on these data cannot generalize well to dynamic wild scenarios. To meet the goal of improving the explicit generalization of ReID models, we develop a new Open-World, Diverse, Cross-Spatial-Temporal dataset named OWD with several distinct features. 1) Diverse collection scenes: multiple independent open-world and highly dynamic collecting scenes, including streets, intersections, shopping malls, etc. 2) Diverse lighting variations: long time spans from daytime to nighttime with abundant illumination changes. 3) Diverse person status: multiple camera networks in all seasons with normal/adverse weather conditions and diverse pedestrian appearances (e.g., clothes, personal belongings, poses, etc.). 4) Protected privacy: invisible faces for privacy critical applications. To improve the implicit generalization of ReID, we further propose a Latent Domain Expansion (LDE) method to develop the potential of source data, which decouples discriminative identity-relevant and trustworthy domain-relevant features and implicitly enforces domain-randomized identity feature space expansion with richer domain diversity to facilitate domain invariant representations. Our comprehensive evaluations with most benchmark datasets in the community are crucial for progress, although this work is far from the grand goal toward open-world and dynamic wild applications.

📄 PDF Abstract BibTeX arXiv:2403.15119

Code (1)

fxw13/OWD 공식 구현

Tasks

DiversityPerson Re-Identification

Similar Papers 제목 키워드 기반

UrbanDiT: A Foundation Model for Open-World Urban Spatio-Temporal Learning

2024-11-19 · Yuan Yuan, Chonghua Han, Jingtao Ding, Depeng Jin 외

The urban environment is characterized by complex spatio-temporal dynamics arising from diverse human activities and interactions. Effectively modeling these dynamics is essential for understanding and optimizing urban s…

ImputationMulti-Task LearningPrompt Learning

SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning

2025-12-18 · Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax arxiv

Autonomous robotic systems require spatio-temporal understanding of dynamic environments to ensure reliable navigation and interaction. While Vision-Language Models (VLMs) provide open-world semantic priors, they lack gr…

Scene UnderstandingPoint Clouds

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?

2026-07-23 · Han Li, Si Liu, Zehao Huang, Dongxin Lyu 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturally develop through continuous observation…

Spatial-Temporal Transformer Networks for Traffic Flow Forecasting

2020-01-09 · Mingxing Xu, Wenrui Dai, Chunmiao Liu, Xing Gao 외

Traffic forecasting has emerged as a core component of intelligent transportation systems. However, timely accurate traffic forecasting, especially long-term forecasting, still remains an open challenge due to the highly…

Traffic Prediction

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

2024-12-13 · Ruijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao 외

Although large vision-language-action (VLA) models pretrained on extensive robot datasets offer promising generalist policies for robotic learning, they still struggle with spatial-temporal dynamics in interactive roboti…

Robot ManipulationVision-Language-Action