paper-with-me

홈 › Papers

APTv2: Benchmarking Animal Pose Estimation and Tracking with a Large-scale Dataset and Beyond

2023-12-25 · Yuxiang Yang, Yingqi Deng, Yufei Xu, Jing Zhang

Animal Pose Estimation and Tracking (APT) is a critical task in detecting and monitoring the keypoints of animals across a series of video frames, which is essential for understanding animal behavior. Past works relating to animals have primarily focused on either animal tracking or single-frame animal pose estimation only, neglecting the integration of both aspects. The absence of comprehensive APT datasets inhibits the progression and evaluation of animal pose estimation and tracking methods based on videos, thereby constraining their real-world applications. To fill this gap, we introduce APTv2, the pioneering large-scale benchmark for animal pose estimation and tracking. APTv2 comprises 2,749 video clips filtered and collected from 30 distinct animal species. Each video clip includes 15 frames, culminating in a total of 41,235 frames. Following meticulous manual annotation and stringent verification, we provide high-quality keypoint and tracking annotations for a total of 84,611 animal instances, split into easy and hard subsets based on the number of instances that exists in the frame. With APTv2 as the foundation, we establish a simple baseline method named \posetrackmethodname and provide benchmarks for representative models across three tracks: (1) single-frame animal pose estimation track to evaluate both intra- and inter-domain transfer learning performance, (2) low-data transfer and generalization track to evaluate the inter-species domain generalization performance, and (3) animal pose tracking track. Our experimental results deliver key empirical insights, demonstrating that APTv2 serves as a valuable benchmark for animal pose estimation and tracking. It also presents new challenges and opportunities for future research. The code and dataset are released at \href{https://github.com/ViTAE-Transformer/APTv2}{https://github.com/ViTAE-Transformer/APTv2}.

📄 PDF Abstract BibTeX arXiv:2312.15612

Code (1)

vitae-transformer/aptv2 공식 구현

Tasks

Animal Pose EstimationBenchmarkingDomain GeneralizationPose EstimationPose TrackingTransfer Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

APT-36K: A Large-scale Benchmark for Animal Pose Estimation and Tracking

2022-06-12 · Yuxiang Yang, Junjie Yang, Yufei Xu, Jing Zhang 외

Animal pose estimation and tracking (APT) is a fundamental task for detecting and tracking animal keypoints from a sequence of video frames. Previous animal-related datasets focus either on animal tracking or single-fram…

Animal Pose EstimationDomain GeneralizationPose EstimationTransfer Learning

SAM 3D Animal: Promptable Animal 3D Reconstruction from Images in the Wild

2026-05-08 · Xuyi Hu, Jin Lyu, Jiuming Liu, Yebin Liu 외 arxiv

3D animal reconstruction in the wild remains challenging due to large species variation, frequent occlusions, and the prevalence of multi-animal scenes, while existing methods predominantly focus on single-animal setting…

3D Reconstruction

Multi-animal pose estimation, identification and tracking with DeepLabCut

2022-04-12 · Nature Methods 2022 4 · Jessy Lauer, Mu Zhou, Shaokai Ye, William Menegas 외

Estimating the pose of multiple animals is a challenging computer vision problem: frequent interactions cause occlusions and complicate the association of detected keypoints to the correct individuals, as well as having …

Animal Pose EstimationPose Estimation

AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition

2025-12-03 · Zichuan Lin, Yicheng Liu, Yang Yang, Lvfang Tao 외 arxiv

Vision-Language Models (VLMs) have achieved remarkable success in visual question answering tasks, but their reliance on large numbers of visual tokens introduces significant computational overhead. While existing effici…

Visual Question AnsweringReinforcement Learning

BuckTales : A multi-UAV dataset for multi-object tracking and re-identification of wild antelopes

2024-11-11 · Hemal Naik, Junran Yang, Dipin Das, Margaret C Crofoot 외

Understanding animal behaviour is central to predicting, understanding, and mitigating impacts of natural and anthropogenic changes on animal populations and ecosystems. However, the challenges of acquiring and processin…

BenchmarkingMulti-Object TrackingObject Tracking