paper-with-me

홈 › Papers

Recurrence over Video Frames (RoVF) for the Re-identification of Meerkats

2024-06-18 · Mitchell Rogers, Kobe Knowles, Gaël Gendron, Shahrokh Heidari, David Arturo Soriano Valdez, Mihailo Azhar, Padriac O'Leary, Simon Eyre, Michael Witbrock, Patrice Delmas

Deep learning approaches for animal re-identification have had a major impact on conservation, significantly reducing the time required for many downstream tasks, such as well-being monitoring. We propose a method called Recurrence over Video Frames (RoVF), which uses a recurrent head based on the Perceiver architecture to iteratively construct an embedding from a video clip. RoVF is trained using triplet loss based on the co-occurrence of individuals in the video frames, where the individual IDs are unavailable. We tested this method and various models based on the DINOv2 transformer architecture on a dataset of meerkats collected at the Wellington Zoo. Our method achieves a top-1 re-identification accuracy of $49\%$, which is higher than that of the best DINOv2 model ($42\%$). We found that the model can match observations of individuals where humans cannot, and our model (RoVF) performs better than the comparisons with minimal fine-tuning. In future work, we plan to improve these models by using pre-text tasks, apply them to animal behaviour classification, and perform a hyperparameter search to optimise the models further.

📄 PDF Abstract BibTeX arXiv:2406.13002

Code (0)

등록된 구현이 없습니다.

Tasks

Triplet

Methods 이 논문이 사용한 방법론

Triplet Loss The goal of Triplet loss, in the context of Siamese Networks, is to maximize the joint probability among all score-pairs i.e. the product of all probabilities. By using its…

Similar Papers 제목 키워드 기반

INPROVF: Leveraging Large Language Models to Repair High-level Robot Controllers from Assumption Violations

2025-03-17 · Qian Meng, Jin Peng Zhou, Kilian Q. Weinberger, Hadas Kress-Gazit

This paper presents INPROVF, an automatic framework that combines large language models (LLMs) and formal methods to speed up the repair process of high-level robot controllers. Previous approaches based solely on formal…

Recurrence-in-Recurrence Networks for Video Deblurring

2022-03-12 · JoonKyu Park, Seungjun Nah, Kyoung Mu Lee

State-of-the-art video deblurring methods often adopt recurrent neural networks to model the temporal dependency between the frames. While the hidden states play key role in delivering information to the next frame, abru…

DeblurringVideo Deblurring

Recurrence without Recurrence: Stable Video Landmark Detection with Deep Equilibrium Models

2023-04-02 · CVPR 2023 1 · Paul Micaelli, Arash Vahdat, Hongxu Yin, Jan Kautz 외

Cascaded computation, whereby predictions are recurrently refined over several stages, has been a persistent theme throughout the development of landmark detection models. In this work, we show that the recently proposed…

Face Alignment

Smoothing Slot Attention Iterations and Recurrences

2025-08-07 · Rongzhen Zhao, Wenyan Yang, Juho Kannala, Joni Pajarinen arxiv

Slot Attention (SA) lies at the heart of mainstream Object-Centric Learning (OCL). Image features can be aggregated into object-level representations by SA \textit{iteratively} refining cold-start query slots. For video,…

Visual Reasoning

Health system learning achieves generalist neuroimaging models

2025-11-23 · Akhil Kondepudi, Akshay Rao, Chenhui Zhao, Yiwei Lyu 외 arxiv

Frontier artificial intelligence (AI) models, such as OpenAI's GPT-5 and Meta's DINOv3, have advanced rapidly through training on internet-scale public data, yet such systems lack access to private clinical data. Neuroim…

Visual Grounding