paper-with-me

홈 › Papers

Follow Anything: Open-set detection, tracking, and following in real-time

2023-08-10 · Alaa Maalouf, Ninad Jadhav, Krishna Murthy Jatavallabhula, Makram Chahine, Daniel M. Vogt, Robert J. Wood, Antonio Torralba, Daniela Rus

Tracking and following objects of interest is critical to several robotics use cases, ranging from industrial automation to logistics and warehousing, to healthcare and security. In this paper, we present a robotic system to detect, track, and follow any object in real-time. Our approach, dubbed ``follow anything'' (FAn), is an open-vocabulary and multimodal model -- it is not restricted to concepts seen at training time and can be applied to novel classes at inference time using text, images, or click queries. Leveraging rich visual descriptors from large-scale pre-trained models (foundation models), FAn can detect and segment objects by matching multimodal queries (text, images, clicks) against an input image sequence. These detected and segmented objects are tracked across image frames, all while accounting for occlusion and object re-emergence. We demonstrate FAn on a real-world robotic system (a micro aerial vehicle) and report its ability to seamlessly follow the objects of interest in a real-time control loop. FAn can be deployed on a laptop with a lightweight (6-8 GB) graphics card, achieving a throughput of 6-20 frames per second. To enable rapid adoption, deployment, and extensibility, we open-source all our code on our project webpage at https://github.com/alaamaalouf/FollowAnything . We also encourage the reader to watch our 5-minutes explainer video in this https://www.youtube.com/watch?v=6Mgt3EPytrw .

📄 PDF Abstract BibTeX arXiv:2308.05737

Code (1)

alaamaalouf/followanything 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

2024-12-20 · Jiaming Ji, Jiayi Zhou, Hantao Lou, Boyuan Chen 외

Reinforcement learning from human feedback (RLHF) has proven effective in enhancing the instruction-following capabilities of large language models; however, it remains underexplored in the cross-modality domain. As the …

AllInstruction Following

ReSurgSAM2: Referring Segment Anything in Surgical Video via Credible Long-term Tracking

2025-05-13 · Haofeng Liu, Mingqi Gao, Xuxiao Luo, Ziyue Wang 외

Surgical scene segmentation is critical in computer-assisted surgery and is vital for enhancing surgical quality and patient outcomes. Recently, referring surgical segmentation is emerging, given its advantage of providi…

DiversityMambaScene SegmentationSegmentation

RewardAnything: Generalizable Principle-Following Reward Models

2025-06-04 · Zhuohao Yu, Jiali Zeng, Weizheng Gu, Yidong Wang 외

Reward Models, essential for guiding Large Language Model optimization, are typically trained on fixed preference datasets, resulting in rigid alignment to single, implicit preference distributions. This prevents adaptat…

Instruction FollowingLarge Language ModelModel Optimization

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

2025-06-14 · Yinghao Ma, Siyou Li, Juntao Yu, Emmanouil Benetos 외

Recent advances in audio-text large language models (LLMs) have opened new possibilities for music understanding and generation. However, existing benchmarks are limited in scope, often relying on simplified tasks or mul…

Beat TrackingGenre classificationInformation RetrievalInstruction Following+5

Vector Fields for Path Following on Lie Groups with Application in Robot Control

2026-02-25 · Felipe Bartelt, Luciano C. A. Pimenta, Weijia Yao, Vinicius M. Gonçalves arxiv

Many robotic systems allow independent control of position and orientation (pose), including omnidirectional aerial vehicles, underwater robots, and manipulator end-effectors. In many applications, these systems must fol…