paper-with-me

홈 › Papers

DeepSea MOT: A benchmark dataset for multi-object tracking on deep-sea video

2025-09-03 · Kevin Barnard, Elaine Liu, Kristine Walz, Brian Schlining, Nancy Jacobsen Stout, Lonny Lundsten arxiv

Benchmarking multi-object tracking and object detection model performance is an essential step in machine learning model development, as it allows researchers to evaluate model detection and tracker performance on human-generated 'test' data, facilitating consistent comparisons between models and trackers and aiding performance optimization. In this study, a novel benchmark video dataset was developed and used to assess the performance of several Monterey Bay Aquarium Research Institute object detection models and a FathomNet single-class object detection model together with several trackers. The dataset consists of four video sequences representing midwater and benthic deep-sea habitats. Performance was evaluated using Higher Order Tracking Accuracy, a metric that balances detection, localization, and association accuracy. To the best of our knowledge, this is the first publicly available benchmark for multi-object tracking in deep-sea video footage. We provide the benchmark data, a clearly documented workflow for generating additional benchmark videos, as well as example Python notebooks for computing metrics.

📄 PDF Abstract BibTeX arXiv:2509.03499

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Object TrackingObject Detection

Similar Papers 제목 키워드 기반

VistaHop: Benchmarking Long-Horizon Visual DeepSearch

2026-06-02 · Hang He, Chuhuai Yue, Chengqi Dong, Chengcheng Wan 외 arxiv

Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evidence, and connecting fine-grained clues…

Question AnsweringVisual GroundingVisual ReasoningImage Cropping

DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents

2026-01-28 · Nikita Gupta, Riju Chatterjee, Lukas Haas, Connie Tao 외 arxiv

We introduce DeepSearchQA, a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional benchmarks that target single answer retrieval or b…

Entity Resolution

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

2026-05-09 · Tao Yu, yiming ding, Shenghua Chai, Minghui Zhang 외 arxiv

Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start from audio alone and actively search for cross-modal evidence remains …

Towards Less Biased Data-driven Scoring with Deep Learning-Based End-to-end Database Search in Tandem Mass Spectrometry

2024-05-08 · Yonghan Yu, Ming Li

Peptide identification in mass spectrometry-based proteomics is crucial for understanding protein function and dynamics. Traditional database search methods, though widely used, rely on heuristic scoring functions and st…

Contrastive LearningDecoder

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

2026-07-08 · Xinyu Geng, Xuanhua He, Sixiang Chen, Yanjing Xiao 외 arxiv

Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak superv…

Reinforcement Learning