paper-with-me

홈 › Papers

NDD20: A large-scale few-shot dolphin dataset for coarse and fine-grained categorisation

2020-05-27 · Cameron Trotter, Georgia Atkinson, Matt Sharpe, Kirsten Richardson, A. Stephen McGough, Nick Wright, Ben Burville, Per Berggren

We introduce the Northumberland Dolphin Dataset 2020 (NDD20), a challenging image dataset annotated for both coarse and fine-grained instance segmentation and categorisation. This dataset, the first release of the NDD, was created in response to the rapid expansion of computer vision into conservation research and the production of field-deployable systems suited to extreme environmental conditions -- an area with few open source datasets. NDD20 contains a large collection of above and below water images of two different dolphin species for traditional coarse and fine-grained segmentation. All data contained in NDD20 was obtained via manual collection in the North Sea around the Northumberland coastline, UK. We present experimentation using standard deep learning network architecture trained using NDD20 and report baselines results.

📄 PDF Abstract BibTeX arXiv:2005.13359

Code (1)

robustsam/RobustSAM pytorch

Tasks

Instance SegmentationSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

Dolphin: A Large-Scale Automatic Speech Recognition Model for Eastern Languages

2025-03-26 · Yangyang Meng, Jinpeng Li, Guodong Lin, Yu Pu 외

This report introduces Dolphin, a large-scale multilingual automatic speech recognition (ASR) model that extends the Whisper architecture to support a wider range of languages. Our approach integrates in-house proprietar…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Dolphin v1.0 Technical Report

2025-09-30 · Taohan Weng, Kaibing Hu, Henan Liu, Siya Liu 외 arxiv

Ultrasound is crucial in modern medicine but faces challenges like operator dependence, image noise, and real-time scanning, hindering AI integration. While large multimodal models excel in other medical imaging areas, t…

Reinforcement Learning

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting

2025-05-20 · Hao Feng, Shu Wei, Xiang Fei, Wei Shi 외

Document image parsing is challenging due to its complexly intertwined elements such as text paragraphs, figures, formulas, and tables. Current approaches either assemble specialized expert models or directly generate pa…

DOLPHINS: Dataset for Collaborative Perception enabled Harmonious and Interconnected Self-driving

2022-07-15 · Ruiqing Mao, Jingyu Guo, Yukuan Jia, Yuxuan Sun 외

Vehicle-to-Everything (V2X) network has enabled collaborative perception in autonomous driving, which is a promising solution to the fundamental defect of stand-alone intelligence including blind zones and long-range per…

Autonomous DrivingObject Detection

Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention

2025-09-28 · Kai Li, Kejun Gao, Xiaolin Hu arxiv

Audio-visual speech separation (AVSS) methods leverage visual cues to extract target speech and have demonstrated strong separation quality in noisy acoustic environments. However, these methods usually involve a large n…

Speech Separation