paper-with-me

Papers

KOLOMVERSE: Korea open large-scale image dataset for object detection in the maritime universe

2022-06-20 · Abhilasha Nanda, Sung Won Cho, Hyeopwoo Lee, Jin Hyoung Park

Over the years, datasets have been developed for various object detection tasks. Object detection in the maritime domain is essential for the safety and navigation of ships. However, there is still a lack of publicly available large-scale datasets in the maritime domain. To overcome this challenge, we present KOLOMVERSE, an open large-scale image dataset for object detection in the maritime domain by KRISO (Korea Research Institute of Ships and Ocean Engineering). We collected 5,845 hours of video data captured from 21 territorial waters of South Korea. Through an elaborate data quality assessment process, we gathered around 2,151,470 4K resolution images from the video data. This dataset considers various environments: weather, time, illumination, occlusion, viewpoint, background, wind speed, and visibility. The KOLOMVERSE consists of five classes (ship, buoy, fishnet buoy, lighthouse and wind farm) for maritime object detection. The dataset has images of 3840$\times$2160 pixels and to our knowledge, it is by far the largest publicly available dataset for object detection in the maritime domain. We performed object detection experiments and evaluated our dataset on several pre-trained state-of-the-art architectures to show the effectiveness and usefulness of our dataset. The dataset is available at: \url{https://github.com/MaritimeDataset/KOLOMVERSE}.

📄 PDF Abstract BibTeX arXiv:2206.09885

Code (0)

등록된 구현이 없습니다.

Tasks

4kObjectobject-detectionObject Detection

Similar Papers 제목 키워드 기반

Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain Texts

2025-10-28 · Seyoung Song, Nawon Kim, Songeun Chae, Kiwoong Park 외 arxiv

The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remaine…

KORMo: Korean Open Reasoning Model for Everyone

2025-10-10 · Minjun Kim, Hyeonseok Lim, Hangyeol Yoo, Inho Won 외 arxiv

This work presents the first large-scale investigation into constructing a fully open bilingual large language model (LLM) for a non-English language, specifically Korean, trained predominantly on synthetic data. We intr…

OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset

2023-01-16 · Jeongkyun Park, Jung-Wook Hwang, Kwanghee Choi, Seung-Hyun Lee 외

Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, induce dependencies with various prediction models d…

Audio-Visual Speech RecognitionLip ReadingSpeaker Recognitionspeech-recognition+2

Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs

2024-10-16 · Hyeonwoo Kim, Dahyun Kim, Jihoo Kim, Sukyung Lee 외

The Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models (LLMs), yet it has certain limitations. Notably, the disconnect between quantitative improvements on the overly academic lead…

Benchmarking

Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study

2025-09-04 · Junghwan Lim, Gangwon Jo, Sungmin Lee, Jiyoung Park 외 arxiv

We introduce Llama-3-Motif, a language model consisting of 102 billion parameters, specifically designed to enhance Korean capabilities while retaining strong performance in English. Developed on the Llama 3 architecture…