paper-with-me

홈 › Papers

Discriminative Spatial-Semantic VOS Solution: 1st Place Solution for 6th LSVOS

2024-08-29 · Deshui Miao, Yameng Gu, Xin Li, Zhenyu He, YaoWei Wang, Ming-Hsuan Yang

Video object segmentation (VOS) is a crucial task in computer vision, but current VOS methods struggle with complex scenes and prolonged object motions. To address these challenges, the MOSE dataset aims to enhance object recognition and differentiation in complex environments, while the LVOS dataset focuses on segmenting objects exhibiting long-term, intricate movements. This report introduces a discriminative spatial-temporal VOS model that utilizes discriminative object features as query representations. The semantic understanding of spatial-semantic modules enables it to recognize object parts, while salient features highlight more distinctive object characteristics. Our model, trained on extensive VOS datasets, achieved first place (\textbf{80.90\%} $\mathcal{J \& F}$) on the test set of the 6th LSVOS challenge in the VOS Track, demonstrating its effectiveness in tackling the aforementioned challenges. The code will be available at \href{https://github.com/yahooo-m/VOS-Solution}{code}.

📄 PDF Abstract BibTeX arXiv:2408.16431

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectObject RecognitionSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
VOS VOS is a type of video object segmentation model consisting of two network components. The target appearance model consists of a light-weight module, which is learned during…

Similar Papers 제목 키워드 기반

Distance Based Image Classification: A solution to generative classification's conundrum?

2022-10-04 · Wen-Yan Lin, Siying Liu, Bing Tian Dai, Hongdong Li

Most classifiers rely on discriminative boundaries that separate instances of each class from everything else. We argue that discriminative boundaries are counter-intuitive as they define semantics by what-they-are-not; …

image-classificationImage Classification

Bidirectional Cross-Attention Fusion of High-Resolution RGB and Low-Resolution Hyperspectral Inputs for Multimodal Semantic Segmentation

2026-03-14 · Jonas V. Funk, Lukas Roming, Andreas Michel, Paul Bäcker 외 arxiv

Multimodal semantic segmentation with heterogeneous sensors must reconcile complementary information across modalities that differ in spatial resolution and channel dimensionality. In particular, high-resolution RGB imag…

Semantic Segmentation

SimpleMatch: A Simple and Strong Baseline for Semantic Correspondence

2026-01-18 · Hailing Jin, Huiying Li arxiv

Recent advances in semantic correspondence have been largely driven by the use of pre-trained large-scale models. However, a limitation of these approaches is their dependence on high-resolution input images to achieve o…

Semantic correspondence

Dense semantic labeling of sub-decimeter resolution images with convolutional neural networks

2016-08-02 · Michele Volpi, Devis Tuia

Semantic labeling (or pixel-level land-cover classification) in ultra-high resolution imagery (< 10cm) requires statistical models able to learn high level concepts from spatial data, with large appearance variations. Co…

General ClassificationLand Cover ClassificationSuperpixels

VoxSegNet: Volumetric CNNs for Semantic Part Segmentation of 3D Shapes

2018-09-01 · Zongji Wang, Feng Lu

Voxel is an important format to represent geometric data, which has been widely used for 3D deep learning in shape analysis due to its generalization ability and regular data format. However, fine-grained tasks like part…

Segmentation