Scene Segmentation
6개 벤치마크 · 논문 322편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation
SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation
Fully Convolutional Networks for Semantic Segmentation
Point Transformer
Dual Attention Network for Scene Segmentation
KPConv: Flexible and Deformable Convolution for Point Clouds
Papers
SurgAM: Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy
Surgical automation is being increasingly studied, yet bridging visual scene understanding with autonomous action planning remains a fundamental challenge. While much research effort has been made on scene perception (e.…
Scene UnderstandingScene SegmentationGeoSAM-3D: Geodesic Prompt Propagation for Open-Vocabulary 3D Scene Segmentation from Monocular Video
Open-vocabulary 3D scene segmentation usually assumes RGB-D video, calibrated multi-view imagery, or a reconstructed mesh. GeoSAM-3D studies a lighter setting: a user uploads a short monocular video, clicks or names an o…
Scene SegmentationMPerS: Dynamic MLLM MixExperts Perception-Guided Remote Sensing Scene Segmentation
The multimodal fusion of images and scene captions has been extensively explored and applied in various fields. However, when dealing with complex remote sensing (RS) scenes, existing studies have predominantly concentra…
Semantic SegmentationScene SegmentationQwen3.5-Omni Technical Report
In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and suppor…
Scene SegmentationVisual GroundingSpeech SynthesisTreeGaussian: Tree-Guided Cascaded Contrastive Learning for Hierarchical Consistent 3D Gaussian Scene Segmentation and Understanding
3D Gaussian Splatting (3DGS) has emerged as a real-time, differentiable representation for neural scene understanding. However, existing 3DGS-based methods struggle to represent hierarchical 3D semantic structures and ca…
Contrastive LearningScene UnderstandingScene SegmentationV-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
Video large language models (VideoLLMs) show strong capability in video understanding, yet long-context inference is still dominated by massive redundant visual tokens in the prefill stage. We revisit token compression f…
Scene Segmentation