paper-with-me

Scene Segmentation

6개 벤치마크 · 논문 322편 · 이 태스크의 논문 보기 →

Benchmarks

SUN-RGBD

결과 15개

ScanNet

결과 9개

StreetHazards

결과 9개

MovieNet

결과 6개

NYU Depth v2

결과 3개

UAVid

결과 3개

Most implemented

Point Transformer

2020-12-16 · 구현 24개

Papers

SurgAM: Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy

2026-07-05 · Lei Song, Yonghao Long, Mengya Xu, Jiayi Geng 외 arxiv

Surgical automation is being increasingly studied, yet bridging visual scene understanding with autonomous action planning remains a fundamental challenge. While much research effort has been made on scene perception (e.…

Scene UnderstandingScene Segmentation

GeoSAM-3D: Geodesic Prompt Propagation for Open-Vocabulary 3D Scene Segmentation from Monocular Video

2026-05-30 · Arun Sharma arxiv

Open-vocabulary 3D scene segmentation usually assumes RGB-D video, calibrated multi-view imagery, or a reconstructed mesh. GeoSAM-3D studies a lighter setting: a user uploads a short monocular video, clicks or names an o…

Scene Segmentation

MPerS: Dynamic MLLM MixExperts Perception-Guided Remote Sensing Scene Segmentation

2026-05-11 · Ziyi Wang, Xianping Ma, Ziyao Wang, Hongyang Zhang 외 arxiv

The multimodal fusion of images and scene captions has been extensively explored and applied in various fields. However, when dealing with complex remote sensing (RS) scenes, existing studies have predominantly concentra…

Semantic SegmentationScene Segmentation

Qwen3.5-Omni Technical Report

2026-04-17 · Qwen Team arxiv

In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and suppor…

Scene SegmentationVisual GroundingSpeech Synthesis

TreeGaussian: Tree-Guided Cascaded Contrastive Learning for Hierarchical Consistent 3D Gaussian Scene Segmentation and Understanding

2026-03-31 · Jingbin You, Zehao Li, Hao Jiang, Xinzhu Ma 외 arxiv

3D Gaussian Splatting (3DGS) has emerged as a real-time, differentiable representation for neural scene understanding. However, existing 3DGS-based methods struggle to represent hierarchical 3D semantic structures and ca…

Contrastive LearningScene UnderstandingScene Segmentation

V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models

2026-03-29 · Xinying Lin, Xuyang Liu, Yiyu Wang, Teng Ma 외 arxiv

Video large language models (VideoLLMs) show strong capability in video understanding, yet long-context inference is still dominated by massive redundant visual tokens in the prefill stage. We revisit token compression f…

Scene Segmentation

전체 322편 보기 →