paper-with-me

Panoptic Segmentation

27개 벤치마크 · 논문 505편 · 이 태스크의 논문 보기 →

Benchmarks

COCO test-dev

결과 114개

Cityscapes val

결과 111개

COCO minival

결과 93개

ADE20K val

결과 75개

Mapillary val

결과 39개

Cityscapes test

결과 30개

LaRS

결과 24개

S3DIS Area5

결과 15개

ScanNetV2

결과 15개

Indian Driving Dataset

결과 12개

PanNuke

결과 12개

ScanNet

결과 12개

PASTIS

결과 9개

COCO panoptic

결과 6개

NYU Depth v2

결과 6개

SemanticKITTI

결과 6개

ADE20K

결과 4개

DALES

결과 3개

Hypersim

결과 3개

KITTI-360

결과 3개

PASTIS-R

결과 3개

Panoptic nuScenes val

결과 3개

S3DIS

결과 3개

SUN-RGBD

결과 3개

Most implemented

Mask R-CNN

2017-03-20 · 구현 179개

ResNeSt: Split-Attention Networks

2020-04-19 · 구현 36개

Visual Attention Network

2022-02-20 · 구현 21개

Papers

GhostPoint: Self-Supervised Representation Learning by Hallucinating Occluded LiDAR Structure

2026-08-14 · Mohamed Abdelsamad, Bin Yang, Michael Ulrich, Miao Zhang 외 arxiv

3D object detection from LiDAR point clouds is a core problem in autonomous driving. Recent advances in self-supervised learning (SSL) enable scalable pretraining and transfers well to per-point tasks such as semantic an…

Self-Supervised LearningRepresentation LearningPanoptic Segmentation3D Object Detection

Extending a Large View Synthesis Model for Multi-view Panoptic Segmentation

2026-07-22 · Kwonyoung Ryu, In-Jae Lee, Jonghyun Jin, Hyunjee Lee 외 arxiv

Large view synthesis models synthesize novel views through cross-view attention without explicit 3D representations, and recent studies have shown that they learn accurate spatial correspondence from RGB supervision alon…

Panoptic SegmentationNovel View SynthesisScene Understanding3D Reconstruction

Instance-Enriched Semantic Maps for Visual Language Navigation

2026-07-14 · Jiho Hong, Eunae Kang, Sanghyun Kim, Young-Sik Shin arxiv

Visual Language Navigation (VLN) aims to enable an embodied agent to navigate complex environments by following natural language instructions. Recent approaches build semantic spatial maps and leverage Large Language Mod…

Panoptic SegmentationDecision Making

DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery

2026-07-14 · Xinyue Xu, Zheng Zhang, Kunyang Ma, Ge Zhu 외 arxiv

As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cognition capability on street view imagery for complex logical reasonin…

Visual Question AnsweringPanoptic SegmentationScene UnderstandingSpatial Reasoning

Pano3D: Unified 3D Reconstruction and Panoptic Segmentation

2026-06-12 · Victor Barberteguy, Ahmet Iscen, Mathilde Caron, Alireza Fathi 외 arxiv

Recent advances in 3D feedforward reconstruction neural networks have achieved remarkable success in dense reconstruction from images without any camera parameters. Yet, equipping these models with robust semantic unders…

Panoptic Segmentation3D Reconstruction

EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation

2026-06-08 · Runsong Zhu, Jiaxin Guo, Xiaoyang Guo, Zhengzhe Liu 외 arxiv

This paper introduces EPS3D, a new end-to-end feed-forward framework for open-vocabulary 3D panoptic segmentation. Unlike existing methods relying on additional preprocessing, we design an end-to-end architecture, with a…

Panoptic SegmentationScene Understanding3D scene Editing

전체 505편 보기 →