paper-with-me

Papers

MUVA: A New Large-Scale Benchmark for Multi-View Amodal Instance Segmentation in the Shopping Scenario

2023-01-01 · ICCV 2023 1 · Zhixuan Li, Weining Ye, Juan Terven, Zachary Bennett, Ying Zheng, Tingting Jiang, Tiejun Huang

Amodal Instance Segmentation (AIS) endeavors to accurately deduce complete object shapes that are partially or fully occluded. However, the inherent ill-posed nature of single-view datasets poses challenges in determining occluded shapes. A multi-view framework may help alleviate this problem, as humans often adjust their perspective when encountering occluded objects. At present, this approach has not yet been explored by existing methods and datasets. To bridge this gap, we propose a new task called Multi-view Amodal Instance Segmentation (MAIS) and introduce the MUVA dataset, the first MUlti-View AIS dataset that takes the shopping scenario as instantiation. MUVA provides comprehensive annotations, including multi-view amodal/visible segmentation masks, 3D models, and depth maps, making it the largest image-level AIS dataset in terms of both the number of images and instances. Additionally, we propose a new method for aggregating representative features across different instances and views, which demonstrates promising results in accurately predicting occluded objects from one viewpoint by leveraging information from other viewpoints. Besides, we also demonstrate that MUVA can benefit the AIS task in real-world scenarios.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Amodal Instance SegmentationInstance SegmentationSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models

2024-08-31 · Yuxiang Guo, Faizan Siddiqui, Yang Zhao, Rama Chellappa 외

Predicting and reasoning how a video would make a human feel is crucial for developing socially intelligent systems. Although Multimodal Large Language Models (MLLMs) have shown impressive video understanding capabilitie…

Video Understanding

MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild

2026-06-15 · Haotian Qi, Gabriel Skantze arxiv

Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human-robot interaction scenarios. We introduce MuVAP, a causal multimodal framework t…

MuVAM: A Multi-View Attention-based Model for Medical Visual Question Answering

2021-07-07 · Haiwei Pan, Shuning He, Kejia Zhang, Bo Qu 외

Medical Visual Question Answering (VQA) is a multi-modal challenging task widely considered by research communities of the computer vision and natural language processing. Since most current medical VQA models focus on v…

Medical Visual Question AnsweringMissing LabelsQuestion AnsweringVisual Question Answering+1

FTDMamba: Frequency-Assisted Temporal Dilation Mamba for Unmanned Aerial Vehicle Video Anomaly Detection

2026-01-16 · Cheng-Zhuang Liu, Si-Bao Chen, Qing-Ling Shu, Chris Ding 외 arxiv

Recent advances in video anomaly detection (VAD) mainly focus on ground-based surveillance or unmanned aerial vehicle (UAV) videos with static backgrounds, whereas research on UAV videos with dynamic backgrounds remains …

Video Anomaly Detection

Multi-View Attention Learning for Residual Disease Prediction of Ovarian Cancer

2023-06-26 · Xiangneng Gao, Shulan Ruan, Jun Shi, Guoqing Hu 외

In the treatment of ovarian cancer, precise residual disease prediction is significant for clinical and surgical decision-making. However, traditional methods are either invasive (e.g., laparoscopy) or time-consuming (e.…

Computed Tomography (CT)Decision MakingDisease PredictionPrediction