paper-with-me

홈 › Papers

Audit After Segmentation: Reference-Free Mask Quality Assessment for Language-Referred Audio-Visual Segmentation

2026-02-03 · Jinxing Zhou, Yanghao Zhou, Yaoting Wang, Zongyan Han, Jiaqi Ma, Henghui Ding, Rao Muhammad Anwer, Hisham Cholakkal arxiv

Language-referred audio-visual segmentation (Ref-AVS) aims to segment target objects described by natural language by jointly reasoning over video, audio, and text. Beyond generating segmentation masks, providing rich and interpretable diagnoses of mask quality remains largely underexplored. In this work, we introduce Mask Quality Assessment in the Ref-AVS context (MQA-RefAVS), a new task that evaluates the quality of candidate segmentation masks without relying on ground-truth annotations as references at inference time. Given audio-visual-language inputs and each provided segmentation mask, the task requires estimating its IoU with the unobserved ground truth, identifying the corresponding error type, and recommending an actionable quality-control decision. To support this task, we construct MQ-RAVSBench, a benchmark featuring diverse and representative mask error modes that span both geometric and semantic issues. We further propose MQ-Auditor, a multimodal large language model (MLLM)-based auditor that explicitly reasons over multimodal cues and mask information to produce quantitative and qualitative mask quality assessments. Extensive experiments demonstrate that MQ-Auditor outperforms strong open-source and commercial MLLMs and can be integrated with existing Ref-AVS systems to detect segmentation failures and support downstream segmentation improvement. Data and codes will be released at https://github.com/jasongief/MQA-RefAVS.

📄 PDF Abstract BibTeX arXiv:2602.03892

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reference Traces for Auditing Invisible Weight Updates and Guiding Exact-Budget Protection

2026-07-09 · Zekai Shang arxiv

Direct low-precision write-back can erase optimizer proposals, while aggregate update visibility need not identify parameters worth protecting. We study two uses of high-precision reference traces: candidate-matched audi…

Boosting SAM for Cross-Domain Few-Shot Segmentation via Conditional Point Sparsification

2026-02-05 · Jiahao Nie, Yun Xing, Wenbin An, Qingsong Zhao 외 arxiv

Motivated by the success of the Segment Anything Model (SAM) in promptable segmentation, recent studies leverage SAM to develop training-free solutions for few-shot segmentation, which aims to predict object masks in the…

Cross-Domain Few-Shot

An Automatic Method for Complete Brain Matter Segmentation from Multislice CT scan

2018-09-11 · Soumi Ray, Vinod Kumar, Chirag Ahuja, Niranjan Khandelwal

Computed tomography imaging is well accepted for its imaging speed, image contrast & resolution and cost. Thus it has wide use in detection and diagnosis of brain diseases. But unfortunately reported works on CT segmenta…

Segmentation

Leveraging Foundation models for Unsupervised Audio-Visual Segmentation

2023-09-13 · Swapnil Bhosale, Haosen Yang, Diptesh Kanojia, Xiatian Zhu

Audio-Visual Segmentation (AVS) aims to precisely outline audible objects in a visual scene at the pixel level. Existing AVS methods require fine-grained annotations of audio-mask pairs in supervised learning fashion. Th…

Segmentation

Understanding Model Behavior in Monocular Polyp Sizing

2026-05-19 · Xinqi Xiong, Andrea Dunn Beltran, Junmyeong Choi, Sarah K. McGill 외 arxiv

Accurate polyp size stratification guides surveillance decisions, with lesions larger than 5 mm typically requiring closer follow-up. However, monocular colonoscopy lacks a reliable metric reference. We present a diagnos…

Depth Estimation