paper-with-me

홈 › Papers

FLOAT: Factorized Learning of Object Attributes for Improved Multi-object Multi-part Scene Parsing

2022-03-30 · CVPR 2022 1 · Rishubh Singh, Pranav Gupta, Pradeep Shenoy, Ravikiran Sarvadevabhatla

Multi-object multi-part scene parsing is a challenging task which requires detecting multiple object classes in a scene and segmenting the semantic parts within each object. In this paper, we propose FLOAT, a factorized label space framework for scalable multi-object multi-part parsing. Our framework involves independent dense prediction of object category and part attributes which increases scalability and reduces task complexity compared to the monolithic label space counterpart. In addition, we propose an inference-time 'zoom' refinement technique which significantly improves segmentation quality, especially for smaller objects/parts. Compared to state of the art, FLOAT obtains an absolute improvement of 2.0% for mean IOU (mIOU) and 4.8% for segmentation quality IOU (sqIOU) on the Pascal-Part-58 dataset. For the larger Pascal-Part-108 dataset, the improvements are 2.1% for mIOU and 3.9% for sqIOU. We incorporate previously excluded part attributes and other minor parts of the Pascal-Part dataset to create the most comprehensive and challenging version which we dub Pascal-Part-201. FLOAT obtains improvements of 8.6% for mIOU and 7.5% for sqIOU on the new dataset, demonstrating its parsing effectiveness across a challenging diversity of objects and parts. The code and datasets are available at floatseg.github.io.

📄 PDF Abstract BibTeX arXiv:2203.16168

Code (1)

floatseg/floatseg.github.io 공식 구현

Tasks

2D Semantic SegmentationObjectScene Parsing

Similar Papers 제목 키워드 기반

Unsupervised Primitive Discovery for Improved 3D Generative Modeling

2019-06-09 · CVPR 2019 6 · Salman H. Khan, Yulan Guo, Munawar Hayat, Nick Barnes

3D shape generation is a challenging problem due to the high-dimensional output space and complex part configurations of real-world objects. As a result, existing algorithms experience difficulties in accurate generative…

3D Shape Generation

NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

2024-03-05 · Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan 외

While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intricately encompasses various attributes (e.…

QuantizationSpeech Synthesistext-to-speechText to Speech

A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model

2025-05-20 · Yang Xiang, Canan Huang, Desheng Hu, Jingguang Tian 외

Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as seman…

Speech Enhancement

Content-Context Factorized Representations for Automated Speech Recognition

2022-05-19 · David M. Chan, Shalini Ghosh

Deep neural networks have largely demonstrated their ability to perform automated speech recognition (ASR) by extracting meaningful features from input audio frames. Such features, however, may consist not only of inform…

speech-recognitionSpeech Recognition

Multi-scale Attributed Node Embedding

2019-09-28 · Benedek Rozemberczki, Carl Allen, Rik Sarkar

We present network embedding algorithms that capture information about a node from the local distribution over node attributes around it, as observed over random walks following an approach similar to Skip-gram. Observat…

AttributeNetwork Embedding