The Devil is in the Decoder: Classification, Regression and GANs
Many machine vision applications, such as semantic segmentation and depth prediction, require predictions for every pixel of the input image. Models for such problems usually consist of encoders which decrease spatial resolution while learning a high-dimensional representation, followed by decoders who recover the original input resolution and result in low-dimensional predictions. While encoders have been studied rigorously, relatively few studies address the decoder side. This paper presents an extensive comparison of a variety of decoders for a variety of pixel-wise tasks ranging from classification, regression to synthesis. Our contributions are: (1) Decoders matter: we observe significant variance in results between different types of decoders on various problems. (2) We introduce new residual-like connections for decoders. (3) We introduce a novel decoder: bilinear additive upsampling. (4) We explore prediction artifacts.
Code (1)
Tasks
Boundary DetectionDecoderDepth EstimationDepth PredictionGeneral ClassificationregressionSemantic SegmentationSimilar Papers 제목 키워드 기반
Fuzzy Generative Adversarial Networks
Generative Adversarial Networks (GANs) are well-known tools for data generation and semi-supervised classification. GANs, with less labeled data, outperform Deep Neural Networks (DNNs) and Convolutional Neural Networks (…
regressionGeneralizing semi-supervised generative adversarial networks to regression using feature contrasting
In this work, we generalize semi-supervised generative adversarial networks (GANs) from classification problems to regression problems. In the last few years, the importance of improving the training of neural networks u…
Age EstimationCrowd CountingGeneral ClassificationregressionDetector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
Multimodal large language models (MLLMs) are rapidly expanding from general video understanding to finer-grained understanding such as spatio-temporal video grounding (STVG) and reasoning. In these tasks, an MLLM must lo…
Spatio-Temporal Video GroundingPropeller Motion of a Devil-Stick using Normal Forcing
The problem of realizing rotary propeller motion of a devil-stick in the vertical plane using forces purely normal to the stick is considered. This problem represents a nonprehensile manipulation task of an underactuated…
AgeFlow: Conditional Age Progression and Regression with Normalizing Flows
Age progression and regression aim to synthesize photorealistic appearance of a given face image with aging and rejuvenation effects, respectively. Existing generative adversarial networks (GANs) based methods suffer fro…
AttributeKnowledge Distillationregression