paper-with-me

홈 › Papers

Zoom Better to See Clearer: Human and Object Parsing with Hierarchical Auto-Zoom Net

2015-11-21 · Fangting Xia, Peng Wang, Liang-Chieh Chen, Alan L. Yuille

Parsing articulated objects, e.g. humans and animals, into semantic parts (e.g. body, head and arms, etc.) from natural images is a challenging and fundamental problem for computer vision. A big difficulty is the large variability of scale and location for objects and their corresponding parts. Even limited mistakes in estimating scale and location will degrade the parsing output and cause errors in boundary details. To tackle these difficulties, we propose a "Hierarchical Auto-Zoom Net" (HAZN) for object part parsing which adapts to the local scales of objects and parts. HAZN is a sequence of two "Auto-Zoom Net" (AZNs), each employing fully convolutional networks that perform two tasks: (1) predict the locations and scales of object instances (the first AZN) or their parts (the second AZN); (2) estimate the part scores for predicted object instance or part regions. Our model can adaptively "zoom" (resize) predicted image regions into their proper scales to refine the parsing. We conduct extensive experiments over the PASCAL part datasets on humans, horses, and cows. For humans, our approach significantly outperforms the state-of-the-arts by 5% mIOU and is especially better at segmenting small instances and small parts. We obtain similar improvements for parsing cows and horses over alternative methods. In summary, our strategy of first zooming into objects and then zooming into parts is very effective. It also enables us to process different regions of the image at different scales adaptively so that, for example, we do not need to waste computational resources scaling the entire image.

📄 PDF Abstract BibTeX arXiv:1511.06881

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive Image Zoom-in with Bounding Box Transformation for UAV Object Detection

2026-02-07 · Tao Wang, Chenyu Lin, Chenwei Tang, Jizhe Zhou 외 arxiv

Detecting objects from UAV-captured images is challenging due to the small object size. In this work, a simple and efficient adaptive zoom-in framework is explored for object detection on UAV images. The main motivation …

Object Detection

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

2024-11-25 · Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu 외

An image, especially with high-resolution, typically consists of numerous visual elements, ranging from dominant large objects to fine-grained detailed objects. When perceiving such images, multimodal large language mode…

AI AgentVisual Question AnsweringVisual Question Answering (VQA)

FLOAT: Factorized Learning of Object Attributes for Improved Multi-object Multi-part Scene Parsing

2022-03-30 · CVPR 2022 1 · Rishubh Singh, Pranav Gupta, Pradeep Shenoy, Ravikiran Sarvadevabhatla

Multi-object multi-part scene parsing is a challenging task which requires detecting multiple object classes in a scene and segmenting the semantic parts within each object. In this paper, we propose FLOAT, a factorized …

2D Semantic SegmentationObjectScene Parsing

ZoomNet: Part-Aware Adaptive Zooming Neural Network for 3D Object Detection

2020-03-01 · Zhenbo Xu, Wei zhang, Xiaoqing Ye, Xiao Tan 외

3D object detection is an essential task in autonomous driving and robotics. Though great progress has been made, challenges remain in estimating 3D pose for distant and occluded objects. In this paper, we present a nove…

2D Object Detection3D Object DetectionAutonomous DrivingDisparity Estimation+2

Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object Detection

2022-03-05 · CVPR 2022 1 · Pang Youwei, Zhao Xiaoqi, Xiang Tian-Zhu, Zhang Lihe 외

The recently proposed camouflaged object detection (COD) attempts to segment objects that are visually blended into their surroundings, which is extremely complex and difficult in real-world scenarios. Apart from high in…

Camouflaged Object SegmentationImage Segmentationobject-detectionObject Detection+1