paper-with-me

홈 › Papers

Non-local RoI for Cross-Object Perception

2018-11-25 · Shou-Yao Roy Tseng, Hwann-Tzong Chen, Shao-Heng Tai, Tyng-Luh Liu

We present a generic and flexible module that encodes region proposals by both their intrinsic features and the extrinsic correlations to the others. The proposed non-local region of interest (NL-RoI) can be seamlessly adapted into different generalized R-CNN architectures to better address various perception tasks. Observe that existing techniques from R-CNN treat RoIs independently and perform the prediction solely based on image features within each region proposal. However, the pairwise relationships between proposals could further provide useful information for detection and segmentation. NL-RoI is thus formulated to enrich each RoI representation with the information from all other RoIs, and yield a simple, low-cost, yet effective module for region-based convolutional networks. Our experimental results show that NL-RoI can improve the performance of Faster/Mask R-CNN for object detection and instance segmentation.

📄 PDF Abstract BibTeX arXiv:1811.10002

Code (0)

등록된 구현이 없습니다.

Tasks

Instance SegmentationObjectobject-detectionObject DetectionRegion ProposalSegmentationSemantic Segmentation

Methods 이 논문이 사용한 방법론

RoIAlign Region of Interest Align, or RoIAlign, is an operation for extracting a small feature map from each RoI in detection and segmentation based tasks. It removes the harsh…
Mask R-CNN Mask R-CNN extends Faster R-CNN to solve instance segmentation tasks. It achieves this by adding a branch for predicting an…
RPN A Region Proposal Network, or RPN, is a fully convolutional network that simultaneously predicts object bounds and objectness scores at each position. The RPN is trained…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
RoIPool 설명 없음
Faster R-CNN Faster R-CNN is an object detection model that improves on Fast R-CNN by utilising a region proposal network…

Similar Papers 제목 키워드 기반

Towards Global Localization using Multi-Modal Object-Instance Re-Identification

2024-09-18 · Aneesh Chavan, Vaibhav Agrawal, Vineeth Bhat, Sarthak Chittawar 외

Re-identification (ReID) is a critical challenge in computer vision, predominantly studied in the context of pedestrians and vehicles. However, robust object-instance ReID, which has significant implications for tasks su…

Camera LocalizationObjectScene Understanding

Self-Localized Collaborative Perception

2024-06-18 · Zhenyang Ni, Zixing Lei, Yifan Lu, Dingju Wang 외

Collaborative perception has garnered considerable attention due to its capacity to address several inherent challenges in single-agent perception, including occlusion and out-of-range issues. However, existing collabora…

UniHead: Unifying Multi-Perception for Detection Heads

2023-09-23 · Hantao Zhou, Rui Yang, Yachao Zhang, Haoran Duan 외

The detection head constitutes a pivotal component within object detectors, tasked with executing both classification and localization functions. Regrettably, the commonly used parallel head often lacks omni perceptual c…

Perception Test 2023: A Summary of the First Challenge And Outcome

2023-12-20 · Joseph Heyward, João Carreira, Dima Damen, Andrew Zisserman 외

The First Perception Test challenge was held as a half-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2023, with the goal of benchmarking state-of-the-art video models on the recen…

BenchmarkingGrounded Video Question AnsweringMultiple-choiceObject Tracking+3

Distilling Cross-Temporal Contexts for Continuous Sign Language Recognition

2023-01-01 · CVPR 2023 1 · Leming Guo, Wanli Xue, Qing Guo, Bo Liu 외

Continuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-the-art methods typically have two modules, a spatial perception module and a temporal aggregation module, whic…

Knowledge DistillationSign Language Recognition