paper-with-me

CoVA

Context-aware Visual Attention-based (CoVA) webpage object detection pipeline

2000년 도입 · 논문 2편에서 사용

Context-Aware Visual Attention-based end-to-end pipeline for Webpage Object Detection (_CoVA_) aims to learn function _f_ to predict labels _y = [$y_1, y_2, ..., y_N$]_ for a webpage containing _N_ elements. The input to CoVA consists of: 1. a screenshot of a webpage, 2. list of bounding boxes _[x, y, w, h]_ of the web elements, and 3. neighborhood information for each element obtained from the DOM tree. This information is processed in four stages: 1. the graph representation extraction for the webpage, 2. the Representation Network (_RN_), 3. the Graph Attention Network (_GAT_), and 4. a fully connected (_FC_) layer. The graph representation extraction computes for every web element _i_ its set of _K_ neighboring web elements _$N_i$_. The _RN_ consists of a Convolutional Neural Net (_CNN_) and a positional encoder aimed to learn a visual representation _$v_i$_ for each web element _i ∈ {1, ..., N}_. The _GAT_ combines the visual representation _$v_i$_ of the web element _i_ to be classified and those of its neighbors, i.e., _$v_k$ ∀k ∈ $N_i$_ to compute the contextual representation _$c_i$_ for web element _i_. Finally, the visual and contextual representations of the web element are concatenated and passed through the _FC_ layer to obtain the classification output.

출처: CoVA: Context-aware Visual Attention for Webpage Information Extraction

소개 논문: CoVA: Context-aware Visual Attention for Webpage Information Extraction

Webpage Object Detection Pipeline · Computer VisionObject Detection Models · Computer Vision