A Two-stage Framework for Compound Figure Separation
Scientific literature contains large volumes of complex, unstructured figures that are compound in nature (i.e. composed of multiple images, graphs, and drawings). Separation of these compound figures is critical for information retrieval from these figures. In this paper, we propose a new strategy for compound figure separation, which decomposes the compound figures into constituent subfigures while preserving the association between the subfigures and their respective caption components. We propose a two-stage framework to address the proposed compound figure separation problem. In particular, the subfigure label detection module detects all subfigure labels in the first stage. Then, in the subfigure detection module, the detected subfigure labels help to detect the subfigures by optimizing the feature selection process and providing the global layout information as extra features. Extensive experiments are conducted to validate the effectiveness and superiority of the proposed framework, which improves the detection precision by 9%.
Code (0)
등록된 구현이 없습니다.
Tasks
feature selectionInformation RetrievalRetrievalVocal Bursts Valence PredictionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Automatic Separation of Compound Figures in Scientific Articles
Content-based analysis and retrieval of digital images found in scientific articles is often hindered by images consisting of multiple subfigures (compound figures). We address this problem by proposing a method to autom…
ArticlesRetrievalCompound Figure Separation of Biomedical Images: Mining Large Datasets for Self-supervised Learning
With the rapid development of self-supervised learning (e.g., contrastive learning), the importance of having large-scale images (even without annotations) for training a more generalizable AI model has been widely recog…
Contrastive LearningImage Augmentationimage-classificationImage Classification+2Compound Figure Separation of Biomedical Images with Side Loss
Unsupervised learning algorithms (e.g., self-supervised learning, auto-encoder, contrastive learning) allow deep learning models to learn effective image representations from large-scale unlabeled data. In medical image …
Contrastive LearningImage AugmentationMedical Image AnalysisSelf-Supervised LearningA Data Driven Approach for Compound Figure Separation Using Convolutional Neural Networks
A key problem in automatic analysis and understanding of scientific papers is to extract semantic information from non-textual paper components like figures, diagrams, tables, etc. Much of this work requires a very first…
Transfer LearningSemantic Segmentation for Compound figures
Scientific literature contains large volumes of unstructured data,with over 30\% of figures constructed as a combination of multiple images, these compound figures cannot be analyzed directly with existing information re…
Information RetrievalRetrievalSegmentationSemantic Segmentation