paper-with-me

Papers

CatalogBank: A Structured and Interoperable Catalog Dataset with a Semi-Automatic Annotation Tool (DocumentLabeler) for Engineering System Design

2024-08-15 · Hasan Sinan Bank, Daniel R. Herber

In the realm of document engineering and Natural Language Processing (NLP), the integration of digitally born catalogs into product design processes presents a novel avenue for enhancing information extraction and interoperability. This paper introduces CatalogBank, a dataset developed to bridge the gap between textual descriptions and other data modalities related to engineering design catalogs. We utilized existing information extraction methodologies to extract product information from PDF-based catalogs to use in downstream tasks to generate a baseline metric. Our approach not only supports the potential automation of design workflows but also overcomes the limitations of manual data entry and non-standard metadata structures that have historically impeded the seamless integration of textual and other data modalities. Through the use of DocumentLabeler, an open-source annotation tool adapted for our dataset, we demonstrated the potential of CatalogBank in supporting diverse document-based tasks such as layout analysis and knowledge extraction. Our findings suggest that CatalogBank can contribute to document engineering and NLP by providing a robust dataset for training models capable of understanding and processing complex document formats with relatively less effort using the semi-automated annotation tool DocumentLabeler.

📄 PDF Abstract BibTeX arXiv:2408.08238

Code (2)

bankh/DocumentLabeler 공식 구현 pytorch
bankh/catalogbank 공식 구현

Similar Papers 제목 키워드 기반

FAIRification of MLC data

2022-11-23 · Ana Kostovska, Jasmin Bogatinovski, Andrej Treven, Sašo Džeroski 외

The multi-label classification (MLC) task has increasingly been receiving interest from the machine learning (ML) community, as evidenced by the growing number of papers and methods that appear in the literature. Hence, …

BenchmarkingManagementMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding

2026-08-21 · Rohan Kumar, Steven Xu, Kyle MacDonald, Matthew Long 외 arxiv

Product catalogs underpin search, discovery, and recommendation in e-commerce, yet they are often attribute-sparse: the attributes shoppers and downstream systems rely on are either buried in unstructured content such as…

Semantic Product Search for Matching Structured Product Catalogs in E-Commerce

2020-08-18 · Jason Ingyu Choi, Surya Kallumadi, Bhaskar Mitra, Eugene Agichtein 외

Retrieving all semantically relevant products from the product catalog is an important problem in E-commerce. Compared to web documents, product catalogs are more structured and sparse due to multi-instance fields that e…

Construction of a Battery Research Knowledge Graph using a Global Open Catalog

2026-04-22 · Luca Foppiano, Sae Dieb, Malik Zain, Kazuki Kasama 외 arxiv

Battery research is a rapidly growing and highly interdisciplinary field, making it increasingly difficult to track relevant expertise and identify potential collaborators across institutional boundaries. In this work, w…

Community Detection

Musical Audio Similarity with Self-supervised Convolutional Neural Networks

2022-02-04 · Carl Thomé, Sebastian Piwell, Oscar Utterbäck

We have built a music similarity search engine that lets video producers search by listenable music excerpts, as a complement to traditional full-text search. Our system suggests similar sounding track segments in a larg…

Triplet