paper-with-me

홈 › Papers

MyFixit: An Annotated Dataset, Annotation Tool, and Baseline Methods for Information Extraction from Repair Manuals

2020-05-01 · LREC 2020 5 · Nima Nabizadeh, Dorothea Kolossa, Martin Heckmann

Text instructions are among the most widely used media for learning and teaching. Hence, to create assistance systems that are capable of supporting humans autonomously in new tasks, it would be immensely productive, if machines were enabled to extract task knowledge from such text instructions. In this paper, we, therefore, focus on information extraction (IE) from the instructional text in repair manuals. This brings with it the multiple challenges of information extraction from the situated and technical language in relatively long and often complex instructions. To tackle these challenges, we introduce a semi-structured dataset of repair manuals. The dataset is annotated in a large category of devices, with information that we consider most valuable for an automated repair assistant, including the required tools and the disassembled parts at each step of the repair progress. We then propose methods that can serve as baselines for this IE task: an unsupervised method based on a bags-of-n-grams similarity for extracting the needed tools in each repair step, and a deep-learning-based sequence labeling model for extracting the identity of disassembled parts. These baseline methods are integrated into a semi-automatic web-based annotator application that is also available along with the dataset.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Repair 설명 없음

Similar Papers 제목 키워드 기반

MSNER: A Multilingual Speech Dataset for Named Entity Recognition

2024-05-19 · Quentin Meeus, Marie-Francine Moens, Hugo Van hamme

While extensively explored in text-based tasks, Named Entity Recognition (NER) remains largely neglected in spoken language understanding. Existing resources are limited to a single, English-only dataset. This paper addr…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Tools Impact on the Quality of Annotations for Chat Untangling

2021-08-01 · ACL 2021 5 · Jhonny Cerezo, Felipe Bravo-Marquez, Alexandre Henri Bergel

The quality of the annotated data directly influences in the success of supervised NLP models. However, creating annotated datasets is often time-consuming and expensive. Although the annotation tool takes an important r…

NarrativeTime: Dense Temporal Annotation on a Timeline

2019-08-29 · Anna Rogers, Marzena Karpinska, Ankita Gupta, Vladislav Lialin 외

For the past decade, temporal annotation has been sparse: only a small portion of event pairs in a text was annotated. We present NarrativeTime, the first timeline-based annotation framework that achieves full coverage o…

Chunking

SALMA: Arabic Sense-Annotated Corpus and WSD Benchmarks

2023-10-29 · Mustafa Jarrar, Sanad Malaysha, Tymaa Hammouda, Mohammed Khalilia

SALMA, the first Arabic sense-annotated corpus, consists of ~34K tokens, which are all sense-annotated. The corpus is annotated using two different sense inventories simultaneously (Modern and Ghani). SALMA novelty lies …

Word Sense Disambiguation

ROBUST-MIPS: A Combined Skeletal Pose and Instance Segmentation Dataset for Laparoscopic Surgical Instruments

2025-08-27 · Zhe Han, Charlie Budd, Gongyu Zhang, Huanyu Tian 외 arxiv

Localisation of surgical tools constitutes a foundational building block for computer-assisted interventional technologies. Works in this field typically focus on training deep learning models to perform segmentation tas…

Instance SegmentationPose Estimation