Building Dataset for Grounding of Formulae — Annotating Coreference Relations Among Math Identifiers
Grounding the meaning of each symbol in math formulae is important for automated understanding of scientific documents. Generally speaking, the meanings of math symbols are not necessarily constant, and the same symbol is used in multiple meanings. Therefore, coreference relations between symbols need to be identified for grounding, and the task has aspects of both description alignment and coreference analysis. In this study, we annotated 15 papers selected from arXiv.org with the grounding information. In total, 12,352 occurrences of math identifiers in these papers were annotated, and all coreference relations between them were made explicit in each paper. The constructed dataset shows that regardless of the ambiguity of symbols in math formulae, coreference relations can be labeled with a high inter-annotator agreement. The constructed dataset enables us to achieve automation of formula grounding, and in turn, make deeper use of the knowledge in scientific documents using techniques such as math information extraction. The built grounding dataset is available at https://sigmathling.kwarc.info/resources/grounding- dataset/.
Code (1)
Tasks
MathSimilar Papers 제목 키워드 기반
Mention Annotations Alone Enable Efficient Domain Adaptation for Coreference Resolution
Although recent neural models for coreference resolution have led to substantial improvements on benchmark datasets, transferring these models to new target domains containing out-of-vocabulary spans and requiring differ…
coreference-resolutionCoreference ResolutionDomain AdaptationExtending Phrase Grounding with Pronouns in Visual Dialogues
Conventional phrase grounding aims to localize noun phrases mentioned in a given caption to their corresponding image regions, which has achieved great success recently. Apparently, sole noun phrase grounding is not enou…
Phrase GroundingParallel Data Helps Neural Entity Coreference Resolution
Coreference resolution is the task of finding expressions that refer to the same entity in a text. Coreference models are generally trained on monolingual annotated data but annotating coreference is expensive and challe…
coreference-resolutionCoreference ResolutionQualitative and Quantitative Analysis of Diversity in Cross-document Coreference Resolution Datasets
Established cross-document coreference resolution (CDCR) datasets contain manually annotated event-centric mentions of events and entities that form coreference chains with identity relations. In this paper, we qualitati…
coreference-resolutionCoreference ResolutionCross Document Coreference ResolutionDiversityGRAVL-BERT: Graphical Visual-Linguistic Representations for Multimodal Coreference Resolution
Learning from multimodal data has become a popular research topic in recent years. Multimodal coreference resolution (MCR) is an important task in this area. MCR involves resolving the references across different modalit…
coreference-resolutionCoreference ResolutionVisual Grounding