MathAlign: Linking Formula Identifiers to their Contextual Natural Language Descriptions
Extending machine reading approaches to extract mathematical concepts and their descriptions is useful for a variety of tasks, ranging from mathematical information retrieval to increasing accessibility of scientific documents for the visually impaired. This entails segmenting mathematical formulae into identifiers and linking them to their natural language descriptions. We propose a rule-based approach for this task, which extracts LaTeX representations of formula identifiers and links them to their in-text descriptions, given only the original PDF and the location of the formula of interest. We also present a novel evaluation dataset for this task, as well as the tool used to create it.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalReading ComprehensionRetrievalSimilar Papers 제목 키워드 기반
Query Brand Entity Linking in E-Commerce Search
Associating user search queries with the correct brand entity is critical for e-commerce product retrieval, yet remains challenging due to the brevity of queries (three to four words on average), their lack of grammatica…
Predicting Failures of LLMs to Link Biomedical Ontology Terms to Identifiers Evidence Across Models and Ontologies
Large language models often perform well on biomedical NLP tasks but may fail to link ontology terms to their correct identifiers. We investigate why these failures occur by analyzing predictions across two major ontolog…
Contextualizing Hate Speech Classifiers with Post-hoc Explanation
Hate speech classifiers trained on imbalanced datasets struggle to determine if group identifiers like "gay" or "black" are used in offensive or prejudiced ways. Such biases manifest in false positives when these identif…
Connecting people digitally - a semantic web based approach to linking heterogeneous data sets
In this paper we present a semantic enrichment approach for linking two distinct data sets: the {\"O}BL (Austrian Biographical Dictionary) and the DB{\"O} (Database of Bavarian Dialects in Austria). Although the data set…
Entity LinkingWord Sense DisambiguationCross-lingual Linking of Multi-word Entities and their corresponding Acronyms
This paper reports on an approach and experiments to automatically build a cross-lingual multi-word entity resource. Starting from a collection of millions of acronym/expansion pairs for 22 languages where expansion vari…
Translation