paper-with-me

Papers

SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions

2024-06-20 · Huitong Pan, Qi Zhang, Cornelia Caragea, Eduard Dragut, Longin Jan Latecki

We present SciDMT, an enhanced and expanded corpus for scientific mention detection, offering a significant advancement over existing related resources. SciDMT contains annotated scientific documents for datasets (D), methods (M), and tasks (T). The corpus consists of two components: 1) the SciDMT main corpus, which includes 48 thousand scientific articles with over 1.8 million weakly annotated mention annotations in the format of in-text span, and 2) an evaluation set, which comprises 100 scientific articles manually annotated for evaluation purposes. To the best of our knowledge, SciDMT is the largest corpus for scientific entity mention detection. The corpus's scale and diversity are instrumental in developing and refining models for tasks such as indexing scientific papers, enhancing information retrieval, and improving the accessibility of scientific knowledge. We demonstrate the corpus's utility through experiments with advanced deep learning architectures like SciBERT and GPT-3.5. Our findings establish performance baselines and highlight unresolved challenges in scientific mention detection. SciDMT serves as a robust benchmark for the research community, encouraging the development of innovative models to further the field of scientific information extraction.

📄 PDF Abstract BibTeX arXiv:2406.14756

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesDiversityInformation Retrieval

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Multi-Head Attention 설명 없음
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

The diffusion of scientific terms – tracing individuals’ influence in the history of science for English

2021-11-01 · EMNLP (LaTeCHCLfL, CLFL, LaTeCH) 2021 11 · Yuri Bizzoni, Stefania Degaetano-Ortlieb, Katrin Menzel, Elke Teich

Tracing the influence of individuals or groups in social networks is an increasingly popular task in sociolinguistic studies. While methods to determine someone’s influence in shortterm contexts (e.g., social media, on-l…

Astronomy

ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation Networks

2019-09-04 · Michihiro Yasunaga, Jungo Kasai, Rui Zhang, Alexander R. Fabbri 외

Scientific article summarization is challenging: large, annotated corpora are not available, and the summary should ideally include the article's impacts on research community. This paper provides novel solutions to thes…

Scientific Document SummarizationText Summarization

CSL: A Large-scale Chinese Scientific Literature Dataset

2022-09-12 · COLING 2022 10 · Yudong Li, Yuqing Zhang, Zhe Zhao, Linlin Shen 외

Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development …

text-classificationText Classification

SciBERT: A Pretrained Language Model for Scientific Text

2019-03-26 · IJCNLP 2019 11 · Iz Beltagy, Kyle Lo, Arman Cohan

Obtaining large-scale annotated data for NLP tasks in the scientific domain is challenging and expensive. We release SciBERT, a pretrained language model based on BERT (Devlin et al., 2018) to address the lack of high-qu…

Citation Intent ClassificationDependency ParsingGeneral ClassificationLanguage Modeling+8

ACL-Fig: A Dataset for Scientific Figure Classification

2023-01-28 · Zeba Karishma, Shaurya Rohatgi, Kavya Shrinivas Puranik, Jian Wu 외

Most existing large-scale academic search engines are built to retrieve text-based information. However, there are no large-scale retrieval services for scientific figures and tables. One challenge for such services is u…

ClassificationQuestion AnsweringRetrieval