paper-with-me

홈 › Papers

Falcon 7b for Software Mention Detection in Scholarly Documents

2024-05-14 · AmeerAli Khan, Qusai Ramadan, Cong Yang, Zeyd Boukhers

This paper aims to tackle the challenge posed by the increasing integration of software tools in research across various disciplines by investigating the application of Falcon-7b for the detection and classification of software mentions within scholarly texts. Specifically, the study focuses on solving Subtask I of the Software Mention Detection in Scholarly Publications (SOMD), which entails identifying and categorizing software mentions from academic literature. Through comprehensive experimentation, the paper explores different training strategies, including a dual-classifier approach, adaptive sampling, and weighted loss scaling, to enhance detection accuracy while overcoming the complexities of class imbalance and the nuanced syntax of scholarly writing. The findings highlight the benefits of selective labelling and adaptive sampling in improving the model's performance. However, they also indicate that integrating multiple strategies does not necessarily result in cumulative improvements. This research offers insights into the effective application of large language models for specific tasks such as SOMD, underlining the importance of tailored approaches to address the unique challenges presented by academic text analysis.

📄 PDF Abstract BibTeX arXiv:2405.08514

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Software Mention Recognition with a Three-Stage Framework Based on BERTology Models at SOMD 2024

2024-04-23 · Thuy Nguyen Thi, Anh Nguyen Viet, Thin Dang Van, Ngan Nguyen Luu Thuy

This paper describes our systems for the sub-task I in the Software Mention Detection in Scholarly Publications shared-task. We propose three approaches leveraging different pre-trained language models (BERT, SciBERT, an…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+3

SecureFalcon: Are We There Yet in Automated Software Vulnerability Detection with LLMs?

2023-07-13 · Mohamed Amine Ferrag, Ammar Battah, Norbert Tihanyi, Ridhi Jain 외

Software vulnerabilities can cause numerous problems, including crashes, data loss, and security breaches. These issues greatly compromise quality and can negatively impact the market adoption of software applications an…

Binary ClassificationBug fixingC++ codeCode Completion+3

SoMeSci- A 5 Star Open Data Gold Standard Knowledge Graph of Software Mentions in Scientific Articles

2021-08-20 · David Schindler, Felix Bensmann, Stefan Dietze, Frank Krüger

Knowledge about software used in scientific investigations is important for several reasons, for instance, to enable an understanding of provenance and methods involved in data handling. However, software is usually not …

ArticlesEntity DisambiguationEntity Linkingnamed-entity-recognition+4

Overview of the SV-Ident 2022 Shared Task on Survey Variable Identification in Social Science Publications

2022-09-19 · sdp (COLING) 2022 10 · Tornike Tsereteli, Yavuz Selim Kartal, Simone Paolo Ponzetto, Andrea Zielinski 외

In this paper, we provide an overview of the SV-Ident shared task as part of the 3rd Workshop on Scholarly Document Processing (SDP) at COLING 2022. In the shared task, participants were provided with a sentence and a vo…

SentenceVariable DetectionVariable Disambiguation

Deduplication of Scholarly Documents using Locality Sensitive Hashing and Word Embeddings

2020-05-01 · LREC 2020 5 · Bikash Gyawali, Lucas Anastasiou, Petr Knoth

Deduplication is the task of identifying near and exact duplicate data items in a collection. In this paper, we present a novel method for deduplication of scholarly documents. We develop a hybrid model which uses struct…

Word Embeddings