paper-with-me

Papers

Generative AI for Software Metadata: Overview of the Information Retrieval in Software Engineering Track at FIRE 2023

2023-10-27 · Srijoni Majumdar, Soumen Paul, Debjyoti Paul, Ayan Bandyopadhyay, Samiran Chattopadhyay, Partha Pratim Das, Paul D Clough, Prasenjit Majumder

The Information Retrieval in Software Engineering (IRSE) track aims to develop solutions for automated evaluation of code comments in a machine learning framework based on human and large language model generated labels. In this track, there is a binary classification task to classify comments as useful and not useful. The dataset consists of 9048 code comments and surrounding code snippet pairs extracted from open source github C based projects and an additional dataset generated individually by teams using large language models. Overall 56 experiments have been submitted by 17 teams from various universities and software companies. The submissions have been evaluated quantitatively using the F1-Score and qualitatively based on the type of features developed, the supervised learning model used and their corresponding hyper-parameters. The labels generated from large language models increase the bias in the prediction model but lead to less over-fitted results.

📄 PDF Abstract BibTeX arXiv:2311.03374

Code (0)

등록된 구현이 없습니다.

Tasks

Binary ClassificationInformation RetrievalLanguage ModelingLanguage ModellingLarge Language ModelRetrieval

Similar Papers 제목 키워드 기반

Building a research-software catalog with a coding agent: from hackathon prototype to public deployment

2026-09-04 · Kazuyoshi Yoshimi, Satoshi Terasaki, Gotai Yamada arxiv

Generative AI and coding agents can accelerate research software development, but they also increase the need for efficient software discovery and maintenance. We developed a repository catalog during a three-day hackath…

S3LLM: Large-Scale Scientific Software Understanding with LLMs using Source, Metadata, and Document

2024-03-15 · Kareem Shaik, Dali Wang, Weijian Zheng, Qinglei Cao 외

The understanding of large-scale scientific software poses significant challenges due to its diverse codebase, extensive code length, and target computing architectures. The emergence of generative AI, specifically large…

Natural Language QueriesRAGRetrieval-augmented Generation

User Manual of Automatic Data Curation Tool(ADCT): A bulk data curator software in Library and Information Science

2022-10-31 · A. Banerjee, B. Sutradhar

In library and information science, document storage and user-specific document retrieval are the main aspects of digital library services. To preserve the cultural heritage, documents, and literature, we need a common p…

DescriptiveRetrieval

TREC 2020 Podcasts Track Overview

2021-03-29 · Rosie Jones, Ben Carterette, Ann Clifton, Maria Eskevich 외

The Podcast Track is new at the Text Retrieval Conference (TREC) in 2020. The podcast track was designed to encourage research into podcasts in the information retrieval and NLP research communities. The track consisted …

Information RetrievalRetrievalText Retrieval

Overview of the TREC 2023 Product Product Search Track

2023-11-14 · Daniel Campos, Surya Kallumadi, Corby Rosset, Cheng Xiang Zhai 외

This is the first year of the TREC Product search track. The focus this year was the creation of a reusable collection and evaluation of the impact of the use of metadata and multi-modal data on retrieval accuracy. This …

Retrieval