paper-with-me

홈 › Papers

Effective Context Selection in LLM-based Leaderboard Generation: An Empirical Study

2024-06-06 · Salomon Kabongo, Jennifer D'Souza, Sören Auer

This paper explores the impact of context selection on the efficiency of Large Language Models (LLMs) in generating Artificial Intelligence (AI) research leaderboards, a task defined as the extraction of (Task, Dataset, Metric, Score) quadruples from scholarly articles. By framing this challenge as a text generation objective and employing instruction finetuning with the FLAN-T5 collection, we introduce a novel method that surpasses traditional Natural Language Inference (NLI) approaches in adapting to new developments without a predefined taxonomy. Through experimentation with three distinct context types of varying selectivity and length, our study demonstrates the importance of effective context selection in enhancing LLM accuracy and reducing hallucinations, providing a new pathway for the reliable and efficient generation of AI leaderboards. This contribution not only advances the state of the art in leaderboard generation but also sheds light on strategies to mitigate common challenges in LLM-based information extraction.

📄 PDF Abstract BibTeX arXiv:2407.02409

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesNatural Language InferenceText Generation

Similar Papers 제목 키워드 기반

Qwen Goes Brrr: Off-the-Shelf RAG for Ukrainian Multi-Domain Document Understanding

2026-05-11 · Anton Bazdyrev, Ivan Bashtovyi, Ivan Havlytskyi, Oleksandr Kharytonov 외 arxiv

We participated in the Fifth UNLP shared task on multi-domain document understanding, where systems must answer Ukrainian multiple-choice questions from PDF collections and localize the supporting document and page. We p…

Answer GenerationAnswer SelectionPassage Ranking

Automated Mining of Leaderboards for Empirical AI Research

2021-08-31 · Salomon Kabongo, Jennifer D'Souza, Sören Auer

With the rapid growth of research publications, empowering scientists to keep oversight over the scientific progress is of paramount importance. In this regard, the Leaderboards facet of information organization provides…

Knowledge GraphsScientific Results Extraction

Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL

2026-08-15 · Yiyun Su, Zujun Peng, Yu Tian, Yuting Liu 외 arxiv

LLM-based Text-to-SQL progress is reported across heterogeneous benchmarks, backbones, and inference protocols, making cross-system comparison fragile. We reframe the field as a leaderboard aggregation: we collect the me…

The Trust Paradox: How CS Researchers Engage LLM Leaderboards

2026-05-27 · Pouya Sadeghi, Anamaria Crisan, Jimmy Lin arxiv

Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite known limitations in their reliability and robustness. Yet how they sha…

On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards

2024-07-04 · Zhimin Zhao, Abdul Ali Bangash, Filipe Roseiro Côgo, Bram Adams 외

Foundation models (FM), such as large language models (LLMs), which are large-scale machine learning (ML) models, have demonstrated remarkable adaptability in various downstream software engineering (SE) tasks, such as c…

Code Completion