paper-with-me

Papers

ArxivBench: Can LLMs Assist Researchers in Conducting Research?

2025-04-06 · Ning li, Jingran Zhang, Justin Cui

Large language models (LLMs) have demonstrated remarkable effectiveness in completing various tasks such as reasoning, translation, and question answering. However the issue of factual incorrect content in LLM-generated responses remains a persistent challenge. In this study, we evaluate both proprietary and open-source LLMs on their ability to respond with relevant research papers and accurate links to articles hosted on the arXiv platform, based on high level prompts. To facilitate this evaluation, we introduce arXivBench, a benchmark specifically designed to assess LLM performance across eight major subject categories on arXiv and five subfields within computer science, one of the most popular categories among them. Our findings reveal a concerning accuracy of LLM-generated responses depending on the subject, with some subjects experiencing significantly lower accuracy than others. Notably, Claude-3.5-Sonnet exhibits a substantial advantage in generating both relevant and accurate responses. And interestingly, most LLMs achieve a much higher accuracy in the Artificial Intelligence sub-field than other sub-fields. This benchmark provides a standardized tool for evaluating the reliability of LLM-generated scientific responses, promoting more dependable use of LLMs in academic and research environments. Our code is open-sourced at https://github.com/arxivBenchLLM/arXivBench and our dataset is available on huggingface at https://huggingface.co/datasets/arXivBenchLLM/arXivBench.

📄 PDF Abstract BibTeX arXiv:2504.10496

Code (1)

arxivbenchllm/arxivbench 공식 구현

Tasks

ArticlesQuestion Answering

Similar Papers 제목 키워드 기반

AAAR-1.0: Assessing AI's Potential to Assist Research

2024-10-29 · Renze Lou, Hanzi Xu, Sijia Wang, Jiangshu Du 외

Numerous studies have assessed the proficiency of AI systems, particularly large language models (LLMs), in facilitating everyday tasks such as email writing, question answering, and creative content generation. However,…

Question Answering

LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing

2024-06-24 · Jiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng 외

This work is motivated by two key trends. On one hand, large language models (LLMs) have shown remarkable versatility in various generative tasks such as writing, drawing, and question answering, significantly reducing t…

Question Answering

Empowering Computing Education Researchers Through LLM-Assisted Content Analysis

2025-08-26 · Laurie Gale, Sebastian Mateos Nicolajsen arxiv

Computing education research (CER) is often instigated by practitioners wanting to improve both their own and the wider discipline's teaching practice. However, the latter is often difficult as many researchers lack the …

Acceleron: A Tool to Accelerate Research Ideation

2024-03-07 · Harshit Nigam, Manasi Patwardhan, Lovekesh Vig, Gautam Shroff

Several tools have recently been proposed for assisting researchers during various stages of the research life-cycle. However, these primarily concentrate on tasks such as retrieving and recommending relevant literature,…

AI-Empowered Human Research Integrating Brain Science and Social Sciences Insights

2024-11-16 · Feng Xiong, Xinguo Yu, Hon Wai Leong

This paper explores the transformative role of artificial intelligence (AI) in enhancing scientific research, particularly in the fields of brain science and social sciences. We analyze the fundamental aspects of human r…