paper-with-me

Papers

Designing an Evaluation Framework for Large Language Models in Astronomy Research

2024-05-30 · John F. Wu, Alina Hyk, Kiera McCormick, Christine Ye, Simone Astarita, Elina Baral, Jo Ciuca, Jesse Cranney, Anjalie Field, Kartheik Iyer, Philipp Koehn, Jenn Kotler, Sandor Kruk, Michelle Ntampaka, Charles O'Neill, Joshua E. G. Peek, Sanjib Sharma, Mikaeel Yunus

Large Language Models (LLMs) are shifting how scientific research is done. It is imperative to understand how researchers interact with these models and how scientific sub-communities like astronomy might benefit from them. However, there is currently no standard for evaluating the use of LLMs in astronomy. Therefore, we present the experimental design for an evaluation study on how astronomy researchers interact with LLMs. We deploy a Slack chatbot that can answer queries from users via Retrieval-Augmented Generation (RAG); these responses are grounded in astronomy papers from arXiv. We record and anonymize user questions and chatbot answers, user upvotes and downvotes to LLM responses, user feedback to the LLM, and retrieved documents and similarity scores with the query. Our data collection method will enable future dynamic evaluations of LLM tools for astronomy.

📄 PDF Abstract BibTeX arXiv:2405.20389

Code (1)

jsalt2024-evaluating-llms-for-astronomy/astro-arxiv-bot 공식 구현 pytorch

Tasks

AstronomyChatbotExperimental DesignRAGRetrievalRetrieval-augmented Generation

Similar Papers 제목 키워드 기반

ORBIT: Cost-Effective Dataset Curation for Large Language Model Domain Adaptation with an Astronomy Case Study

2024-12-19 · Eric Modesitt, Ke Yang, Spencer Hulsey, ChengXiang Zhai 외

Recent advances in language modeling demonstrate the need for high-quality domain-specific training data, especially for tasks that require specialized knowledge. General-purpose models, while versatile, often lack the d…

AstronomyDomain AdaptationLanguage ModelingLanguage Modelling+2

AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy

2025-05-26 · Sebastian Antony Joseph, Syed Murtaza Husain, Stella S. R. Offner, Stéphanie Juneau 외

Large Language Models (LLMs) are being explored for applications in scientific research, including their capabilities to synthesize literature, answer research questions, generate research ideas, and even conduct computa…

Astronomy

OmniSpectra: A Unified Foundation Model for Native Resolution Astronomical Spectra

2026-01-21 · Md Khairul Islam, Judy Fox arxiv

We present OmniSpectra, the first native-resolution foundation model for astronomy spectra. Unlike traditional models, which are limited to fixed-length input sizes or configurations, OmniSpectra handles spectra of any l…

Zero-shot GeneralizationTransfer Learning

AstroMLab 2: AstroLLaMA-2-70B Model and Benchmarking Specialised LLMs for Astronomy

2024-09-29 · Rui Pan, Tuan Dung Nguyen, Hardik Arora, Alberto Accomazzi 외

Continual pretraining of large language models on domain-specific data has been proposed to enhance performance on downstream tasks. In astronomy, the previous absence of astronomy-focused benchmarks has hindered objecti…

AstronomyBenchmarkingContinual Pretraining

AstroLLaMA: Towards Specialized Foundation Models in Astronomy

2023-09-12 · Tuan Dung Nguyen, Yuan-Sen Ting, Ioana Ciucă, Charlie O'Neill 외

Large language models excel in many human-language tasks but often falter in highly specialized domains like scholarly astronomy. To bridge this gap, we introduce AstroLLaMA, a 7-billion-parameter model fine-tuned from L…

AstronomyCausal Language ModelingDomain AdaptationLanguage Modeling+1