Designing an Evaluation Framework for Large Language Models in Astronomy Research
Large Language Models (LLMs) are shifting how scientific research is done. It is imperative to understand how researchers interact with these models and how scientific sub-communities like astronomy might benefit from them. However, there is currently no standard for evaluating the use of LLMs in astronomy. Therefore, we present the experimental design for an evaluation study on how astronomy researchers interact with LLMs. We deploy a Slack chatbot that can answer queries from users via Retrieval-Augmented Generation (RAG); these responses are grounded in astronomy papers from arXiv. We record and anonymize user questions and chatbot answers, user upvotes and downvotes to LLM responses, user feedback to the LLM, and retrieved documents and similarity scores with the query. Our data collection method will enable future dynamic evaluations of LLM tools for astronomy.
Code (1)
Tasks
AstronomyChatbotExperimental DesignRAGRetrievalRetrieval-augmented GenerationSimilar Papers 제목 키워드 기반
ORBIT: Cost-Effective Dataset Curation for Large Language Model Domain Adaptation with an Astronomy Case Study
Recent advances in language modeling demonstrate the need for high-quality domain-specific training data, especially for tasks that require specialized knowledge. General-purpose models, while versatile, often lack the d…
AstronomyDomain AdaptationLanguage ModelingLanguage Modelling+2AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy
Large Language Models (LLMs) are being explored for applications in scientific research, including their capabilities to synthesize literature, answer research questions, generate research ideas, and even conduct computa…
AstronomyOmniSpectra: A Unified Foundation Model for Native Resolution Astronomical Spectra
We present OmniSpectra, the first native-resolution foundation model for astronomy spectra. Unlike traditional models, which are limited to fixed-length input sizes or configurations, OmniSpectra handles spectra of any l…
Zero-shot GeneralizationTransfer LearningAstroMLab 2: AstroLLaMA-2-70B Model and Benchmarking Specialised LLMs for Astronomy
Continual pretraining of large language models on domain-specific data has been proposed to enhance performance on downstream tasks. In astronomy, the previous absence of astronomy-focused benchmarks has hindered objecti…
AstronomyBenchmarkingContinual PretrainingAstroLLaMA: Towards Specialized Foundation Models in Astronomy
Large language models excel in many human-language tasks but often falter in highly specialized domains like scholarly astronomy. To bridge this gap, we introduce AstroLLaMA, a 7-billion-parameter model fine-tuned from L…
AstronomyCausal Language ModelingDomain AdaptationLanguage Modeling+1