paper-with-me

홈 › Papers

Are Frontier Large Language Models Suitable for Q&A in Science Centres?

2024-12-06 · Jacob Watson, Fabrício Góes, Marco Volpe, Talles Medeiros

This paper investigates the suitability of frontier Large Language Models (LLMs) for Q&A interactions in science centres, with the aim of boosting visitor engagement while maintaining factual accuracy. Using a dataset of questions collected from the National Space Centre in Leicester (UK), we evaluated responses generated by three leading models: OpenAI's GPT-4, Claude 3.5 Sonnet, and Google Gemini 1.5. Each model was prompted for both standard and creative responses tailored to an 8-year-old audience, and these responses were assessed by space science experts based on accuracy, engagement, clarity, novelty, and deviation from expected answers. The results revealed a trade-off between creativity and accuracy, with Claude outperforming GPT and Gemini in both maintaining clarity and engaging young audiences, even when asked to generate more creative responses. Nonetheless, experts observed that higher novelty was generally associated with reduced factual reliability across all models. This study highlights the potential of LLMs in educational settings, emphasizing the need for careful prompt engineering to balance engagement with scientific rigor.

📄 PDF Abstract BibTeX arXiv:2412.05200

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Residual Connection 설명 없음
Transformer A Transformer is a model architecture that eschews recurrence and instead relies entirely on an [attention…
Adam 설명 없음

Similar Papers 제목 키워드 기반

The Language Archive --- a new hub for language resources

2012-05-01 · LREC 2012 5 · Sebastian Drude, Daan Broeder, Paul Trilsbeek, Peter Wittenburg

This contribution presents “The Language Archive” (TLA), a new unit at the MPI for Psycholinguistics, discussing the current developments in management of scientific data, considering the need for new data research infra…

Language AcquisitionManagement

Humanity's Last Exam

2025-01-24 · Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li 외

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve over 90\% accuracy on popular benchm…

Humanity's Last ExamLanguage ModelingLanguage ModellingLarge Language Model+2

FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks

2026-01-29 · Miles Wang, Robi Lin, Kat Hu, Joy Jiao 외 arxiv

We introduce FrontierScience, a benchmark evaluating expert-level scientific reasoning in frontier language models. Recent model progress has nearly saturated existing science benchmarks, which often rely on multiple-cho…

Local Search Yields a PTAS for k-Means in Doubling Metrics

2016-03-29 · Zachary Friggstad, Mohsen Rezapour, Mohammad R. Salavatipour

The most well known and ubiquitous clustering problem encountered in nearly every branch of science is undoubtedly $k$-means: given a set of data points and a parameter $k$, select $k$ centres and partition the data poin…

Clustering

Does Spatial Cognition Emerge in Frontier Models?

2024-10-09 · Santhosh Kumar Ramakrishnan, Erik Wijmans, Philipp Kraehenbuehl, Vladlen Koltun

Not yet. We present SPACE, a benchmark that systematically evaluates spatial cognition in frontier models. Our benchmark builds on decades of research in cognitive science. It evaluates large-scale mapping abilities that…