paper-with-me

Papers

A Report on the llms evaluating the high school questions

2025-04-30 · Zhu Jiawei, Chen Wei

This report aims to evaluate the performance of large language models (LLMs) in solving high school science questions and to explore their potential applications in the educational field. With the rapid development of LLMs in the field of natural language processing, their application in education has attracted widespread attention. This study selected mathematics exam questions from the college entrance examinations (2019-2023) as evaluation data and utilized at least eight LLM APIs to provide answers. A comprehensive assessment was conducted based on metrics such as accuracy, response time, logical reasoning, and creativity. Through an in-depth analysis of the evaluation results, this report reveals the strengths and weaknesses of LLMs in handling high school science questions and discusses their implications for educational practice. The findings indicate that although LLMs perform excellently in certain aspects, there is still room for improvement in logical reasoning and creative problem-solving. This report provides an empirical foundation for further research and application of LLMs in the educational field and offers suggestions for improvement.

📄 PDF Abstract BibTeX arXiv:2505.00057

Code (0)

등록된 구현이 없습니다.

Tasks

Logical Reasoning

Similar Papers 제목 키워드 기반

IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge

2026-03-24 · Ali Abdelaal, Mohammed Nader Al Haffar, Mahmoud Fawzi, Walid Magdy arxiv

Large language models are increasingly consulted for Islamic knowledge, yet no comprehensive benchmark evaluates their performance across core Islamic disciplines. We introduce IslamicMMLU, a benchmark of 10,013 multiple…

Bias Detection

Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions

2025-08-04 · Farah Atif, Nursultan Askarbekuly, Kareem Darwish, Monojit Choudhury arxiv

Despite the increasing usage of Large Language Models (LLMs) in answering questions in a variety of domains, their reliability and accuracy remain unexamined for a plethora of domains including the religious domains. In …

From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation

2025-08-05 · Chengliang Zhou, Mei Wang, Ting Zhang, Qiannan Zhu 외 arxiv

Large Language Models (LLMs) have demonstrated remarkable capabilities in mathematical problem-solving. However, the transition from providing answers to generating high-quality educational questions presents significant…

Question Generation

VisScience: An Extensive Benchmark for Evaluating K12 Educational Multi-modal Scientific Reasoning

2024-09-10 · Zhihuan Jiang, Zhen Yang, Jinhao Chen, Zhengxiao Du 외

Multi-modal large language models (MLLMs) have demonstrated promising capabilities across various tasks by integrating textual and visual information to achieve visual understanding in complex scenarios. Despite the avai…

Question AnsweringVisual Question Answering

VNHSGE: VietNamese High School Graduation Examination Dataset for Large Language Models

2023-05-20 · Dao Xuan-Quy, Le Ngoc-Bich, Vo The-Duy, Phan Xuan-Dung 외

The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article. The dataset, which covers nine subjects, was generated…

Multiple-choiceQuestion AnsweringQuestion RewritingReading Comprehension+2