paper-with-me

홈 › Papers

Fine-tuning ChatGPT for Automatic Scoring of Written Scientific Explanations in Chinese

2025-01-12 · Jie Yang, Ehsan Latif, Yuze He, Xiaoming Zhai

The development of explanations for scientific phenomena is essential in science assessment, but scoring student-written explanations remains challenging and resource-intensive. Large language models (LLMs) have shown promise in addressing this issue, particularly in alphabetic languages like English. However, their applicability to logographic languages is less explored. This study investigates the potential of fine-tuning ChatGPT, a leading LLM, to automatically score scientific explanations written in Chinese. Student responses to seven scientific explanation tasks were collected and automatically scored, with scoring accuracy examined in relation to reasoning complexity using the Kendall correlation. A qualitative analysis explored how linguistic features influenced scoring accuracy. The results show that domain-specific adaptation enables ChatGPT to score Chinese scientific explanations with accuracy. However, scoring accuracy correlates with reasoning complexity: a negative correlation for lower-level responses and a positive one for higher-level responses. The model overrates complex reasoning in low-level responses with intricate sentence structures and underrates high-level responses using concise causal reasoning. These correlations stem from linguistic features--simplicity and clarity enhance accuracy for lower-level responses, while comprehensiveness improves accuracy for higher-level ones. Simpler, shorter responses tend to score more accurately at lower levels, whereas longer, information-rich responses yield better accuracy at higher levels. These findings demonstrate the effectiveness of LLMs in automatic scoring within a Chinese context and emphasize the importance of linguistic features and reasoning complexity in fine-tuning scoring models for educational assessments.

📄 PDF Abstract BibTeX arXiv:2501.06704

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fine-tuning ChatGPT for Automatic Scoring

2023-10-16 · Ehsan Latif, Xiaoming Zhai

This study highlights the potential of fine-tuned ChatGPT (GPT-3.5) for automatically scoring student written constructed responses using example assessment tasks in science education. Recent studies on OpenAI's generati…

Language Modelling

Using ChatGPT to Score Essays and Short-Form Constructed Responses

2024-08-18 · Mark D. Shermis

This study aimed to determine if ChatGPT's large language models could match the scoring accuracy of human and machine scores from the ASAP competition. The investigation focused on various prediction models, including l…

FairnessForm

Empirical Study of Large Language Models as Automated Essay Scoring Tools in English Composition__Taking TOEFL Independent Writing Task for Example

2024-01-07 · Wei Xia, Shaoguang Mao, Chanjing Zheng

Large language models have demonstrated exceptional capabilities in tasks involving natural language generation, reasoning, and comprehension. This study aims to construct prompts and comments grounded in the diverse sco…

Automated Essay ScoringPrompt LearningText Generation

Seventeenth-Century Spanish American Notary Records for Fine-Tuning Spanish Large Language Models

2024-06-09 · Shraboni Sarker, Ahmad Tamim Hamad, Hulayyil Alshammari, Viviana Grieco 외

Large language models have gained tremendous popularity in domains such as e-commerce, finance, healthcare, and education. Fine-tuning is a common approach to customize an LLM on a domain-specific dataset for a desired d…

Language ModelingLanguage ModellingMasked Language Modeling

Matching Exemplar as Next Sentence Prediction (MeNSP): Zero-shot Prompt Learning for Automatic Scoring in Science Education

2023-01-20 · Xuansheng Wu, Xinyu He, Tianming Liu, Ninghao Liu 외

Developing models to automatically score students' written responses to science problems is critical for science education. However, collecting and labeling sufficient student responses for training models is time and co…

Prompt LearningSentence