paper-with-me

홈 › Papers

SemScore: Automated Evaluation of Instruction-Tuned LLMs based on Semantic Textual Similarity

2024-01-30 · Ansar Aynetdinov, Alan Akbik

Instruction-tuned Large Language Models (LLMs) have recently showcased remarkable advancements in their ability to generate fitting responses to natural language instructions. However, many current works rely on manual evaluation to judge the quality of generated responses. Since such manual evaluation is time-consuming, it does not easily scale to the evaluation of multiple models and model variants. In this short paper, we propose a straightforward but remarkably effective evaluation metric called SemScore, in which we directly compare model outputs to gold target responses using semantic textual similarity (STS). We conduct a comparative evaluation of the model outputs of 12 prominent instruction-tuned LLMs using 8 widely-used evaluation metrics for text generation. We find that our proposed SemScore metric outperforms all other, in many cases more complex, evaluation metrics in terms of correlation to human evaluation. These findings indicate the utility of our proposed metric for the evaluation of instruction-tuned LLMs.

📄 PDF Abstract BibTeX arXiv:2401.17072

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Textual SimilaritySTSText Generation

Similar Papers 제목 키워드 기반

Optimizing Psychological Counseling with Instruction-Tuned Large Language Models

2024-06-19 · Wenjie Li, Tianyu Sun, Kun Qian, Wenhong Wang

The advent of large language models (LLMs) has significantly advanced various fields, including natural language processing and automated dialogue systems. This paper explores the application of LLMs in psychological cou…

Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation

2025-10-10 · Fanwei Zhu, Jinke Yu, Zulong Chen, Ying Zhou 외 arxiv

Automated resume information extraction is critical for scaling talent acquisition, yet its real-world deployment faces three major challenges: the extreme heterogeneity of resume layouts and content, the high cost and l…

Information Extraction

Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback

2023-07-29 · Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen 외

A key technology for the development of large language models (LLMs) involves instruction tuning that helps align the models' responses with human expectations to realize impressive learning abilities. Two major approach…

EvalYaks: Instruction Tuning Datasets and LoRA Fine-tuned Models for Automated Scoring of CEFR B2 Speaking Assessment Transcripts

2024-08-22 · Nicy Scaria, Silvester John Joseph Kennedy, Thomas Latinovich, Deepak Subramani

Relying on human experts to evaluate CEFR speaking assessments in an e-learning environment creates scalability challenges, as it limits how quickly and widely assessments can be conducted. We aim to automate the evaluat…

Instruction-Tuned Video-Audio Models Elucidate Functional Specialization in the Brain

2025-06-09 · Subba Reddy Oota, Khushbu Pahwa, Prachi Jindal, Satya Sai Srinath Namburi 외

Recent voxel-wise multimodal brain encoding studies have shown that multimodal large language models (MLLMs) exhibit a higher degree of brain alignment compared to unimodal models in both unimodal and multimodal stimulus…

Disentanglement