paper-with-me

홈 › Papers

Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding

2024-12-17 · Tadesse Destaw Belay, Israel Abebe Azime, Abinew Ali Ayele, Grigori Sidorov, Dietrich Klakow, Philipp Slusallek, Olga Kolesnikova, Seid Muhie Yimam

Large Language Models (LLMs) show promising learning and reasoning abilities. Compared to other NLP tasks, multilingual and multi-label emotion evaluation tasks are under-explored in LLMs. In this paper, we present EthioEmo, a multi-label emotion classification dataset for four Ethiopian languages, namely, Amharic (amh), Afan Oromo (orm), Somali (som), and Tigrinya (tir). We perform extensive experiments with an additional English multi-label emotion dataset from SemEval 2018 Task 1. Our evaluation includes encoder-only, encoder-decoder, and decoder-only language models. We compare zero and few-shot approaches of LLMs to fine-tuning smaller language models. The results show that accurate multi-label emotion classification is still insufficient even for high-resource languages such as English, and there is a large gap between the performance of high-resource and low-resource languages. The results also show varying performance levels depending on the language and model type. EthioEmo is available publicly to further improve the understanding of emotions in language models and how people convey emotions through various languages.

📄 PDF Abstract BibTeX arXiv:2412.17837

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderEmotion Classification

Similar Papers 제목 키워드 기반

VL-Taboo: An Analysis of Attribute-based Zero-shot Capabilities of Vision-Language Models

2022-09-12 · Felix Vogel, Nina Shvetsova, Leonid Karlinsky, Hilde Kuehne

Vision-language models trained on large, randomly collected data had significant impact in many areas since they appeared. But as they show great performance in various fields, such as image-text-retrieval, their inner w…

AttributeImage-text RetrievalRetrievalText Retrieval+1

MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models

2026-05-14 · Tianwei Chen, Takuya Furusawa, Yuki Hirakawa, Ryotaro Shimizu 외 arxiv

This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language models (MLLMs) to predict the emotions evoked by images. Recent user s…

M2G-Eval: Enhancing and Evaluating Multi-granularity Multilingual Code Generation

2025-12-27 · Fanglin Xu, Wei Zhang, Jian Yang, Guo Chen 외 arxiv

The rapid advancement of code large language models (LLMs) has sparked significant research interest in systematically evaluating their code generation capabilities, yet existing benchmarks predominantly assess models at…

Code Generation

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

2026-03-14 · Quoc-Huy Trinh, Xi Ding, Yang Liu, Zhenyue Qin 외 arxiv

Visual spatial intelligence is critical for medical image interpretation, yet remains largely unexplored in Multimodal Large Language Models (MLLMs) for 3D imaging. This gap persists due to a systemic lack of datasets fe…

Spatial Reasoning

MMRel: A Relation Understanding Benchmark in the MLLM Era

2024-06-13 · Jiahao Nie, Gongjie Zhang, Wenbin An, Yap-Peng Tan 외

Though Multi-modal Large Language Models (MLLMs) have recently achieved significant progress, they often face various problems while handling inter-object relations, i.e., the interaction or association among distinct ob…

DiversityHallucinationObjectRelation+1