paper-with-me

홈 › Papers

Evaluation of large language models using an Indian language LGBTI+ lexicon

2023-10-26 · Aditya Joshi, Shruta Rawat, Alpana Dange

Large language models (LLMs) are typically evaluated on the basis of task-based benchmarks such as MMLU. Such benchmarks do not examine responsible behaviour of LLMs in specific contexts. This is particularly true in the LGBTI+ context where social stereotypes may result in variation in LGBTI+ terminology. Therefore, domain-specific lexicons or dictionaries may be useful as a representative list of words against which the LLM's behaviour needs to be evaluated. This paper presents a methodology for evaluation of LLMs using an LGBTI+ lexicon in Indian languages. The methodology consists of four steps: formulating NLP tasks relevant to the expected behaviour, creating prompts that test LLMs, using the LLMs to obtain the output and, finally, manually evaluating the results. Our qualitative analysis shows that the three LLMs we experiment on are unable to detect underlying hateful content. Similarly, we observe limitations in using machine translation as means to evaluate natural language understanding in languages other than English. The methodology presented in this paper can be useful for LGBTI+ lexicons in other languages as well as other domain-specific lexicons. The work done in this paper opens avenues for responsible behaviour of LLMs, as demonstrated in the context of prevalent social perception of the LGBTI+ community.

📄 PDF Abstract BibTeX arXiv:2310.17787

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationMMLUNatural Language Understanding

Similar Papers 제목 키워드 기반

EMBEDDIA hackathon report: Automatic sentiment and viewpoint analysis of Slovenian news corpus on the topic of LGBTIQ+

2021-04-01 · EACL (Hackashop) 2021 4 · Matej Martinc, Nina Perger, Andraž Pelicon, Matej Ulčar 외

We conduct automatic sentiment and viewpoint analysis of the newly created Slovenian news corpus containing articles related to the topic of LGBTIQ+ by employing the state-of-the-art news sentiment classifier and a syste…

ArticlesChange Detection

Towards Large Language Model driven Reference-less Translation Evaluation for English and Indian Languages

2024-04-03 · Vandan Mujadia, Pruthwik Mishra, Arafat Ahsan, Dipti Misra Sharma

With the primary focus on evaluating the effectiveness of large language models for automatic reference-less translation assessment, this work presents our experiments on mimicking human direct assessment to evaluate the…

Language ModelingLanguage ModellingLarge Language ModelTranslation+1

Crosslingual Optimized Metric for Translation Assessment of Indian Languages

2025-09-22 · Arafat Ahsan, Vandan Mujadia, Pruthwik Mishra, Yash Bhaskar 외 arxiv

Automatic evaluation of translation remains a challenging task owing to the orthographic, morphological, syntactic and semantic richness and divergence observed across languages. String-based metrics such as BLEU have pr…

Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages

2025-10-08 · Amir Hossein Yari, Kalmit Kulkarni, Ahmad Raza Khan, Fajri Koto arxiv

While automatic metrics drive progress in Machine Translation (MT) and Text Summarization (TS), existing metrics have been developed and validated almost exclusively for English and other high-resource languages. This na…

Machine TranslationText Summarization

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations

2025-05-24 · Ashwin Sankar, Yoach Lacombe, Sherry Thomas, Praveen Srinivasa Varadhan 외

We introduce RASMALAI, a large-scale speech dataset with rich text descriptions, designed to advance controllable and expressive text-to-speech (TTS) synthesis for 23 Indian languages and English. It comprises 13,000 hou…

Expressive Speech SynthesisSpeech Synthesistext-to-speechText to Speech