paper-with-me

홈 › Papers

Using Contextually Aligned Online Reviews to Measure LLMs' Performance Disparities Across Language Varieties

2025-02-10 · Zixin Tang, Chieh-Yang Huang, Tsung-Che Li, Ho Yin Sam Ng, Hen-Hsen Huang, Ting-Hao 'Kenneth' Huang

A language can have different varieties. These varieties can affect the performance of natural language processing (NLP) models, including large language models (LLMs), which are often trained on data from widely spoken varieties. This paper introduces a novel and cost-effective approach to benchmark model performance across language varieties. We argue that international online review platforms, such as Booking.com, can serve as effective data sources for constructing datasets that capture comments in different language varieties from similar real-world scenarios, like reviews for the same hotel with the same rating using the same language (e.g., Mandarin Chinese) but different language varieties (e.g., Taiwan Mandarin, Mainland Mandarin). To prove this concept, we constructed a contextually aligned dataset comprising reviews in Taiwan Mandarin and Mainland Mandarin and tested six LLMs in a sentiment analysis task. Our results show that LLMs consistently underperform in Taiwan Mandarin.

📄 PDF Abstract BibTeX arXiv:2502.07058

Code (0)

등록된 구현이 없습니다.

Tasks

Sentiment Analysis

Similar Papers 제목 키워드 기반

CPR: Leveraging LLMs for Topic and Phrase Suggestion to Facilitate Comprehensive Product Reviews

2025-04-18 · Ekta Gujral, Apurva Sinha, Lishi Ji, Bijayani Sanghamitra Mishra

Consumers often heavily rely on online product reviews, analyzing both quantitative ratings and textual descriptions to assess product quality. However, existing research hasn't adequately addressed how to systematically…

PANORAMA: A synthetic PII-laced dataset for studying sensitive data memorization in LLMs

2025-05-18 · Sriram Selvam, Anneswa Ghosh

The memorization of sensitive and personally identifiable information (PII) by large language models (LLMs) poses growing privacy risks as models scale and are increasingly deployed in real-world applications. Existing e…

ArticlesAttributeMemorizationPrivacy Preserving

eC-Tab2Text: Aspect-Based Text Generation from e-Commerce Product Tables

2025-02-20 · Luis Antonio Gutiérrez Guanilo, Mir Tafseer Nayeem, Cristian López, Davood Rafiei

Large Language Models (LLMs) have demonstrated exceptional versatility across diverse domains, yet their application in e-commerce remains underexplored due to a lack of domain-specific datasets. To address this gap, we …

AttributeText Generation

Appraising the Potential Uses and Harms of LLMs for Medical Systematic Reviews

2023-05-19 · Hye Sun Yun, Iain J. Marshall, Thomas A. Trikalinos, Byron C. Wallace

Medical systematic reviews play a vital role in healthcare decision making and policy. However, their production is time-consuming, limiting the availability of high-quality and up-to-date evidence summaries. Recent adva…

Decision MakingHallucination

PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality

2026-06-18 · Zeyuan Chen, Ziqing Yang, Yihan Ma, Michael Backes 외 arxiv

As academic submissions grow, the traditional peer review process struggles to keep up, raising concerns about quality and fairness. A trend of using large language models (LLMs) for assistance has emerged. In this work,…

Prompt Engineering