paper-with-me

홈 › Papers

Trusting CHATGPT: how minor tweaks in the prompts lead to major differences in sentiment classification

2025-04-16 · Jaime E. Cuellar, Oscar Moreno-Martinez, Paula Sofia Torres-Rodriguez, Jaime Andres Pavlich-Mariscal, Andres Felipe Mican-Castiblanco, Juan Guillermo Torres-Hurtado

One fundamental question for the social sciences today is: how much can we trust highly complex predictive models like ChatGPT? This study tests the hypothesis that subtle changes in the structure of prompts do not produce significant variations in the classification results of sentiment polarity analysis generated by the Large Language Model GPT-4o mini. Using a dataset of 100.000 comments in Spanish on four Latin American presidents, the model classified the comments as positive, negative, or neutral on 10 occasions, varying the prompts slightly each time. The experimental methodology included exploratory and confirmatory analyses to identify significant discrepancies among classifications. The results reveal that even minor modifications to prompts such as lexical, syntactic, or modal changes, or even their lack of structure impact the classifications. In certain cases, the model produced inconsistent responses, such as mixing categories, providing unsolicited explanations, or using languages other than Spanish. Statistical analysis using Chi-square tests confirmed significant differences in most comparisons between prompts, except in one case where linguistic structures were highly similar. These findings challenge the robustness and trust of Large Language Models for classification tasks, highlighting their vulnerability to variations in instructions. Moreover, it was evident that the lack of structured grammar in prompts increases the frequency of hallucinations. The discussion underscores that trust in Large Language Models is based not only on technical performance but also on the social and institutional relationships underpinning their use.

📄 PDF Abstract BibTeX arXiv:2504.12180

Code (0)

등록된 구현이 없습니다.

Tasks

Large Language ModelSentiment AnalysisSentiment Classification

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

Testing the Reliability of ChatGPT for Text Annotation and Classification: A Cautionary Remark

2023-04-17 · Michael V. Reiss

Recent studies have demonstrated promising potential of ChatGPT for various text annotation and classification tasks. However, ChatGPT is non-deterministic which means that, as with human coders, identical input can lead…

Classificationtext annotation

Is ChatGPT A Good Keyphrase Generator? A Preliminary Study

2023-03-23 · Mingyang Song, Haiyun Jiang, Shuming Shi, Songfang Yao 외

The emergence of ChatGPT has recently garnered significant attention from the computational linguistics community. To demonstrate its capabilities as a keyphrase generator, we conduct a preliminary evaluation of ChatGPT …

Diversitydocument understandingKeyphrase Generation

Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine

2023-01-20 · Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang 외

This report provides a preliminary evaluation of ChatGPT for machine translation, including translation prompt, multilingual translation, and translation robustness. We adopt the prompts advised by ChatGPT to trigger its…

Machine TranslationSentenceTranslation

Towards Making the Most of ChatGPT for Machine Translation

2023-03-24 · Keqin Peng, Liang Ding, Qihuang Zhong, Li Shen 외

ChatGPT shows remarkable capabilities for machine translation (MT). Several prior studies have shown that it achieves comparable results to commercial systems for high-resource languages, but lags behind in complex tasks…

In-Context LearningMachine TranslationTranslationWord Translation

Personality over Precision: Exploring the Influence of Human-Likeness on ChatGPT Use for Search

2025-11-09 · Mert Yazan, Frederik Bungaran Ishak Situmeang, Suzan Verberne arxiv

Conversational search interfaces, like ChatGPT, offer an interactive, personalized, and engaging user experience compared to traditional search. On the downside, they are prone to cause overtrust issues where users rely …