Assessing how hyperparameters impact Large Language Models' sarcasm detection performance
Sarcasm detection is challenging for both humans and machines. This work explores how model characteristics impact sarcasm detection in OpenAI's GPT, and Meta's Llama-2 models, given their strong natural language understanding, and popularity. We evaluate fine-tuned and zero-shot models across various sizes, releases, and hyperparameters. Experiments were conducted on the political and balanced (pol-bal) portion of the popular Self-Annotated Reddit Corpus (SARC2.0) sarcasm dataset. Fine-tuned performance improves monotonically with model size within a model family, while hyperparameter tuning also impacts performance. In the fine-tuning scenario, full precision Llama-2-13b achieves state-of-the-art accuracy and $F_1$-score, both measured at 0.83, comparable to average human performance. In the zero-shot setting, one GPT-4 model achieves competitive performance to prior attempts, yielding an accuracy of 0.70 and an $F_1$-score of 0.75. Furthermore, a model's performance may increase or decline with each release, highlighting the need to reassess performance after each release.
Code (0)
등록된 구현이 없습니다.
Tasks
Natural Language UnderstandingSarcasm DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Impact of emoji exclusion on the performance of Arabic sarcasm detection models
The complex challenge of detecting sarcasm in Arabic speech on social media is increased by the language diversity and the nature of sarcastic expressions. There is a significant gap in the capability of existing models …
NavigateSarcasm DetectionCofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal Models
Social media abounds with multimodal sarcasm, and identifying sarcasm targets is particularly challenging due to the implicit incongruity not directly evident in the text and image modalities. Current methods for Multimo…
Language ModelingLanguage ModellingMultimodal ReasoningSarcasm Detection+1Sarcasm Detection in Twitter -- Performance Impact while using Data Augmentation: Word Embeddings
Sarcasm is the use of words usually used to either mock or annoy someone, or for humorous purposes. Sarcasm is largely used in social networks and microblogging websites, where people mock or censure in a way that makes …
Data AugmentationOpinion MiningSarcasm DetectionSentiment Analysis+1Towards Assessing the Impact of Bayesian Optimization's Own Hyperparameters
Bayesian Optimization (BO) is a common approach for hyperparameter optimization (HPO) in automated machine learning. Although it is well-accepted that HPO is crucial to obtain well-performing machine learning models, tun…
Bayesian OptimizationBIG-bench Machine LearningHyperparameter OptimizationNeural Architecture SearchPresence of informal language, such as emoticons, hashtags, and slang, impact the performance of sentiment analysis models on social media text?
This study aimed to investigate the influence of the presence of informal language, such as emoticons and slang, on the performance of sentiment analysis models applied to social media text. A convolutional neural networ…
Sentiment Analysis