paper-with-me

Papers

ChatGPT Prompting Cannot Estimate Predictive Uncertainty in High-Resource Languages

2023-11-10 · Martino Pelucchi, Matias Valdenegro-Toro

ChatGPT took the world by storm for its impressive abilities. Due to its release without documentation, scientists immediately attempted to identify its limits, mainly through its performance in natural language processing (NLP) tasks. This paper aims to join the growing literature regarding ChatGPT's abilities by focusing on its performance in high-resource languages and on its capacity to predict its answers' accuracy by giving a confidence level. The analysis of high-resource languages is of interest as studies have shown that low-resource languages perform worse than English in NLP tasks, but no study so far has analysed whether high-resource languages perform as well as English. The analysis of ChatGPT's confidence calibration has not been carried out before either and is critical to learn about ChatGPT's trustworthiness. In order to study these two aspects, five high-resource languages and two NLP tasks were chosen. ChatGPT was asked to perform both tasks in the five languages and to give a numerical confidence value for each answer. The results show that all the selected high-resource languages perform similarly and that ChatGPT does not have a good confidence calibration, often being overconfident and never giving low confidence values.

📄 PDF Abstract BibTeX arXiv:2311.06427

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enhancing Medical Support in the Arabic Language Through Personalized ChatGPT Assistance

2024-03-21 · Mohamed Issa, Ahmed Abdelwahed

This Paper discusses the growing popularity of online medical diagnosis as an alternative to traditional doctor visits. It highlights the limitations of existing tools and emphasizes the advantages of using ChatGPT, whic…

Medical Diagnosis

ChatGPT and Deepseek: Can They Predict the Stock Market and Macroeconomy?

2025-02-14 · Jian Chen, Guohao Tang, Guofu Zhou, Wu Zhu

We study whether ChatGPT and DeepSeek can extract information from the Wall Street Journal to predict the stock market and the macroeconomy. We find that ChatGPT has predictive power. DeepSeek underperforms ChatGPT, whic…

Evaluating language models as risk scores

2024-07-19 · André F. Cruz, Moritz Hardt, Celestine Mendler-Dünner

Current question-answering benchmarks predominantly focus on accuracy in realizable prediction tasks. Conditioned on a question and answer-key, does the most likely token match the ground truth? Such benchmarks necessari…

Multiple-choiceQuestion Answering

Can Base ChatGPT be Used for Forecasting without Additional Optimization?

2024-04-11 · Van Pham, Scott Cunningham

This study investigates whether OpenAI's ChatGPT-3.5 and ChatGPT-4 can forecast future events. To evaluate the accuracy of the predictions, we take advantage of the fact that the training data at the time of our experime…

Distractor generation for multiple-choice questions with predictive prompting and large language models

2023-07-30 · Semere Kiros Bitew, Johannes Deleu, Chris Develder, Thomas Demeester

Large Language Models (LLMs) such as ChatGPT have demonstrated remarkable performance across various tasks and have garnered significant attention from both researchers and practitioners. However, in an educational conte…

Distractor GenerationMultiple-choice