paper-with-me

홈 › Papers

ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time

2023-04-27 · Shangqing Tu, Chunyang Li, Jifan Yu, Xiaozhi Wang, Lei Hou, Juanzi Li

ChatGPT has achieved great success and can be considered to have acquired an infrastructural status. There are abundant works for evaluating ChatGPT on benchmarks. However, existing benchmarks encounter two challenges: (1) Disregard for periodical evaluation and (2) Lack of fine-grained features. In this paper, we construct ChatLog, an ever-updating dataset with large-scale records of diverse long-form ChatGPT responses for 21 NLP benchmarks from March, 2023 to now. We conduct a comprehensive performance evaluation to find that most capabilities of ChatGPT improve over time except for some abilities, and there exists a step-wise evolving pattern of ChatGPT. We further analyze the inherent characteristics of ChatGPT by extracting the knowledge and linguistic features. We find some stable features that stay unchanged and apply them on the detection of ChatGPT-generated texts to improve the robustness of cross-version detection. We will continuously maintain our project at \url{https://github.com/THU-KEG/ChatLog/}.

📄 PDF Abstract BibTeX arXiv:2304.14106

Code (1)

thu-keg/chatlog 공식 구현

Tasks

Natural Language Understanding

Similar Papers 제목 키워드 기반

ChatLogic: Integrating Logic Programming with Large Language Models for Multi-Step Reasoning

2024-07-14 · Zhongsheng Wang, Jiamou Liu, Qiming Bao, Hongfei Rong 외

Large language models (LLMs) such as ChatGPT and GPT-4 have demonstrated impressive capabilities in various generative tasks. However, their performance is often hampered by limitations in accessing and leveraging long-t…

Language ModelingLanguage Modelling

SmartSales: Sales Script Extraction and Analysis from Sales Chatlog

2022-04-19 · Hua Liang, Tianyu Liu, Peiyi Wang, Mengliang Rao 외

In modern sales applications, automatic script extraction and management greatly decrease the need for human labor to collect the winning sales scripts, which largely boost the success rate for sales and can be shared ac…

Management

DialogQAE: N-to-N Question Answer Pair Extraction from Customer Service Chatlog

2022-12-14 · Xin Zheng, Tianyu Liu, Haoran Meng, Xu Wang 외

Harvesting question-answer (QA) pairs from customer service chatlog in the wild is an efficient way to enrich the knowledge base for customer service chatbots in the cold start or continuous integration scenarios. Prior …

Retrieval

HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats

2026-04-30 · Rebecca Soskin Hicks, Mikhail Trofimov, Dominick Lim, Rahul K. Arora 외 arxiv

Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited. We introduce HealthBench Professional, an open benchmark for evaluat…

From ChatGPT to DeepSeek AI: A Comprehensive Analysis of Evolution, Deviation, and Future Implications in AI-Language Models

2025-04-04 · Simrandeep Singh, Shreya Bansal, Abdulmotaleb El Saddik, Mukesh Saini

The rapid advancement of artificial intelligence (AI) has reshaped the field of natural language processing (NLP), with models like OpenAI ChatGPT and DeepSeek AI. Although ChatGPT established a strong foundation for con…

Multiple-choice