Adaptive Testing for LLM-Based Applications: A Diversity-based Approach
The recent surge of building software systems powered by Large Language Models (LLMs) has led to the development of various testing frameworks, primarily focused on treating prompt templates as the unit of testing. Despite the significant costs associated with test input execution and output assessment, the curation of optimized test suites is yet overlooked in these tools, which calls for tailored test selection or prioritization strategies. In this paper, we show that diversity-based testing techniques, such as Adaptive Random Testing (ART) with appropriate string distance metrics, can be effectively applied to the testing of prompt templates. Our proposed adaptive testing approach adjusts the conventional ART process to this context by selecting new test inputs based on scores derived from existing test suite and their labelling results. Our results, obtained using various implementations that explore several string-based distances, confirm that our approach enables the discovery of failures with reduced testing budgets and promotes the generation of more varied outputs.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversitySimilar Papers 제목 키워드 기반
DATTA: Towards Diversity Adaptive Test-Time Adaptation in Dynamic Wild World
Test-time adaptation (TTA) effectively addresses distribution shifts between training and testing data by adjusting models on test samples, which is crucial for improving model inference in real-world applications. Howev…
DiversityTest-time AdaptationQuality meets Diversity: A Model-Agnostic Framework for Computerized Adaptive Testing
Computerized Adaptive Testing (CAT) is emerging as a promising testing application in many scenarios, such as education, game and recruitment, which targets at diagnosing the knowledge mastery levels of examinees on requ…
Active LearningDiversityPEOAT: Personalization-Guided Evolutionary Question Assembly for One-Shot Adaptive Testing
With the rapid advancement of intelligent education, Computerized Adaptive Testing (CAT) has attracted increasing attention by integrating educational psychology with deep learning technologies. Unlike traditional paper-…
DiCriTest: Testing Scenario Generation for Decision-Making Agents Considering Diversity and Criticality
The growing deployment of decision-making agents in dynamic environments increases the demand for safety verification. While critical testing scenario generation has emerged as an appealing verification methodology, effe…
Dimensionality ReductionGMOCAT: A Graph-Enhanced Multi-Objective Method for Computerized Adaptive Testing
Computerized Adaptive Testing(CAT) refers to an online system that adaptively selects the best-suited question for students with various abilities based on their historical response records. Most CAT methods only focus o…
DiversityGraph Neural NetworkMulti-Objective Reinforcement Learning