Papers Overall - Test
“Overall - Test” 태그가 달린 논문 34편 · 필터 해제
Underage Detection through a Multi-Task and MultiAge Approach for Screening Minors in Unconstrained Imagery
Accurate automatic screening of minors in unconstrained images demands models that are robust to distribution shift and resilient to the children under-representation in publicly available data. To overcome these issues,…
Age EstimationOverall - TestStochastic OptimizationAI5GTest: AI-Driven Specification-Aware Automated Testing and Validation of 5G O-RAN Components
The advent of Open Radio Access Networks (O-RAN) has transformed the telecommunications industry by promoting interoperability, vendor diversity, and rapid innovation. However, its disaggregated architecture introduces c…
Overall - TestDeep Modeling and Optimization of Medical Image Classification
Deep models, such as convolutional neural networks (CNNs) and vision transformer (ViT), demonstrate remarkable performance in image classification. However, those deep models require large data to fine-tune, which is imp…
AvgClassificationFederated Learningimage-classification+3Modeling speech emotion with label variance and analyzing performance across speakers and unseen acoustic conditions
Spontaneous speech emotion data usually contain perceptual grades where graders assign emotion score after listening to the speech files. Such perceptual grades introduce uncertainty in labels due to grader opinion varia…
Emotion RecognitionOverall - TestUnify and Triumph: Polyglot, Diverse, and Self-Consistent Generation of Unit Tests with LLMs
Large language model (LLM)-based test generation has gained attention in software engineering, yet most studies evaluate LLMs' ability to generate unit tests in a single attempt for a given language, missing the opportun…
DiversityLarge Language ModelOverall - TestCost-Saving LLM Cascades with Early Abstention
LLM cascades deploy small LLMs to answer most queries, limiting the use of large and expensive LLMs to difficult queries. This approach can significantly reduce costs without impacting performance. However, risk-sensitiv…
GSM8KMMLUOverall - TestTriviaQA+1Classifier Enhanced Deep Learning Model for Erythroblast Differentiation with Limited Data
Hematological disorders, which involve a variety of malignant conditions and genetic diseases affecting blood formation, present significant diagnostic challenges. One such major challenge in clinical settings is differe…
DiagnosticOverall - TestGradual Learning: Optimizing Fine-Tuning with Partially Mastered Knowledge in Large Language Models
During the pretraining phase, large language models (LLMs) acquire vast amounts of knowledge from extensive text corpora. Nevertheless, in later stages such as fine-tuning and inference, the model may encounter knowledge…
HallucinationOverall - TestArtificial Data Point Generation in Clustered Latent Space for Small Medical Datasets
One of the growing trends in machine learning is the use of data generation techniques, since the performance of machine learning models is dependent on the quantity of the training dataset. However, in many medical appl…
Overall - TestSynthetic Data GenerationEfficient Training of Deep Neural Operator Networks via Randomized Sampling
Neural operators (NOs) employ deep neural networks to learn mappings between infinite-dimensional function spaces. Deep operator network (DeepONet), a popular NO architecture, has demonstrated success in the real-time pr…
Overall - TestThe Future of Software Testing: AI-Powered Test Case Generation and Validation
Software testing is a crucial phase in the software development lifecycle (SDLC), ensuring that products meet necessary functional, performance, and quality benchmarks before release. Despite advancements in automation, …
Overall - Testsoftware testingOptimal Layer Selection for Latent Data Augmentation
While data augmentation (DA) is generally applied to input data, several studies have reported that applying DA to hidden layers in neural networks, i.e., feature augmentation, can improve performance. However, in previo…
Data Augmentationimage-classificationImage ClassificationOverall - Test+1WATT: Weight Average Test-Time Adaptation of CLIP
Vision-Language Models (VLMs) such as CLIP have yielded unprecedented performance for zero-shot image classification, yet their generalization capability may still be seriously challenged when confronted to domain shifts…
image-classificationImage ClassificationOverall - TestTest-time Adaptation+1Network two-sample test for block models
We consider the two-sample testing problem for networks, where the goal is to determine whether two sets of networks originated from the same stochastic model. Assuming no vertex correspondence and allowing for different…
Graph MatchingOverall - TestStochastic Block ModelTwo-sample testingmmID: High-Resolution mmWave Imaging for Human Identification
Achieving accurate human identification through RF imaging has been a persistent challenge, primarily attributed to the limited aperture size and its consequent impact on imaging resolution. The existing imaging solution…
Activity RecognitionOverall - TestPose EstimationSmall Language Models Fine-tuned to Coordinate Larger Language Models improve Complex Reasoning
Large Language Models (LLMs) prompted to generate chain-of-thought (CoT) exhibit impressive reasoning capabilities. Recent attempts at prompt decomposition toward solving complex, multi-step reasoning problems depend on …
Overall - TestProblem DecompositionTransferable Availability Poisoning Attacks
We consider availability data poisoning attacks, where an adversary aims to degrade the overall test accuracy of a machine learning model by crafting small perturbations to its training data. Existing poisoning strategie…
Contrastive LearningData PoisoningOverall - TestContraction Properties of the Global Workspace Primitive
To push forward the important emerging research field surrounding multi-area recurrent neural networks (RNNs), we expand theoretically and empirically on the provably stable RNNs of RNNs introduced by Kozachkov et al. in…
Overall - TestTargeted Data Generation: Finding and Fixing Model Weaknesses
Even when aggregate accuracy is high, state-of-the-art NLP models often fail systematically on specific subgroups of data, resulting in unfair outcomes and eroding user trust. Additional data collection may not help in a…
Data AugmentationNatural Language InferenceOverall - TestSentiment AnalysisHave LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models
The performance of large language models (LLMs) on existing reasoning benchmarks has significantly improved over the past years. In response, we present JEEBench, a considerably more challenging benchmark dataset for eva…
Overall - Test