Benchmarking
2개 벤치마크 · 논문 5,548편 · 이 태스크의 논문 보기 →
Benchmarks
CloudEval-YAML
Wiki-40B
Most implemented
MMDetection: Open MMLab Detection Toolbox and Benchmark
Learning Transferable Visual Models From Natural Language Supervision
Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
CIDEr: Consensus-based Image Description Evaluation
The StarCraft Multi-Agent Challenge
Benchmarking Graph Neural Networks
Papers
Visual Place Recognition for Large-Scale UAV Applications
Visual Place Recognition (vPR) plays a crucial role in Unmanned Aerial Vehicle (UAV) navigation, enabling robust localization across diverse environments. Despite significant advancements, aerial vPR faces unique challen…
BenchmarkingDiversityVisual Place RecognitionTraining Transformers with Enforced Lipschitz Constants
Neural networks are often highly sensitive to input and weight perturbations. This sensitivity has been linked to pathologies such as vulnerability to adversarial examples, divergent training, and overfitting. To combat …
BenchmarkingDisentangling coincident cell events using deep transfer learning and compressive sensing
Accurate single-cell analysis is critical for diagnostics, immunomonitoring, and cell therapy, but coincident events - where multiple cells overlap in a sensing zone - can severely compromise signal fidelity. We present …
BenchmarkingCompressive SensingTransfer LearningMUPAX: Multidimensional Problem Agnostic eXplainable AI
Robust XAI techniques should ideally be simultaneously deterministic, model agnostic, and guaranteed to converge. We propose MULTIDIMENSIONAL PROBLEM AGNOSTIC EXPLAINABLE AI (MUPAX), a deterministic, model agnostic expla…
Anatomical Landmark DetectionAudio ClassificationBenchmarkingFeature Importance+3DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition
The landscape of video recognition has evolved significantly, shifting from traditional Convolutional Neural Networks (CNNs) to Transformer-based architectures for improved accuracy. While 3D CNNs have been effective at …
BenchmarkingKnowledge DistillationSpatio-temporal Action RecognitionTemporal Action Localization+2DCR: Quantifying Data Contamination in LLMs Evaluation
The rapid advancement of large language models (LLMs) has heightened concerns about benchmark data contamination (BDC), where models inadvertently memorize evaluation data, inflating performance metrics and undermining g…
Arithmetic ReasoningBenchmarkingComputational EfficiencyFake News Detection+1