paper-with-me

Papers

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

2026-09-10 · Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang, Pengqian Han, Junkai Wang, Wanghan Xu hf

Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluation evidence. Daily discovery draws on 37 sources: 13 direct connectors and 24 first-party research and engineering feeds. The catalog contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. We describe collection and retrieval, audit the full catalog, and examine benchmark saturation, adoption trends, and the limits of score comparisons. A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. We release the web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a command-line interface (CLI) for offline queries, and reproducible analysis.

📄 PDF Abstract BibTeX arXiv:2609.11115

Code (2)

Valiant-Cat/hfpaper
ktwu01/benchmark-radar ★ 207

Similar Papers 제목 키워드 기반

Generative Latent Alignment for Interpretable Radar Based Occupancy Detection in Ambient Assisted Living

2026-01-27 · Huy Trinh arxiv

In this work, we study how to make mmWave radar presence detection more interpretable for Ambient Assisted Living (AAL) settings, where camera-based sensing raises privacy concerns. We propose a Generative Latent Alignme…

Living Labs - An Ethical Challenge for Researchers and Platform Providers

2017-06-22 · Schaer Philipp

The infamous Facebook emotion contagion experiment is one of the most prominent and best-known online experiments based on the concept of what we here call "living labs". In these kinds of experiments, real-world applica…

Experimental DesignInformation RetrievalRetrieval

RF Sensing for Continuous Monitoring of Human Activities for Home Consumer Applications

2020-03-21 · Moeness G. Amin, Arun Ravisankar, Ronny G. Guendel

Radar for indoor monitoring is an emerging area of research and development, covering and supporting different health and wellbeing applications of smart homes, assisted living, and medical diagnosis. We report on a succ…

Medical DiagnosisTranslation

Radar Classification of Contiguous Activities of Daily Living

2019-12-17 · Ronny Gerhard Guendel

We consider radar classifications of Activities of Daily Living (ADL) which can prove beneficial in fall detection, analysis of daily routines, and discerning physical and cognitive human conditions. We focus on contiguo…

ClassificationGeneral ClassificationTranslation

Vision-RADAR fusion for Robotics BEV Detections: A Survey

2023-02-13 · Apoorv Singh

Due to the trending need of building autonomous robotic perception system, sensor fusion has attracted a lot of attention amongst researchers and engineers to make best use of cross-modality information. However, in orde…

object-detectionObject DetectionSensor FusionSurvey