paper-with-me

Papers

MDBench: Benchmarking Data-Driven Methods for Model Discovery

2025-09-24 · Amirmohammad Ziaei Bideh, Aleksandra Georgievska, Jonathan Gryak arxiv

Model discovery aims to uncover governing differential equations of dynamical systems directly from experimental data. Benchmarking such methods is essential for tracking progress and understanding trade-offs in the field. While prior efforts have focused mostly on identifying single equations, typically framed as symbolic regression, there remains a lack of comprehensive benchmarks for discovering dynamical models. To address this, we introduce MDBench, an open-source benchmarking framework for evaluating model discovery methods on dynamical systems. MDBench assesses 12 algorithms on 14 partial differential equations (PDEs) and 63 ordinary differential equations (ODEs) under varying levels of noise. Evaluation metrics include derivative prediction accuracy, model complexity, and equation fidelity. We also introduce seven challenging PDE systems from fluid dynamics and thermodynamics, revealing key limitations in current methods. Our findings illustrate that linear methods and genetic programming methods achieve the lowest prediction error for PDEs and ODEs, respectively. Moreover, linear models are in general more robust against noise. MDBench accelerates the advancement of model discovery methods by offering a rigorous, extensible benchmarking framework and a rich, diverse collection of dynamical system datasets, enabling systematic evaluation, comparison, and improvement of equation accuracy and robustness.

📄 PDF Abstract BibTeX arXiv:2509.20529

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation

2025-07-21 · Yibo He, Shuoran Zhao, Jiaming Huang, Yingjie Fu 외 arxiv

SIMD (Single Instruction Multiple Data) instructions and their compiler intrinsics are widely supported by modern processors to accelerate performance-critical tasks. SIMD intrinsic programming, a trade-off between codin…

Code Generation

3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark

2025-03-26 · Ivan Sviridov, Amina Miftakhova, Artemiy Tereshchenko, Galina Zubkova 외

Though Large Vision-Language Models (LVLMs) are being actively explored in medicine, their ability to conduct telemedicine consultations combining accurate diagnosis with professional dialogue remains underexplored. In t…

DiagnosticMultimodal Reasoning

CMDBench: A Benchmark for Coarse-to-fine Multimodal Data Discovery in Compound AI Systems

2024-06-02 · Yanlin Feng, Sajjadur Rahman, Aaron Feng, Vincent Chen 외

Compound AI systems (CASs) that employ LLMs as agents to accomplish knowledge-intensive tasks via interactions with tools and data retrievers have garnered significant interest within database and AI communities. While t…

Question Answering

WelQrate: Defining the Gold Standard in Small Molecule Drug Discovery Benchmarking

2024-11-14 · Yunchao, Liu, Ha Dong, Xin Wang 외

While deep learning has revolutionized computer-aided drug discovery, the AI community has predominantly focused on model innovation and placed less emphasis on establishing best benchmarking practices. We posit that wit…

BenchmarkingDrug Discovery

CIPCaD-Bench: Continuous Industrial Process datasets for benchmarking Causal Discovery methods

2022-08-02 · Giovanni Menegozzo, Diego Dall'Alba, Paolo Fiorini

Causal relationships are commonly examined in manufacturing processes to support faults investigations, perform interventions, and make strategic decisions. Industry 4.0 has made available an increasing amount of data th…

BenchmarkingCausal DiscoveryFault Detection