paper-with-me

홈 › Papers

Croissant: A Metadata Format for ML-Ready Datasets

2024-03-28 · Mubashara Akhtar, Omar Benjelloun, Costanza Conforti, Luca Foschini, Joan Giner-Miguelez, Pieter Gijsbers, Sujata Goswami, Nitisha Jain, Michalis Karamousadakis, Michael Kuchnik, Satyapriya Krishna, Sylvain Lesage, Quentin Lhoest, Pierre Marcenac, Manil Maskey, Peter Mattson, Luis Oala, Hamidah Oderinwale, Pierre Ruyssen, Tim Santos, Rajat Shinde, Elena Simperl, Arjun Suresh, Goeffry Thomas, Slava Tykhonov, Joaquin Vanschoren, Susheel Varma, Jos van der Velde, Steffen Vogler, Carole-Jean Wu, Luyao Zhang

Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that creates a shared representation across ML tools, frameworks, and platforms. Croissant makes datasets more discoverable, portable, and interoperable, thereby addressing significant challenges in ML data management. Croissant is already supported by several popular dataset repositories, spanning hundreds of thousands of datasets, enabling easy loading into the most commonly-used ML frameworks, regardless of where the data is stored. Our initial evaluation by human raters shows that Croissant metadata is readable, understandable, complete, yet concise.

📄 PDF Abstract BibTeX arXiv:2403.19546

Code (1)

mlcommons/croissant 공식 구현 tf

Tasks

FrictionManagement

Similar Papers 제목 키워드 기반

Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

2026-05-14 · Rafi Al Attrach, Rajna Fani, Sebastian Lobentanzer, Joan Giner-Miguelez 외 arxiv

Croissant has emerged as the metadata standard for machine learning datasets, providing a structured, JSON-LD-based format that makes dataset discovery, automated ingestion, and reproducible analysis machine-checkable ac…

Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations

2026-05-28 · Omar Benjelloun, Leonardo Martins Bianco, Isabelle Guyon, Thanh Gia Hieu Khuong 외 arxiv

Reproducibility is fundamental to the scientific method, yet remains a critical challenge in machine learning. Contributing factors include underspecified execution details and brittle software environments. Human-centri…

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation

2025-05-17 · Vincent Koc

Tiny QA Benchmark++ (TQB++) presents an ultra-lightweight, multilingual smoke-test suite designed to give large-language-model (LLM) pipelines a unit-test style safety net dataset that runs in seconds with minimal cost. …

Dataset GenerationGPULarge Language ModelMMLU+2

An Ontology for Machine Learning Interatomic Potentials

2026-07-25 · Daniel Hernández, Jong Hyun Jung, Yuji Ikeda, Yongliang Ou 외 arxiv

Machine learning interatomic potentials (MLIPs) approximate quantum-mechanical energies and forces---conventionally computed by density functional theory (DFT) or wave-function methods---at a fraction of the cost. The fi…

MRM3: Machine Readable ML Model Metadata

2025-05-19 · Andrej Čop, Blaž Bertalanič, Marko Grobelnik, Carolina Fortuna

As the complexity and number of machine learning (ML) models grows, well-documented ML models are essential for developers and companies to use or adapt them to their specific use cases. Model metadata, already present i…

model