paper-with-me

Papers

Exploring Differences Between Tabular Enterprise Data and Public Benchmarks

2026-06-29 · Myung Jun Kim, Maximilian Schambach, Frank Essenberger, Andre Sres, Johannes Höhne arxiv

Tabular data dominate the landscape of data science, increasingly attracting innovative machine learning models and tailored benchmarks. Yet, little is known for enterprise data, where tables constitute the backbone of business operations. To broaden the benchmarking landscape for business applications, this work aims to actualize the characteristics of enterprise data by providing an analysis of data statistics and performance measurements of tabular models such as TabPFN, TabICL and ConTextTab. Through our analysis, we find enterprise data markedly differ from tabular benchmarks and we demonstrate that a tabular model that performs well on typical tabular benchmarks may perform poorly on real world enterprise data -- and vice versa. This lack of generalization underlines the need for additional benchmarks with enterprise-grade characteristics.

📄 PDF Abstract BibTeX arXiv:2606.30452

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SALT-KG: A Benchmark for Semantics-Aware Learning on Enterprise Tables

2026-01-12 · Isaiah Onando Mulang, Felix Sasaki, Tassilo Klein, Jonas Kolk 외 arxiv

Building upon the SALT benchmark for relational prediction (Klein et al., 2024), we introduce SALT-KG, a benchmark for semantics-aware learning on enterprise tables. SALT-KG extends SALT by linking its multi-table transa…

CASPR: Customer Activity Sequence-based Prediction and Representation

2022-11-16 · Pin-Jung Chen, Sahil Bhatnagar, Sagar Goyal, Damian Konrad Kowalczyk 외

Tasks critical to enterprise profitability, such as customer churn prediction, fraudulent account detection or customer lifetime value estimation, are often tackled by models trained on features engineered from customer …

Feature EngineeringPredictionRepresentation Learning

A fuzzy-rough uncertainty measure to discover bias encoded explicitly or implicitly in features of structured pattern classification datasets

2021-08-20 · Gonzalo Nápoles, Lisa Koutsoviti Koumeri

The need to measure bias encoded in tabular data that are used to solve pattern recognition problems is widely recognized by academia, legislators and enterprises alike. In previous work, we proposed a bias quantificatio…

Advancing Retrieval-Augmented Generation for Structured Enterprise and Internal Data

2025-07-16 · Chandana Cheerla arxiv

Organizations increasingly rely on proprietary enterprise data, including HR records, structured reports, and tabular documents, for critical decision-making. While Large Language Models (LLMs) have strong generative cap…

Domain Adaptation for Enterprise Email Search

2019-06-19 · Brandon Tran, Maryam Karimzadehgan, Rama Kumar Pasumarthi, Michael Bendersky 외

In the enterprise email search setting, the same search engine often powers multiple enterprises from various industries: technology, education, manufacturing, etc. However, using the same global ranking model across dif…

Domain AdaptationInformation RetrievalRetrieval