paper-with-me

홈 › Papers

Something's Fishy In The Data Lake: A Critical Re-evaluation of Table Union Search Benchmarks

2025-05-27 · Allaa Boutaleb, Bernd Amann, Hubert Naacke, Rafael Angarita

Recent table representation learning and data discovery methods tackle table union search (TUS) within data lakes, which involves identifying tables that can be unioned with a given query table to enrich its content. These methods are commonly evaluated using benchmarks that aim to assess semantic understanding in real-world TUS tasks. However, our analysis of prominent TUS benchmarks reveals several limitations that allow simple baselines to perform surprisingly well, often outperforming more sophisticated approaches. This suggests that current benchmark scores are heavily influenced by dataset-specific characteristics and fail to effectively isolate the gains from semantic understanding. To address this, we propose essential criteria for future benchmarks to enable a more realistic and reliable evaluation of progress in semantic table union search.

📄 PDF Abstract BibTeX arXiv:2505.21329

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

LakeFM: Toward a Foundation Model for Aquatic Ecosystems Using Irregular Multivariate Multi-depth Time Series Data

2026-06-09 · Abhilash Neog, Sepideh Fatemi, Medha Sawhney, Kazi Sajeed Mehrab 외 arxiv

Understanding and forecasting lake dynamics is critical for monitoring water quality and ecosystem health across lakes and reservoirs. While machine learning methods have been recently applied to ecological time-series d…

The Fishyscapes Benchmark: Measuring Blind Spots in Semantic Segmentation

2019-04-05 · Hermann Blum, Paul-Edouard Sarlin, Juan Nieto, Roland Siegwart 외

Deep learning has enabled impressive progress in the accuracy of semantic segmentation. Yet, the ability to estimate uncertainty and detect failure is key for safety-critical applications like autonomous driving. Existin…

Anomaly DetectionAutonomous DrivingSegmentationSemantic Segmentation

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes

2026-07-14 · Michael Solodko, Steven Gong, Guangwei Yu, Satya Krishna Gorti 외 arxiv

While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Answering questions over enterprise and scientific data lakes requires systems to na…

Question Answering

Deep Lake: a Lakehouse for Deep Learning

2022-09-22 · Sasun Hambardzumyan, Abhinav Tuli, Levon Ghukasyan, Fariz Rahman 외

Traditional data lakes provide critical data infrastructure for analytical workloads by enabling time travel, running SQL queries, ingesting data with ACID transactions, and visualizing petabyte-scale datasets on cloud s…

Decision MakingDeep LearningGPU

SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering

2021-02-18 · Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma 외

Medical visual question answering (Med-VQA) has tremendous potential in healthcare. However, the development of this technology is hindered by the lacking of publicly-available and high-quality labeled datasets for train…

Medical Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)