paper-with-me

홈 › Papers

LUCAS-MEGA: A Large-Scale Multimodal Dataset for Representation Learning in Soil-Environment Systems

2026-05-05 · Kuangdai Leng, Simon Jeffery, Panos Panagos, Tarje Nissen-Meyer arxiv

Understanding soil is fundamental to agriculture, carbon cycling, and environmental sustainability, yet progress is limited by fragmented and heterogeneous datasets that constrain modeling to small-scale predictive settings rather than high-dimensional representation learning. We introduce LUCAS-MEGA, a large-scale multimodal dataset constructed through systematic data fusion of European soil-environment observations, with the LUCAS survey as its backbone. The fused dataset comprises over 70,000 samples and more than 1,000 features spanning physical, chemical, environmental, biological, and visual attributes, aggregated from 68 source datasets. To enable integration at scale, we develop SoilFuser, a multi-agent, human-in-the-loop data fusion pipeline that standardizes heterogeneous data formats and measurement protocols, resolves inconsistencies and invalid entries (e.g., unit inconsistencies, codebook mismatches, and erroneous values), incorporates natural language annotations, and harmonizes multimodal attributes and metadata into a unified, machine learning-ready feature space. The resulting dataset captures key characteristics of real-world soil observations, including multimodality, uneven feature coverage, and heterogeneous uncertainty. To demonstrate the usability of LUCAS-MEGA for data-driven modeling, we pretrain a multimodal tabular transformer (SoilFormer) using a self-supervised objective based on feature masking, achieving stable training, strong predictive performance, and representations that support uncertainty-aware prediction. We further show that the learned representations recover relationships consistent with established soil processes. LUCAS-MEGA is released with open access and is accompanied by composable, agent-friendly APIs that support structured querying and data-driven workflows.

📄 PDF Abstract BibTeX arXiv:2605.04323

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Deep Lucas-Kanade Homography for Multimodal Image Alignment

2021-04-22 · CVPR 2021 1 · Yiming Zhao, Xinming Huang, Ziming Zhang

Estimating homography to align image pairs captured by different sensors or image pairs with large appearance changes is an important and general challenge for many computer vision applications. In contrast to others, we…

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models

2024-11-30 · Yutong Zhou, Masahiro Ryo

We introduce AgriBench, the first agriculture benchmark designed to evaluate MultiModal Large Language Models (MM-LLMs) for agriculture applications. To further address the agriculture knowledge-based dataset limitation …

DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection

2025-10-26 · Kangran Zhao, Yupeng Chen, Xiaoyu Zhang, Yize Chen 외 arxiv

The misuse of advanced generative AI models has resulted in the widespread proliferation of falsified data, particularly forged human-centric audiovisual content, which poses substantial societal risks (e.g., financial f…

DeepFake Detection

Soil Texture Classification with 1D Convolutional Neural Networks based on Hyperspectral Data

2019-01-15 · Felix M. Riese, Sina Keller

Soil texture is important for many environmental processes. In this paper, we study the classification of soil texture based on hyperspectral data. We develop and implement three 1-dimensional (1D) convolutional neural n…

General ClassificationTexture Classification

MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

2024-12-19 · Junjie Zhou, Zheng Liu, Ze Liu, Shitao Xiao 외

Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that lever…

Image RetrievalRetrievalZero-Shot Composed Image Retrieval (ZS-CIR)