paper-with-me

홈 › Papers

A foundation model enpowered by a multi-modal prompt engine for universal seismic geobody interpretation across surveys

2024-09-08 · Hang Gao, Xinming Wu, Luming Liang, Hanlin Sheng, Xu Si, Gao Hui, Yaxing Li

Seismic geobody interpretation is crucial for structural geology studies and various engineering applications. Existing deep learning methods show promise but lack support for multi-modal inputs and struggle to generalize to different geobody types or surveys. We introduce a promptable foundation model for interpreting any geobodies across seismic surveys. This model integrates a pre-trained vision foundation model (VFM) with a sophisticated multi-modal prompt engine. The VFM, pre-trained on massive natural images and fine-tuned on seismic data, provides robust feature extraction for cross-survey generalization. The prompt engine incorporates multi-modal prior information to iteratively refine geobody delineation. Extensive experiments demonstrate the model's superior accuracy, scalability from 2D to 3D, and generalizability to various geobody types, including those unseen during training. To our knowledge, this is the first highly scalable and versatile multi-modal foundation model capable of interpreting any geobodies across surveys while supporting real-time interactions. Our approach establishes a new paradigm for geoscientific data interpretation, with broad potential for transfer to other tasks.

📄 PDF Abstract BibTeX arXiv:2409.04962

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Survey of Automatic Prompt Engineering: An Optimization Perspective

2025-02-17 · Wenwu Li, Xiangfeng Wang, Wenhao Li, Bo Jin

The rise of foundation models has shifted focus from resource-intensive fine-tuning to prompt engineering, a paradigm that steers model behavior through input design rather than weight updates. While manual prompt engine…

cross-modal alignmentPrompt EngineeringSurvey

Sub-Region-Aware Modality Fusion and Adaptive Prompting for Multi-Modal Brain Tumor Segmentation

2026-01-22 · Shadi Alijani, Fereshteh Aghaee Meibodi, Homayoun Najjaran arxiv

The successful adaptation of foundation models to multi-modal medical imaging is a critical yet unresolved challenge. Existing models often struggle to effectively fuse information from multiple sources and adapt to the …

Brain Tumor SegmentationPrompt Engineering

PF3Det: A Prompted Foundation Feature Assisted Visual LiDAR 3D Detector

2025-04-04 · Kaidong Li, Tianxiao Zhang, Kuan-Chuan Peng, Guanghui Wang

3D object detection is crucial for autonomous driving, leveraging both LiDAR point clouds for precise depth information and camera images for rich semantic information. Therefore, the multi-modal methods that combine bot…

3D Object DetectionAutonomous Drivingobject-detectionObject Detection+1

Apollo: Zero-shot MultiModal Reasoning with Multiple Experts

2023-10-25 · Daniela Ben-David, Tzuf Paz-Argaman, Reut Tsarfaty

We propose a modular framework that leverages the expertise of different foundation models over different modalities and domains in order to perform a single, complex, multi-modal task, without relying on prompt engineer…

Image CaptioningMultimodal ReasoningPrompt Engineering

SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials

2025-09-25 · Qixin Wan, Zilong Wang, Jingwen Zhou, Wanting Wang 외 arxiv

Foundation models have shown remarkable capabilities in various domains, but their performance on complex, multimodal engineering problems remains largely unexplored. We introduce SoM-1K, the first large-scale multimodal…

Multimodal Reasoning