paper-with-me

홈 › Papers

OMNIFLOW: A Physics-Grounded Multimodal Agent for Generalized Scientific Reasoning

2026-03-16 · Hao Wu, Yongheng Zhang, Yuan Gao, Fan Xu, Fan Zhang, Ruobing Xie, Ruijian Gou, Yuxuan Liang, Xiaomeng Huang, Xian Wu arxiv

Large Language Models (LLMs) have demonstrated exceptional logical reasoning capabilities but frequently struggle with the continuous spatiotemporal dynamics governed by Partial Differential Equations (PDEs), often resulting in non-physical hallucinations. Existing approaches typically resort to costly, domain-specific fine-tuning, which severely limits cross-domain generalization and interpretability. To bridge this gap, we propose OMNIFLOW, a neuro-symbolic architecture designed to ground frozen multimodal LLMs in fundamental physical laws without requiring domain-specific parameter updates. OMNIFLOW introduces a novel \textit{Semantic-Symbolic Alignment} mechanism that projects high-dimensional flow tensors into topological linguistic descriptors, enabling the model to perceive physical structures rather than raw pixel values. Furthermore, we construct a Physics-Guided Chain-of-Thought (PG-CoT) workflow that orchestrates reasoning through dynamic constraint injection (e.g., mass conservation) and iterative reflexive verification. We evaluate OMNIFLOW on a comprehensive benchmark spanning microscopic turbulence, theoretical Navier-Stokes equations, and macroscopic global weather forecasting. Empirical results demonstrate that OMNIFLOW significantly outperforms traditional deep learning baselines in zero-shot generalization and few-shot adaptation tasks. Crucially, it offers transparent, physically consistent reasoning reports, marking a paradigm shift from black-box fitting to interpretable scientific reasoning.

📄 PDF Abstract BibTeX arXiv:2603.15797

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationDomain GeneralizationWeather ForecastingLogical Reasoning

Similar Papers 제목 키워드 기반

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

2024-12-02 · CVPR 2025 1 · Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Zichun Liao 외

We introduce OmniFlow, a novel generative model designed for any-to-any generation tasks such as text-to-image, text-to-audio, and audio-to-image synthesis. OmniFlow advances the rectified flow (RF) framework used in tex…

Audio SynthesisImage GenerationText Generation

OmniFlow: Human Omnidirectional Optical Flow

2021-04-16 · Roman Seidel, André Apitzsch, Gangolf Hirtz

Optical flow is the motion of a pixel between at least two consecutive video frames and can be estimated through an end-to-end trainable convolutional neural network. To this end, large training datasets are required to …

Optical Flow Estimation

PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation

2025-12-05 · Shima Imani, Seungwhan Moon, Adel Ahmadyan, Lu Zhang 외 arxiv

Evaluating vision-language models (VLMs) in scientific domains like mathematics and physics poses unique challenges that go far beyond predicting final answers. These domains demand conceptual understanding, symbolic rea…

Program Synthesis

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

2026-09-09 · Yizhan Li, Jianxin You, Mengyang Xiong, Yinhuan Chen 외 hf

Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as…

Question Answering

Grounded Gesture Generation: Language, Motion, and Space

2025-07-06 · Anna Deichler, Jim O'Regan, Teo Guichoux, David Johansson 외 arxiv

Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing models typically specialize either in de…

Synthetic Data GenerationGesture Generation