paper-with-me

Papers

EyeWorld: A Generative World Model of Ocular State and Dynamics

2026-03-14 · Ziyu Gao, Xinyuan Wu, Xiaolan Chen, Zhuoran Liu, Ruoyu Chen, Bowen Liu, Bingjie Yan, Zhenhan Wang, Kai Jin, Jiancheng Yang, Yih Chung Tham, Mingguang He, Danli Shi arxiv

Ophthalmic decision-making depends on subtle lesion-scale cues interpreted across multimodal imaging and over time, yet most medical foundation models remain static and degrade under modality and acquisition shifts. Here we introduce EyeWorld, a generative world model that conceptualizes the eye as a partially observed dynamical system grounded in clinical imaging. EyeWorld learns an observation-stable latent ocular state shared across modalities, unifying fine-grained parsing, structure-preserving cross-modality translation and quality-robust enhancement within a single framework. Longitudinal supervision further enables time-conditioned state transitions, supporting forecasting of clinically meaningful progression while preserving stable anatomy. By moving from static representation learning to explicit dynamical modeling, EyeWorld provides a unified approach to robust multimodal interpretation and prognosis-oriented simulation in medicine.

📄 PDF Abstract BibTeX arXiv:2603.14039

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

DyST: Towards Dynamic Neural Scene Representations on Real-World Videos

2023-10-09 · Maximilian Seitzer, Sjoerd van Steenkiste, Thomas Kipf, Klaus Greff 외

Visual understanding of the world goes beyond the semantics and flat structure of individual images. In this work, we aim to capture both the 3D structure and dynamics of real-world scenes from monocular real-world video…

Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection

2026-06-24 · Zican Wang, Niloy Mitra arxiv

We present a self-supervised framework for learning implicit 3D physical dynamics directly from video-derived supervisory signals. While current generative video models achieve high visual fidelity, they lack a 3D geomet…

Natural Human Motion Recovery by Aligning High-Order Temporal Dynamics from Monocular Videos

2026-05-26 · Dingkun Wei, Zehong Shen, Yan Xia, Georgios Pavlakos 외 arxiv

Human motion recovered from monocular videos often appears overly smooth or dynamically inconsistent, even when joint positions are numerically accurate. We observe that this limitation stems from the absence of reliable…

Gaussian Sequences with Multi-Scale Dynamics for 4D Reconstruction from Monocular Casual Videos

2026-02-14 · Can Li, Jie Gu, Jingmin Chen, Fangzhou Qiu 외 arxiv

Understanding dynamic scenes from casual videos is critical for scalable robot learning, yet four-dimensional (4D) reconstruction under strictly monocular settings remains highly ill-posed. To address this challenge, our…

SeeClear: Reliable Transparent Object Depth Estimation via Generative Opacification

2026-03-20 · Xiaoying Wang, Yumeng He, Jingkai Shi, Jiayin Lu 외 arxiv

Monocular depth estimation remains challenging for transparent objects, where refraction and transmission are difficult to model and break the appearance assumptions used by depth networks. As a result, state-of-the-art …

Transparent Object Depth EstimationMonocular Depth Estimation