paper-with-me

홈 › Papers

DGFamba: Learning Flow Factorized State Space for Visual Domain Generalization

2025-04-10 · Qi Bi, Jingjun Yi, Hao Zheng, Haolan Zhan, Wei Ji, Yawen Huang, Yuexiang Li

Domain generalization aims to learn a representation from the source domain, which can be generalized to arbitrary unseen target domains. A fundamental challenge for visual domain generalization is the domain gap caused by the dramatic style variation whereas the image content is stable. The realm of selective state space, exemplified by VMamba, demonstrates its global receptive field in representing the content. However, the way exploiting the domain-invariant property for selective state space is rarely explored. In this paper, we propose a novel Flow Factorized State Space model, dubbed as DG-Famba, for visual domain generalization. To maintain domain consistency, we innovatively map the style-augmented and the original state embeddings by flow factorization. In this latent flow space, each state embedding from a certain style is specified by a latent probability path. By aligning these probability paths in the latent space, the state embeddings are able to represent the same content distribution regardless of the style differences. Extensive experiments conducted on various visual domain generalization settings show its state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2504.08019

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Generalization

Similar Papers 제목 키워드 기반

LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow

2026-07-12 · Hang Long, Tianhao Zhao, Junkai Lin, Youjia Zhang 외 arxiv

Flow matching over carefully designed latent representations has recently emerged as a powerful paradigm for topology-aware mesh generation. Existing approaches, however, model vertices and connectivity jointly in a join…

EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning

2026-09-08 · Jingpu Yang, Fengxian Ji, Mingxuan Cui, Yilin Sun 외 arxiv

UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter t…

Video Question AnsweringSpatial Reasoning

Investigation of Factorized Optical Flows as Mid-Level Representations

2022-03-09 · Hsuan-Kung Yang, Tsu-Ching Hsiao, Ting-Hsuan Liao, Hsu-Shen Liu 외

In this paper, we introduce a new concept of incorporating factorized flow maps as mid-level representations, for bridging the perception and the control modules in modular learning based robotic frameworks. To investiga…

Deep Reinforcement LearningOptical Flow Estimationreinforcement-learningReinforcement Learning (RL)

Flow Factorized Representation Learning

2023-09-22 · NeurIPS 2023 11 · Yue Song, T. Anderson Keller, Nicu Sebe, Max Welling

A prominent goal of representation learning research is to achieve representations which are factorized in a useful manner with respect to the ground truth factors of variation. The fields of disentangled and equivariant…

DisentanglementRepresentation Learning

FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery

2026-05-14 · Patrick Kwon, Chen Chen arxiv

Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evidence. This ambiguity is not uniform across the body, as torso pose and root stru…

Human Mesh Recovery