paper-with-me

Papers

Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations

2025-10-20 · Tal Barami, Nimrod Berman, Ilan Naiman, Amos H. Hason, Rotem Ezra, Omri Azencot arxiv

Learning disentangled representations in sequential data is a key goal in deep learning, with broad applications in vision, audio, and time series. While real-world data involves multiple interacting semantic factors over time, prior work has mostly focused on simpler two-factor static and dynamic settings, primarily because such settings make data collection easier, thereby overlooking the inherently multi-factor nature of real-world data. We introduce the first standardized benchmark for evaluating multi-factor sequential disentanglement across six diverse datasets spanning video, audio, and time series. Our benchmark includes modular tools for dataset integration, model development, and evaluation metrics tailored to multi-factor analysis. We additionally propose a post-hoc Latent Exploration Stage to automatically align latent dimensions with semantic factors, and introduce a Koopman-inspired model that achieves state-of-the-art results. Moreover, we show that Vision-Language Models can automate dataset annotation and serve as zero-shot disentanglement evaluators, removing the need for manual labels and human intervention. Together, these contributions provide a robust and scalable foundation for advancing multi-factor sequential disentanglement. Our code is available on GitHub, and the datasets and trained models are available on Hugging Face.

📄 PDF Abstract BibTeX arXiv:2510.17313

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sequential Disentanglement by Extracting Static Information From A Single Sequence Element

2024-06-26 · Nimrod Berman, Ilan Naiman, Idan Arbiv, Gal Fadlon 외

One of the fundamental representation learning tasks is unsupervised sequential disentanglement, where latent codes of inputs are decomposed to a single static factor and a sequence of dynamic factors. To extract this la…

Data AugmentationDisentanglementInductive BiasRepresentation Learning

DiViD: Disentangled Video Diffusion for Static-Dynamic Factorization

2025-07-18 · Marzieh Gheisari, Auguste Genovesio arxiv

Unsupervised disentanglement of static appearance and dynamic motion in video remains a fundamental challenge, often hindered by information leakage and blurry reconstructions in existing VAE- and GAN-based approaches. W…

DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across Modalities

2025-10-07 · Hedi Zisling, Ilan Naiman, Nimrod Berman, Supasorn Suwajanakorn 외 arxiv

Unsupervised representation learning, particularly sequential disentanglement, aims to separate static and dynamic factors of variation in data without relying on labels. This remains a challenging problem, as existing a…

Representation Learning

Towards Robust Unsupervised Disentanglement of Sequential Data -- A Case Study Using Music Audio

2022-05-12 · Yin-Jyun Luo, Sebastian Ewert, Simon Dixon

Disentangled sequential autoencoders (DSAEs) represent a class of probabilistic graphical models that describes an observed sequence with dynamic latent variables and a static latent variable. The former encode informati…

Data AugmentationDisentanglementInductive Bias

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness

2026-05-27 · Manuel Frank, Haithem Afli arxiv

Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding robustness is multidimensional, since models respond differently to dif…