Papers Disentanglement
“Disentanglement” 태그가 달린 논문 1,854편 · 필터 해제
Instruction-Tuned Video-Audio Models Elucidate Functional Specialization in the Brain
Recent voxel-wise multimodal brain encoding studies have shown that multimodal large language models (MLLMs) exhibit a higher degree of brain alignment compared to unimodal models in both unimodal and multimodal stimulus…
DisentanglementLearning Speaker-Invariant Visual Features for Lipreading
Lipreading is a challenging cross-modal task that aims to convert visual lip movements into spoken text. Existing lipreading methods often extract visual features that include speaker-specific lip attributes (e.g., shape…
DisentanglementLipreadingSpeaker RecognitionHalf-AVAE: Adversarial-Enhanced Factorized and Structured Encoder-Free VAE for Underdetermined Independent Component Analysis
This study advances the Variational Autoencoder (VAE) framework by addressing challenges in Independent Component Analysis (ICA) under both determined and underdetermined conditions, focusing on enhancing the independenc…
Causal InferenceDecoderDisentanglementVariational InferenceSLAC: Simulation-Pretrained Latent Action Space for Whole-Body Real-World RL
Building capable household and industrial robots requires mastering the control of versatile, high-degree-of-freedom (DoF) systems such as mobile manipulators. While reinforcement learning (RL) holds promise for autonomo…
DisentanglementIndustrial RobotsReinforcement Learning (RL)Safe ExplorationTowards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
Expressive voice conversion aims to transfer both speaker identity and expressive attributes from a target speech to a given source speech. In this work, we improve over a self-supervised, non-autoregressive framework wi…
DisentanglementStyle TransferVoice ConversionPartComposer: Learning and Composing Part-Level Concepts from Single-Image Examples
We present PartComposer: a framework for part-level concept learning from single-image examples that enables text-to-image diffusion models to compose novel objects from meaningful components. Existing methods either str…
DisentanglementLASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
Speaker recognition models face challenges in multi-lingual settings due to the entanglement of linguistic information within speaker embeddings. The overlap between vocal traits such as accent, vocal anatomy, and a lang…
AnatomyDisentanglementSpeaker RecognitionCoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consiste…
Dialogue GenerationDisentanglementSentenceHASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
Effective speech representations for spoken language models must balance semantic relevance with acoustic fidelity for high-quality reconstruction. However, existing approaches struggle to achieve both simultaneously. To…
DisentanglementSelf-Supervised LearningDisentangling Granularity: An Implicit Inductive Bias in Factorized VAEs
Despite the success in learning semantically meaningful, unsupervised disentangled representations, variational autoencoders (VAEs) and their variants face a fundamental theoretical challenge: substantial evidence indica…
DisentanglementInductive BiasRevisiting Cross-Modal Knowledge Distillation: A Disentanglement Approach for RGBD Semantic Segmentation
Multi-modal RGB and Depth (RGBD) data are predominant in many domains such as robotics, autonomous driving and remote sensing. The combination of these multi-modal data enhances environmental perception by providing 3D s…
Autonomous DrivingContrastive LearningData AugmentationDisentanglement+3REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
Speech time reversal refers to the process of reversing the entire speech signal in time, causing it to play backward. Such signals are completely unintelligible since the fundamental structures of phonemes and syllables…
DisentanglementSpeaker IdentificationVoice ConversionErasing Concepts, Steering Generations: A Comprehensive Survey of Concept Suppression
Text-to-Image (T2I) models have demonstrated impressive capabilities in generating high-quality and diverse visual content from natural language prompts. However, uncontrolled reproduction of sensitive, copyrighted, or h…
Adversarial RobustnessDisentanglementSpecificityCausal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models
Multimodal Large Language Models (MLLMs) have demonstrated strong performance in visual understanding tasks, yet they often suffer from object hallucinations--generating descriptions of objects that are inconsistent with…
DisentanglementHallucinationLanguage ModelingLanguage Modelling+1Causality and "In-the-Wild" Video-Based Person Re-ID: A Survey
Video-based person re-identification (Re-ID) remains brittle in real-world deployments despite impressive benchmark performance. Most existing models rely on superficial correlations such as clothing, background, or ligh…
counterfactualCounterfactual ReasoningDisentanglementFairness+4BrainStratify: Coarse-to-Fine Disentanglement of Intracranial Neural Dynamics
Decoding speech directly from neural activity is a central goal in brain-computer interface (BCI) research. In recent years, exciting advances have been made through the growing use of intracranial field potential record…
Brain Computer InterfaceDisentanglementQuantizationEta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
Self-supervised learning (SSL) has reduced the reliance on expensive labeling in speech technologies by learning meaningful representations from unannotated data. Since most SSL-based downstream tasks prioritize content …
DisentanglementSelf-Supervised LearningVoice ConversionDisentangled Human Body Representation Based on Unsupervised Semantic-Aware Learning
In recent years, more and more attention has been paid to the learning of 3D human representation. However, the complexity of lots of hand-defined human body constraints and the absence of supervision data limit that the…
DisentanglementPose TransferAn Interpretable Representation Learning Approach for Diffusion Tensor Imaging
Diffusion Tensor Imaging (DTI) tractography offers detailed insights into the structural connectivity of the brain, but presents challenges in effective representation and interpretation in deep learning models. In this …
Contrastive LearningDecoderDisentanglementRepresentation Learning+1JEDI: The Force of Jensen-Shannon Divergence in Disentangling Diffusion Models
We introduce JEDI, a test-time adaptation method that enhances subject separation and compositional alignment in diffusion models without requiring retraining or external supervision. JEDI operates by minimizing semantic…
DisentanglementTest-time Adaptation