Papers Disentanglement
“Disentanglement” 태그가 달린 논문 1,854편 · 필터 해제
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
Disentangling content and style from a single image, known as content-style decomposition (CSD), enables recontextualization of extracted content and stylization of extracted styles, offering greater creative flexibility…
DisentanglementTowards Imperceptible JPEG Image Hiding: Multi-range Representations-driven Adversarial Stego Generation
Deep hiding has been exploring the hiding capability of deep learning-based models, aiming to conceal image-level messages into cover images and reveal them from generated stego images. Existing schemes are easily detect…
DisentanglementSteganalysisGenerative Head-Mounted Camera Captures for Photorealistic Avatars
Enabling photorealistic avatar animations in virtual and augmented reality (VR/AR) has been challenging because of the difficulty of obtaining ground truth state of faces. It is physically impossible to obtain synchroniz…
DisentanglementReflections Unlock: Geometry-Aware Reflection Disentanglement in 3D Gaussian Splatting for Photorealistic Scenes Rendering
Accurately rendering scenes with reflective surfaces remains a significant challenge in novel view synthesis, as existing methods like Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) often misinterpret ref…
3DGSDisentanglementNeRFNovel View Synthesis+1Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
Domain Generalization (DG) aims to enhance model robustness in unseen or distributionally shifted target domains through training exclusively on source domains. Although existing DG techniques, such as data manipulation,…
DisentanglementDomain GeneralizationCausal-SAM-LLM: Large Language Models as Causal Reasoners for Robust Medical Segmentation
The clinical utility of deep learning models for medical image segmentation is severely constrained by their inability to generalize to unseen domains. This failure is often rooted in the models learning spurious correla…
AnatomyDisentanglementImage SegmentationMedical Image Segmentation+1Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization
Domain Generalization (DG) seeks to develop a versatile model capable of performing effectively on unseen target domains. Notably, recent advances in pre-trained Visual Foundation Models (VFMs), such as CLIP, have demons…
DescriptiveDisentanglementDomain GeneralizationLarge Language Model+1SemFaceEdit: Semantic Face Editing on Generative Radiance Manifolds
Despite multiple view consistency offered by 3D-aware GAN techniques, the resulting images often lack the capacity for localized editing. In response, generative radiance manifolds emerge as an efficient approach for con…
DisentanglementWordCon: Word-level Typography Control in Scene Text Rendering
Achieving precise word-level typography control within generated images remains a persistent challenge. To address it, we newly construct a word-level controlled scene text dataset and introduce the Text-Image Alignment …
Disentanglementparameter-efficient fine-tuningDisentangled representations of microscopy images
Microscopy image analysis is fundamental for different applications, from diagnosis to synthetic engineering and environmental monitoring. Modern acquisition systems have granted the possibility to acquire an escalating …
ClassificationDisentanglementimage-classificationImage Classification+2Optimizing Multilingual Text-To-Speech with Accents & Emotions
State-of-the-art text-to-speech (TTS) systems realize high naturalness in monolingual environments, synthesizing speech with correct multilingual accents (especially for Indic languages) and context-relevant emotions sti…
DisentanglementEmotion Recognitiontext-to-speechText to Speech+1Factorized RVQ-GAN For Disentangled Speech Tokenization
We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge dist…
DisentanglementKnowledge DistillationSpeech TokenizationDreamLight: Towards Harmonious and Consistent Image Relighting
We introduce a model named DreamLight for universal image relighting in this work, which can seamlessly composite subjects into a new background while maintaining aesthetic uniformity in terms of lighting and color tone.…
DisentanglementImage RelightingDisProtEdit: Exploring Disentangled Representations for Multi-Attribute Protein Editing
We introduce DisProtEdit, a controllable protein editing framework that leverages dual-channel natural language supervision to learn disentangled representations of structural and functional properties. Unlike prior appr…
AttributeDisentanglementLarge Language ModelRepresentation LearningAdversarial Disentanglement by Backpropagation with Physics-Informed Variational Autoencoder
Inference and prediction under partial knowledge of a physical system is challenging, particularly when multiple confounding sources influence the measured response. Explicitly accounting for these influences in physics-…
DisentanglementDual-View Disentangled Multi-Intent Learning for Enhanced Collaborative Filtering
Disentangling user intentions from implicit feedback has become a promising strategy to enhance recommendation accuracy and interpretability. Prior methods often model intentions independently and lack explicit supervisi…
Collaborative FilteringDisentanglementTask Adaptation from Skills: Information Geometry, Disentanglement, and New Objectives for Unsupervised Reinforcement Learning
Unsupervised reinforcement learning (URL) aims to learn general skills for unseen downstream tasks. Mutual Information Skill Learning (MISL) addresses URL by maximizing the mutual information between states and skills bu…
DisentanglementDiversityUnsupervised Reinforcement LearningDisentangling Dual-Encoder Masked Autoencoder for Respiratory Sound Classification
Deep neural networks have been applied to audio spectrograms for respiratory sound classification, but it remains challenging to achieve satisfactory performance due to the scarcity of available data. Moreover, domain mi…
DisentanglementSound ClassificationTowards Cross-Subject EMG Pattern Recognition via Dual-Branch Adversarial Feature Disentanglement
Cross-subject electromyography (EMG) pattern recognition faces significant challenges due to inter-subject variability in muscle anatomy, electrode placement, and signal characteristics. Traditional methods rely on subje…
AnatomyDisentanglementElectromyography (EMG)Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement (ICCV-25 🥳)
Clothes-Changing Re-Identification (CC-ReID) aims to recognize individuals across different locations and times, irrespective of clothing. Existing methods often rely on additional models or annotations to learn robust, …
DisentanglementPerson Re-Identification