multimodal generation
1개 벤치마크 · 논문 187편 · 이 태스크의 논문 보기 →
Benchmarks
Multi-Modal CelebA-HQ
Most implemented
PMG : Personalized Multimodal Generation with Large Language Models
Retrieval-Augmented Generation for AI-Generated Content: A Survey
Finite Scalar Quantization: VQ-VAE Made Simple
Emerging Properties in Unified Multimodal Pretraining
Vision-to-Music Generation: A Survey
Papers
FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding
Multimodal speculative decoding accelerates vision-language models by allowing a lightweight draft model to propose candidate tokens for parallel verification by a larger target model. Existing methods typically conditio…
multimodal generationVisual GroundingG-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation
Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communica…
Reinforcement Learningmultimodal generationSPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilitie…
multimodal generationSpatial Reasoning3D ReconstructionM$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data
Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for comprehensive disease understanding. However, gene expression data acquisition remains…
multimodal generationContrastive LearningVQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios
Tactile image generation significantly reduces the dependency on expensive and wear-prone sensors by synthesizing high-fidelity tactile data, offering an efficient solution for tactile information acquisition in robotic …
multimodal generationImage GenerationConcurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Mask…
multimodal generationVisual Reasoning