paper-with-me

multimodal generation

1개 벤치마크 · 논문 187편 · 이 태스크의 논문 보기 →

Benchmarks

Multi-Modal CelebA-HQ

결과 1개

Most implemented

Vision-to-Music Generation: A Survey

2025-03-27 · 구현 2개

Papers

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

2026-08-24 · Hengjie Zhu, Dayan Wu, Zihao Zhang, Xinze Liu 외 arxiv

Multimodal speculative decoding accelerates vision-language models by allowing a lightweight draft model to propose candidate tokens for parallel verification by a larger target model. Existing methods typically conditio…

multimodal generationVisual Grounding

G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

2026-08-20 · Shiao Xie, Siyu Chen, Jianwei Lv, Bo Yuan 외 arxiv

Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communica…

Reinforcement Learningmultimodal generation

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

2026-08-14 · Jinsheng Quan, Jianhua Li, Siyi Xie, Xuanke Shi 외 arxiv

Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilitie…

multimodal generationSpatial Reasoning3D Reconstruction

M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data

2026-07-23 · Francesca Pia Panaccione, Carlo Sgaravatti, Marco Venere arxiv

Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for comprehensive disease understanding. However, gene expression data acquisition remains…

multimodal generationContrastive Learning

VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios

2026-07-16 · Kailin Lyu, Long Xiao, Jianing Zeng, Di Wu 외 arxiv

Tactile image generation significantly reduces the dependency on expensive and wear-prone sensors by synthesizing high-fidelity tactile data, offering an efficient solution for tactile information acquisition in robotic …

multimodal generationImage Generation

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

2026-07-14 · Minh-Quan Le, Armand Comas, Alexandros Lattas, Stylianos Moschoglou 외 hf

Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Mask…

multimodal generationVisual Reasoning

전체 187편 보기 →