paper-with-me

홈 › Papers

MLLM-Guided Semantic Correction for Text-to-Video Generation

2026-08-17 · Junhao Chen, Zheqi Lv, Keting Yin, Shengyu Zhang, Zhou Zhao, Feiyang Chen, Xinyu Duan, Baoxing Huai, Fei Wu arxiv

Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from semantic errors such as missing objects, incorrect attributes, or mismatched actions. Although some semantic correction methods perform optimization before sampling or refinement after sampling, how to detect and correct semantic deviations during the video generation process remains underexplored. In this paper, we introduce a training-free, interpretable mid-generation correction framework that integrates multimodal large language model (MLLM) feedback directly into the diffusion sampling loop. Our framework achieves diffusion trajectory correction by injecting semantic evaluation signals during video synthesis, enabling the model to optimize the generated content through continuous self-reflection. We propose two key modules: a Semantic Assessment Supervisor that generates intermediate preview frames for semantic evaluations and deviation diagnostics, and a Semantic Modification Assistant that corrects semantic drift during inference via a controllable latent trajectory intervention. Our method improves semantic alignment, visual fidelity, and temporal consistency without modifying model parameters. We validate the effectiveness of our approach through extensive experiments across multiple benchmarks.

📄 PDF Abstract BibTeX arXiv:2608.16513

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video Generation

Similar Papers 제목 키워드 기반

FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation

2025-07-09 · Liqiang Jing, Viet Lai, Seunghyun Yoon, Trung Bui 외

Video Multimodal Large Language Models (VideoMLLMs) have achieved remarkable progress in both Video-to-Text and Text-to-Video tasks. However, they often suffer fro hallucinations, generating content that contradicts the …

DescriptiveText GenerationVideo Generation

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

2026-04-28 · Minghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng 외 arxiv

Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale and semantic diversity, causing performance gaps between common and r…

Reinforcement LearningDense Captioning

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion

2025-05-26 · Zheqi Lv, JunHao Chen, Qi Tian, Keting Yin 외

Diffusion models have become the mainstream architecture for text-to-image generation, achieving remarkable progress in visual quality and prompt controllability. However, current inference pipelines generally lack inter…

DenoisingImage GenerationLarge Language ModelMultimodal Large Language Model+2

TiFRe: Text-guided Video Frame Reduction for Efficient Video Multi-modal Large Language Models

2026-02-09 · Xiangtian Zheng, Zishuo Wang, Yuxin Peng arxiv

With the rapid development of Large Language Models (LLMs), Video Multi-Modal Large Language Models (Video MLLMs) have achieved remarkable performance in video-language tasks such as video understanding and question answ…

Semantic SimilarityQuestion Answering

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

2026-08-26 · Eunjee Choi, JungHoon Sung, Seongwhan Cho, Chu Xin 외 arxiv

Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss annotations, provid…

Sign Language RecognitionRepresentation Learning