paper-with-me

홈 › Papers

Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories

2025-12-17 · Chayan Jain, Rishant Sharma, Archit Garg, Ishan Bhanuka, Pratik Narang, Dhruv Kumar arxiv

Generating long, cohesive video stories with consistent characters is a significant challenge for current text-to-video AI. We introduce a method that approaches video generation in a filmmaker-like manner. Instead of creating a video in one step, our proposed pipeline first uses a large language model to generate a detailed production script. This script guides a text-to-image model in creating consistent visuals for each character, which then serve as anchors for a video generation model to synthesize each scene individually. Our baseline comparisons validate the necessity of this multi-stage decomposition; specifically, we observe that removing the visual anchoring mechanism results in a catastrophic drop in character consistency scores (from 7.99 to 0.55), confirming that visual priors are essential for identity preservation. Furthermore, we analyze cultural disparities in current models, revealing distinct biases in subject consistency and dynamic degree between Indian vs Western-themed generations.

📄 PDF Abstract BibTeX arXiv:2512.16954

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Modeling and Optimization of Transistor Voltage Amplifiers Based on Stochastic Thermodynamics

2025-04-27 · Xiaoxuan Peng, Xiaohu Ge

As transistor sizes reach the mesoscopic scale, the limitations of traditional methods in ensuring thermodynamic consistency have made power dissipation optimization in transistor amplifiers a critical challenge. Based o…

Seeing Temporal Modulation of Lights From Standard Cameras

2018-06-01 · CVPR 2018 6 · Naoki Sakakibara, Fumihiko Sakaue, Jun Sato

In this paper, we propose a novel method for measuring the temporal modulation of lights by using off-the-shelf cameras. In particular, we show that the invisible flicker patterns of various lights such as fluorescent li…

Deblurring

Exploiting Multilingualism through Multistage Fine-Tuning for Low-Resource Neural Machine Translation

2019-11-01 · IJCNLP 2019 11 · Raj Dabre, Atsushi Fujita, Chenhui Chu

This paper highlights the impressive utility of multi-parallel corpora for transfer learning in a one-to-many low-resource neural machine translation (NMT) setting. We report on a systematic comparison of multistage fine…

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationNMT+2

A hybrid deep learning framework for integrated segmentation and registration: evaluation on longitudinal white matter tract changes

2019-08-26 · Bo Li, Wiro Niessen, Stefan Klein, Marius de Groot 외

To accurately analyze changes of anatomical structures in longitudinal imaging studies, consistent segmentation across multiple time-points is required. Existing solutions often involve independent registration and segme…

Segmentation

Automating Video Thumbnails Selection and Generation with Multimodal and Multistage Analysis

2024-10-18 · Elia Fantini

This thesis presents an innovative approach to automate video thumbnail selection for traditional broadcast content. Our methodology establishes stringent criteria for diverse, representative, and aesthetically pleasing …

Face RecognitionImage Matting