paper-with-me

Papers

Cut2Next: Generating Next Shot via In-Context Tuning

2025-08-11 · Jingwen He, Hongbo Liu, Jiajun Li, Ziqi Huang, Yu Qiao, Wanli Ouyang, Ziwei Liu arxiv

Effective multi-shot generation demands purposeful, film-like transitions and strict cinematic continuity. Current methods, however, often prioritize basic visual consistency, neglecting crucial editing patterns (e.g., shot/reverse shot, cutaways) that drive narrative flow for compelling storytelling. This yields outputs that may be visually coherent but lack narrative sophistication and true cinematic integrity. To bridge this, we introduce Next Shot Generation (NSG): synthesizing a subsequent, high-quality shot that critically conforms to professional editing patterns while upholding rigorous cinematic continuity. Our framework, Cut2Next, leverages a Diffusion Transformer (DiT). It employs in-context tuning guided by a novel Hierarchical Multi-Prompting strategy. This strategy uses Relational Prompts to define overall context and inter-shot editing styles. Individual Prompts then specify per-shot content and cinematographic attributes. Together, these guide Cut2Next to generate cinematically appropriate next shots. Architectural innovations, Context-Aware Condition Injection (CACI) and Hierarchical Attention Mask (HAM), further integrate these diverse signals without introducing new parameters. We construct RawCuts (large-scale) and CuratedCuts (refined) datasets, both with hierarchical prompts, and introduce CutBench for evaluation. Experiments show Cut2Next excels in visual consistency and text fidelity. Crucially, user studies reveal a strong preference for Cut2Next, particularly for its adherence to intended editing patterns and overall cinematic continuity, validating its ability to generate high-quality, narratively expressive, and cinematically coherent subsequent shots.

📄 PDF Abstract BibTeX arXiv:2508.08244

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Better Few-Shot and Finetuning Performance with Forgetful Causal Language Models

2022-10-24 · Hao liu, Xinyang Geng, Lisa Lee, Igor Mordatch 외

Large language models (LLM) trained using the next-token-prediction objective, such as GPT3 and PaLM, have revolutionized natural language processing in recent years by showing impressive zero-shot and few-shot capabilit…

Language ModelingLanguage ModellingNatural Language InferenceRepresentation Learning

Contextual Biasing of Named-Entities with Large Language Models

2023-09-01 · Chuanneng Sun, Zeeshan Ahmed, Yingyi Ma, Zhe Liu 외

This paper studies contextual biasing with Large Language Models (LLMs), where during second-pass rescoring additional contextual information is provided to a LLM to boost Automatic Speech Recognition (ASR) performance. …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Prompt Learningspeech-recognition+2

Learning Next Action Predictors from Human-Computer Interaction

2026-03-06 · Omar Shaikh, Valentin Teutschbein, Kanishk Gandhi, Yikun Chi 외 arxiv

Truly proactive AI systems must anticipate what we will do next. This foresight demands far richer information than the sparse signals we type into our prompts -- it demands reasoning over the entire context of what we s…

Comparison of ConvNeXt and Vision-Language Models for Breast Density Assessment in Screening Mammography

2025-06-16 · Yusdivia Molina-Román, David Gómez-Ortiz, Ernestina Menasalvas-Ruiz, José Gerardo Tamez-Peña 외

Mammographic breast density classification is essential for cancer risk assessment but remains challenging due to subjective interpretation and inter-observer variability. This study compares multimodal and CNN-based met…

breast density classificationClassificationzero-shot-classificationZero-Shot Learning

ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling

2026-03-26 · Yawen Luo, Xiaoyu Shi, Junhao Zhuang, Yutian Chen 외 arxiv

Multi-shot video generation is crucial for long narrative storytelling, yet current bidirectional architectures suffer from limited interactivity and high latency. We propose ShotStream, a novel causal multi-shot archite…

Video Generation