paper-with-me

Papers

Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video

2025-11-05 · Alexander Htet Kyaw, Lenin Ravindranath Sivalingam arxiv

We present a node-based storytelling system for multimodal content generation. The system represents stories as graphs of nodes that can be expanded, edited, and iteratively refined through direct user edits and natural-language prompts. Each node can integrate text, images, audio, and video, allowing creators to compose multimodal narratives. A task selection agent routes between specialized generative tasks that handle story generation, node structure reasoning, node diagram formatting, and context generation. The interface supports targeted editing of individual nodes, automatic branching for parallel storylines, and node-based iterative refinement. Our results demonstrate that node-based editing supports control over narrative structure and iterative generation of text, images, audio, and video. We report quantitative outcomes on automatic story outline generation and qualitative observations of editing workflows. Finally, we discuss current limitations such as scalability to longer narratives and consistency across multiple nodes, and outline future work toward human-in-the-loop and user-centered creative AI tools.

📄 PDF Abstract BibTeX arXiv:2511.03227

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal generationStory Generation

Similar Papers 제목 키워드 기반

AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control

2025-11-26 · Xinyue Guo, Xiaoran Yang, Lipan Zhang, Jianxuan Yang 외 arxiv

Sound effect editing-modifying audio by adding, removing, or replacing elements-remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limite…

Audio Generation

Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing

2026-04-12 · Zeyue Tian, Binxin Yang, Zhaoyang Liu, Jiexuan Zhang 외 arxiv

Recent progress in multimodal models has spurred rapid advances in audio understanding, generation, and editing. However, these capabilities are typically addressed by specialized models, leaving the development of a tru…

Audio Generation

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

2025-06-26 · Huadai Liu, Jialei Wang, Kaicheng Luo, Wen Wang 외

While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries,…

Audio GenerationLarge Language ModelMultimodal Large Language Model

CoDi-2: In-Context Interleaved and Interactive Any-to-Any Generation

2024-01-01 · CVPR 2024 1 · Zineng Tang, ZiYi Yang, Mahmoud Khademi, Yang Liu 외

We present CoDi-2 a Multimodal Large Language Model (MLLM) for learning in-context interleaved multimodal representations. By aligning modalities with language for both encoding and generation CoDi-2 empowers Large L…

Image GenerationLanguage ModelingLanguage ModellingLarge Language Model+1

SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model

2026-02-25 · Guibin Chen, Dixuan Lin, Jiangping Yang, Youqiang Zhang 외 arxiv

SkyReels V4 is a unified multi modal video foundation model for joint video audio generation, inpainting, and editing. The model adopts a dual stream Multimodal Diffusion Transformer (MMDiT) architecture, where one branc…

Instruction FollowingAudio GenerationVideo Generation