paper-with-me

홈 › Papers

The Lost Melody: Empirical Observations on Text-to-Video Generation From A Storytelling Perspective

2024-05-13 · Andrew Shin, Yusuke Mori, Kunitake Kaneko

Text-to-video generation task has witnessed a notable progress, with the generated outcomes reflecting the text prompts with high fidelity and impressive visual qualities. However, current text-to-video generation models are invariably focused on conveying the visual elements of a single scene, and have so far been indifferent to another important potential of the medium, namely a storytelling. In this paper, we examine text-to-video generation from a storytelling perspective, which has been hardly investigated, and make empirical remarks that spotlight the limitations of current text-to-video generation scheme. We also propose an evaluation framework for storytelling aspects of videos, and discuss the potential future directions.

📄 PDF Abstract BibTeX arXiv:2405.08720

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video GenerationVideo Generation

Similar Papers 제목 키워드 기반

Accompanied Singing Voice Synthesis with Fully Text-controlled Melody

2024-07-02 · RuiQi Li, Zhiqing Hong, Yongqi Wang, Lichao Zhang 외

Text-to-song (TTSong) is a music generation task that synthesizes accompanied singing voices. Current TTSong methods, inherited from singing voice synthesis (SVS), require melody-related information that can sometimes be…

Language ModelingLanguage ModellingMusic GenerationSinging Voice Synthesis

REFFLY: Melody-Constrained Lyrics Editing Model

2024-08-30 · Songyan Zhao, Bingxuan Li, Yufei Tian, Nanyun Peng

Automatic melody-to-lyric (M2L) generation aims to create lyrics that align with a given melody. While most previous approaches generate lyrics from scratch, revision, editing plain text draft to fit it into the melody, …

modelStyle TransferTranslation

Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer

2024-10-07 · Siyuan Hou, Shansong Liu, Ruibin Yuan, Wei Xue 외

Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structu…

Music GenerationMusic Style TransferStyle TransferText-to-Music Generation

MelodyGLM: Multi-task Pre-training for Symbolic Melody Generation

2023-09-19 · Xinda Wu, Zhijie Huang, Kejun Zhang, Jiaxing Yu 외

Pre-trained language models have achieved impressive results in various music understanding and generation tasks. However, existing pre-training methods for symbolic melody generation struggle to capture multi-scale, mul…

Rhythm

SongMASS: Automatic Song Writing with Pre-training and Alignment Constraint

2020-12-09 · Zhonghao Sheng, Kaitao Song, Xu Tan, Yi Ren 외

Automatic song writing aims to compose a song (lyric and/or melody) by machine, which is an interesting topic in both academia and industry. In automatic song writing, lyric-to-melody generation and melody-to-lyric gener…

Sentence