paper-with-me

Papers

SAME: A Semantically-Aligned Music Autoencoder

2026-05-18 · Julian D. Parker, Zach Evans, CJ Carr, Zachary Zukowski, Josiah Taylor, Matthew Rice, Jordi Pons arxiv

Latent representations are at the heart of the majority of modern generative models. In the audio domain they are typically produced by a neural-audio-codec autoencoder. In this work we introduce SAME (Semantically-Aligned Music autoEncoder), an autoencoder for stereo music and general audio that reaches a 4096$\times$ temporal compression ratio while maintaining reconstruction quality and downstream generative performance. We achieve this by combining a tranformer-based backbone with set of semantic regularisation approaches, phase-aware reconstruction losses and improved discriminator designs. The architecture delivers substantial computational cost benefits, through both its high compression ratio and its reliance on well-optimised transformer primitives. Two variants (a large SAME-L and a CPU-deployable SAME-S) are released in open-weights form.

📄 PDF Abstract BibTeX arXiv:2605.18613

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling

2024-06-06 · CVPR 2025 1 · Zeyue Tian, Zhaoyang Liu, Ruibin Yuan, Jiahao Pan 외

In this work, we systematically study music generation conditioned solely on the video. First, we present a large-scale dataset comprising 360K video-music pairs, including various genres such as movie trailers, advertis…

DiversityMusic Generation

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

2026-04-19 · Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj, Gouthaman KV 외 arxiv

Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditioning alone and provide limited semantic…

Music Generation

COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations

2020-06-15 · Xavier Favory, Konstantinos Drossos, Tuomas Virtanen, Xavier Serra

Audio representation learning based on deep neural networks (DNNs) emerged as an alternative approach to hand-crafted features. For achieving high performance, DNNs often need a large amount of annotated data which can b…

Representation Learning

Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment

2025-11-27 · Jiaying Hong, Ting Zhu, Thanet Markchom, Huizhi Liang arxiv

With the rise of AI-generated content (AIGC), generating perceptually natural and feeling-aligned music from multimodal inputs has become a central challenge. Existing approaches often rely on explicit emotion labels tha…

Audio GenerationMusic Generation

PIANOTREE VAE: Structured Representation Learning for Polyphonic Music

2020-08-17 · Ziyu Wang, Yiyi Zhang, Yixiao Zhang, Junyan Jiang 외

The dominant approach for music representation learning involves the deep unsupervised model family variational autoencoder (VAE). However, most, if not all, viable attempts on this problem have largely been limited to m…

Music GenerationRepresentation Learning