paper-with-me

홈 › Papers

AlignNet: A Unifying Approach to Audio-Visual Alignment

2020-02-12 · Jianren Wang, Zhaoyuan Fang, Hang Zhao

We present AlignNet, a model that synchronizes videos with reference audios under non-uniform and irregular misalignments. AlignNet learns the end-to-end dense correspondence between each frame of a video and an audio. Our method is designed according to simple and well-established principles: attention, pyramidal processing, warping, and affinity function. Together with the model, we release a dancing dataset Dance50 for training and evaluation. Qualitative, quantitative and subjective evaluation results on dance-music alignment and speech-lip alignment demonstrate that our method far outperforms the state-of-the-art methods. Project video and code are available at https://jianrenw.github.io/AlignNet.

📄 PDF Abstract BibTeX arXiv:2002.05070

Code (1)

zfang399/AlignNet pytorch

Similar Papers 제목 키워드 기반

AlignNet: Learning dataset score alignment functions to enable better training of speech quality estimators

2024-06-14 · Jaden Pieper, Stephen D. Voran

We develop two complementary advances for training no-reference (NR) speech quality estimators with independent datasets. Multi-dataset finetuning (MDF) pretrains an NR estimator on a single dataset and then finetunes it…

Frequency-Based Alignment of EEG and Audio Signals Using Contrastive Learning and SincNet for Auditory Attention Detection

2025-03-06 · Yuan Liao, Yuhong Zhang, Qiushi Han, Yuhang Yang 외

Humans exhibit a remarkable ability to focus auditory attention in complex acoustic environments, such as cocktail parties. Auditory attention detection (AAD) aims to identify the attended speaker by analyzing brain sign…

Contrastive LearningEEG

AlignNet: Unsupervised Entity Alignment

2020-07-17 · Antonia Creswell, Kyriacos Nikiforou, Oriol Vinyals, Andre Saraiva 외

Recently developed deep learning models are able to learn to segment scenes into component objects without supervision. This opens many new and exciting avenues of research, allowing agents to take objects (or entities) …

Entity Alignment

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

2025-10-17 · Hanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang 외 arxiv

Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM.…

ResAlignNet: A Data-Driven Approach for INS/DVL Alignment

2025-11-17 · Guy Damari, Itzik Klein arxiv

Autonomous underwater vehicles rely on precise navigation systems that combine the inertial navigation system and the Doppler velocity log for successful missions in challenging environments where satellite navigation is…