paper-with-me

Papers

Enhancing Visual Forced Alignment with Local Context-Aware Feature Extraction and Multi-Task Learning

2025-03-05 · Yi He, Lei Yang, Shilin Wang

This paper introduces a novel approach to Visual Forced Alignment (VFA), aiming to accurately synchronize utterances with corresponding lip movements, without relying on audio cues. We propose a novel VFA approach that integrates a local context-aware feature extractor and employs multi-task learning to refine both global and local context features, enhancing sensitivity to subtle lip movements for precise word-level and phoneme-level alignment. Incorporating the improved Viterbi algorithm for post-processing, our method significantly reduces misalignments. Experimental results show our approach outperforms existing methods, achieving a 6% accuracy improvement at the word-level and 27% improvement at the phoneme-level in LRS2 dataset. These improvements offer new potential for applications in automatically subtitling TV shows or user-generated content platforms like TikTok and YouTube Shorts.

📄 PDF Abstract BibTeX arXiv:2503.03286

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task Learning

Similar Papers 제목 키워드 기반

Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video

2023-02-27 · Minsu Kim, Chae Won Kim, Yong Man Ro

Forced alignment refers to a technology that time-aligns a given transcription with a corresponding speech. However, as the forced alignment technologies have developed using speech audio, they might fail in alignment wh…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+2

ContextGuard-LVLM: Enhancing News Veracity through Fine-grained Cross-modal Contextual Consistency Verification

2025-08-08 · Sihan Ma, Qiming Wu, Ruotong Jiang, Frank Burns arxiv

The proliferation of digital news media necessitates robust methods for verifying content veracity, particularly regarding the consistency between visual and textual information. Traditional approaches often fall short i…

Logical Reasoning

PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment

2026-03-19 · Tianci Luo, Jinpeng Wang, Shiyu Qin, Niu Lian 외 arxiv

Visual In-Context Learning (VICL) aims to complete vision tasks by imitating pixel demonstrations. Recent work pioneered prompt fusion that combines the advantages of various demonstrations, which shows a promising way t…

Data Augmentation

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding

2025-08-23 · Leilei Guo, Antonio Carlos Rivera, Peiyu Tang, Haoxuan Ren 외 arxiv

Large Language Models (LLMs) and Vision-Language Large Models (LVLMs) have achieved remarkable progress in natural language processing and multimodal understanding. Despite their impressive generalization capabilities, c…

Referring ExpressionVisual Reasoning

Multilingual Word-Level Forced Alignment with Self-Supervised Representations and Learned Dynamic Programming

2026-06-09 · Roy Weber, Meidan Zehavi, Rotem Rousso, Joseph Keshet arxiv

We present a method for accurate multilingual word-level forced alignment, consisting of an alignment encoder and a learned alignment decoder. The encoder integrates two representations: one from the Massively Multilingu…