paper-with-me

홈 › Papers

Jointly Modeling Embedding and Translation to Bridge Video and Language

2015-05-07 · CVPR 2016 6 · Yingwei Pan, Tao Mei, Ting Yao, Houqiang Li, Yong Rui

Automatically describing video content with natural language is a fundamental challenge of multimedia. Recurrent Neural Networks (RNN), which models sequence dynamics, has attracted increasing attention on visual interpretation. However, most existing approaches generate a word locally with given previous words and the visual content, while the relationship between sentence semantics and visual content is not holistically exploited. As a result, the generated sentences may be contextually correct but the semantics (e.g., subjects, verbs or objects) are not true. This paper presents a novel unified framework, named Long Short-Term Memory with visual-semantic Embedding (LSTM-E), which can simultaneously explore the learning of LSTM and visual-semantic embedding. The former aims to locally maximize the probability of generating the next word given previous words and visual content, while the latter is to create a visual-semantic embedding space for enforcing the relationship between the semantics of the entire sentence and visual content. Our proposed LSTM-E consists of three components: a 2-D and/or 3-D deep convolutional neural networks for learning powerful video representation, a deep RNN for generating sentences, and a joint embedding model for exploring the relationships between visual content and sentence semantics. The experiments on YouTube2Text dataset show that our proposed LSTM-E achieves to-date the best reported performance in generating natural sentences: 45.3% and 31.0% in terms of BLEU@4 and METEOR, respectively. We also demonstrate that LSTM-E is superior in predicting Subject-Verb-Object (SVO) triplets to several state-of-the-art techniques.

📄 PDF Abstract BibTeX arXiv:1505.01861

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceTranslation

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation

2026-01-17 · Zijie Lou, Xiangwei Feng, Jiaxin Wang, Jiangtao Yao 외 arxiv

Existing video object removal methods predominantly rely on diffusion models following a noise-to-data paradigm, where generation starts from uninformative Gaussian noise. This approach discards the rich structural and c…

Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment

2025-12-08 · Youngjoon Jang, Liliane Momeni, Zifan Jiang, Joon Son Chung 외 arxiv

Our aim is to develop a unified model for sign language understanding, that performs sign language translation (SLT) and sign-subtitle alignment (SSA). Together, these two tasks enable the conversion of continuous signin…

Sign Language Translation

Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval

2026-07-23 · Kyeongmo Chae, Jihoon Lee, Sangtae Ahn arxiv

This paper proposes the Distribution-Alignment Bridge (DAB), a framework that reconceptualizes text-to-video retrieval as a distribution alignment task rather than traditional deterministic point matching. By modeling bo…

Video Retrieval

Vision Bridge Transformer at Scale

2025-11-28 · Zhenxiong Tan, Zeqing Wang, Xingyi Yang, Songhua Liu 외 arxiv

We introduce Vision Bridge Transformer (ViBT), a large-scale instantiation of Brownian Bridge Models designed for conditional generation. Unlike traditional diffusion models that transform noise into data, Bridge Models …

Image Editing

Visual Grounding in Video for Unsupervised Word Translation

2020-03-11 · CVPR 2020 6 · Gunnar A. Sigurdsson, Jean-Baptiste Alayrac, Aida Nematzadeh, Lucas Smaira 외

There are thousands of actively spoken languages on Earth, but a single visual world. Grounding in this visual world has the potential to bridge the gap between all these languages. Our goal is to use visual grounding to…

TranslationVisual GroundingWord Translation