paper-with-me

홈 › Papers

Extended Self-Critical Pipeline for Transforming Videos to Text (TRECVID-VTT Task 2021) -- Team: MMCUniAugsburg

2021-12-28 · Philipp Harzig, Moritz Einfalt, Katja Ludwig, Rainer Lienhart

The Multimedia and Computer Vision Lab of the University of Augsburg participated in the VTT task only. We use the VATEX and TRECVID-VTT datasets for training our VTT models. We base our model on the Transformer approach for both of our submitted runs. For our second model, we adapt the X-Linear Attention Networks for Image Captioning which does not yield the desired bump in scores. For both models, we train on the complete VATEX dataset and 90% of the TRECVID-VTT dataset for pretraining while using the remaining 10% for validation. We finetune both models with self-critical sequence training, which boosts the validation performance significantly. Overall, we find that training a Video-to-Text system on traditional Image Captioning pipelines delivers very poor performance. When switching to a Transformer-based architecture our results greatly improve and the generated captions match better with the corresponding video.

📄 PDF Abstract BibTeX arXiv:2112.14100

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

PresentAgent: Multimodal Agent for Presentation Video Generation

2025-07-05 · Jingwei Shi, Zeyu Zhang, Biao Wu, Yanjie Liang 외

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances …

text-to-speechText to SpeechVideo Generation

Φ-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation

2026-05-23 · Ofir Abramovich, Nadav Z. Cohen, Adi Rosenthal, Ariel Shamir arxiv

Latent video diffusion models generate videos by progressively transforming Gaussian noise into realistic samples conditioned on text or visual inputs. However, existing conditioning methods often require additional trai…

Video Generation

Video-Infinity: Distributed Long Video Generation

2024-06-24 · Zhenxiong Tan, Xingyi Yang, Songhua Liu, Xinchao Wang

Diffusion models have recently achieved remarkable results for video generation. Despite the encouraging performances, the generated videos are typically constrained to a small number of frames, resulting in clips lastin…

GPUVideo Generation

GenCompositor: Generative Video Compositing with Diffusion Transformer

2025-09-02 · Shuzhou Yang, Xiaoyu Li, Xiaodong Cun, Guangzhi Wang 외 arxiv

Video compositing combines live-action footage to create video production, serving as a crucial technique in video creation and film production. Traditional pipelines require intensive labor efforts and expert collaborat…

HIDRO-VQA: High Dynamic Range Oracle for Video Quality Assessment

2023-11-18 · Shreshth Saini, Avinab Saha, Alan C. Bovik

We introduce HIDRO-VQA, a no-reference (NR) video quality assessment model designed to provide precise quality evaluations of High Dynamic Range (HDR) videos. HDR videos exhibit a broader spectrum of luminance, detail, a…

Video Quality AssessmentVisual Question Answering (VQA)