paper-with-me

홈 › Papers

Pre-training for Video Captioning Challenge 2020 Summary

2020-07-27 · Yingwei Pan, Jun Xu, Yehao Li, Ting Yao, Tao Mei

The Pre-training for Video Captioning Challenge 2020 Summary: results and challenge participants' technical reports.

📄 PDF Abstract BibTeX arXiv:2008.00947

Code (0)

등록된 구현이 없습니다.

Tasks

Video Captioning

Similar Papers 제목 키워드 기반

MMSummary: Multimodal Summary Generation for Fetal Ultrasound Video

2024-08-07 · Xiaoqing Guo, Qianhui Men, J. Alison Noble

We present the first automated multimodal summary generation system, MMSummary, for medical imaging video, particularly with a focus on fetal ultrasound analysis. Imitating the examination process performed by a human so…

AnatomyLanguage ModelingLanguage ModellingLarge Language Model

ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning

2026-01-14 · Po-han Li, Shenghui Chen, Ufuk Topcu, Sandeep Chinchali arxiv

Multimodal video captioning condenses dense footage into a structured format of keyframes and natural language. By creating a cohesive multimodal summary, this approach anchors generative AI in rich semantic evidence and…

Video Question AnsweringVideo Captioning

Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and Summarization

2024-05-31 · Richard Luo, Austin Peng, Adithya Vasudev, Rishabh Jain

Video is an increasingly prominent and information-dense medium, yet it poses substantial challenges for language models. A typical video consists of a sequence of shorter segments, or shots, that collectively form a coh…

SentenceVideo CaptioningVideo Summarization

Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction

2025-12-12 · Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller 외 arxiv

Vision-Language Models (VLMs) are able to process increasingly longer videos. Yet, important visual information is easily lost throughout the entire context and missed by VLMs. Also, it is important to design tools that …

Video SummarizationVideo Captioning

VATEX Captioning Challenge 2019: Multi-modal Information Fusion and Multi-stage Training Strategy for Video Captioning

2019-10-13 · Ziqi Zhang, Yaya Shi, Jiutong Wei, Chunfeng Yuan 외

Multi-modal information is essential to describe what has happened in a video. In this work, we represent videos by various appearance, motion and audio information guided with video topic. By following multi-stage train…

Video Captioning