paper-with-me

Papers

Between Flexibility and Consistency: Joint Generation of Captions and Subtitles

2021-07-13 · ACL (IWSLT) 2021 8 · Alina Karakanta, Marco Gaido, Matteo Negri, Marco Turchi

Speech translation (ST) has lately received growing interest for the generation of subtitles without the need for an intermediate source language transcription and timing (i.e. captions). However, the joint generation of source captions and target subtitles does not only bring potential output quality advantages when the two decoding processes inform each other, but it is also often required in multilingual scenarios. In this work, we focus on ST models which generate consistent captions-subtitles in terms of structure and lexical content. We further introduce new metrics for evaluating subtitling consistency. Our findings show that joint decoding leads to increased performance and consistency between the generated captions and subtitles while still allowing for sufficient flexibility to produce subtitles conforming to language-specific needs and norms.

📄 PDF Abstract BibTeX arXiv:2107.06246

Code (1)

mgaido91/FBK-fairseq-ST 공식 구현 pytorch

Tasks

Translation

Similar Papers 제목 키워드 기반

Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models

2024-07-02 · Fei Shen, Hu Ye, Sibo Liu, Jun Zhang 외

Recent research showcases the considerable potential of conditional diffusion models for generating consistent stories. However, current methods, which predominantly generate stories in an autoregressive and excessively …

Story Visualization

Joint Generation of Captions and Subtitles with Dual Decoding

2022-05-13 · IWSLT (ACL) 2022 5 · Jitao Xu, François Buet, Josep Crego, Elise Bertin-Lemée 외

As the amount of audio-visual content increases, the need to develop automatic captioning and subtitling solutions to match the expectations of a growing international audience appears as the only viable way to boost thr…

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation

2025-02-04 · Xueqing Deng, Qihang Yu, Ali Athar, Chenglin Yang 외

This paper introduces the COCONut-PanCap dataset, created to enhance panoptic segmentation and grounded image captioning. Building upon the COCO dataset with advanced COCONut panoptic masks, this dataset aims to overcome…

Image CaptioningPanoptic SegmentationSegmentation

Unified Dense Prediction of Video Diffusion

2025-03-12 · CVPR 2025 1 · Lehan Yang, Lu Qi, Xiangtai Li, Sheng Li 외

We present a unified network for simultaneously generating videos and their corresponding entity segmentation and depth maps from text prompts. We utilize colormap to represent entity masks and depth maps, tightly integr…

PredictionVideo Generation

AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization

2026-07-15 · Yiyang Yao, Shanglin Liu, Jianming Lv, Chengjun Wang 외 arxiv

Contrastive Language-Image Pretraining learns a shared representation space through large-scale contrastive learning. However, existing methods that enforce global consistency regularization overlook a key challenge: the…

Contrastive Learning