paper-with-me

Papers

UniT3D: A Unified Transformer for 3D Dense Captioning and Visual Grounding

2022-12-01 · ICCV 2023 1 · Dave Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner, Angel X. Chang

Performing 3D dense captioning and visual grounding requires a common and shared understanding of the underlying multimodal relationships. However, despite some previous attempts on connecting these two related tasks with highly task-specific neural modules, it remains understudied how to explicitly depict their shared nature to learn them simultaneously. In this work, we propose UniT3D, a simple yet effective fully unified transformer-based architecture for jointly solving 3D visual grounding and dense captioning. UniT3D enables learning a strong multimodal representation across the two tasks through a supervised joint pre-training scheme with bidirectional and seq-to-seq objectives. With a generic architecture design, UniT3D allows expanding the pre-training scope to more various training sources such as the synthesized data from 2D prior knowledge to benefit 3D vision-language tasks. Extensive experiments and analysis demonstrate that UniT3D obtains significant gains for 3D dense captioning and visual grounding.

📄 PDF Abstract BibTeX arXiv:2212.00836

Code (0)

등록된 구현이 없습니다.

Tasks

3D dense captioning3D visual groundingDense CaptioningVisual Grounding

Similar Papers 제목 키워드 기반

D3Net: A Unified Speaker-Listener Architecture for 3D Dense Captioning and Visual Grounding

2021-12-02 · Dave Zhenyu Chen, Qirui Wu, Matthias Nießner, Angel X. Chang

Recent studies on dense captioning and visual grounding in 3D have achieved impressive results. Despite developments in both areas, the limited amount of available 3D vision-language data causes overfitting issues for 3D…

3D dense captioning3D visual groundingCaption GenerationDense Captioning+1

One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework

2025-10-03 · Lorenzo Bianchi, Giacomo Pacini, Fabio Carrara, Nicola Messina 외 arxiv

Zero-shot captioners are recently proposed models that utilize common-space vision-language representations to caption images without relying on paired image-text data. To caption an image, they proceed by textually deco…

Dense Captioning

A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer

2020-05-17 · Vladimir Iashin, Esa Rahtu

Dense video captioning aims to localize and describe important events in untrimmed videos. Existing methods mainly tackle this task by exploiting only visual features, while completely neglecting the audio track. Only a …

Dense Video CaptioningTemporal Action Proposal GenerationVideo Captioning

GRiT: A Generative Region-to-text Transformer for Object Understanding

2022-12-01 · Jialian Wu, JianFeng Wang, Zhengyuan Yang, Zhe Gan 외

This paper presents a Generative RegIon-to-Text transformer, GRiT, for object understanding. The spirit of GRiT is to formulate object understanding as <region, text> pairs, where region locates objects and text describe…

DecoderDense CaptioningDescriptiveObject+2

Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning

2023-02-27 · CVPR 2023 1 · Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech 외

In this work, we introduce Vid2Seq, a multi-modal single-stage dense event captioning model pretrained on narrated videos which are readily-available at scale. The Vid2Seq architecture augments a language model with spec…

Dense Video CaptioningLanguage ModelingLanguage ModellingSentence+1