paper-with-me

Papers

PIC 4th Challenge: Semantic-Assisted Multi-Feature Encoding and Multi-Head Decoding for Dense Video Captioning

2022-07-06 · Yifan Lu, Ziqi Zhang, Yuxin Chen, Chunfeng Yuan, Bing Li, Weiming Hu

The task of Dense Video Captioning (DVC) aims to generate captions with timestamps for multiple events in one video. Semantic information plays an important role for both localization and description of DVC. We present a semantic-assisted dense video captioning model based on the encoding-decoding framework. In the encoding stage, we design a concept detector to extract semantic information, which is then fused with multi-modal visual features to sufficiently represent the input video. In the decoding stage, we design a classification head, paralleled with the localization and captioning heads, to provide semantic supervision. Our method achieves significant improvements on the YouMakeup dataset under DVC evaluation metrics and achieves high performance in the Makeup Dense Video Captioning (MDVC) task of PIC 4th Challenge.

📄 PDF Abstract BibTeX arXiv:2207.02583

Code (0)

등록된 구현이 없습니다.

Tasks

Dense Video CaptioningVideo Captioning

Similar Papers 제목 키워드 기반

Learning Joint Source-Channel Encoding in IRS-assisted Multi-User Semantic Communications

2025-04-10 · Haidong Wang, Songhan Zhao, Lanhua Li, Bo Gu 외

In this paper, we investigate a joint source-channel encoding (JSCE) scheme in an intelligent reflecting surface (IRS)-assisted multi-user semantic communication system. Semantic encoding not only compresses redundant in…

Deep Reinforcement LearningSchedulingSemantic Communication

Learning to Optimize Joint Source and RIS-assisted Channel Encoding for Multi-User Semantic Communication Systems

2026-03-22 · Haidong Wang, Songhan Zhao, Bo Gu, Shimin Gong 외 arxiv

In this paper, we explore a joint source and reconfigurable intelligent surface (RIS)-assisted channel encoding (JSRE) framework for multi-user semantic communications, where a deep neural network (DNN) extracts semantic…

Reinforcement LearningSemantic CommunicationSemantic Similarity

Large Generative Model Assisted 3D Semantic Communication

2024-03-09 · Feibo Jiang, Yubo Peng, Li Dong, Kezhi Wang 외

Semantic Communication (SC) is a novel paradigm for data transmission in 6G. However, there are several challenges posed when performing SC in 3D scenarios: 1) 3D semantic extraction; 2) Latent semantic redundancy; and 3…

Generative Adversarial NetworkmodelNeRFSemantic Communication+1

AS3D: 2D-Assisted Cross-Modal Understanding with Semantic-Spatial Scene Graphs for 3D Visual Grounding

2025-05-07 · Feng Xiao, Hongbin Xu, Guocan Zhao, Wenxiong Kang

3D visual grounding aims to localize the unique target described by natural languages in 3D scenes. The significant gap between 3D and language modalities makes it a notable challenge to distinguish multiple similar obje…

3D visual groundingGraph AttentionObjectRelational Reasoning+1

TopicFM+: Boosting Accuracy and Efficiency of Topic-Assisted Feature Matching

2023-07-02 · Khang Truong Giang, Soohwan Song, Sungho Jo

This study tackles the challenge of image matching in difficult scenarios, such as scenes with significant variations or limited texture, with a strong emphasis on computational efficiency. Previous studies have attempte…

Computational Efficiency