paper-with-me

홈 › Papers

A Thousand Frames in Just a Few Words: Lingual Description of Videos through Latent Topics and Sparse Object Stitching

2013-06-01 · CVPR 2013 6 · Pradipto Das, Chenliang Xu, Richard F. Doell, Jason J. Corso

The problem of describing images through natural language has gained importance in the computer vision community. Solutions to image description have either focused on a top-down approach of generating language through combinations of object detections and language models or bottom-up propagation of keyword tags from training images to test images through probabilistic or nearest neighbor techniques. In contrast, describing videos with natural language is a less studied problem. In this paper, we combine ideas from the bottom-up and top-down approaches to image description and propose a method for video description that captures the most relevant contents of a video in a natural language description. We propose a hybrid system consisting of a low level multimodal latent topic model for initial keyword annotation, a middle level of concept detectors and a high level module to produce final lingual descriptions. We compare the results of our system to human descriptions in both short and long forms on two datasets, and demonstrate that final system output has greater agreement with the human descriptions than any single level.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image DescriptionVideo Description

Similar Papers 제목 키워드 기반

Cross-lingual Linking of Automatically Constructed Frames and FrameNet

2022-06-01 · LREC 2022 6 · Ryohei Sasano

A semantic frame is a conceptual structure describing an event, relation, or object along with its participants. Several semantic frame resources have been manually elaborated, and there has been much interest in the pos…

Cross-Lingual Word EmbeddingsWord Embeddings

The Impact of Cross-Lingual Adjustment of Contextual Word Representations on Zero-Shot Transfer

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Large pre-trained multilingual models such as mBERT and XLM-R enabled effective cross-lingual zero-shot transfer in many NLP tasks. A cross-lingual adjustment of these models using a small parallel corpus can further imp…

Machine TranslationNERXLM-R

Query Brand Entity Linking in E-Commerce Search

2025-02-03 · Dong Liu, Sreyashi Nag arxiv

Associating user search queries with the correct brand entity is critical for e-commerce product retrieval, yet remains challenging due to the brevity of queries (three to four words on average), their lack of grammatica…

Corpus-based Check-up for Thesaurus

2019-07-01 · ACL 2019 7 · Natalia Loukachevitch

In this paper we discuss the usefulness of applying a checking procedure to existing thesauri. The procedure is based on the analysis of discrepancies of corpus-based and thesaurus-based word similarities. We applied the…

Multi-lingual Common Semantic Space Construction via Cluster-consistent Word Embedding

2018-04-21 · EMNLP 2018 10 · Lifu Huang, Kyunghyun Cho, Boliang Zhang, Heng Ji 외

We construct a multilingual common semantic space based on distributional semantics, where words from multiple languages are projected into a shared space to enable knowledge and resource transfer across languages. Beyon…

ClusteringWord Alignment