paper-with-me

홈 › Papers

TennisVid2Text: Fine-grained Descriptions for Domain Specific Videos

2015-11-26 · Mohak Sukhwani, C. V. Jawahar

Automatically describing videos has ever been fascinating. In this work, we attempt to describe videos from a specific domain - broadcast videos of lawn tennis matches. Given a video shot from a tennis match, we intend to generate a textual commentary similar to what a human expert would write on a sports website. Unlike many recent works that focus on generating short captions, we are interested in generating semantically richer descriptions. This demands a detailed low-level analysis of the video content, specially the actions and interactions among subjects. We address this by limiting our domain to the game of lawn tennis. Rich descriptions are generated by leveraging a large corpus of human created descriptions harvested from Internet. We evaluate our method on a newly created tennis video data set. Extensive analysis demonstrate that our approach addresses both semantic correctness as well as readability aspects involved in the task.

📄 PDF Abstract BibTeX arXiv:1511.08522

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GIST: Generating Image-Specific Text for Fine-grained Object Classification

2023-07-21 · Kathleen M. Lewis, Emily Mu, Adrian V. Dalca, John Guttag

Recent vision-language models outperform vision-only models on many image classification tasks. However, because of the absence of paired text/image descriptions, it remains difficult to fine-tune these models for fine-g…

ClassificationFine-Grained Image Classificationimage-classificationImage Classification+6

Motion Generation from Fine-grained Textual Descriptions

2024-03-20 · Kunhang Li, Yansong Feng

The task of text2motion is to generate human motion sequences from given textual descriptions, where the model explores diverse mappings from natural language instructions to human body movements. While most existing wor…

Motion Generation

Tell and Predict: Kernel Classifier Prediction for Unseen Visual Classes from Unstructured Text Descriptions

2015-06-29 · Mohamed Elhoseiny, Ahmed Elgammal, Babak Saleh

In this paper we propose a framework for predicting kernelized classifiers in the visual domain for categories with no training images where the knowledge comes from textual description about these categories. Through ou…

Zero-Shot Learning

Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging

2025-05-28 · Runze Xia, Shuo Feng, Renzhi Wang, Congchi Yin 외

Brain-to-Image reconstruction aims to recover visual stimuli perceived by humans from brain activity. However, the reconstructed visual stimuli often missing details and semantic inconsistencies, which may be attributed …

Image ReconstructionLanguage ModelingLanguage ModellingSemantic Similarity+1

Think Twice Before Recognizing: Large Multimodal Models for General Fine-grained Traffic Sign Recognition

2024-09-03 · Yaozong Gan, Guang Li, Ren Togo, Keisuke Maeda 외

We propose a new strategy called think twice before recognizing to improve fine-grained traffic sign recognition (TSR). Fine-grained TSR in the wild is difficult due to the complex road conditions, and existing approache…

In-Context LearningTraffic Sign Recognition