Defining Visually Descriptive Language
Code (0)
등록된 구현이 없습니다.
Tasks
DescriptiveImage CaptioningImage RetrievalSimilar Papers 제목 키워드 기반
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applications increasingly demand visually ground…
Continual PretrainingAn Analytical Model of Language Resource Sustainability
This paper elaborates on a sustainability model for Language Resources, both at a descriptive and analytical level. The first part, devoted to the descriptive model, elaborates on the definition of this concept both from…
DescriptiveInformation RetrievalMachine TranslationmodelUsing Descriptive Video Services to Create a Large Data Source for Video Annotation Research
In this work, we introduce a dataset of video annotated with high quality natural language phrases describing the visual content in a given segment of time. Our dataset is based on the Descriptive Video Service (DVS) tha…
DescriptiveVideo DescriptionLanguage as a Label: Zero-Shot Multimodal Classification of Everyday Postures under Data Scarcity
Recent Vision-Language Models (VLMs) enable zero-shot classification by aligning images and text in a shared space, a promising approach for data-scarce conditions. However, the influence of prompt design on recognizing …
SemStyle: Learning to Generate Stylised Image Captions using Unaligned Text
Linguistic style is an essential part of written communication, with the power to affect both clarity and attractiveness. With recent advances in vision and language, we can start to tackle the problem of generating imag…
DescriptiveImage CaptioningLanguage ModelingLanguage Modelling