paper-with-me

홈 › Papers

TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning

2024-09-30 · Joshua Feinglass, Yezhou Yang

Zero-shot inference, where pre-trained models perform tasks without specific training data, is an exciting emergent ability of large models like CLIP. Although there has been considerable exploration into enhancing zero-shot abilities in image captioning (IC) for popular datasets such as MSCOCO and Flickr8k, these approaches fall short with fine-grained datasets like CUB, FLO, UCM-Captions, and Sydney-Captions. These datasets require captions to discern between visually and semantically similar classes, focusing on detailed object parts and their attributes. To overcome this challenge, we introduce TRaining-Free Object-Part Enhancement (TROPE). TROPE enriches a base caption with additional object-part details using object detector proposals and Natural Language Processing techniques. It complements rather than alters the base caption, allowing seamless integration with other captioning methods and offering users enhanced flexibility. Our evaluations show that TROPE consistently boosts performance across all tested zero-shot IC approaches and achieves state-of-the-art results on fine-grained IC datasets.

📄 PDF Abstract BibTeX arXiv:2409.19960

Code (1)

JoshuaFeinglass/TROPE 공식 구현

Tasks

Image CaptioningObject

Methods 이 논문이 사용한 방법론

BASE 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Overview of PicTropes, a film trope dataset

2018-09-28 · Rubén H. García-Ortega, Juan J. Merelo-Guervós, Pablo García Sánchez, Gad Pitaru

From the database DBTropes.org, we have created a dataset of films and the tropes that they use, which we have called PicTropes. In this report we provide the descriptive analysis and a further discussion on the dataset …

Descriptive

Analyzing Gender Bias within Narrative Tropes

2020-10-30 · EMNLP (NLP+CSS) 2020 11 · Dhruvil Gala, Mohammad Omar Khursheed, Hannah Lerner, Brendan O'Connor 외

Popular media reflects and reinforces societal biases through the use of tropes, which are narrative elements, such as archetypal characters and plot arcs, that occur frequently across media. In this paper, we specifical…

TrUMAn: Trope Understanding in Movies and Animations

2021-08-10 · Hung-Ting Su, Po-Wei Shen, Bing-Chen Tsai, Wen-Feng Cheng 외

Understanding and comprehending video content is crucial for many real-world applications such as search and recommendation systems. While recent progress of deep learning has boosted performance on various tasks using v…

Recommendation Systems

Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models

2026-03-25 · Shengli Zhou, Minghang Zheng, Feng Zheng, Yang Liu arxiv

Spatial reasoning focuses on locating target objects based on spatial relations in 3D scenes, which plays a crucial role in developing intelligent embodied agents. Due to the limited availability of 3D scene-language pai…

Spatial Reasoning

A Study on the Performance of U-Net Modifications in Retroperitoneal Tumor Segmentation

2025-02-01 · Moein Heidari, Ehsan Khodapanah Aghdam, Alexander Manzella, Daniel Hsu 외

The retroperitoneum hosts a variety of tumors, including rare benign and malignant types, which pose diagnostic and treatment challenges due to their infrequency and proximity to vital structures. Estimating tumor volume…

DiagnosticMambaOrgan SegmentationSegmentation+1