paper-with-me

홈 › Papers

LARE: Latent Augmentation using Regional Embedding with Vision-Language Model

2024-09-19 · Kosuke Sakurai, Tatsuya Ishii, Ryotaro Shimizu, Linxin Song, Masayuki Goto

In recent years, considerable research has been conducted on vision-language models that handle both image and text data; these models are being applied to diverse downstream tasks, such as "image-related chat," "image recognition by instruction," and "answering visual questions." Vision-language models (VLMs), such as Contrastive Language-Image Pre-training (CLIP), are also high-performance image classifiers that are being developed into domain adaptation methods that can utilize language information to extend into unseen domains. However, because these VLMs embed images as a single point in a unified embedding space, there is room for improvement in the classification accuracy. Therefore, in this study, we proposed the Latent Augmentation using Regional Embedding (LARE), which embeds the image as a region in the unified embedding space learned by the VLM. By sampling the augmented image embeddings from within this latent region, LARE enables data augmentation to various unseen domains, not just to specific unseen domains. LARE achieves robust image classification for domains in and out using augmented image embeddings to fine-tune VLMs. We demonstrate that LARE outperforms previous fine-tuning models in terms of image classification accuracy on three benchmarks. We also demonstrate that LARE is a more robust and general model that is valid under multiple conditions, such as unseen domains, small amounts of data, and imbalanced data.

📄 PDF Abstract BibTeX arXiv:2409.12597

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDomain Adaptationimage-classificationImage ClassificationLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

LaRe: Latent Refocusing for Multimodal Reasoning

2025-11-04 · Jizheng Ma, Xiaofei Zhou, Geyuan Zhang, Yanlong Song 외 arxiv

Chain of Thought (CoT) reasoning enhances logical performance by decomposing complex tasks, yet its multimodal extension faces a trade-off. The prevailing Thinking with Images paradigm achieves visual refocusing by expli…

Multimodal Reasoning

Enhancing Semantic Segmentation with Continual Self-Supervised Pre-training

2025-09-22 · Brown Ebouky, Ajad Chhatkuli, Cristiano Malossi, Christoph Studer 외 arxiv

Self-supervised learning (SSL) has emerged as a central paradigm for training foundation models by leveraging large-scale unlabeled datasets, often producing representations with strong generalization capabilities. These…

Self-Supervised LearningSemantic Segmentation

FLARE: Robot Learning with Implicit World Modeling

2025-05-21 · Ruijie Zheng, Jing Wang, Scott Reed, Johan Bjorck 외

We introduce $\textbf{F}$uture $\textbf{LA}$tent $\textbf{RE}$presentation Alignment ($\textbf{FLARE}$), a novel framework that integrates predictive latent world modeling into robot policy learning. By aligning features…

Imitation LearningVision-Language-Action

Difflare: Removing Image Lens Flare with Latent Diffusion Model

2024-07-20 · Tianwen Zhou, Qihao Duan, Zitong Yu

The recovery of high-quality images from images corrupted by lens flare presents a significant challenge in low-level vision. Contemporary deep learning methods frequently entail training a lens flare removing model from…

Flare Removal

Direction-Oriented Visual-semantic Embedding Model for Remote Sensing Image-text Retrieval

2023-10-12 · Qing Ma, Jiancheng Pan, Cong Bai

Image-text retrieval has developed rapidly in recent years. However, it is still a challenge in remote sensing due to visual-semantic imbalance, which leads to incorrect matching of non-semantic visual and textual featur…

Cross-Modal RetrievalImage-text RetrievalRetrievalText Retrieval