paper-with-me

Papers

FedVLM: Scalable Personalized Vision-Language Models through Federated Learning

2025-07-23 · Arkajyoti Mitra, Afia Anjum, Paul Agbaje, Mert Pesé, Habeeb Olufowobi arxiv

Vision-language models (VLMs) demonstrate impressive zero-shot and few-shot learning capabilities, making them essential for several downstream tasks. However, fine-tuning these models at scale remains challenging, particularly in federated environments where data is decentralized and non-iid across clients. Existing parameter-efficient tuning methods like LoRA (Low-Rank Adaptation) reduce computational overhead but struggle with heterogeneous client data, leading to suboptimal generalization. To address these challenges, we propose FedVLM, a federated LoRA fine-tuning framework that enables decentralized adaptation of VLMs while preserving model privacy and reducing reliance on centralized training. To further tackle data heterogeneity, we introduce personalized LoRA (pLoRA), which dynamically adapts LoRA parameters to each client's unique data distribution, significantly improving local adaptation while maintaining global model aggregation. Experiments on the RLAIF-V dataset show that pLoRA improves client-specific performance by 24.5% over standard LoRA, demonstrating superior adaptation in non-iid settings. FedVLM provides a scalable and efficient solution for fine-tuning VLMs in federated settings, advancing personalized adaptation in distributed learning scenarios.

📄 PDF Abstract BibTeX arXiv:2507.17088

Code (0)

등록된 구현이 없습니다.

Tasks

Federated LearningFew-Shot Learning

Similar Papers 제목 키워드 기반

FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models

2025-06-11 · Weiying Zheng, Ziyue Lin, Pengxin Guo, Yuyin Zhou 외

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding and generation by integrating visual and textual information. While instruction tuning and parameter-efficient fine-tun…

BenchmarkingFederated Learningparameter-efficient fine-tuningPrivacy Preserving

Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation

2026-04-10 · Lingfeng Huang, Huizhong Guo, Tianjun Wei, Yingpeng Du 외 arxiv

Large language model (LLM) agents are increasingly deployed as scalable user simulators for recommender system evaluation. Yet existing simulators perceive recommendations through text or structured metadata rather than …

FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model

2025-04-24 · Kaicheng Pang, Xingxing Zou, Waikeung Wong

Fashion styling and personalized recommendations are pivotal in modern retail, contributing substantial economic value in the fashion industry. With the advent of vision-language models (VLM), new opportunities have emer…

Image GenerationLanguage ModelingLanguage ModellingVirtual Try-on

On-Board Vision-Language Models for Personalized Autonomous Vehicle Motion Control: System Design and Real-World Validation

2024-11-17 · Can Cui, Zichong Yang, Yupeng Zhou, Juntong Peng 외

Personalized driving refers to an autonomous vehicle's ability to adapt its driving behavior or control strategies to match individual users' preferences and driving styles while maintaining safety and comfort standards.…

Autonomous VehiclesNatural Language UnderstandingRAGRetrieval-augmented Generation

ZoRRO: A Zero-Weight Personalized Recommender System for Scalable News Recommendation

2026-07-12 · Johannes Kruse, Ryotaro Shimizu, Kasper Lindskow, Jon Tofteskov 외 arxiv

We present ZoRRO (Zero-Weight Personalized Recommender System), a zero-weight, training-free framework for personalized news recommendation designed for scalable real-world deployment. ZoRRO outperforms strong neural bas…