paper-with-me

홈 › Papers

Visual Large Language Models for Generalized and Specialized Applications

2025-01-06 · YiFan Li, Zhixin Lai, Wentao Bao, Zhen Tan, Anh Dao, Kewei Sui, Jiayi Shen, Dong Liu, Huan Liu, Yu Kong

Visual-language models (VLM) have emerged as a powerful tool for learning a unified embedding space for vision and language. Inspired by large language models, which have demonstrated strong reasoning and multi-task capabilities, visual large language models (VLLMs) are gaining increasing attention for building general-purpose VLMs. Despite the significant progress made in VLLMs, the related literature remains limited, particularly from a comprehensive application perspective, encompassing generalized and specialized applications across vision (image, video, depth), action, and language modalities. In this survey, we focus on the diverse applications of VLLMs, examining their using scenarios, identifying ethics consideration and challenges, and discussing future directions for their development. By synthesizing these contents, we aim to provide a comprehensive guide that will pave the way for future innovations and broader applications of VLLMs. The paper list repository is available: https://github.com/JackYFL/awesome-VLLMs.

📄 PDF Abstract BibTeX arXiv:2501.02765

Code (1)

jackyfl/awesome-vllms 공식 구현

Tasks

Ethics

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts

2026-05-07 · Yuhua Jiang, Junjie Lu, Xinyao Qin, Xiaoyu Chen 외 arxiv

Vision-language-action (VLA) models inherit rich visual-semantic priors from pre-trained vision-language backbones, but adapting them to robotic control remains challenging. Full fine-tuning (FFT) is prone to overfitting…

parameter-efficient fine-tuning

Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs

2025-10-21 · Jiaao Yu, Shenwei Li, Mingjie Han, Yifei Yin 외 arxiv

Recent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet, a significant gap persists in their ad…

Reinforcement Learning

Evaluating and Enhancing Large Language Models Performance in Domain-specific Medicine: Osteoarthritis Management with DocOA

2024-01-20 · Xi Chen, MingKe You, Li Wang, Weizhi Liu 외

The efficacy of large language models (LLMs) in domain-specific medicine, particularly for managing complex diseases such as osteoarthritis (OA), remains largely unexplored. This study focused on evaluating and enhancing…

ManagementRAGRetrievalRetrieval-augmented Generation

Semi-Supervised Medical Image Segmentation via Knowledge Mining from Large Models

2025-03-10 · Yuchen Mao, Hongwei Li, Yinyi Lai, Giorgos Papanastasiou 외

Large-scale vision models like SAM have extensive visual knowledge, yet their general nature and computational demands limit their use in specialized tasks like medical image segmentation. In contrast, task-specific mode…

General KnowledgeImage SegmentationMedical Image SegmentationSemantic Segmentation+1

Domain Prompt Learning with Quaternion Networks

2023-12-12 · CVPR 2024 1 · Qinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma 외

Prompt learning has emerged as an effective and data-efficient technique in large Vision-Language Models (VLMs). However, when adapting VLMs to specialized domains such as remote sensing and medical imaging, domain promp…

Contrastive LearningPrompt Learning