paper-with-me

홈 › Papers

RLLaVA: An RL-central Framework for Language and Vision Assistants

2025-12-25 · Lei Zhao, Zihao Ma, Boyu Lin, Yuhe Liu, Wenjun Wu, Lei Huang arxiv

We present an RL-central framework for Language and Vision Assistants (RLLaVA) with its formulation of Markov decision process (MDP). RLLaVA decouples RL algorithmic logic from model architecture and distributed execution, supporting researchers in implementing new RL algorithms with minimal code, and to plug in a broad family of RL methods and vision-language models (VLMs) while remaining agnostic to specific training and inference engines. RLLaVA makes resource-efficient training of 1B--7B models feasible on common GPUs; notably, 4B-scale models can be trained end-to-end with full-parameter updates on a single 24GB GPU. Experiments on multi-modal and agentic tasks demonstrate that RLLaVA has task extensibility, and the models trained with it consistently improve performance over base models, competitive with other specially engineered RL frameworks. The code is available at https://github.com/TinyLoopX/RLLaVA.

📄 PDF Abstract BibTeX arXiv:2512.21450

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants

2026-03-27 · Mahesh Bhosale, Abdul Wasi, Shantam Srivastava, Shifa Latif 외 arxiv

While powerful in image-conditioned generation, multimodal large language models (MLLMs) can display uneven performance across demographic groups, highlighting fairness risks. In safety-critical clinical settings, such d…

parameter-efficient fine-tuningvisual instruction followingVisual Question Answering

CarLLaVA: Vision language models for camera-only closed-loop driving

2024-06-14 · Katrin Renz, Long Chen, Ana-Maria Marcu, Jan Hünermann 외

In this technical report, we present CarLLaVA, a Vision Language Model (VLM) for autonomous driving, developed for the CARLA Autonomous Driving Challenge 2.0. CarLLaVA uses the vision encoder of the LLaVA VLM and the LLa…

Autonomous DrivingBench2DriveCARLA Leaderboard 2.0Language Modeling+1

HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models

2024-03-20 · Wenqiao Zhang, Tianwei Lin, Jiang Liu, Fangxun Shu 외

Recent advancements indicate that scaling up Multimodal Large Language Models (MLLMs) effectively enhances performance on downstream multimodal tasks. The prevailing MLLM paradigm, \emph{e.g.}, LLaVA, transforms visual f…

MMEVisual Question Answering

PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLM

2024-06-05 · Tao Yang, Yingmin Luo, Zhongang Qi, Yang Wu 외

Layout generation is the keystone in achieving automated graphic design, requiring arranging the position and size of various multi-modal design elements in a visually pleasing and constraint-following manner. Previous a…

Language ModellingLarge Language ModelLayout Generation

Multimodal Foundation Models: From Specialists to General-Purpose Assistants

2023-09-18 · Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang 외

This paper presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to gene…

Image GenerationSurveyText to Image GenerationText-to-Image Generation