paper-with-me

홈 › Papers

Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning

2025-08-11 · Bao Li, Xiaomei Zhang, Miao Xu, Zhaoxin Fan, Xiangyu Zhu, Zhen Lei arxiv

Generating 3D human poses from multimodal inputs such as images or text requires models to capture both rich spatial and semantic correspondences. While pose-specific multimodal large language models (MLLMs) have shown promise in this task, they are typically trained with supervised objectives such as SMPL parameter regression or token-level prediction, which struggle to model the inherent ambiguity and achieve task-specific alignment required for accurate 3D pose generation. To address these limitations, we propose Pose-RFT, a reinforcement fine-tuning framework tailored for 3D human pose generation in MLLMs. We formulate the task as a hybrid action reinforcement learning problem that jointly optimizes discrete language prediction and continuous pose generation. To this end, we introduce HyGRPO, a hybrid reinforcement learning algorithm that performs group-wise reward normalization over sampled responses to guide joint optimization of discrete and continuous actions. Pose-RFT further incorporates task-specific reward functions to guide optimization towards spatial alignment in image-to-pose generation and semantic consistency in text-to-pose generation. Extensive experiments on multiple pose generation benchmarks demonstrate that Pose-RFT significantly improves performance over existing pose-specific MLLMs, validating the effectiveness of hybrid action reinforcement fine-tuning for 3D pose generation.

📄 PDF Abstract BibTeX arXiv:2508.07804

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal

2024-12-15 · Yuhao Wang, Zhiyuan Zhu, Heyang Liu, Yusheng Liao 외

Multimodal large language models (MLLMs) excel at multimodal perception and understanding, yet their tendency to generate hallucinated or inaccurate responses undermines their trustworthiness. Existing methods have large…

Hybrid RAG-empowered Multi-modal LLM for Secure Data Management in Internet of Medical Things: A Diffusion-based Contract Approach

2024-07-01 · Cheng Su, Jinbo Wen, Jiawen Kang, Yonghua Wang 외

Secure data management and effective data sharing have become paramount in the rapidly evolving healthcare landscape, especially with the growing integration of the Internet of Medical Things (IoMT). The rise of generati…

Deep Reinforcement LearningManagementModel-based Reinforcement LearningRAG+2

HyViLM: Enhancing Fine-Grained Recognition with a Hybrid Encoder for Vision-Language Models

2024-12-11 · Shiding Zhu, Wenhui Dong, Jun Song, Yingbo Wang 외

Recently, there has been growing interest in the capability of multimodal large language models (MLLMs) to process high-resolution images. A common approach currently involves dynamically cropping the original high-resol…

TextVQA

MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm

2025-12-02 · Wei Chen, Chaoqun Du, Feng Gu, Wei He 외 arxiv

We present MindGPT-4ov, a multimodal large language model (MLLM) that introduces a general post-training paradigm spanning data production, model training, and efficient deployment. It achieves state-of-the-art performan…

Reinforcement LearningDomain Adaptation

A Hybrid Swarm Intelligence Approach for Optimizing Multimodal Large Language Models Deployment in Edge-Cloud-based Federated Learning Environments

2025-02-04 · Gaith Rjouba, Hanae Elmekki, Saidul Islam, Jamal Bentahar 외

The combination of Federated Learning (FL), Multimodal Large Language Models (MLLMs), and edge-cloud computing enables distributed and real- time data processing while preserving privacy across edge devices and cloud inf…

Cloud ComputingFederated Learning