paper-with-me

Papers

Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies

2026-05-12 · Alberta Longhini, David Emukpere, Jean-Michel Renders, Seungsu Kim arxiv

We address the problem of fine-tuning pre-trained generative policies with reinforcement learning (RL) while preserving the multimodality of their action distributions. Existing methods for RL fine-tuning of generative policies (e.g., diffusion policies) improve task performance but often collapse diverse behaviors into a single reward-maximizing mode. To mitigate this issue, we propose an unsupervised mode discovery framework that uncovers latent behavioral modes within generative policies. The discovered modes enable the use of mutual information as an intrinsic reward, regularizing RL fine-tuning to enhance task success while maintaining behavioral diversity. Experiments on robotic manipulation tasks demonstrate that our method consistently outperforms conventional fine-tuning approaches, achieving higher success rates and preserving richer multimodal action distributions.

📄 PDF Abstract BibTeX arXiv:2605.11387

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

BehaviorVLM: Unified Finetuning-Free Behavioral Understanding with Vision-Language Reasoning

2026-03-12 · Jingyang Ke, Weihan Li, Amartya Pradhan, Jeffrey Markowitz 외 arxiv

Understanding freely moving animal behavior is central to neuroscience, where pose estimation and behavioral understanding form the foundation for linking neural activity to natural actions. Yet both tasks still depend h…

Video CaptioningPose Estimation

Fine-tuning a multimodal large language model for clinician-grade autism behavioral scoring from short home videos

2026-06-25 · Mohammadmahdi Honarmand, Parnian Azizian, Aaron Kline, Kae Nurge 외 arxiv

Autism spectrum disorder (ASD) affects 1 in 31 US children, yet median age at diagnosis exceeds four years. Artificial intelligence pipelines that provide quantified diagnosis using easy to access observational data (e.g…

Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment

2025-08-27 · Julian Arnold, Niels Lörch arxiv

Fine-tuning LLMs on narrowly harmful datasets can lead to behavior that is broadly misaligned with respect to human values. To understand when and how this emergent misalignment occurs, we develop a comprehensive framewo…

Change Detection

MMQ: Multimodal Mixture-of-Quantization Tokenization for Semantic ID Generation and User Behavioral Adaptation

2025-08-21 · Yi Xu, Moyu Zhang, Chenxuan Li, Zhihao Liao 외 arxiv

Recommender systems traditionally represent items using unique identifiers (ItemIDs), but this approach struggles with large, dynamic item corpora and sparse long-tail data, limiting scalability and generalization. Seman…

Delivery Optimized Discovery in Behavioral User Segmentation under Budget Constraint

2024-02-04 · Harshita Chopra, Atanu R. Sinha, Sunav Choudhary, Ryan A. Rossi 외

Users' behavioral footprints online enable firms to discover behavior-based user segments (or, segments) and deliver segment specific messages to users. Following the discovery of segments, delivery of messages to users …

Stochastic Optimization