paper-with-me

Papers

Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models

2025-08-10 · Xuanhan Wang, Huimin Deng, Ke Liu, Jun Wang, Lianli Gao, Jingkuan Song arxiv

Human-centric vision models (HVMs) have achieved remarkable generalization due to large-scale pretraining on massive person images. However, their dependence on large neural architectures and the restricted accessibility of pretraining data significantly limits their practicality in real-world applications. To address this limitation, we propose Dynamic Pattern Alignment Learning (DPAL), a novel distillation-based pretraining framework that efficiently trains lightweight HVMs to acquire strong generalization from large HVMs. In particular, human-centric visual perception are highly dependent on three typical visual patterns, including global identity pattern, local shape pattern and multi-person interaction pattern. To achieve generalizable lightweight HVMs, we firstly design a dynamic pattern decoder (D-PaDe), acting as a dynamic Mixture of Expert (MoE) model. It incorporates three specialized experts dedicated to adaptively extract typical visual patterns, conditioned on both input image and pattern queries. And then, we present three levels of alignment objectives, which aims to minimize generalization gap between lightweight HVMs and large HVMs at global image level, local pixel level, and instance relation level. With these two deliberate designs, the DPAL effectively guides lightweight model to learn all typical human visual patterns from large HVMs, which can generalize to various human-centric vision tasks. Extensive experiments conducted on 15 challenging datasets demonstrate the effectiveness of the DPAL. Remarkably, when employing PATH-B as the teacher, DPAL-ViT/Ti (5M parameters) achieves surprising generalizability similar to existing large HVMs such as PATH-B (84M) and Sapiens-L (307M), and outperforms previous distillation-based pretraining methods including Proteus-ViT/Ti (5M) and TinyMiM-ViT/Ti (5M) by a large margin.

📄 PDF Abstract BibTeX arXiv:2508.07144

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Gravity-Aware Hierarchical Routing for Lightweight SensorLLM on Human Activity Recognition

2026-06-01 · Hao Li, Mingrui Zheng, Yasuyuki Tahara, Yuichi Sei arxiv

Recent studies on sensor-language alignment have shown that two-stage frameworks can improve the semantic modeling ability of wearable-sensor human activity recognition (HAR), where SensorLLM-style methods first perform …

Human Activity Recognition

Decoding Dynamic Visual Experience from Calcium Imaging via Cell-Pattern-Aware Pretraining

2025-10-21 · Sangyoon Bae, Mehdi Azabou, Blake Richards, Jiook Cha arxiv

Neural recordings exhibit a distinctive form of heterogeneity rooted in differences in cell types, intrinsic circuit dynamics, and stochastic stimulus-response variability that goes beyond ordinary dataset variability, m…

Self-Supervised LearningRepresentation Learning

Toward Generalizable Deblurring: Leveraging Massive Blur Priors with Linear Attention for Real-World Scenarios

2026-01-10 · Yuanting Gao, Shuo Cao, Xiaohui Li, Yuandong Pu 외 arxiv

Image deblurring has advanced rapidly with deep learning, yet most methods exhibit poor generalization beyond their training datasets, with performance dropping significantly in real-world scenarios. Our analysis shows t…

Image Deblurring

Thin Bridges for Drug Text Alignment: Lightweight Contrastive Learning for Target Specific Drug Retrieval

2025-09-30 · Mallikarjuna Tupakula arxiv

Multimodal foundation models hold promise for drug discovery and biomedical applications, but most existing approaches rely on heavy pretraining or large scale multimodal corpora. We investigate whether thin contrastive …

Contrastive LearningDrug Discovery

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model

2026-05-29 · Xiang Zhu, Puzhen Yuan, Yichen Liu, Jianyu Chen arxiv

Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observations and executable actions. While latent…

Representation Learning