paper-with-me

홈 › Papers

PoseLLM: Enhancing Language-Guided Human Pose Estimation with MLP Alignment

2025-07-12 · Dewen Zhang, Tahir Hussain, Wangpeng An, Hayaru Shouno

Human pose estimation traditionally relies on architectures that encode keypoint priors, limiting their generalization to novel poses or unseen keypoints. Recent language-guided approaches like LocLLM reformulate keypoint localization as a vision-language task, enabling zero-shot generalization through textual descriptions. However, LocLLM's linear projector fails to capture complex spatial-textual interactions critical for high-precision localization. To address this, we propose PoseLLM, the first Large Language Model (LLM)-based pose estimation framework that replaces the linear projector with a nonlinear MLP vision-language connector. This lightweight two-layer MLP with GELU activation enables hierarchical cross-modal feature transformation, enhancing the fusion of visual patches and textual keypoint descriptions. Trained exclusively on COCO data, PoseLLM achieves 77.8 AP on the COCO validation set, outperforming LocLLM by +0.4 AP, while maintaining strong zero-shot generalization on Human-Art and MPII. Our work demonstrates that a simple yet powerful nonlinear connector significantly boosts localization accuracy without sacrificing generalization, advancing the state-of-the-art in language-guided pose estimation. Code is available at https://github.com/Ody-trek/PoseLLM.

📄 PDF Abstract BibTeX arXiv:2507.09139

Code (0)

등록된 구현이 없습니다.

Tasks

Large Language ModelPose EstimationZero-shot Generalization

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…

Similar Papers 제목 키워드 기반

Text-guided Diffusion Model for 3D Molecule Generation

2024-10-04 · Yanchen Luo, Junfeng Fang, Sihang Li, Zhiyuan Liu 외

The de novo generation of molecules with targeted properties is crucial in biology, chemistry, and drug discovery. Current generative models are limited to using single property values as conditions, struggling with comp…

3D Molecule GenerationDiversityDrug Discoverymodel

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

2025-09-16 · Ali Salamatian, Amirhossein Abaskohi, Wan-Cyuan Fan, Mir Rayat Imtiaz Hossain 외 arxiv

Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particular…

Chart Question Answering

Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts

2024-10-31 · Xiang Deng, Youxin Pang, Xiaochen Zhao, Chao Xu 외

This paper introduces Stereo-Talker, a novel one-shot audio-driven human video synthesis system that generates 3D talking videos with precise lip synchronization, expressive body gestures, temporally consistent photo-rea…

Language ModelingLanguage ModellingLarge Language ModelMixture-of-Experts+1

Generative AI for Industrial Contour Detection: A Language-Guided Vision System

2025-08-29 · Liang Gong, Tommy, Wang, Sara Chaker 외 arxiv

Industrial computer vision systems often struggle with noise, material variability, and uncontrolled imaging conditions, limiting the effectiveness of classical edge detectors and handcrafted pipelines. In this work, we …

Language-Guided Reinforcement Learning for Hard Attention in Few-Shot Learning

2023-10-11 · Bahareh Nikpour, Narges Armanfard

Attention mechanisms have demonstrated significant potential in enhancing learning models by identifying key portions of input data, particularly in scenarios with limited training samples. Inspired by human perception, …

Deep Reinforcement LearningFew-Shot LearningHard Attention