paper-with-me

Papers

On-device Large Multi-modal Agent for Human Activity Recognition

2025-12-17 · Md Shakhrul Iman Siam, Ishtiaque Ahmed Showmik, Guanqun Song, Ting Zhu arxiv

Human Activity Recognition (HAR) has been an active area of research, with applications ranging from healthcare to smart environments. The recent advancements in Large Language Models (LLMs) have opened new possibilities to leverage their capabilities in HAR, enabling not just activity classification but also interpretability and human-like interaction. In this paper, we present a Large Multi-Modal Agent designed for HAR, which integrates the power of LLMs to enhance both performance and user engagement. The proposed framework not only delivers activity classification but also bridges the gap between technical outputs and user-friendly insights through its reasoning and question-answering capabilities. We conduct extensive evaluations using widely adopted HAR datasets, including HHAR, Shoaib, Motionsense to assess the performance of our framework. The results demonstrate that our model achieves high classification accuracy comparable to state-of-the-art methods while significantly improving interpretability through its reasoning and Q&A capabilities.

📄 PDF Abstract BibTeX arXiv:2512.19742

Code (0)

등록된 구현이 없습니다.

Tasks

Human Activity Recognition

Similar Papers 제목 키워드 기반

YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks

2025-01-16 · Saptarashmi Bandyopadhyay, Vikas Bahirwani, Lavisha Aggarwal, Bhanu Guda 외

Multimodal AI Agents are AI models that have the capability of interactively and cooperatively assisting human users to solve day-to-day tasks. Augmented Reality (AR) head worn devices can uniquely improve the user exper…

AI AgentScene UnderstandingSSIM

Benchmarking Mobile Device Control Agents across Diverse Configurations

2024-04-25 · Juyong Lee, Taywon Min, Minyong An, Dongyoon Hahm 외

Mobile device control agents can largely enhance user interactions and productivity by automating daily tasks. However, despite growing interest in developing practical agents, the absence of a commonly adopted benchmark…

BenchmarkingImitation Learning

Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

2024-01-29 · Junyang Wang, Haiyang Xu, Jiabo Ye, Ming Yan 외

Mobile device agent based on Multimodal Large Language Models (MLLM) is becoming a popular application. In this paper, we introduce Mobile-Agent, an autonomous multi-modal mobile device agent. Mobile-Agent first leverage…

Octopus v3: Technical Report for On-device Sub-billion Multimodal AI Agent

2024-04-17 · Wei Chen, Zhiyuan Li

A multimodal AI agent is characterized by its ability to process and learn from various types of data, including natural language, visual, and audio inputs, to inform its actions. Despite advancements in large language m…

AI Agent

AppAgent v2: Advanced Agent for Flexible Mobile Interactions

2024-08-05 · Yanda Li, Chi Zhang, Wanqi Yang, Bin Fu 외

With the advancement of Multimodal Large Language Models (MLLM), LLM-driven visual agents are increasingly impacting software interfaces, particularly those with graphical user interfaces. This work introduces a novel LL…

RAG