paper-with-me

홈 › Papers

TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints

2026-06-07 · Vinh-Thuan Ly arxiv

Current advancements in Audio Reasoning rely on massive Large Audio-Language Models (LALMs), hindering deployment in resource-constrained environments. We introduce TinyGiantALM, a compact 1.5B efficiency-oriented alternative. Instead of brute-force scaling, we propose an Instruction-Aware Feature Refinement framework using a Query-guided Projector and Semantic Gating to filter acoustic signals based on user intent. On the MMAR benchmark, TinyGiantALM achieves 46.4% zero-shot accuracy, significantly outperforming 7B-13B baselines. While a reasoning gap in logical narrative remains versus 30B+ models and certain trade-offs exist in overly dense or spatial scenes, our approach notably surpasses models up to 8x larger in disentangling mixed-modality environments. These findings demonstrate that architectural precision offers a tangible pathway to secure robust perception capabilities on edge-friendly scales.

📄 PDF Abstract BibTeX arXiv:2606.08425

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval

2026-08-17 · Chen-An Li, Hung-yi Lee arxiv

Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for instruction-aware speech retrieval, in which natural-language in…

Semantic Retrieval

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

2026-06-22 · Quanyue Song, Yishan He, Yanfei Zhang, Shihao Cheng 외 arxiv

Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain visual temporal consistency and fail to explicitly perceive us…

Video Generation

Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving

2026-01-17 · Ziang Guo, Feng Yang, Xuefeng Zhang, Jiaqi Guo 외 arxiv

Vision Language Action (VLA) models promise an open-vocabulary interface that can translate perceptual ambiguity into semantically grounded driving decisions, yet they still treat language as a static prior fixed at infe…

Autonomous Driving

Generalized zero-shot audio-to-intent classification

2023-11-04 · Veera Raghavendra Elluru, Devang Kulshreshtha, Rohit Paturi, Sravan Bodapati 외

Spoken language understanding systems using audio-only data are gaining popularity, yet their ability to handle unseen intents remains limited. In this study, we propose a generalized zero-shot audio-to-intent classifica…

ClassificationGoal-Oriented Dialogintent-classificationIntent Classification+4

Learning Discriminative Representations and Decision Boundaries for Open Intent Detection

2022-03-11 · Hanlei Zhang, Hua Xu, Shaojie Zhao, Qianrui Zhou

Open intent detection is a significant problem in natural language understanding, which aims to identify the unseen open intent while ensuring known intent identification performance. However, current methods face two ma…

Intent DetectionNatural Language UnderstandingOpen Intent Detection