paper-with-me

Papers

MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting

2024-06-11 · Zhiqi Ai, Zhiyong Chen, Shugong Xu

In this paper, we propose MM-KWS, a novel approach to user-defined keyword spotting leveraging multi-modal enrollments of text and speech templates. Unlike previous methods that focus solely on either text or speech features, MM-KWS extracts phoneme, text, and speech embeddings from both modalities. These embeddings are then compared with the query speech embedding to detect the target keywords. To ensure the applicability of MM-KWS across diverse languages, we utilize a feature extractor incorporating several multilingual pre-trained models. Subsequently, we validate its effectiveness on Mandarin and English tasks. In addition, we have integrated advanced data augmentation tools for hard case mining to enhance MM-KWS in distinguishing confusable words. Experimental results on the LibriPhrase and WenetPhrase datasets demonstrate that MM-KWS outperforms prior methods significantly.

📄 PDF Abstract BibTeX arXiv:2406.07310

Code (1)

aizhiqi-work/MM-KWS 공식 구현 pytorch

Tasks

Data AugmentationKeyword Spotting

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

LogoDiffuser: Training-Free Multilingual Logo Generation and Stylization via Letter-Aware Attention Control

2026-03-10 · Mingyu Kang, Hyein Seo, Yuna Jeong, Junhyeong Park 외 arxiv

Recent advances in text-to-image generation have been remarkable, but generating multilingual design logos that harmoniously integrate visual and textual elements remains a challenging task. Existing methods often distor…

Text-to-Image GenerationText Generation

Do What I Say: A Spoken Prompt Dataset for Instruction-Following

2026-03-10 · Maike Züfle, Sara Papi, Fabian Retkowski, Szymon Mazurek 외 arxiv

Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompts, which may not reflect real-world scenarios where users interact with …

"Haet Bhasha aur Diskrimineshun": Phonetic Perturbations in Code-Mixed Hinglish to Red-Team LLMs

2025-05-20 · Darpan Aswal, Siddharth D Jaiswal

Large Language Models (LLMs) have become increasingly powerful, with multilingual and multimodal capabilities improving by the day. These models are being evaluated through audits, alignment studies and red-teaming effor…

Image GenerationRed TeamingSafety AlignmentText Generation

Multimodal Conditional 3D Face Geometry Generation

2024-07-01 · Christopher Otto, Prashanth Chandran, Sebastian Weiss, Markus Gross 외

We present a new method for multimodal conditional 3D face geometry generation that allows user-friendly control over the output identity and expression via a number of different conditioning signals. Within a single mod…

3D geometryFace GenerationFace Model

Multilingual Jailbreak Challenges in Large Language Models

2023-10-10 · Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong Bing

While large language models (LLMs) exhibit remarkable capabilities across a wide range of tasks, they pose potential safety concerns, such as the ``jailbreak'' problem, wherein malicious instructions can manipulate LLMs …