paper-with-me

Papers

PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models

2024-06-12 · Runyan Yang, Huibao Yang, Xiqing Zhang, Tiantian Ye, Ying Liu, Yingying Gao, Shilei Zhang, Chao Deng, Junlan Feng

Recently, there have been attempts to integrate various speech processing tasks into a unified model. However, few previous works directly demonstrated that joint optimization of diverse tasks in multitask speech models has positive influence on the performance of individual tasks. In this paper we present a multitask speech model -- PolySpeech, which supports speech recognition, speech synthesis, and two speech classification tasks. PolySpeech takes multi-modal language model as its core structure and uses semantic representations as speech inputs. We introduce semantic speech embedding tokenization and speech reconstruction methods to PolySpeech, enabling efficient generation of high-quality speech for any given speaker. PolySpeech shows competitiveness across various tasks compared to single-task models. In our experiments, multitask optimization achieves performance comparable to single-task optimization and is especially beneficial for specific tasks.

📄 PDF Abstract BibTeX arXiv:2406.07801

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingspeech-recognitionSpeech RecognitionSpeech Synthesis

Similar Papers 제목 키워드 기반

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

2026-05-31 · Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu 외 arxiv

While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription. Existing benchmarks suffer from three critical lim…

SpeechEQ: Speech Emotion Recognition based on Multi-scale Unified Datasets and Multitask Learning

2022-06-27 · Zuheng Kang, Junqing Peng, Jianzong Wang, Jing Xiao

Speech emotion recognition (SER) has many challenges, but one of the main challenges is that each framework does not have a unified standard. In this paper, we propose SpeechEQ, a framework for unifying SER tasks based o…

Emotion RecognitionPhoneme RecognitionSpeech Emotion Recognition

A Survey on Speech Large Language Models

2024-10-24 · Jing Peng, Yucheng Wang, Yangui Fang, Yu Xi 외

Large Language Models (LLMs) exhibit strong contextual understanding and remarkable multitask performance. As a result, researchers have been actively exploring the integration of LLMs into the domain of speech understan…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeech Emotion Recognition+6

Exploring Universal Speech Attributes for Speaker Verification with an Improved Cross-stitch Network

2020-10-13 · Jiajun Qi, Wu Guo, Jingjing Shi, Yafeng Chen 외

The universal speech attributes for x-vector based speaker verification (SV) are addressed in this paper. The manner and place of articulation form the fundamental speech attribute unit (SAU), and then new speech attribu…

AttributeSpeaker Verification

SyntaxMind at BLP-2025 Task 1: Leveraging Attention Fusion of CNN and GRU for Hate Speech Detection

2026-01-09 · Md. Shihab Uddin Riad arxiv

This paper describes our system used in the BLP-2025 Task 1: Hate Speech Detection. We participated in Subtask 1A and Subtask 1B, addressing hate speech classification in Bangla text. Our approach employs a unified archi…

Hate Speech Detection