paper-with-me

홈 › Papers

Synthesizing the Virtual Advocate: A Multi-Persona Speech Generation Framework for Diverse Linguistic Jurisdictions in Indic Languages

2026-01-19 · Aniket Deroy arxiv

Legal advocacy requires a unique combination of authoritative tone, rhythmic pausing for emphasis, and emotional intelligence. This study investigates the performance of the Gemini 2.5 Flash TTS and Gemini 2.5 Pro TTS models in generating synthetic courtroom speeches across five Indic languages: Tamil, Telugu, Bengali, Hindi, and Gujarati. We propose a prompting framework that utilizes Gemini 2.5s native support for 5 languages and its context-aware pacing to produce distinct advocate personas. The evolution of Large Language Models (LLMs) has shifted the focus of TexttoSpeech (TTS) technology from basic intelligibility to context-aware, expressive synthesis. In the legal domain, synthetic speech must convey authority and a specific professional persona a task that becomes significantly more complex in the linguistically diverse landscape of India. The models exhibit a "monotone authority," excelling at procedural information delivery but struggling with the dynamic vocal modulation and emotive gravitas required for persuasive advocacy. Performance dips in Bengali and Gujarati further highlight phonological frontiers for future refinement. This research underscores the readiness of multilingual TTS for procedural legal tasks while identifying the remaining challenges in replicating the persuasive artistry of human legal discourse. The code is available at-https://github.com/naturenurtureelite/Synthesizing-the-Virtual-Advocate/tree/main

📄 PDF Abstract BibTeX arXiv:2602.11172

Code (0)

등록된 구현이 없습니다.

Tasks

Emotional Intelligence

Similar Papers 제목 키워드 기반

Real-time Gesture Animation Generation from Speech for Virtual Human Interaction

2022-08-05 · Manuel Rebol, Christian Gütl, Krzysztof Pietroszek

We propose a real-time system for synthesizing gestures directly from speech. Our data-driven approach is based on Generative Adversarial Neural Networks to model the speech-gesture relationship. We utilize the large amo…

ComedicSpeech: Text To Speech For Stand-up Comedies in Low-Resource Scenarios

2023-05-20 · Yuyue Wang, Huan Xiao, Yihan Wu, Ruihua Song

Text to Speech (TTS) models can generate natural and high-quality speech, but it is not expressive enough when synthesizing speech with dramatic expressiveness, such as stand-up comedies. Considering comedians have diver…

Rhythmtext-to-speechText to Speech

AMII: Adaptive Multimodal Inter-personal and Intra-personal Model for Adapted Behavior Synthesis

2023-05-18 · Jieyeon Woo, Mireille Fares, Catherine Pelachaud, Catherine Achard

Socially Interactive Agents (SIAs) are physical or virtual embodied agents that display similar behavior as human multimodal behavior. Modeling SIAs' non-verbal behavior, such as speech and facial gestures, has always be…

Modeling Spoken Information Queries for Virtual Assistants: Open Problems, Challenges and Opportunities

2023-04-25 · Christophe Van Gysel

Virtual assistants are becoming increasingly important speech-driven Information Retrieval platforms that assist users with various tasks. We discuss open problems and challenges with respect to modeling spoken informati…

domain classificationInformation RetrievalKnowledge GraphsRetrieval+2

Speech-driven Personalized Gesture Synthetics: Harnessing Automatic Fuzzy Feature Inference

2024-03-16 · Fan Zhang, Zhaohan Wang, Xin Lyu, Siyuan Zhao 외

Speech-driven gesture generation is an emerging field within virtual human creation. However, a significant challenge lies in accurately determining and processing the multitude of input features (such as acoustic, seman…

Gesture Generation