paper-with-me

홈 › Papers

Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese

2026-07-08 · Rodrigo de Freitas Lima, Julio Cesar Galdino, Marcos Vinicius Treviso arxiv

Automatic prosodic segmentation identifies boundaries between speech units from acoustic and linguistic evidence. Although recent deep learning approaches have produced strong results for English, automatic segmentation for Brazilian Portuguese (BP) still relies mostly on rule-based or traditional machine-learning methods. This paper presents SAMPA, a Whisper-based segmenter that transcribes BP speech while inserting explicit markers for terminal prosodic boundaries. We fine-tune Whisper large-v3 on manually segmented recordings from the NURC-SP dataset and evaluate different training and test-time filtering configurations, including out-of-distribution testing on the MuPe-Diversidades dataset. SAMPA achieves competitive boundary-detection performance across settings, with the best models reaching F1=0.731 on the held-out test split and F1=0.796 on MuPe-Diversidades. Finally, through n-gram and acoustic-visual analyses, we show that our model follows morphosyntactic, semantic, and prosodic cues for detecting prosodic boundaries.

📄 PDF Abstract BibTeX arXiv:2607.07408

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Brazilian Portuguese Phonological-prosodic Algorithm Applied to Language Acquisition: A Case Study

2014-04-01 · WS 2014 4 · Vera Vasil{\'e}vski, M{\'a}rcio Jos{\'e} Ara{\'u}jo, Helena Ferro Blasi
Language Acquisition

The C-ORAL-BRASIL I: Reference Corpus for Spoken Brazilian Portuguese

2012-05-01 · LREC 2012 5 · Tommaso Raso, Heliana Mello, Maryual{\^e} Malvessi Mittmann

C-ORAL-BRASIL I is a Brazilian Portuguese spontaneous speech corpus compiled following the same architecture adopted by the C-ORAL-ROM resource. The main goal is the documentation of the diaphasic and diastratic variatio…

text-to-speechText to Speech

The Impact of Prosodic Segmentation on Speech Synthesis of Spontaneous Speech

2025-11-06 · Julio Cesar Galdino, Sidney Evaldo Leal, Leticia Gabriella De Souza, Rodrigo de Freitas Lima 외 arxiv

Spontaneous speech presents several challenges for speech synthesis, particularly in capturing the natural flow of conversation, including turn-taking, pauses, and disfluencies. Although speech synthesis systems have mad…

Speech Synthesis

Introducing the SEA\_AP: an Enhanced Tool for Automatic Prosodic Analysis

2016-05-01 · LREC 2016 5 · Marta Mart{\'\i}nez, Roc{\'\i}o Varela, Carmen Garc{\'\i}a Mateo, Elisa Fern{\'a}ndez Rei 외

SEA{\_}AP (Segmentador e Etiquetador Autom{\'a}tico para An{\'a}lise Pros{\'o}dica, Automatic Segmentation and Labelling for Prosodic Analysis) toolkit is an application that performs audio segmentation and labelling to …

Segmentation

Image captioning for Brazilian Portuguese using GRIT model

2024-02-07 · Rafael Silva de Alencar, William Alberto Cruz Castañeda, Marcellus Amadeus

This work presents the early development of a model of image captioning for the Brazilian Portuguese language. We used the GRIT (Grid - and Region-based Image captioning Transformer) model to accomplish this work. GRIT i…

Image Captioningmodel