MYCanCor: A Video Corpus of spoken Malaysian Cantonese
Code (1)
Similar Papers 제목 키워드 기반
CantoMap: a Hong Kong Cantonese MapTask Corpus
This work reports on the construction of a corpus of connected spoken Hong Kong Cantonese. The corpus aims at providing an additional resource for the study of modern (Hong Kong) Cantonese and also involves several contr…
HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation
We introduce HK-LegiCoST, a new three-way parallel corpus of Cantonese-English translations, containing 600+ hours of Cantonese audio, its standard traditional Chinese transcript, and English translation, segmented and a…
Cross-corpusSentencespeech-recognitionSpeech Recognition+1Cifu: a Frequency Lexicon of Hong Kong Cantonese
This paper introduces Cifu, a lexical database for Hong Kong Cantonese (HKC) that offers phonological and orthographic information, frequency measures, and lexical neighborhood information for lexical items in HKC. Cifu …
DiversityWeCanTalk: A New Multi-language, Multi-modal Resource for Speaker Recognition
The WeCanTalk (WCT) Corpus is a new multi-language, multi-modal resource for speaker recognition. The corpus contains Cantonese, Mandarin and English telephony and video speech data from over 200 multilingual speakers lo…
Speaker RecognitionEnriching Linguistic Representation in the Cantonese Wordnet and Building the New Cantonese Wordnet Corpus
This paper reports on the most recent improvements on the Cantonese Wordnet, a wordnet project started in 2019 (Sio and Morgado da Costa, 2019) with the aim of capturing and organizing lexico-semantic information of Hong…