paper-with-me

홈 › Papers

Building and Modelling Multilingual Subjective Corpora

2014-05-01 · LREC 2014 5 · Motaz Saad, David Langlois, Kamel Sma{\"\i}li

Building multilingual opinionated models requires multilingual corpora annotated with opinion labels. Unfortunately, such kind of corpora are rare. We consider opinions in this work as subjective or objective. In this paper, we introduce an annotation method that can be reliably transferred across topic domains and across languages. The method starts by building a classifier that annotates sentences into subjective/objective label using a training data from {``}movie reviews{''} domain which is in English language. The annotation can be transferred to another language by classifying English sentences in parallel corpora and transferring the same annotation to the same sentences of the other language. We also shed the light on the link between opinion mining and statistical language modelling, and how such corpora are useful for domain specific language modelling. We show the distinction between subjective and objective sentences which tends to be stable across domains and languages. Our experiments show that language models trained on objective (respectively subjective) corpus lead to better perplexities on objective (respectively subjective) test.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModellingMachine TranslationOpinion MiningSentiment AnalysisSubjectivity Analysis

Similar Papers 제목 키워드 기반

Conception: Multilingually-Enhanced, Human-Readable Concept Vector Representations

2020-12-01 · COLING 2020 8 · Simone Conia, Roberto Navigli

To date, the most successful word, word sense, and concept modelling techniques have used large corpora and knowledge resources to produce dense vector representations that capture semantic similarities in a relatively l…

Word Sense DisambiguationWord Similarity

Optimizing Multilingual Text-To-Speech with Accents & Emotions

2025-06-19 · Pranav Pawar, Akshansh Dwivedi, Jenish Boricha, Himanshu Gohil 외

State-of-the-art text-to-speech (TTS) systems realize high naturalness in monolingual environments, synthesizing speech with correct multilingual accents (especially for Indic languages) and context-relevant emotions sti…

DisentanglementEmotion Recognitiontext-to-speechText to Speech+1

DHPLT: large-scale multilingual diachronic corpora and word representations for semantic change modelling

2026-02-12 · Mariia Fedorova, Andrey Kutuzov, Khonzoda Umarova arxiv

In this resource paper, we present DHPLT, an open collection of diachronic corpora in 41 diverse languages. DHPLT is based on the web-crawled HPLT datasets; we use web crawl timestamps as the approximate signal of docume…

Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams

2020-06-20 · Huirong Huang, Zhiyong Wu, Shiyin Kang, Dongyang Dai 외

Generating 3D speech-driven talking head has received more and more attention in recent years. Recent approaches mainly have following limitations: 1) most speaker-independent methods need handcrafted features that are t…

Talking Head Generation

Building Open Javanese and Sundanese Corpora for Multilingual Text-to-Speech

2018-05-01 · LREC 2018 5 · Jaka Aris Eko Wibawa, Supheakmungkol Sarin, Chenfang Li, Knot Pipatsrisawat 외
Automatic Speech Recognition (ASR)Speech RecognitionSpeech Synthesistext-to-speech+1