paper-with-me

홈 › Papers

Adapting Language Specific Components of Cross-Media Analysis Frameworks to Less-Resourced Languages: the Case of Amharic

2020-05-01 · LREC 2020 5 · Yonas Woldemariam, Adam Dahlgren

We present an ASR based pipeline for Amharic that orchestrates NLP components within a cross media analysis framework (CMAF). One of the major challenges that are inherently associated with CMAFs is effectively addressing multi-lingual issues. As a result, many languages remain under-resourced and fail to leverage out of available media analysis solutions. Although spoken natively by over 22 million people and there is an ever-increasing amount of Amharic multimedia content on the Web, querying them with simple text search is difficult. Searching for, especially audio/video content with simple key words, is even hard as they exist in their raw form. In this study, we introduce a spoken and textual content processing workflow into a CMAF for Amharic. We design an ASR-named entity recognition (NER) pipeline that includes three main components: ASR, a transliterator and NER. We explore various acoustic modeling techniques and develop an OpenNLP-based NER extractor along with a transliterator that interfaces between ASR and NER. The designed ASR-NER pipeline for Amharic promotes the multi-lingual support of CMAFs. Also, the state-of-the art design principles and techniques employed in this study shed light for other less-resourced languages, particularly the Semitic ones.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Similar Papers 제목 키워드 기반

Least but not Last: Fine-tuning Intermediate Principal Components for Better Performance-Forgetting Trade-Offs

2026-02-03 · Alessio Quercia, Arya Bangun, Ira Assent, Hanno Scharr arxiv

Low-Rank Adaptation (LoRA) methods have emerged as crucial techniques for adapting large pre-trained models to downstream tasks under computational and memory constraints. However, they face a fundamental challenge in ba…

Continual Learning

Cross-lingual Capsule Network for Hate Speech Detection in Social Media

2021-08-06 · Aiqi Jiang, Arkaitz Zubiaga

Most hate speech detection research focuses on a single language, generally English, which limits their generalisability to other languages. In this paper we investigate the cross-lingual hate speech detection task, tack…

Hate Speech Detection

Investigating Gender Bias in Language Models Using Causal Mediation Analysis

2020-12-01 · NeurIPS 2020 12 · Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 외

Many interpretation methods for neural models in natural language processing investigate how information is encoded inside hidden representations. However, these methods can only measure whether the information exists, n…

Identifying and Adapting Transformer-Components Responsible for Gender Bias in an English Language Model

2023-10-19 · Abhijith Chintam, Rahel Beloch, Willem Zuidema, Michael Hanna 외

Language models (LMs) exhibit and amplify many types of undesirable biases learned from the training data, including gender bias. However, we lack tools for effectively and efficiently changing this behavior without hurt…

Causal DiscoveryLanguage ModelingLanguage Modellingparameter-efficient fine-tuning

Visual Affect Around the World: A Large-scale Multilingual Visual Sentiment Ontology

2015-08-16 · Brendan Jou, Tao Chen, Nikolaos Pappas, Miriam Redi 외

Every culture and language is unique. Our work expressly focuses on the uniqueness of culture and language in relation to human affect, specifically sentiment and emotion semantics, and how they manifest in social multim…

Cultural Vocal Bursts Intensity Prediction