paper-with-me

Papers

Using a Language Technology Infrastructure for German in order to Anonymize German Sign Language Corpus Data

2016-05-01 · LREC 2016 5 · Julian Bleicken, Thomas Hanke, Uta Salden, Sven Wagner

For publishing sign language corpus data on the web, anonymization is crucial even if it is impossible to hide the visual appearance of the signers: In a small community, even vague references to third persons may be enough to identify those persons. In the case of the DGS Korpus (German Sign Language corpus) project, we want to publish data as a contribution to the cultural heritage of the sign language community while annotation of the data is still ongoing. This poses the question how well anonymization can be achieved given that no full linguistic analysis of the data is available. Basically, we combine analysis of all data that we have, including named entity recognition on translations into German. For this, we use the WebLicht language technology infrastructure. We report on the reliability of these methods in this special context and also illustrate how the anonymization of the video data is technically achieved in order to minimally disturb the viewer.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Similar Papers 제목 키워드 기반

A9-Dataset: Multi-Sensor Infrastructure-Based Dataset for Mobility Research

2022-04-13 · Christian Creß, Walter Zimmer, Leah Strand, Venkatnarayanan Lakshminarasimhan 외

Data-intensive machine learning based techniques increasingly play a prominent role in the development of future mobility solutions - from driver assistance and automation functions in vehicles, to real-time traffic mana…

Management

Enabling automated driving by ICT infrastructure: A reference architecture

2020-03-11

Information and communication technology (ICT) is an enabler for establishing automated vehicles (AVs) in today's traffic systems. By providing complementary and/or redundant information via radio communication to the AV…

CO-Fun: A German Dataset on Company Outsourcing in Fund Prospectuses for Named Entity Recognition and Relation Extraction

2024-03-22 · Neda Foroutan, Markus Schröder, Andreas Dengel

The process of cyber mapping gives insights in relationships among financial entities and service providers. Centered around the outsourcing practices of companies within fund prospectuses in Germany, we introduce a data…

named-entity-recognitionNamed Entity RecognitionRelationRelation Extraction

AI-assisted German Employment Contract Review: A Benchmark Dataset

2025-01-27 · Oliver Wardas, Florian Matthes

Employment contracts are used to agree upon the working conditions between employers and employees all over the world. Understanding and reviewing contracts for void or unfair clauses requires extensive knowledge of the …

Fairness

The Multilingual Amazon Reviews Corpus

2020-10-06 · EMNLP 2020 11 · Phillip Keung, Yichao Lu, György Szarvas, Noah A. Smith

We present the Multilingual Amazon Reviews Corpus (MARC), a large-scale collection of Amazon reviews for multilingual text classification. The corpus contains reviews in English, Japanese, German, French, Spanish, and Ch…

ClassificationCross-Lingual TransferGeneral ClassificationMultilingual text classification+4