A Corpus for Multilingual Analysis of Online Terms of Service
We present the first annotated corpus for multilingual analysis of potentially unfair clauses in online Terms of Service. The data set comprises a total of 100 contracts, obtained from 25 documents annotated in four different languages: English, German, Italian, and Polish. For each contract, potentially unfair clauses for the consumer are annotated, for nine different unfairness categories. We show how a simple yet efficient annotation projection technique based on sentence embeddings could be used to automatically transfer annotations across languages.
Code (0)
등록된 구현이 없습니다.
Tasks
SentenceUnfairness DetectionSimilar Papers 제목 키워드 기반
A Million Tweets Are Worth a Few Points: Tuning Transformers for Customer Service Tasks
In online domain-specific customer service applications, many companies struggle to deploy advanced NLP models successfully, due to the limited availability of and noise in their datasets. While prior research demonstrat…
On-line Multilingual Linguistic Services
In this demo, we present our free on-line multilingual linguistic services which allow to analyze sentences or to extract collocations from a corpus directly on-line, or by uploading a corpus. They are available for 8 Eu…
Dependency ParsingPOSPOS TaggingTelenor Nordics Customer Service self-help corpus
This paper presents a multilingual customer service self-help corpus comprising 1,122 manually validated documents in Finnish, Danish, Norwegian, and Swedish, totaling 274,599 words and 1,884,833 characters. The document…
Cross-Lingual TransferInformation RetrievalProduct Review Translation: Parallel Corpus Creation and Robustness towards User-generated Noisy Text
Reviews written by the users for a particular product or service play an influencing role for the customers to make an informative decision. Although online e-commerce portals have immensely impacted our lives, available…
Language ModellingMachine TranslationNMTTranslationToward Multilingual Identification of Online Registers
We consider cross- and multilingual text classification approaches to the identification of online registers (genres), i.e. text varieties with specific situational characteristics. Register is the most important predict…
Multilingual text classificationMultilingual Word Embeddingstext-classificationText Classification+1