Cross-lingual Annotation Projection in Legal Texts
We study annotation projection in text classification problems where source documents are published in multiple languages and may not be an exact translation of one another. In particular, we focus on the detection of unfair clauses in privacy policies and terms of service. We present the first English-German parallel asymmetric corpus for the task at hand. We study and compare several language-agnostic sentence-level projection methods. Our results indicate that a combination of word embeddings and dynamic time warping performs best.
Code (1)
Tasks
Cross-Lingual TransferDynamic Time WarpingSentencetext-classificationText ClassificationTranslationWord EmbeddingsMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Legal document retrieval across languages: topic hierarchies based on synsets
Cross-lingual annotations of legislative texts enable us to explore major themes covered in multilingual legal data and are a key facilitator of semantic similarity when searching for similar documents. Multilingual prob…
RetrievalSemantic SimilaritySemantic Textual SimilarityTopic Models+1Multilingual Projection for Parsing Truly Low-Resource Languages
We propose a novel approach to cross-lingual part-of-speech tagging and dependency parsing for truly low-resource languages. Our annotation projection-based approach yields tagging and parsing models for over 100 languag…
Cross-Lingual TransferDependency ParsingPart-Of-Speech TaggingTransfer LearningResolving Legalese: A Multilingual Exploration of Negation Scope Resolution in Legal Documents
Resolving the scope of a negation within a sentence is a challenging NLP task. The complexity of legal texts and the lack of annotated in-domain negation corpora pose challenges for state-of-the-art (SotA) models when pe…
NegationNegation Scope ResolutionSentenceLex Rosetta: Transfer of Predictive Models Across Languages, Jurisdictions, and Legal Domains
In this paper, we examine the use of multi-lingual sentence embeddings to transfer predictive models for functional segmentation of adjudicatory decisions across jurisdictions, legal systems (common and civil law), langu…
SegmentationSentenceSentence EmbeddingsPAXQA: Generating Cross-lingual Question Answering Examples at Training Scale
Existing question answering (QA) systems owe much of their success to large, high-quality training data. Such annotation efforts are costly, and the difficulty compounds in the cross-lingual setting. Therefore, prior cro…
Cross-Lingual Question AnsweringDataset GenerationMachine TranslationQuestion Answering+3