paper-with-me

홈 › Papers

Lex Rosetta: Transfer of Predictive Models Across Languages, Jurisdictions, and Legal Domains

2021-12-15 · Jaromir Savelka, Hannes Westermann, Karim Benyekhlef, Charlotte S. Alexander, Jayla C. Grant, David Restrepo Amariles, Rajaa El Hamdani, Sébastien Meeùs, Michał Araszkiewicz, Kevin D. Ashley, Alexandra Ashley, Karl Branting, Mattia Falduti, Matthias Grabmair, Jakub Harašta, Tereza Novotná, Elizabeth Tippett, Shiwanni Johnson

In this paper, we examine the use of multi-lingual sentence embeddings to transfer predictive models for functional segmentation of adjudicatory decisions across jurisdictions, legal systems (common and civil law), languages, and domains (i.e. contexts). Mechanisms for utilizing linguistic resources outside of their original context have significant potential benefits in AI & Law because differences between legal systems, languages, or traditions often block wider adoption of research outcomes. We analyze the use of Language-Agnostic Sentence Representations in sequence labeling models using Gated Recurrent Units (GRUs) that are transferable across languages. To investigate transfer between different contexts we developed an annotation scheme for functional segmentation of adjudicatory decisions. We found that models generalize beyond the contexts on which they were trained (e.g., a model trained on administrative decisions from the US can be applied to criminal law decisions from Italy). Further, we found that training the models on multiple contexts increases robustness and improves overall performance when evaluating on previously unseen contexts. Finally, we found that pooling the training data from all the contexts enhances the models' in-context performance.

📄 PDF Abstract BibTeX arXiv:2112.07882

Code (1)

lexrosetta/caselaw_functional_segmentation_multilingual 공식 구현 pytorch

Tasks

SegmentationSentenceSentence Embeddings

Similar Papers 제목 키워드 기반

Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions

2026-05-28 · Volodymyr Ovcharov arxiv

Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual comparison impossible. We introduce Multi-Legal-Bench, the first cros…

Legal Reasoning

Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics

2025-08-14 · Carter Blum, Katja Filippova, Ann Yuan, Asma Ghandeharioun 외 arxiv

Large language models (LLMs) struggle with cross-lingual knowledge transfer: they sometimes hallucinate when asked in one language about facts expressed in a different language during training. This work introduces a con…

Cross-Lingual Transfer

CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming

2024-10-27 · Ali TehraniJamsaz, Arijit Bhattacharjee, Le Chen, Nesreen K. Ahmed 외

Recent advancements in Large Language Models (LLMs) have renewed interest in automatic programming language translation. Encoder-decoder transformer models, in particular, have shown promise in translating between differ…

Code TranslationDecoderTranslation

Rosetta at AlexandriaX-2026: LoRA-Adapted NileChat for Context-Aware Dialectal Arabic Dialogue Translation

2026-09-09 · Nada Esmaeil, Fathima Rena, Sibi Subhash, Osama Elgendy 외 arxiv

This paper describes the Rosetta system for Subtask 1 (Context-Aware English-to-Dialectal Arabic Dialogue Translation) of the AlexandriaX shared task, participating in both constrained and unconstrained tracks. The appro…

Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning

2025-03-25 · Shaun Baek, Shaun Esua-Mensah, Cyrus Tsui, Sejan Vigneswaralingam 외

Large Language Models (LLMs) are primarily trained on high-resource natural languages, limiting their effectiveness in low-resource settings and in tasks requiring deep logical reasoning. This research introduces Rosetta…

Language ModelingLanguage ModellingLarge Language ModelLogical Reasoning+1