paper-with-me

홈 › Papers

Simplify Your Law: Using Information Theory to Deduplicate Legal Documents

2021-10-02 · Corinna Coupette, Jyotsna Singh, Holger Spamann

Textual redundancy is one of the main challenges to ensuring that legal texts remain comprehensible and maintainable. Drawing inspiration from the refactoring literature in software engineering, which has developed methods to expose and eliminate duplicated code, we introduce the duplicated phrase detection problem for legal texts and propose the Dupex algorithm to solve it. Leveraging the Minimum Description Length principle from information theory, Dupex identifies a set of duplicated phrases, called patterns, that together best compress a given input text. Through an extensive set of experiments on the Titles of the United States Code, we confirm that our algorithm works well in practice: Dupex will help you simplify your law.

📄 PDF Abstract BibTeX arXiv:2110.00735

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Let MT simplify and speed up your Alignment for TM creation

2020-11-01 · EAMT 2020 11 · Judith Klein, Giorgio Bernardinello

Large quantities of multilingual legal documents are waiting to be regularly aligned and used for future translations. For reasons of time, effort and cost, manual alignment is not an option. Automatically aligned segmen…

Mining Legal Arguments in Court Decisions

2022-08-12 · Ivan Habernal, Daniel Faber, Nicola Recchia, Sebastian Bretthauer 외

Identifying, classifying, and analyzing arguments in legal discourse has been a prominent area of research since the inception of the argument mining field. However, there has been a major discrepancy between the way nat…

Argument Mining

Machine Learners Should Acknowledge the Legal Implications of Large Language Models as Personal Data

2025-03-03 · Henrik Nolte, Michèle Finck, Kristof Meding

Does GPT know you? The answer depends on your level of public recognition; however, if your information was available on a website, the answer could be yes. Most Large Language Models (LLMs) memorize training data to som…

A Mental Trespass? Unveiling Truth, Exposing Thoughts and Threatening Civil Liberties with Non-Invasive AI Lie Detection

2021-02-16 · Taylan Sen, Kurtis Haut, Denis Lomakin, Ehsan Hoque

Imagine an app on your phone or computer that can tell if you are being dishonest, just by processing affective features of your facial expressions, body movements, and voice. People could ask about your political prefer…

Decision Making

Assessing the Reliability of Large Language Models in the Bengali Legal Context: A Comparative Evaluation Using LLM-as-Judge and Legal Experts

2025-11-07 · Sabik Aftahee, A. F. M. Farhad, Arpita Mallik, Ratnajit Dhar 외 arxiv

Accessing legal help in Bangladesh is hard. People face high fees, complex legal language, a shortage of lawyers, and millions of unresolved court cases. Generative AI models like OpenAI GPT-4.1 Mini, Gemini 2.0 Flash, M…