paper-with-me

홈 › Papers

LePaRD: A Large-Scale Dataset of Judges Citing Precedents

2023-11-15 · Robert Mahari, Dominik Stammbach, Elliott Ash, Alex `Sandy' Pentland

We present the Legal Passage Retrieval Dataset LePaRD. LePaRD is a massive collection of U.S. federal judicial citations to precedent in context. The dataset aims to facilitate work on legal passage prediction, a challenging practice-oriented legal retrieval and reasoning task. Legal passage prediction seeks to predict relevant passages from precedential court decisions given the context of a legal argument. We extensively evaluate various retrieval approaches on LePaRD, and find that classification appears to work best. However, we note that legal precedent prediction is a difficult task, and there remains significant room for improvement. We hope that by publishing LePaRD, we will encourage others to engage with a legal NLP task that promises to help expand access to justice by reducing the burden associated with legal research. A subset of the LePaRD dataset is freely available and the whole dataset will be released upon publication.

📄 PDF Abstract BibTeX arXiv:2311.09356

Code (1)

rmahari/lepard 공식 구현 pytorch

Tasks

Passage RetrievalPredictionRetrieval

Similar Papers 제목 키워드 기반

Lepard: Learning partial point cloud matching in rigid and deformable scenes

2021-11-24 · CVPR 2022 1 · Yang Li, Tatsuya Harada

We present Lepard, a Learning based approach for partial point cloud matching in rigid and deformable scenes. The key characteristics are the following techniques that exploit 3D positional knowledge for point cloud matc…

3D Feature Matching3D Point Cloud MatchingPartial Point Cloud MatchingPoint Cloud Registration+1

LEPARD: Learning Explicit Part Discovery for 3D Articulated Shape Reconstruction

2023-09-21 · NeurIPS 2023 11

Reconstructing the 3D articulated shape of an animal from a single in-the-wild image is a challenging task. We propose LEPARD, a learning-based framework that discovers semantically meaningful 3D parts and reconstructs 3…

A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness

2026-02-04 · Leo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami 외 arxiv

Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, these judges are relied upon to evaluate har…

Adversarial Robustness

When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation

2025-10-21 · Abeer Badawi, Elahe Rahimi, Md Tahmid Rahman Laskar, Sheri Grach 외 arxiv

Evaluating Large Language Models (LLMs) for mental health support is challenging due to the emotionally and cognitively complex nature of therapeutic dialogue. Existing benchmarks are limited in scale, reliability, often…

JudgeLM: Fine-tuned Large Language Models are Scalable Judges

2023-10-26 · Lianghui Zhu, Xinggang Wang, Xinlong Wang

Evaluating Large Language Models (LLMs) in open-ended scenarios is challenging because existing benchmarks and metrics can not measure them comprehensively. To address this problem, we propose to fine-tune LLMs as scalab…