paper-with-me

홈 › Papers

FOMO: Topics versus documents in legal eDiscovery

2021-09-16 · Herbert Roitblat

In the United States, the parties to a lawsuit are required to search through their electronically stored information to find documents that are relevant to the specific case and produce them to their opposing party. Negotiations over the scope of these searches often reflect a fear that something will be missed (Fear of Missing Out: FOMO). A Recall level of 80%, for example, means that 20% of the relevant documents will be left unproduced. This paper makes the argument that eDiscovery is the process of identifying responsive information, not identifying documents. Documents are the carriers of the information; they are not the direct targets of the process. A given document may contain one or more topics or factoids and a factoid may appear in more than one document. The coupon collector's problem, Heaps law, and other analyses provide ways to model the problem of finding information from among documents. In eDiscovery, however, the parties do not know how many factoids there might be in a collection or their probabilities. This paper describes a simple model that estimates the confidence that a fact will be omitted from the produced set (the identified set), while being contained in the missed set. Two data sets are then analyzed, a small set involving microaggressions and larger set involving classification of web pages. Both show that it is possible to discover at least one example of each available topic within a relatively small number of documents, meaning the further effort will not return additional novel information. The smaller data set is also used to investigate whether the non-random order of searching for responsive documents commonly used in eDiscovery (called continuous active learning) affects the distribution of topics-it does not.

📄 PDF Abstract BibTeX arXiv:2109.08059

Code (0)

등록된 구현이 없습니다.

Tasks

Active Learning

Similar Papers 제목 키워드 기반

Is there something I'm missing? Topic Modeling in eDiscovery

2020-07-30 · Herbert L. Roitblat

In legal eDiscovery, the parties are required to search through their electronically stored information to find documents that are relevant to a specific case. Negotiations over the scope of these searches are often base…

Active Learning

Probably Reasonable Search in eDiscovery

2022-01-28 · Herbert L. Roitblat

In eDiscovery, a party to a lawsuit or similar action must search through available information to identify those documents and files that are relevant to the suit. Search efforts tend to identify less than 100% of the r…

Learning from Litigation: Graphs and LLMs for Retrieval and Reasoning in eDiscovery

2024-05-29 · Sounak Lahiri, Sumit Pai, Tim Weninger, Sanmitra Bhattacharya

Electronic Discovery (eDiscovery) involves identifying relevant documents from a vast collection based on legal production requests. The integration of artificial intelligence (AI) and natural language processing (NLP) h…

Language ModelingLanguage ModellingLarge Language ModelRetrieval

Annotation of argument structure in Japanese legal documents

2017-09-01 · WS 2017 9 · Hiroaki Yamada, Simone Teufel, Takenobu Tokunaga

We propose a method for the annotation of Japanese civil judgment documents, with the purpose of creating flexible summaries of these. The first step, described in the current paper, concerns content selection, i.e., the…

Argument Mining

D2GCLF: Document-to-Graph Classifier for Legal Document Classification

2022-07-01 · Findings (NAACL) 2022 7 · Qiqi Wang, Kaiqi Zhao, Robert Amor, Benjamin Liu 외

Legal document classification is an essential task in law intelligence to automate the labor-intensive law case filing process. Unlike traditional document classification problems, legal documents should be classified by…

ClassificationDocument ClassificationGraph AttentionRelation