paper-with-me

Papers

A Consolidated System for Robust Multi-Document Entity Risk Extraction and Taxonomy Augmentation

2019-09-23 · Berk Ekmekci, Eleanor Hagerman, Blake Howald

We introduce a hybrid human-automated system that provides scalable entity-risk relation extractions across large data sets. Given an expert-defined keyword taxonomy, entities, and data sources, the system returns text extractions based on bidirectional token distances between entities and keywords and expands taxonomy coverage with word vector encodings. Our system represents a more simplified architecture compared to alerting focused systems - motivated by high coverage use cases in the risk mining space such as due diligence activities and intelligence gathering. We provide an overview of the system and expert evaluations for a range of token distances. We demonstrate that single and multi-sentence distance groups significantly outperform baseline extractions with shorter, single sentences being preferred by analysts. As the taxonomy expands, the amount of relevant information increases and multi-sentence extractions become more preferred, but this is tempered against entity-risk relations become more indirect. We discuss the implications of these observations on users, management of ambiguity and taxonomy expansion, and future system modifications.

📄 PDF Abstract BibTeX arXiv:1909.10368

Code (0)

등록된 구현이 없습니다.

Tasks

ManagementSentenceTaxonomy Expansion

Similar Papers 제목 키워드 기반

Can Generative Models Actually Forge Realistic Identity Documents?

2025-12-25 · Alexander Vinogradov arxiv

Generative image models have recently shown significant progress in image realism, leading to public concerns about their potential misuse for document forgery. This paper explores whether contemporary open-source and pu…

Image Generation

AI-based Identity Fraud Detection: A Systematic Review

2025-01-16 · Chuo Jun Zhang, Asif Q. Gill, Bo Liu, Memoona J. Anwar

With the rapid development of digital services, a large volume of personally identifiable information (PII) is stored online and is subject to cyberattacks such as Identity fraud. Most recently, the use of Artificial Int…

Fraud DetectionSystematic Literature Review

Specificity-Based Sentence Ordering for Multi-Document Extractive Risk Summarization

2019-09-23 · Berk Ekmekci, Eleanor Hagerman, Blake Howald

Risk mining technologies seek to find relevant textual extractions that capture entity-risk relationships. However, when high volume data sets are processed, a multitude of relevant extractions can be returned, shifting …

Extractive SummarizationSentenceSentence OrderingSpecificity

A Consolidated Open Knowledge Representation for Multiple Texts

2017-04-01 · WS 2017 4 · Rachel Wities, Vered Shwartz, Gabriel Stanovsky, Meni Adler 외

We propose to move from Open Information Extraction (OIE) ahead to Open Knowledge Representation (OKR), aiming to represent information conveyed jointly in a set of texts in an open text-based manner. We do so by consoli…

Lexical EntailmentOpen Information Extraction

Why AI Harms Can't Be Fixed One Identity at a Time: What 5300 Incident Reports Reveal About Intersectionality

2026-04-27 · Edyta Bogucka, Sanja Šćepanović, Daniele Quercia arxiv

AI risk assessment is the primary tool for identifying harms caused by AI systems. These include intersectional harms, which arise from the interaction between identity categories (e.g., class and skin tone) and which do…