What Set of Documents to Present to an Analyst?
We describe the human triage scenario envisioned in the Cross-Lingual Information Retrieval (CLIR) problem of the [REDUCT] Program. The overall goal is to maximize the quality of the set of documents that is given to a bilingual analyst, as measured by the AQWV score. The initial set of source documents that are retrieved by the CLIR system is summarized in English and presented to human judges who attempt to remove the irrelevant documents (false alarms); the resulting documents are then presented to the analyst. First, we describe the AQWV performance measure and show that, in our experience, if the acceptance threshold of the CLIR component has been optimized to maximize AQWV, the loss in AQWV due to false alarms is relatively constant across many conditions, which also limits the possible gain that can be achieved by any post filter (such as human judgments) that removes false alarms. Second, we analyze the likely benefits for the triage operation as a function of the initial CLIR AQWV score and the ability of the human judges to remove false alarms without removing relevant documents. Third, we demonstrate that we can increase the benefit for human judgments by combining the human judgment scores with the original document scores returned by the automatic CLIR system.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Lingual Information RetrievalInformation RetrievalRetrievalSimilar Papers 제목 키워드 기반
Do Sell-side Analyst Reports Have Investment Value?
This paper documents new investment value in analyst reports. Analyst narratives embedded with large language models strongly forecast future stock returns, generating significant alpha beyond established analyst-based a…
Where did you get that? Towards Summarization Attribution for Analysts
Analysts require attribution, as nothing can be reported without knowing the source of the information. In this paper, we will focus on automatic methods for attribution, linking each sentence in the summary to a portion…
LLM Augmentations to support Analytical Reasoning over Multiple Documents
Building on their demonstrated ability to perform a variety of tasks, we investigate the application of large language models (LLMs) to enhance in-depth analytical reasoning within the context of intelligence analysis. I…
NetReAct: Interactive Learning for Network Summarization
Generating useful network summaries is a challenging and important problem with several applications like sensemaking, visualization, and compression. However, most of the current work in this space do not take human fee…
Irregularity Detection in Categorized Document Corpora
The paper presents an approach to extract irregularities in document corpora, where the documents originate from different sources and the analyst's interest is to find documents which are atypical for the given source. …
ArticlesDocument ClassificationOutlier DetectionText Categorization