paper-with-me

홈 › Papers

Contextualization for the Organization of Text Documents Streams

2022-05-30 · Rui Portocarrero Sarmento, Douglas O. Cardoso, João Gama, Pavel Brazdil

There has been a significant effort by the research community to address the problem of providing methods to organize documentation with the help of information Retrieval methods. In this report paper, we present several experiments with some stream analysis methods to explore streams of text documents. We use only dynamic algorithms to explore, analyze, and organize the flux of text documents. This document shows a case study with developed architectures of a Text Document Stream Organization, using incremental algorithms like Incremental TextRank, and IS-TFIDF. Both these algorithms are based on the assumption that the mapping of text documents and their document-term matrix in lower-dimensional evolving networks provides faster processing when compared to batch algorithms. With this architecture, and by using FastText Embedding to retrieve similarity between documents, we compare methods with large text datasets and ground truth evaluation of clustering capacities. The datasets used were Reuters and COVID-19 emotions. The results provide a new view for the contextualization of similarity when approaching flux of documents organization tasks, based on the similarity between documents in the flux, and by using mentioned algorithms.

📄 PDF Abstract BibTeX arXiv:2206.02632

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrieval

Methods 이 논문이 사용한 방법론

fastText fastText embeddings exploit subword information to construct word embeddings. Representations are learnt of character $n$-grams, and words represented as the sum of the…

Similar Papers 제목 키워드 기반

LOCALINTEL: Generating Organizational Threat Intelligence from Global and Local Cyber Knowledge

2024-01-18 · Shaswata Mitra, Subash Neupane, Trisha Chakraborty, Sudip Mittal 외

Security Operations Center (SoC) analysts gather threat reports from openly accessible global threat repositories and tailor the information to their organization's needs, such as developing threat intelligence and secur…

Retrieval

SPIRE: Structure-Preserving Interpretable Retrieval of Evidence

2026-02-12 · Mike Rainey, Umut Acar, Muhammed Sezer arxiv

Retrieval-augmented generation over semi-structured sources such as HTML is constrained by a mismatch between document structure and the flat, sequence-based interfaces of today's embedding and generative models. Retriev…

Contextualization with SPLADE for High Recall Retrieval

2024-05-07 · Eugene Yang

High Recall Retrieval (HRR), such as eDiscovery and medical systematic review, is a search problem that optimizes the cost of retrieving most relevant documents in a given collection. Iterative approaches, such as iterat…

RetrievalTAR

SelfDoc: Self-Supervised Document Representation Learning

2021-06-07 · CVPR 2021 1 · Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu 외

We propose SelfDoc, a task-agnostic pre-training framework for document image understanding. Because documents are multimodal and are intended for sequential reading, our framework exploits the positional, textual, and v…

Representation Learning

A Question Answering Framework for Decontextualizing User-facing Snippets from Scientific Documents

2023-05-24 · Benjamin Newman, Luca Soldaini, Raymond Fok, Arman Cohan 외

Many real-world applications (e.g., note taking, search) require extracting a sentence or paragraph from a document and showing that snippet to a human outside of the source document. Yet, users may find snippets difficu…

Question AnsweringQuestion GenerationQuestion-GenerationSentence