paper-with-me

Papers

Accessing accurate documents by mining auxiliary document information

2016-04-15 · Jinju Joby, Jyothi Korra

Earlier techniques of text mining included algorithms like k-means, Naive Bayes, SVM which classify and cluster the text document for mining relevant information about the documents. The need for improving the mining techniques has us searching for techniques using the available algorithms. This paper proposes one technique which uses the auxiliary information that is present inside the text documents to improve the mining. This auxiliary information can be a description to the content. This information can be either useful or completely useless for mining. The user should assess the worth of the auxiliary information before considering this technique for text mining. In this paper, a combination of classical clustering algorithms is used to mine the datasets. The algorithm runs in two stages which carry out mining at different levels of abstraction. The clustered documents would then be classified based on the necessary groups. The proposed technique is aimed at improved results of document clustering.

📄 PDF Abstract BibTeX arXiv:1604.04558

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Methods 이 논문이 사용한 방법론

SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

Extracting Body Text from Academic PDF Documents for Text Mining

2020-10-23 · Changfeng Yu, Cheng Zhang, Jie Wang

Accurate extraction of body text from PDF-formatted academic documents is essential in text-mining applications for deeper semantic understandings. The objective is to extract complete sentences in the body text into a t…

Sentence

Mining both Commonality and Specificity from Multiple Documents for Multi-Document Summarization

2023-03-05 · Bing Ma

The multi-document summarization task requires the designed summarizer to generate a short text that covers the important information of original documents and satisfies content diversity. This paper proposes a multi-doc…

DiversityDocument SummarizationMulti-Document SummarizationSpecificity

Generating an Overview Report over Many Documents

2019-08-17 · Jingwen Wang, Hao Zhang, Cheng Zhang, Wenjing Yang 외

How to efficiently generate an accurate, well-structured overview report (ORPT) over thousands of related documents is challenging. A well-structured ORPT consists of sections of multiple levels (e.g., sections and subse…

AttributeDecision MakingDiversityDocument Summarization+1

EdgeDoc: Hybrid CNN-Transformer Model for Accurate Forgery Detection and Localization in ID Documents

2025-08-22 · Anjith George, Sebastien Marcel arxiv

The widespread availability of tools for manipulating images and documents has made it increasingly easy to forge digital documents, posing a serious threat to Know Your Customer (KYC) processes and remote onboarding sys…

SimDoc: Topic Sequence Alignment based Document Similarity Framework

2016-11-15 · Gaurav Maheshwari, Priyansh Trivedi, Harshita Sahijwani, Kunal Jha 외

Document similarity is the problem of estimating the degree to which a given pair of documents has similar semantic content. An accurate document similarity measure can improve several enterprise relevant tasks such as d…

ClusteringQuestion AnsweringSemantic SimilaritySemantic Textual Similarity