paper-with-me

Papers

A Framework For Refining Text Classification and Object Recognition from Academic Articles

2023-05-27 · Jinghong Li, Koichi Ota, Wen Gu, Shinobu Hasegawa

With the widespread use of the internet, it has become increasingly crucial to extract specific information from vast amounts of academic articles efficiently. Data mining techniques are generally employed to solve this issue. However, data mining for academic articles is challenging since it requires automatically extracting specific patterns in complex and unstructured layout documents. Current data mining methods for academic articles employ rule-based(RB) or machine learning(ML) approaches. However, using rule-based methods incurs a high coding cost for complex typesetting articles. On the other hand, simply using machine learning methods requires annotation work for complex content types within the paper, which can be costly. Furthermore, only using machine learning can lead to cases where patterns easily recognized by rule-based methods are mistakenly extracted. To overcome these issues, from the perspective of analyzing the standard layout and typesetting used in the specified publication, we emphasize implementing specific methods for specific characteristics in academic articles. We have developed a novel Text Block Refinement Framework (TBRF), a machine learning and rule-based scheme hybrid. We used the well-known ACL proceeding articles as experimental data for the validation experiment. The experiment shows that our approach achieved over 95% classification accuracy and 90% detection accuracy for tables and figures.

📄 PDF Abstract BibTeX arXiv:2305.17401

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesObject Recognitiontext-classificationText Classification

Similar Papers 제목 키워드 기반

Deep Attributes from Context-Aware Regional Neural Codes

2015-09-08 · Jianwei Luo, Jianguo Li, Jun Wang, Zhiguo Jiang 외

Recently, many researches employ middle-layer output of convolutional neural network models (CNN) as features for different visual recognition tasks. Although promising results have been achieved in some empirical studie…

AttributeGeneral Classificationimage-classificationImage Classification+1

Improving Multi-label Recognition using Class Co-Occurrence Probabilities

2024-04-24 · Samyak Rawlekar, Shubhang Bhatnagar, Vishnuvardhan Pogunulu Srinivasulu, Narendra Ahuja

Multi-label Recognition (MLR) involves the identification of multiple objects within an image. To address the additional complexity of this problem, recent works have leveraged information from vision-language models (VL…

Useful Blunders: Can Automated Speech Recognition Errors Improve Downstream Dementia Classification?

2024-01-10 · Changye Li, Weizhe Xu, Trevor Cohen, Serguei Pakhomov

\textbf{Objectives}: We aimed to investigate how errors from automatic speech recognition (ASR) systems affect dementia classification accuracy, specifically in the ``Cookie Theft'' picture description task. We aimed to …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Classificationspeech-recognition+1

Listening, Imagining & Refining: A Heuristic Optimized ASR Correction Framework with LLMs

2025-09-18 · Yutong Liu, Ziyue Zhang, Cheng Huang, Yongbin Yu 외 arxiv

Automatic Speech Recognition (ASR) systems remain prone to errors that affect downstream applications. In this paper, we propose LIR-ASR, a heuristic optimized iterative correction framework using LLMs, inspired by human…

Speech Recognition

MMRAG: Multi-Mode Retrieval-Augmented Generation with Large Language Models for Biomedical In-Context Learning

2025-02-21 · Zaifu Zhan, Jun Wang, Shuang Zhou, Jiawen Deng 외

Objective: To optimize in-context learning in biomedical natural language processing by improving example selection. Methods: We introduce a novel multi-mode retrieval-augmented generation (MMRAG) framework, which integr…

DiversityDrug–drug Interaction ExtractionIn-Context Learningnamed-entity-recognition+8