paper-with-me

홈 › Papers

Predicting Themes within Complex Unstructured Texts: A Case Study on Safeguarding Reports

2020-10-27 · Aleksandra Edwards, David Rogers, Jose Camacho-Collados, Hélène de Ribaupierre, Alun Preece

The task of text and sentence classification is associated with the need for large amounts of labelled training data. The acquisition of high volumes of labelled datasets can be expensive or unfeasible, especially for highly-specialised domains for which documents are hard to obtain. Research on the application of supervised classification based on small amounts of training data is limited. In this paper, we address the combination of state-of-the-art deep learning and classification methods and provide an insight into what combination of methods fit the needs of small, domain-specific, and terminologically-rich corpora. We focus on a real-world scenario related to a collection of safeguarding reports comprising learning experiences and reflections on tackling serious incidents involving children and vulnerable adults. The relatively small volume of available reports and their use of highly domain-specific terminology makes the application of automated approaches difficult. We focus on the problem of automatically identifying the main themes in a safeguarding report using supervised classification approaches. Our results show the potential of deep learning models to simulate subject-expert behaviour even for complex tasks with limited labelled data.

📄 PDF Abstract BibTeX arXiv:2010.14584

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationSentenceSentence Classification

Similar Papers 제목 키워드 기반

BEYONDWORDS is All You Need: Agentic Generative AI based Social Media Themes Extractor

2025-02-26 · Mohammed-Khalil Ghali, Abdelrahman Farrag, Sarah Lam, Daehan Won

Thematic analysis of social media posts provides a major understanding of public discourse, yet traditional methods often struggle to capture the complexity and nuance of unstructured, large-scale text data. This study i…

AllDimensionality Reduction

Potential and Perils of Large Language Models as Judges of Unstructured Textual Data

2025-01-14 · Rewina Bedemariam, Natalie Perez, Sreyoshi Bhaduri, Satya Kapoor 외

Rapid advancements in large language models have unlocked remarkable capabilities when it comes to processing and summarizing unstructured text data. This has implications for the analysis of rich, open-ended datasets, s…

BERT-Deep CNN: State-of-the-Art for Sentiment Analysis of COVID-19 Tweets

2022-11-04 · Javad Hassannataj Joloudari, Sadiq Hussain, Mohammad Ali Nematollahi, Rouhollah Bagheri 외

The free flow of information has been accelerated by the rapid development of social media technology. There has been a significant social and psychological impact on the population due to the outbreak of Coronavirus dis…

ArticlesSentiment Analysis

Automating Historical Insight Extraction from Large-Scale Newspaper Archives via Neural Topic Modeling

2025-12-12 · Keerthana Murugaraj, Salima Lamsiyah, Marten During, Martin Theobald arxiv

Extracting coherent and human-understandable themes from large collections of unstructured historical newspaper archives presents significant challenges due to topic evolution, Optical Character Recognition (OCR) noise, …

Autodetection and Classification of Hidden Cultural City Districts from Yelp Reviews

2015-01-12 · Harini Suresh, Nicholas Locascio

Topic models are a way to discover underlying themes in an otherwise unstructured collection of documents. In this study, we specifically used the Latent Dirichlet Allocation (LDA) topic model on a dataset of Yelp review…

ClusteringGeneral ClassificationTopic Models