paper-with-me

Papers

FinRAD: Financial Readability Assessment Dataset - 13,000+ Definitions of Financial Terms for Measuring Readability

2022-06-01 · FNP (LREC) 2022 6 · Sohom Ghosh, Shovon Sengupta, Sudip Naskar, Sunny Kumar Singh

In today’s world, the advancement and spread of the Internet and digitalization have resulted in most information being openly accessible. This holds true for financial services as well. Investors make data driven decisions by analysing publicly available information like annual reports of listed companies, details regarding asset allocation of mutual funds, etc. Many a time these financial documents contain unknown financial terms. In such cases, it becomes important to look at their definitions. However, not all definitions are equally readable. Readability largely depends on the structure, complexity and constituent terms that make up a definition. This brings in the need for automatically evaluating the readability of definitions of financial terms. This paper presents a dataset, FinRAD consisting of financial terms, their definitions and embeddings. In addition to standard readability scores (like “Flesch Reading Index (FRI)”, “Automated Readability Index (ARI)”, “SMOG Index Score (SIS)”,“Dale-Chall formula (DCF)”, etc.), it also contains the readability scores (AR) assigned based on sources from which the terms have been collected. We manually inspect a sample from it to ensure the quality of the assignment. Subsequently, we prove that the rule-based standard readability scores (like “Flesch Reading Index (FRI)”, “Automated Readability Index (ARI)”, “SMOG Index Score (SIS)”,“Dale-Chall formula (DCF)”, etc.) do not correlate well with the manually assigned binary readability scores of definitions of financial terms. Finally, we present a few neural baselines using transformer based architecture to automatically classify these definitions as readable or not. Pre-trained FinBERT model fine-tuned on FinRAD corpus performs the best (AU-ROC = 0.9927, F1 = 0.9610). This corpus can be downloaded from https://github.com/sohomghosh/FinRAD_Financial_Readability_Assessment_Dataset.

📄 PDF Abstract BibTeX

Code (1)

sohomghosh/finrad_financial_readability_assessment_dataset 공식 구현

Similar Papers 제목 키워드 기반

FinRead: A Transfer Learning Based Tool to Assess Readability of Definitions of Financial Terms

2021-12-01 · ICON 2021 12 · Sohom Ghosh, Shovon Sengupta, Sudip Naskar, Sunny Kumar Singh

Simplified definitions of complex terms help learners to understand any content better. Comprehending readability is critical for the simplification of these contents. In most cases, the standard formula based readabilit…

SentenceSentence EmbeddingsTransfer Learning

Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics

2025-10-17 · Catarina G Belem, Parker Glenn, Alfy Samuel, Anoop Kumar 외 arxiv

Automatic readability assessment plays a key role in ensuring effective and accessible written communication. Despite significant progress, the field is hindered by inconsistent definitions of readability and measurement…

A Machine Learning Approach to Persian Text Readability Assessment Using a Crowdsourced Dataset

2018-10-07 · Hamid Mohammadi, Seyed Hossein Khasteh

An automated approach to text readability assessment is essential to a language and can be a powerful tool for improving the understandability of texts written and published in that language. However, the Persian languag…

BIG-bench Machine Learning

Hierarchical Ranking Neural Network for Long Document Readability Assessment

2025-11-26 · Yurui Zheng, Yijun Chen, Shaohong Zhang arxiv

Readability assessment aims to evaluate the reading difficulty of a text. In recent years, while deep learning technology has been gradually applied to readability assessment, most approaches fail to consider either the …

Inter-Rater Agreement Study on Readability Assessment in Bengali

2014-07-08 · Shanta Phani, Shibamouli Lahiri, Arindam Biswas

An inter-rater agreement study is performed for readability assessment in Bengali. A 1-7 rating scale was used to indicate different levels of readability. We obtained moderate to fair agreement among seven independent a…