paper-with-me

홈 › Papers

Outperforming Good-Turing: Preliminary Report

2018-07-06 · Amichai Painsky, Meir Feder

Estimating a large alphabet probability distribution from a limited number of samples is a fundamental problem in machine learning and statistics. A variety of estimation schemes have been proposed over the years, mostly inspired by the early work of Laplace and the seminal contribution of Good and Turing. One of the basic assumptions shared by most commonly-used estimators is the unique correspondence between the symbol's sample frequency and its estimated probability. In this work we tackle this paradigmatic assumption; we claim that symbols with "similar" frequencies shall be assigned the same estimated probability value. This way we regulate the number of parameters and improve generalization. In this preliminary report we show that by applying an ensemble of such regulated estimators, we introduce a dramatic enhancement in the estimation accuracy (typically up to 50%), compared to currently known methods. An implementation of our suggested method is publicly available at the first author's web-page.

📄 PDF Abstract BibTeX arXiv:1807.02287

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Semantic Similarity in Radiology Reports via LLMs and NER

2025-10-03 · Beth Pearson, Ahmed Adnan, Zahraa S. Abdallah arxiv

Radiology report evaluation is a crucial part of radiologists' training and plays a key role in ensuring diagnostic accuracy. As part of the standard reporting workflow, a junior radiologist typically prepares a prelimin…

Semantic SimilarityClinical Knowledge

Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation

2026-08-18 · Iryna Hartsock, Cesar Lam, Christopher Otteni, Aliya Qayyum 외 arxiv

Purpose: To develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance. Materials and Methods: This retrospective study included 638 radiology reports from CT exa…

An interpretable Good--Turing restart criterion for k-means++

2026-07-09 · Renato Cordeiro de Amorim arxiv

The k-means++ algorithm is commonly restarted multiple times to avoid poor local optima, yet the number of restarts is almost always chosen arbitrarily and applied uniformly regardless of data set difficulty. This underm…

Is ChatGPT a Good NLG Evaluator? A Preliminary Study

2023-03-07 · Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun 외

Recently, the emergence of ChatGPT has attracted wide attention from the computational linguistics community. Many prior studies have shown that ChatGPT achieves remarkable performance on various NLP tasks in terms of au…

nlg evaluationStory GenerationText Generation

Is ChatGPT A Good Keyphrase Generator? A Preliminary Study

2023-03-23 · Mingyang Song, Haiyun Jiang, Shuming Shi, Songfang Yao 외

The emergence of ChatGPT has recently garnered significant attention from the computational linguistics community. To demonstrate its capabilities as a keyphrase generator, we conduct a preliminary evaluation of ChatGPT …

Diversitydocument understandingKeyphrase Generation