paper-with-me

홈 › Papers

CEScore: Simple and Efficient Confidence Estimation Model for Evaluating Split and Rephrase

2023-12-03 · AlMotasem Bellah Al Ajlouni, Jinlong Li

The split and rephrase (SR) task aims to divide a long, complex sentence into a set of shorter, simpler sentences that convey the same meaning. This challenging problem in NLP has gained increased attention recently because of its benefits as a pre-processing step in other NLP tasks. Evaluating quality of SR is challenging, as there no automatic metric fit to evaluate this task. In this work, we introduce CEScore, as novel statistical model to automatically evaluate SR task. By mimicking the way humans evaluate SR, CEScore provides 4 metrics (Sscore, Gscore, Mscore, and CEscore) to assess simplicity, grammaticality, meaning preservation, and overall quality, respectively. In experiments with 26 models, CEScore correlates strongly with human evaluations, achieving 0.98 in Spearman correlations at model-level. This underscores the potential of CEScore as a simple and effective metric for assessing the overall quality of SR models.

📄 PDF Abstract BibTeX arXiv:2312.01356

Code (1)

motasemajlouni/cescore 공식 구현

Tasks

SentenceSplit and Rephrase

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs

2026-06-01 · Abhishek Aich, Sparsh Garg, Vijay Kumar BG, Turgun Yusuf Kashgari 외 arxiv

Driving vision-language models (VLMs) must accurately understand scenes across diverse conditions defined by Operational Design Domains (ODDs), yet verification remains sparse: many slices are missing, making empirical f…

FaceScore: Benchmarking and Enhancing Face Quality in Human Generation

2024-06-24 · Zhenyi Liao, Qingsong Xie, Chen Chen, Hannan Lu 외

Diffusion models (DMs) have achieved significant success in generating imaginative images given textual descriptions. However, they are likely to fall short when it comes to real-life scenarios with intricate details. Th…

BenchmarkingDenoisingFace GenerationImage Generation+2

Large Language Model Confidence Estimation via Black-Box Access

2024-06-01 · Tejaswini Pedapati, Amit Dhurandhar, Soumya Ghosh, Soham Dan 외

Estimating uncertainty or confidence in the responses of a model can be significant in evaluating trust not only in the responses, but also in the model as a whole. In this paper, we explore the problem of estimating con…

Language ModelingLanguage ModellingLarge Language Model

FlowPainter: Inpainting Optical Flow via Confidence-Guided Completion

2026-07-11 · Yuang Meng, Chenyang Wu, Xianshun Liu, Chun-Le Guo 외 arxiv

Existing optical flow methods broadly follow two paradigms: iterative optimization and diffusion-based estimation. Iterative methods, exemplified by RAFT, achieve high accuracy through recurrent refinement, but remain ch…

Optical Flow Estimation

A Confidence Machine for Sparse High-Order Interaction Model

2022-05-28 · Diptesh Das, Eugene Ndiaye, Ichiro Takeuchi

In predictive modeling for high-stake decision-making, predictors must be not only accurate but also reliable. Conformal prediction (CP) is a promising approach for obtaining the confidence of prediction results with few…

Conformal PredictionDecision MakingPredictionVocal Bursts Intensity Prediction