arXiv · 2312.01356
CEScore: Simple and Efficient Confidence Estimation Model for Evaluating Split and Rephrase
Abstract
The split and rephrase (SR) task aims to divide a long, complex sentence into a set of shorter, simpler sentences that convey the same meaning. This challenging problem in NLP has gained increased attention recently because of its benefits as a pre-processing step in other NLP tasks. Evaluating quality of SR is challenging, as there no automatic metric fit to evaluate this task. In this work, we introduce CEScore, as novel statistical model to automatically evaluate SR task. By mimicking the way humans evaluate SR, CEScore provides 4 metrics (Sscore, Gscore, Mscore, and CEscore) to assess simplicity, grammaticality, meaning preservation, and overall quality, respectively. In experiments with 26 models, CEScore correlates strongly with human evaluations, achieving 0.98 in Spearman correlations at model-level. This underscores the potential of CEScore as a simple and effective metric for assessing the overall quality of SR models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
AlMotasem Bellah Al Ajlouni, Jinlong Li. 2023-12-03. CEScore: Simple and Efficient Confidence Estimation Model for Evaluating Split and Rephrase. https://arxiv.org/abs/2312.01356
Cite the original work for its findings. Save a collection to share your selection of sources.