Robust Summarization Evaluation Benchmark
홈페이지 · 논문 1편
Robust Summarization Evaluation Benchmark is a large human evaluation dataset consisting of over 22k summary-level annotations over state-of-the-art systems on three datasets. Source: Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation
Texts