paper-with-me

Papers

Bidimensional Leaderboards: Generate and Evaluate Language Hand in Hand

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Natural language processing researchers have identified limitations of evaluation methodology for generation tasks, with new questions raised about the validity of automatic metrics and of crowdworker judgments. Meanwhile, efforts to improve generation models tend to focus on simple n-gram overlap metrics (e.g., BLEU, ROUGE). We argue that new advances on models and metrics should each more directly benefit and inform the other. We therefore propose a generalization of leaderboards, bidimensional leaderboards (Billboards), that simultaneously tracks progress in language generation tasks and metrics for their evaluation. Unlike conventional unidimensional leaderboards that sort submitted systems by predetermined metrics, a Billboard accepts both generators and evaluation metrics as competing entries. A Billboard automatically creates an ensemble metric that selects and linearly combines a few metrics based on a global analysis across generators. Further, metrics are ranked based on their correlation with human judgments. We release four Billboards for machine translation, summarization, and image captioning. We demonstrate that a linear ensemble of a few diverse metrics sometimes substantially outperforms existing metrics in isolation. Our mixed-effects model analysis shows that most automatic metrics, especially the reference-based ones, overrate machine over human generation, demonstrating the importance of updating metrics as generation models become stronger (and perhaps more similar to humans) in the future.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningMachine TranslationText Generation

Similar Papers 제목 키워드 기반

Bidimensional Leaderboards: Generate and Evaluate Language Hand in Hand

2021-12-08 · NAACL 2022 7 · Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan 외

Natural language processing researchers have identified limitations of evaluation methodology for generation tasks, with new questions raised about the validity of automatic metrics and of crowdworker judgments. Meanwhil…

Image CaptioningMachine TranslationText Generation

BattRAE: Bidimensional Attention-Based Recursive Autoencoders for Learning Bilingual Phrase Embeddings

2016-05-25 · Biao Zhang, Deyi Xiong, Jinsong Su

In this paper, we propose a bidimensional attention based recursive autoencoder (BattRAE) to integrate clues and sourcetarget interactions at multiple levels of granularity into bilingual phrase representations. We emplo…

Semantic SimilaritySemantic Textual Similarity

ScAlN-on-SiC Ku-Band Solidly-Mounted Bidimensional Mode Resonators

2024-11-20 · Luca Colombo, Luca Spagnuolo, Kapil Saha, Gabriel Giribaldi 외

This letter reports on Solidly-Mounted Bidimensional Mode Resonators (S2MRs) exploiting a highly-optimized Sezawa mode in 30% Scandium-doped Aluminum Nitride (ScAlN) on Silicon Carbide (SiC) and operating near 16 GHz. Ex…

LEGOBench: Scientific Leaderboard Generation Benchmark

2024-01-11 · Shruti Singh, Shoaib Alam, Husain Malwat, Mayank Singh

The ever-increasing volume of paper submissions makes it difficult to stay informed about the latest state-of-the-art research. To address this challenge, we introduce LEGOBench, a benchmark for evaluating systems that g…

DecoderLanguage ModelingLanguage Modelling

Non-Homogeneous Haze Removal via Artificial Scene Prior and Bidimensional Graph Reasoning

2021-04-05 · Haoran Wei, Qingbo Wu, Hui Li, King Ngi Ngan 외

Due to the lack of natural scene and haze prior information, it is greatly challenging to completely remove the haze from a single image without distorting its visual content. Fortunately, the real-world haze usually pre…

Image DehazingSingle Image Dehazing