paper-with-me

홈 › Papers

AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising

2024-08-12 · Peinan Zhang, Yusuke Sakai, Masato Mita, Hiroki Ouchi, Taro Watanabe

With the increase in the more fluent ad texts automatically created by natural language generation technology, it is in the high demand to verify the quality of these creatives in a real-world setting. We propose AdTEC, the first public benchmark to evaluate ad texts in multiple aspects from the perspective of practical advertising operations. Our contributions are: (i) Defining five tasks for evaluating the quality of ad texts and building a dataset based on the actual operational experience of advertising agencies, which is typically kept in-house. (ii) Validating the performance of existing pre-trained language models (PLMs) and human evaluators on the dataset. (iii) Analyzing the characteristics and providing challenges of the benchmark. The results show that while PLMs have already reached the practical usage level in several tasks, human still outperforms in certain domains, implying that there is significant room for improvement in such area.

📄 PDF Abstract BibTeX arXiv:2408.05906

Code (1)

cyberagentailab/adtec 공식 구현

Tasks

Text Generation

Similar Papers 제목 키워드 기반

CoSy: Evaluating Textual Explanations of Neurons

2024-05-30 · Laura Kopf, Philine Lou Bommer, Anna Hedström, Sebastian Lapuschkin 외

A crucial aspect of understanding the complex nature of Deep Neural Networks (DNNs) is the ability to explain learned concepts within their latent representations. While methods exist to connect neurons to human-understa…

Benchmarking

LCTG Bench: LLM Controlled Text Generation Benchmark

2025-01-27 · Kentaro Kurihara, Masato Mita, Peinan Zhang, Shota Sasaki 외

The rise of large language models (LLMs) has led to more diverse and higher-quality machine-generated text. However, their high expressive power makes it difficult to control outputs based on specific business instructio…

Text Generation

Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content

2025-03-04 · CVPR 2025 1 · ZiCheng Zhang, Tengchuan Kou, Shushi Wang, Chunyi Li 외

Evaluating text-to-vision content hinges on two crucial aspects: visual quality and alignment. While significant progress has been made in developing objective models to assess these dimensions, the performance of such m…

A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models

2024-02-21 · Ashutosh Sathe, Prachi Jain, Sunayana Sitaram

Vision-language models (VLMs) have gained widespread adoption in both industry and academia. In this study, we propose a unified framework for systematically evaluating gender, race, and age biases in VLMs with respect t…

BenchmarkingImage to text

Fundamental Challenges in Evaluating Text2SQL Solutions and Detecting Their Limitations

2025-01-30 · Cedric Renggli, Ihab F. Ilyas, Theodoros Rekatsinas

In this work, we dive into the fundamental challenges of evaluating Text2SQL solutions and highlight potential failure causes and the potential risks of relying on aggregate metrics in existing benchmarks. We identify tw…