paper-with-me

홈 › Papers

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

2025-06-06 · Tian Lan, Yang-Hao Zhou, Zi-Ao Ma, Fanshu Sun, Rui-Qing Sun, Junyu Luo, Rong-Cheng Tu, Heyan Huang, Chen Xu, Zhijing Wu, Xian-Ling Mao

Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these generated outputs presents ongoing challenges. Although numerous automatic evaluation methods exist, current research lacks a systematic framework that comprehensively organizes these methods across text, visual, and audio modalities. To address this issue, we present a comprehensive review and a unified taxonomy of automatic evaluation methods for generated content across all three modalities; We identify five fundamental paradigms that characterize existing evaluation approaches across these domains. Our analysis begins by examining evaluation methods for text generation, where techniques are most mature. We then extend this framework to image and audio generation, demonstrating its broad applicability. Finally, we discuss promising directions for future research in cross-modal evaluation methodologies.

📄 PDF Abstract BibTeX arXiv:2506.10019

Code (0)

등록된 구현이 없습니다.

Tasks

Audio GenerationText Generation

Similar Papers 제목 키워드 기반

What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation

2024-08-26 · Dingyi Yang, Qin Jin

With the development of artificial intelligence, particularly the success of Large Language Models (LLMs), the quantity and quality of automatically generated stories have significantly increased. This has led to the nee…

Machine Translation

Evaluation of Text Generation: A Survey

2020-06-26 · Asli Celikyilmaz, Elizabeth Clark, Jianfeng Gao

The paper surveys evaluation methods of natural language generation (NLG) systems that have been developed in the last few years. We group NLG evaluation methods into three categories: (1) human-centric evaluation metric…

nlg evaluationSurveyText GenerationText Summarization

Vision-Language Models under Cultural and Inclusive Considerations

2024-07-08 · Antonia Karamolegkou, Phillip Rust, Yong Cao, Ruixiang Cui 외

Large vision-language models (VLMs) can assist visually impaired people by describing images from their daily lives. Current evaluation datasets may not reflect diverse cultural user backgrounds or the situational contex…

HallucinationSurvey

SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys

2025-12-02 · Jiahao Zhao, Shuaixing Zhang, Nan Xu, Lei Wang arxiv

LLM-based automatic survey systems are transforming how users acquire information from the web by integrating retrieval, organization, and content synthesis into end-to-end generation pipelines. While recent works focus …

Automatic Scene Generation: State-of-the-Art Techniques, Models, Datasets, Challenges, and Future Prospects

2025-05-28 · IEEE Access 2025 5 · Awal Ahmed Fime, Saifuddin Mahmud, Arpita Das, Md. Sunzidul Islam 외

Automatic scene generation is an essential area of research with applications in robotics, recreation, visual representation, training and simulation, education, and more. This survey provides a comprehensive review of t…

3D GenerationImage to 3DScene GenerationSurvey+1