paper-with-me

Papers

Evaluating Generative AI-Enhanced Content: A Conceptual Framework Using Qualitative, Quantitative, and Mixed-Methods Approaches

2024-11-26 · Saman Sarraf

Generative AI (GenAI) has revolutionized content generation, offering transformative capabilities for improving language coherence, readability, and overall quality. This manuscript explores the application of qualitative, quantitative, and mixed-methods research approaches to evaluate the performance of GenAI models in enhancing scientific writing. Using a hypothetical use case involving a collaborative medical imaging manuscript, we demonstrate how each method provides unique insights into the impact of GenAI. Qualitative methods gather in-depth feedback from expert reviewers, analyzing their responses using thematic analysis tools to capture nuanced improvements and identify limitations. Quantitative approaches employ automated metrics such as BLEU, ROUGE, and readability scores, as well as user surveys, to objectively measure improvements in coherence, fluency, and structure. Mixed-methods research integrates these strengths, combining statistical evaluations with detailed qualitative insights to provide a comprehensive assessment. These research methods enable quantifying improvement levels in GenAI-generated content, addressing critical aspects of linguistic quality and technical accuracy. They also offer a robust framework for benchmarking GenAI tools against traditional editing processes, ensuring the reliability and effectiveness of these technologies. By leveraging these methodologies, researchers can evaluate the performance boost driven by GenAI, refine its applications, and guide its responsible adoption in high-stakes domains like healthcare and scientific research. This work underscores the importance of rigorous evaluation frameworks for advancing trust and innovation in GenAI.

📄 PDF Abstract BibTeX arXiv:2411.17943

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Safety and Fairness for Content Moderation in Generative Models

2023-06-09 · Susan Hao, Piyush Kumar, Sarah Laszlo, Shivani Poddar 외

With significant advances in generative AI, new technologies are rapidly being deployed with generative components. Generative models are typically trained on large datasets, resulting in model behaviors that can mimic t…

Fairness

LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output

2024-11-09 · Elise Karinshak, Amanda Hu, Kewen Kong, Vishwanatha Rao 외

Immense effort has been dedicated to minimizing the presence of harmful or biased generative content and better aligning AI output to human intention; however, research investigating the cultural values of LLMs is still …

The Narrative Continuity Test: A Conceptual Framework for Evaluating Identity Persistence in AI Systems

2025-10-28 · Stefano Natangelo arxiv

Artificial intelligence systems based on large language models (LLMs) can now generate coherent text, music, and images, yet they operate without a persistent state: each inference reconstructs context from scratch. This…

Bridging the Intent Gap: Knowledge-Enhanced Visual Generation

2024-05-21 · Yi Cheng, Ziwei Xu, Dongyun Lin, Harry Cheng 외

For visual content generation, discrepancies between user intentions and the generated content have been a longstanding problem. This discrepancy arises from two main factors. First, user intentions are inherently comple…

World Knowledge

Evaluating Machine Expertise: How Graduate Students Develop Frameworks for Assessing GenAI Content

2025-04-24 · Celia Chen, Alex Leitch

This paper examines how graduate students develop frameworks for evaluating machine-generated expertise in web-based interactions with large language models (LLMs). Through a qualitative study combining surveys, LLM inte…