paper-with-me

Papers

A State-of-the-practice Release-readiness Checklist for Generative AI-based Software Products

2024-03-27 · Harsh Patel, Dominique Boucher, Emad Fallahzadeh, Ahmed E. Hassan, Bram Adams

This paper investigates the complexities of integrating Large Language Models (LLMs) into software products, with a focus on the challenges encountered for determining their readiness for release. Our systematic review of grey literature identifies common challenges in deploying LLMs, ranging from pre-training and fine-tuning to user experience considerations. The study introduces a comprehensive checklist designed to guide practitioners in evaluating key release readiness aspects such as performance, monitoring, and deployment strategies, aiming to enhance the reliability and effectiveness of LLM-based applications in real-world settings.

📄 PDF Abstract BibTeX arXiv:2403.18958

Code (1)

SAILResearch/replication-24-harsh-generative-ai-release-readiness-checklist 공식 구현

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

The Minimum Information about CLinical Artificial Intelligence Checklist for Generative Modeling Research (MI-CLAIM-GEN)

2024-03-05 · Brenda Y. Miao, Irene Y. Chen, Christopher YK Williams, Jaysón Davidson 외

Recent advances in generative models, including large language models (LLMs), vision language models (VLMs), and diffusion models, have accelerated the field of natural language and image processing in medicine and marke…

ACL Ready: RAG Based Assistant for the ACL Checklist

2024-08-07 · Michael Galarnyk, Rutwik Routu, Kosha Bheda, Priyanshu Mehta 외

The ARR Responsible NLP Research checklist website states that the "checklist is designed to encourage best practices for responsible research, addressing issues of research ethics, societal impact and reproducibility." …

EthicsLanguage ModelingLanguage ModellingRAG

Sphinx: Benchmarking and Modeling for LLM-Driven Pull Request Review

2026-01-06 · Daoan Zhang, Shuo Zhang, Zijian Jin, Jiebo Luo 외 arxiv

Pull request (PR) review is essential for ensuring software quality, yet automating this task remains challenging due to noisy supervision, limited contextual understanding, and inadequate evaluation metrics. We present …

Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness

2026-07-30 · Fouad Bousetouane arxiv

AI agents are moving into production workflows where they retrieve information, call tools, maintain state, and act on behalf of users or organizations, but many release decisions still rely on capability signals, demos,…

Exploring Human-AI Collaboration in Agile: Customised LLM Meeting Assistants

2024-04-23 · Beatriz Cabrero-Daniel, Tomas Herda, Victoria Pichler, Martin Eder

This action research study focuses on the integration of "AI assistants" in two Agile software development meetings: the Daily Scrum and a feature refinement, a planning meeting that is part of an in-house Scaled Agile f…