paper-with-me

홈 › Papers

Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic Study

2026-02-16 · Annalisa Szymanski, Oghenemaro Anuyah, Toby Jia-Jun Li, Ronald A. Metoyer arxiv

Large Language Models (LLMs) are increasingly developed for use in complex professional domains, yet little is known about how teams design and evaluate these systems in practice. This paper examines the challenges and trade-offs in LLM development through a 12-week ethnographic study of a team building a pedagogical chatbot. The researcher observed design and evaluation activities and conducted interviews with both developers and domain experts. Analysis revealed four key practices: creating workarounds for data collection, turning to augmentation when expert input was limited, co-developing evaluation criteria with experts, and adopting hybrid expert-developer-LLM evaluation strategies. These practices show how teams made strategic decisions under constraints and demonstrate the central role of domain expertise in shaping the system. Challenges included expert motivation and trust, difficulties structuring participatory design, and questions around ownership and integration of expert knowledge. We propose design opportunities for future LLM development workflows that emphasize AI literacy, transparent consent, and frameworks recognizing evolving expert roles.

📄 PDF Abstract BibTeX arXiv:2602.14357

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

T-FIX: Text-Based Explanations with Features Interpretable to eXperts

2025-11-06 · Shreya Havaldar, Weiqiu You, Chaehyeon Kim, Anton Xue 외 arxiv

As LLMs are deployed in knowledge-intensive settings (e.g., surgery, astronomy, therapy), users are often domain experts who expect not just answers, but explanations that mirror professional reasoning. Yet evaluating wh…

Benchmarking Agents in Insurance Underwriting Environments

2026-01-31 · Amanda Dsouza, Ramya Ramakrishnan, Charles Dickens, Bhavishya Pohani 외 arxiv

As AI agents integrate into enterprise applications, their evaluation demands benchmarks that reflect the complexity of real-world operations. Instead, existing benchmarks overemphasize open-domains such as code, use nar…

Practical Considerations for Agentic LLM Systems

2024-12-05 · Chris Sypherd, Vaishak Belle

As the strength of Large Language Models (LLMs) has grown over recent years, so too has interest in their use as the underlying models for autonomous agents. Although LLMs demonstrate emergent abilities and broad experti…

When AI Bends Metal: AI-Assisted Optimization of Design Parameters in Sheet Metal Forming

2025-11-27 · Ahmad Tarraf, Koutaiba Kassem-Manthey, Seyed Ali Mohammadi, Philipp Martin 외 arxiv

Numerical simulations have revolutionized the industrial design process by reducing prototyping costs, design iterations, and enabling product engineers to explore the design space more efficiently. However, the growing …

Active Learning

Designing forecasting software for forecast users: Empowering non-experts to create and understand their own forecasts

2024-04-22 · Richard Stromer, Oskar Triebe, Chad Zanocco, Ram Rajagopal

Forecasts inform decision-making in nearly every domain. Forecasts are often produced by experts with rare or hard to acquire skills. In practice, forecasts are often used by domain experts and managers with little forec…

Decision Making