paper-with-me

Papers

Orchestrating LLM Agents for Scientific Research: A Pilot Study of Multiple Choice Question (MCQ) Generation and Evaluation

2026-02-21 · Yuan An arxiv

Advances in large language models (LLMs) are rapidly transforming scientific work, yet empirical evidence on how these systems reshape research activities remains limited. We report a mixed-methods pilot evaluation of an AI-orchestrated research workflow in which a human researcher coordinated multiple LLM-based agents to perform data extraction, corpus construction, artifact generation, and artifact evaluation. Using the generation and assessment of multiple-choice questions (MCQs) as a testbed, we collected 1,071 SAT Math MCQs and employed LLM agents to extract questions from PDFs, retrieve and convert open textbooks into structured representations, align each MCQ with relevant textbook content, generate new MCQs under specified difficulty and cognitive levels, and evaluate both original and generated MCQs using a 24-criterion quality framework. Across all evaluations, average MCQ quality was high. However, criterion-level analysis and equivalence testing show that generated MCQs are not fully comparable to expert-vetted baseline questions. Strict similarity (24/24 criteria equivalent) was never achieved. Persistent gaps concentrated in skill\ depth, cognitive engagement, difficulty calibration, and metadata alignment, while surface-level qualities, such as {grammar fluency}, {clarity options}, {no duplicates}, were consistently strong. Beyond MCQ outcomes, the study documents a labor shift. The researcher's work moved from `authoring items'' toward {specification, orchestration, verification}, and {governance}. Formalizing constraints, designing rubrics, building validation loops, recovering from tool failures, and auditing provenance constituted the primary activities. We discuss implications for the future of scientific work, including emerging `AI research operations'' skills required for AI-empowered research pipelines.

📄 PDF Abstract BibTeX arXiv:2602.18891

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics

2025-10-10 · Lianhao Zhou, Hongyi Ling, Cong Fu, Yepeng Huang 외 arxiv

Computing has long served as a cornerstone of scientific discovery. Recently, a paradigm shift has emerged with the rise of large language models (LLMs), introducing autonomous systems, referred to as agents, that accele…

BrainPilot: Automating Brain Discovery with Agentic Research

2026-07-16 · Haoxuan Li, Tianci Gao, Jianhe Li, Yang Fan 외 arxiv

Understanding the brain increasingly depends on integrating evidence across scales, modalities, and disciplines. Addressing a single research question therefore requires a coordinated sequence of operations, from surveyi…

Multi-Turn Agentic Scientific Literature Search via Workflow Induction

2026-07-01 · Jisen Li, Bingxuan Li, Nanyi Jiang, Xuying Ning 외 hf

Scientific literature search often requires more than retrieving papers from a single query: users' intents are underspecified, preference-dependent, and evolve through interaction. Existing search agents typically rely …

Autonomous LLM-driven research from data to human-verifiable research papers

2024-04-24 · Tal Ifargan, Lukas Hafner, Maor Kern, Ori Alcalay 외

As AI promises to accelerate scientific discovery, it remains unclear whether fully AI-driven research is possible and whether it can adhere to key scientific values, such as transparency, traceability and verifiability.…

scientific discovery

MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration

2024-11-10 · Ziqi Ni, Yahao Li, Kaijia Hu, Kunyuan Han 외

The rapid evolution of artificial intelligence, particularly large language models, presents unprecedented opportunities for materials science research. We proposed and developed an AI materials scientist named MatPilot,…