paper-with-me

홈 › Papers

Theory Under Construction: Orchestrating Language Models for Research Software Where the Specification Evolves

2026-04-29 · Halley Young, Nikolaj Björner arxiv

Large language models can now generate substantial code and draft research text, but research-software projects require more than either artifact alone. The mathematical thesis, executable system, benchmark surface, and public claims must mature together, yet often drift apart. We identify two LM-specific failure modes: hallucination accumulation, in which claims exceed what code or theory supports and unsupported assertions propagate across sessions; and desynchronization, in which code, theory, or the model's own world model fall out of alignment. We propose Comet-H, an iterative prompt automaton that orchestrates ideation, implementation, evaluation, grounding, and paper-writing as coupled coordinates of a single workspace state. At each step, a controller selects the next prompt by scoring it against what the workspace currently lacks, carries unfinished follow-up work forward with a half-life, and re-checks the paper and README against the code and benchmarks whenever documentation changes. We frame prompt selection as a small contextual bandit problem over prompt families, with prompts as arms, workspace deficits as context, and a hand-weighted linear score. This transparent scorer, paired with a fading record of unfinished work, bounds long-horizon follow-ups, requires no learned policy, and makes each prompt choice legible from the workspace. We created a portfolio of 46 research-software repositories across two dozen domains. We study A3 in depth, a Python static-analysis tool built entirely within the loop, which reaches (F1 = 0.768) on a 90-case benchmark, compared with a next-best baseline of 0.364. Across approximately 400 commits, we find that audit-and-contraction passes dominate the later phases of every successful trajectory.

📄 PDF Abstract BibTeX arXiv:2604.27209

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

InnerPond: Fostering Inter-Self Dialogue with a Multi-Agent Approach for Introspection

2026-03-29 · Hayeon Jeon, Dakyeom Ahn, Sunyu Pang, Yunseo Choi 외 arxiv

Introspection is central to identity construction and future planning, yet most digital tools approach the self as a unified entity. In contrast, Dialogical Self Theory (DST) views the self as composed of multiple intern…

Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory

2026-07-07 · Jihao Liu, Guoxiong Gao, Zeming Sun, Bin Wu 외 arxiv

Recent LLM-based mathematical reasoning agents have begun to tackle research-level problems and, in several cases, have contributed to the resolution of open problems. However, scaling and orchestrating such agents effec…

Mathematical Reasoning

'Layer su Layer': Identifying and Disambiguating the Italian NPN Construction in BERT's family

2026-04-04 · Greta Gorzoni, Ludovica Pannitto, Francesca Masini arxiv

Interpretability research has highlighted the importance of evaluating Pretrained Language Models (PLMs) and in particular contextual embeddings against explicit linguistic theories to determine what linguistic informati…

Language Modelling

Orchestrating NLP Services for the Legal Domain

2020-03-28 · LREC 2020 5 · Julián Moreno-Schneider, Georg Rehm, Elena Montiel-Ponsoda, Víctor Rodriguez-Doncel 외

Legal technology is currently receiving a lot of attention from various angles. In this contribution we describe the main technical components of a system that is currently under development in the European innovation pr…

Analysis of AI Techniques for Orchestrating Edge-Cloud Application Migration

2025-07-14 · Sadig Gojayev, Ahmad Anaqreh, Carolina Fortuna arxiv

Application migration in edge-cloud system enables high QoS and cost effective service delivery. However, automatically orchestrating such migration is typically solved with heuristic approaches. Starting from the Markov…

Reinforcement Learning