paper-with-me

Papers

Read, Grep, and Synthesize: Diagnosing Cross-Domain Seed Exposure for LLM Research Ideation

2026-05-12 · Yunju Choi, Min Song arxiv

The discovery of novel methodologies for emerging problems is a continuing cycle in ML, often driven by the migration of techniques across domains. Building on this observation, we ask whether current LLM ideation systems benefit from targeted cross-domain retrieval or simply from exposure to diverse mechanisms. We study this question through PaperGym, a three-stage pipeline: (1) tool-augmented seed extraction via read, grep, and bash over an isolated paper environment, (2) cross-domain seed retrieval via paraphrasing across seven ML domains, and (3) method synthesis from retrieved seeds, each scored by rubric-based judges. Tool-augmented extraction improves specificity, and paraphrase-based retrieval broadens domain coverage. In synthesis, cross-domain retrieval receives more pairwise novelty wins than no-retrieval and same-domain baselines, but shows no significant difference from a random diverse-seed control. These findings suggest LLM ideation systems benefit from diverse seed exposure, but do not yet reliably exploit the semantic reason particular seeds were retrieved. We release the seed library, rubric prompts, and run scripts at https://github.com/yunjoochoi/PaperGym

📄 PDF Abstract BibTeX arXiv:2605.11532

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents

2026-08-06 · Wuya Chen, Yihao yang, Yang Cao, Yue Lin arxiv

Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands age…

GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization

2026-02-14 · Juntong Wang, Libin Chen, Xiyuan Wang, Shijia Kang 외 arxiv

Repository-level bug localization-the task of identifying where code must be modified to fix a bug-is a critical software engineering challenge. Standard Large Language Modles (LLMs) are often unsuitable for this task du…

Information Retrieval

Is Grep All You Need? How Agent Harnesses Reshape Agentic Search

2026-05-14 · Sahil Sen, Akhil Kasturi, Elias Lumer, Anmol Gulati 외 arxiv

Recent advances in Large Language Model (LLM) agents have enabled complex agentic workflows where models autonomously retrieve information, call tools, and reason over large corpora to complete tasks on behalf of users. …

LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis

2026-05-26 · Bowen Qin arxiv

CI failure logs are large (median 5k lines, max 200k in this corpus) and noisy. Coding agents that try to debug them depend on an upstream tool to reduce the log to a manageable context, but the field has had no public e…

Careful Selection and Thoughtful Discarding: Graph Explicit Pooling Utilizing Discarded Nodes

2023-11-21 · Chuang Liu, Wenhang Yu, Kuang Gao, Xueqi Ma 외

Graph pooling has been increasingly recognized as crucial for Graph Neural Networks (GNNs) to facilitate hierarchical graph representation learning. Existing graph pooling methods commonly consist of two stages: selectin…

Graph Representation LearningRepresentation Learning