paper-with-me

홈 › Papers

DOLOMITES: Domain-Specific Long-Form Methodical Tasks

2024-05-09 · Chaitanya Malaviya, Priyanka Agrawal, Kuzman Ganchev, Pranesh Srinivasan, Fantine Huot, Jonathan Berant, Mark Yatskar, Dipanjan Das, Mirella Lapata, Chris Alberti

Experts in various fields routinely perform methodical writing tasks to plan, organize, and report their work. From a clinician writing a differential diagnosis for a patient, to a teacher writing a lesson plan for students, these tasks are pervasive, requiring to methodically generate structured long-form output for a given input. We develop a typology of methodical tasks structured in the form of a task objective, procedure, input, and output, and introduce DoLoMiTes, a novel benchmark with specifications for 519 such tasks elicited from hundreds of experts from across 25 fields. Our benchmark further contains specific instantiations of methodical tasks with concrete input and output examples (1,857 in total) which we obtain by collecting expert revisions of up to 10 model-generated examples of each task. We use these examples to evaluate contemporary language models highlighting that automating methodical tasks is a challenging long-form generation problem, as it requires performing complex inferences, while drawing upon the given context as well as domain knowledge.

📄 PDF Abstract BibTeX arXiv:2405.05938

Code (0)

등록된 구현이 없습니다.

Tasks

Form

Similar Papers 제목 키워드 기반

Wild herbivores in forests: four case studies

2015-12-12

A three population system with a top predator population, i.e. the herbivores, and two prey populations, grass and trees, is considered to model the interaction of herbivores with natural resources. We apply the model fo…

Comparative Analysis of Open-Source Language Models in Summarizing Medical Text Data

2024-05-25 · Yuhao Chen, Zhimu Wang, Bo Wen, Farhana Zulkernine

Unstructured text in medical notes and dialogues contains rich information. Recent advancements in Large Language Models (LLMs) have demonstrated superior performance in question answering and summarization tasks on unst…

Question Answering

Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

2025-02-22 · Patrick Tser Jern Kon, Jiachen Liu, Qiuyi Ding, Yiming Qiu 외

Scientific experimentation, a cornerstone of human progress, demands rigor in reliability, methodical control, and interpretability to yield meaningful results. Despite the growing capabilities of large language models (…

AI Agent

Detection of Bark Beetle Attacks using Hyperspectral PRISMA Data and Few-Shot Learning

2025-11-14 · Mattia Ferrari, Giancarlo Papitto, Giorgio Deligios, Lorenzo Bruzzone arxiv

Bark beetle infestations represent a serious challenge for maintaining the health of coniferous forests. This paper proposes a few-shot learning approach leveraging contrastive learning to detect bark beetle infestations…

Contrastive LearningFew-Shot Learning

Scientific Large Language Models: A Survey on Biological & Chemical Domains

2024-01-26 · Qiang Zhang, Keyang Ding, Tianwen Lyv, Xinda Wang 외

Large Language Models (LLMs) have emerged as a transformative power in enhancing natural language comprehension, representing a significant stride toward artificial general intelligence. The application of LLMs extends b…

scientific discoverySurvey