paper-with-me

홈 › Papers

Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning

2025-03-25 · Shaun Baek, Shaun Esua-Mensah, Cyrus Tsui, Sejan Vigneswaralingam, Abdullah Alali, Michael Lu, Vasu Sharma, Sean O'Brien, Kevin Zhu

Large Language Models (LLMs) are primarily trained on high-resource natural languages, limiting their effectiveness in low-resource settings and in tasks requiring deep logical reasoning. This research introduces Rosetta-PL, a benchmark designed to evaluate LLMs' logical reasoning and generalization capabilities in a controlled environment. We construct Rosetta-PL by translating a dataset of logical propositions from Lean into a custom logical language, which is then used to fine-tune an LLM (e.g., GPT-4o). Our experiments analyze the impact of the size of the dataset and the translation methodology on the performance of the model. Our results indicate that preserving logical relationships in the translation process significantly boosts precision, with accuracy plateauing beyond roughly 20,000 training samples. These insights provide valuable guidelines for optimizing LLM training in formal reasoning tasks and improving performance in various low-resource language applications.

📄 PDF Abstract BibTeX arXiv:2505.00001

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelLogical ReasoningTranslation

Similar Papers 제목 키워드 기반

LogicPrpBank: A Corpus for Logical Implication and Equivalence

2024-02-14 · Zhexiong Liu, Jing Zhang, Jiaying Lu, Wenjing Ma 외

Logic reasoning has been critically needed in problem-solving and decision-making. Although Language Models (LMs) have demonstrated capabilities of handling multiple reasoning tasks (e.g., commonsense reasoning), their a…

Decision Making

From Rosetta to Match-Up: A Paired Corpus of Linguistic Puzzles with Human and LLM Benchmarks

2026-05-13 · Neh Majmudar, Anne Huang, Jinfan Frank Hu, Elena Filatova arxiv

In this paper, we examine linguistic puzzles used in high school linguistics competitions, focusing on two common formats: Rosetta Stone and Match-Up. We propose a systematic procedure for converting existing Rosetta Sto…

RosettaSearch: Multi-Objective Inference-Time Search for Protein Sequence Design

2026-04-19 · Meghana Kshirsagar, Allen Nie, Ching-An Cheng, Fanglei Xue 외 arxiv

We introduce RosettaSearch, an inference-time multi-objective optimization approach for backbone conditioned protein sequence design. We use large language models (LLMs) as a generative optimizer within a search algorith…

Protein Design with Agent Rosetta: A Case Study for Specialized Scientific Agents

2026-03-16 · Jacopo Teneggi, S. M. Bargeen A. Turzo, Tanya Marwah, Alberto Bietti 외 arxiv

Large language models (LLMs) are capable of emulating reasoning and using tools, creating opportunities for autonomous agents that execute complex scientific tasks. Protein design provides a natural testbed: although mac…

Prompt EngineeringProtein Design

The Rosetta Paradox: Domain-Specific Performance Inversions in Large Language Models

2024-12-09 · Basab Jha, Ujjwal Puri

While large language models, such as GPT and BERT, have already demonstrated unprecedented skills in everything from natural language processing to domain-specific applications, there came an unexplored phenomenon we ter…

Common Sense ReasoningSpecificity