paper-with-me

홈 › Papers

From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options

2026-08-13 · Obed Junias, Maria Leonor Pacheco hf

Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they judge the individual atoms correctly. We study compound options connected by AND, OR, and NEITHER/NOR, introducing a framework that decomposes each option into atomic answers and scores contrastive hypotheses about each one, so the model never sees a compound option. An operator-constrained integer linear program then composes the calibrated scores into a single prediction. We evaluate on LOGICAL-COMMONSENSEQA and introduce LOGICAL-SATA, a reading-comprehension benchmark derived from SATA-Bench. Our framework improves Macro-F1 from 48.3 to 77.0 on the human-validated LOGICAL-COMMONSENSEQA split and from 47.0 to 75.6 on LOGICAL-SATA, with the largest gains on NEITHER/NOR.

📄 PDF Abstract BibTeX arXiv:2608.12836

Code (1)

obedjunias19/structured-compositional-reasoning

Similar Papers 제목 키워드 기반

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning

2026-06-29 · Eric Peh, Debaditya Roy, Basura Fernando arxiv

Vision-Language Models (VLMs) often achieve high performance on benchmarks while remaining "black boxes", yet they remain prone to hallucination or rely on superficial shortcuts. In this work, we propose a framework desi…

Reinforcement LearningVisual Reasoning

ATOM-Bench: A Real-World Benchmark for Atomic Skills and Compositional Generalization in Manipulation Policies

2026-06-15 · Zenan Wu, Bingqing Wei, Lu Liu, Zheqi He 외 arxiv

Generalist manipulation policies are increasingly presented as foundation models for robotic control, but their real-world generalization remains difficult to diagnose. A policy may succeed on demonstrated tasks while st…

CSyMR: Benchmarking Compositional Music Information Retrieval in Symbolic Music Reasoning

2025-12-16 · Boyang Wang, Yash Vishe, Xin Xu, Zachary Novack 외 arxiv

Natural language information needs over symbolic music scores rarely reduce to a single step lookup. Many queries require compositional Music Information Retrieval (MIR) that extracts multiple pieces of evidence from str…

Information Retrieval

State Space Models are Effective Sign Language Learners: Exploiting Phonological Compositionality for Vocabulary-Scale Recognition

2026-04-09 · Bryan Cheng, Austin Jin, Jasper Zhang arxiv

Sign language recognition suffers from catastrophic scaling failure: models achieving high accuracy on small vocabularies collapse at realistic sizes. Existing architectures treat signs as atomic visual patterns, learnin…

Sign Language RecognitionRepresentation Learning

Detection Accuracy for Evaluating Compositional Explanations of Units

2021-09-16 · Sayo M. Makinwa, Biagio La Rosa, Roberto Capobianco

The recent success of deep learning models in solving complex problems and in different domains has increased interest in understanding what they learn. Therefore, different approaches have been employed to explain these…