paper-with-me

홈 › Papers

COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-Negatives

2026-06-26 · David Steinmann, Antonia Wüst, Kristian Kersting, Wolfgang Stammer arxiv

While interpretable models such as concept bottleneck models (CBMs) and program synthesis methods enable verification of model decisions, their evaluation is typically limited to simple tasks, leaving complex reasoning on real-world images largely unexplored. We introduce COCOLogic-V2, an object-centric dataset for visual inductive reasoning on real-world images covering a broad subset of first-order logic. By categorizing samples into positive variants, near-boundary (NB), and far-from-boundary (FB) negatives, COCOLogic-V2 enables fine-grained diagnosis of model accountability. Our evaluations show that models tend to separate positive and FB samples well but fail on NB samples, while perceptual noise and large rule-induced search spaces pose additional challenges in few-shot settings. Together, these results highlight that visual inductive reasoning remains an open challenge and COCOLogic-V2 provides a concrete foundation for advancing methods in this direction.

📄 PDF Abstract BibTeX arXiv:2606.28194

Code (0)

등록된 구현이 없습니다.

Tasks

Program Synthesis

Similar Papers 제목 키워드 기반

Defining Knowledge: Bridging Epistemology and Large Language Models

2024-10-03 · Constanza Fierro, Ruchira Dhar, Filippos Stamatiou, Nicolas Garneau 외

Knowledge claims are abundant in the literature on large language models (LLMs); but can we say that GPT-4 truly "knows" the Earth is round? To address this question, we review standard definitions of knowledge in episte…

Identifying and Handling Cross-Treebank Inconsistencies in UD: A Pilot Study

2020-12-01 · UDW (COLING) 2020 12 · Tillmann Dönicke, Xiang Yu, Jonas Kuhn

The Universal Dependencies treebanks are a still-growing collection of treebanks for a wide range of languages, all annotated with a common inventory of dependency relations. Yet, the usages of the relations can be categ…

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks

2026-04-20 · Taylor Lundy, Narun K. Raman, Kevin Leyton-Brown arxiv

LLM benchmarks are increasingly dynamic: instead of containing a fixed set of questions, they define templates and parameters that can generate an effectively unlimited number of question variants. This flexibility is va…

Detecting Lip-Syncing Deepfakes: Vision Temporal Transformer for Analyzing Mouth Inconsistencies

2025-04-02 · Soumyya Kanti Datta, Shan Jia, Siwei Lyu

Deepfakes are AI-generated media in which the original content is digitally altered to create convincing but manipulated images, videos, or audio. Among the various types of deepfakes, lip-syncing deepfakes are one of th…

Face Swapping

Tackling the Low-resource Challenge for Canonical Segmentation

2020-10-06 · EMNLP 2020 11 · Manuel Mager, Özlem Çetinoğlu, Katharina Kann

Canonical morphological segmentation consists of dividing words into their standardized morphemes. Here, we are interested in approaches for the task when training data is limited. We compare model performance in a simul…

Imitation LearningSegmentation