paper-with-me

홈 › Papers

Probing Cross-Modal Representations in Multi-Step Relational Reasoning

2021-08-01 · ACL (RepL4NLP) 2021 8 · Iuliia Parfenova, Desmond Elliott, Raquel Fernández, Sandro Pezzelle

We investigate the representations learned by vision and language models in tasks that require relational reasoning. Focusing on the problem of assessing the relative size of objects in abstract visual contexts, we analyse both one-step and two-step reasoning. For the latter, we construct a new dataset of three-image scenes and define a task that requires reasoning at the level of the individual images and across images in a scene. We probe the learned model representations using diagnostic classifiers. Our experiments show that pretrained multimodal transformer-based architectures can perform higher-level relational reasoning, and are able to learn representations for novel tasks and data that are very different from what was seen in pretraining.

📄 PDF Abstract BibTeX

Code (1)

jig-san/multi-step-size-reasoning 공식 구현 pytorch

Tasks

DiagnosticRelational Reasoning

Similar Papers 제목 키워드 기반

Disentangled Shared Representations Improve Morpho-Transcriptomic Integration

2026-08-14 · Julian Ostermaier, Swann Ruyter, Reuben Dorent, Daniel Racoceanu arxiv

Spatial transcriptomics (ST) enables the simultaneous profiling of gene expression and tissue morphology, creating an opportunity to learn multimodal representations capturing shared morpho-transcriptomic structure. Howe…

Representation Learning

Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis

2026-06-15 · Jingyu Hu, Giuseppe Tripodi, Reed Naidoo, Sarah F. McGough 외 arxiv

Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under distribution shift remains underexplored. This work systematically evaluates FM-ba…

Probing-RAG: Self-Probing to Guide Language Models in Selective Document Retrieval

2024-10-17 · Ingeol Baek, Hwan Chang, Byeongjeong Kim, JiMin Lee 외

Retrieval-Augmented Generation (RAG) enhances language models by retrieving and incorporating relevant external knowledge. However, traditional retrieve-and-generate processes may not be optimized for real-world scenario…

Decision MakingRAGRetrievalRetrieval-augmented Generation

Beyond Pattern Recognition: Probing Mental Representations of LMs

2025-02-23 · Moritz Miller, Kumar Shridhar

Language Models (LMs) have demonstrated impressive capabilities in solving complex reasoning tasks, particularly when prompted to generate intermediate explanations. However, it remains an open question whether these int…

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning

2026-06-19 · Awais Rauf, Ahmed Hasssan, Greg Slabaugh arxiv

Understanding long videos requires fine-grained perception and multi-step, higher-order reasoning over complex, long-range spatio-temporal dynamics. Vision-language models (VLMs) encode video frames into visual tokens an…

Relational ReasoningSemantic Retrieval