paper-with-me

홈 › Papers

TraVLR: Now You See It, Now You Don't! A Bimodal Dataset for Evaluating Visio-Linguistic Reasoning

2021-11-21 · Keng Ji Chow, Samson Tan, Min-Yen Kan

Numerous visio-linguistic (V+L) representation learning methods have been developed, yet existing datasets do not adequately evaluate the extent to which they represent visual and linguistic concepts in a unified space. We propose several novel evaluation settings for V+L models, including cross-modal transfer. Furthermore, existing V+L benchmarks often report global accuracy scores on the entire dataset, making it difficult to pinpoint the specific reasoning tasks that models fail and succeed at. We present TraVLR, a synthetic dataset comprising four V+L reasoning tasks. TraVLR's synthetic nature allows us to constrain its training and testing distributions along task-relevant dimensions, enabling the evaluation of out-of-distribution generalisation. Each example in TraVLR redundantly encodes the scene in two modalities, allowing either to be dropped or added during training or testing without losing relevant information. We compare the performance of four state-of-the-art V+L models, finding that while they perform well on test examples from the same modality, they all fail at cross-modal transfer and have limited success accommodating the addition or deletion of one modality. We release TraVLR as an open challenge for the research community.

📄 PDF Abstract BibTeX arXiv:2111.10756

Code (1)

kengjichow/travlr 공식 구현

Tasks

Representation Learning

Methods 이 논문이 사용한 방법론

Best Ways to Get in Touch with Expedia When You’re Dealing with Travel Issues or Booking Errors 설명 없음
30 Ways Contact Expedia Customer Service Using Phone, Email, Chat, and App-Based Support Options 설명 없음

Similar Papers 제목 키워드 기반

Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute Recognition

2025-01-01 · CVPR 2025 1 · Junyi Wu, Yan Huang, Min Gao, Yuzhen Niu 외

Pedestrian attribute recognition (PAR) seeks to predict multiple semantic attributes associated with a specific pedestrian. There are two types of approaches for PAR: unimodal framework and bimodal framework. The for…

AttributePedestrian Attribute RecognitionPerson Re-IdentificationPrompt Learning

Training Vision-Language Models with Less Bimodal Supervision

2022-11-01 · Elad Segal, Ben Bogin, Jonathan Berant

Standard practice in pretraining multimodal models, such as vision-language models, is to rely on pairs of aligned inputs from both modalities, for example, aligned image-text pairs. However, such pairs can be difficult …

Language Modelling

Evaluating How Fine-tuning on Bimodal Data Effects Code Generation

2022-11-15 · Gabriel Orlanski, Seonhye Yang, Michael Healy

Despite the increase in popularity of language models for code generation, it is still unknown how training on bimodal coding forums affects a model's code generation performance and reliability. We, therefore, collect a…

Code GenerationHumanEval

BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models

2025-11-24 · Juncheng Li, Yige Li, Hanxun Huang, Yunhao Chen 외 arxiv

Backdoor attacks undermine the reliability and trustworthiness of machine learning systems by injecting hidden behaviors that can be maliciously activated at inference time. While such threats have been extensively studi…

Visual Question AnsweringImage Captioning

Discriminative Bimodal Networks for Visual Localization and Detection with Natural Language Queries

2017-04-12 · CVPR 2017 7 · Yuting Zhang, Luyao Yuan, Yijie Guo, Zhiyuan He 외

Associating image regions with text queries has been recently explored as a new way to bridge visual and linguistic representations. A few pioneering approaches have been proposed based on recurrent neural language model…

Natural Language QueriesVisual Localization