paper-with-me

Papers

Non-Natural Image Understanding with Advancing Frequency-based Vision Encoders

2025-01-01 · CVPR 2025 1 · Wang Lin, Qingsong Wang, Yueying Feng, Shulei Wang, Tao Jin, Zhou Zhao, Fei Wu, Chang Yao, Jingyuan Chen

Large language models (LLMs) have significantly enhanced cross-modal understanding capabilities by integrating visual encoders with textual embeddings, giving rise to multimodal large language models (MLLMs). However, these models struggle with non-natural images such as geometric and charts, particularly in fields like education and finance. Despite efforts to collect datasets and fine-tune the MLLMs, the gap with natural image understanding is still evident, and the cost of collecting large and diverse non-natural image datasets is high. To address this, we analyzed the limitations of transformer-based vision encoders(ViT) within existing MLLMs from a frequency perspective. Studies have shown that ViT models are less effective at capturing high-frequency information, impairing their ability to capture elements like points, lines, and angles in non-natural images. In response, we introduced FM-ViT, a frequency-modulated vision encoder that utilizes Fourier decomposition to extract high and low frequency components from self-attention features and re-weight them during tuning to non-natural images. In addition, we combine the features of CNN models with FM-ViT and propose EDGE, an MLLM with enhanced graphical encoders tailored for understanding non-natural images. Extensive experiments have confirmed the effectiveness of our FM-ViT and EDGE in 4 types.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Adversarial Robustness of Vision Transformers

2021-03-29 · Rulin Shao, Zhouxing Shi, JinFeng Yi, Pin-Yu Chen 외

Following the success in advancing natural language processing and understanding, transformers are expected to bring revolutionary changes to computer vision. This work provides a comprehensive study on the robustness of…

Adversarial Robustness

Enhancing Image Retrieval : A Comprehensive Study on Photo Search using the CLIP Mode

2024-01-24 · Naresh Kumar Lahajal, Harini S

Photo search, the task of retrieving images based on textual queries, has witnessed significant advancements with the introduction of CLIP (Contrastive Language-Image Pretraining) model. CLIP leverages a vision-language …

Image RetrievalInformation RetrievalNatural Language QueriesNatural Language Understanding+2

A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others

2022-12-09 · CVPR 2023 1 · Zhiheng Li, Ivan Evtimov, Albert Gordo, Caner Hazirbas 외

Machine learning models have been found to learn shortcuts -- unintended decision rules that are unable to generalize -- undermining models' reliability. Previous works address this problem under the tenuous assumption t…

Domain GeneralizationImage ClassificationOut-of-Distribution Generalization

Natural Language Generation from Visual Sequences: Challenges and Future Directions

2025-02-18 · Aditya K Surikuchi, Raquel Fernández, Sandro Pezzelle

The ability to use natural language to talk about visual content is at the core of human intelligence and a crucial feature of any artificial intelligence system. Various studies have focused on generating text for singl…

Image to textText Generation

Language (Re)modelling: Towards Embodied Language Understanding

2020-05-01 · ACL 2020 6 · Ronen Tamari, Chen Shani, Tom Hope, Miriam R. L. Petruck 외

While natural language understanding (NLU) is advancing rapidly, today's technology differs from human-like language understanding in fundamental ways, notably in its inferior efficiency, interpretability, and generaliza…

Natural Language UnderstandingPosition