paper-with-me

Papers

Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding

2026-02-03 · Byeongju Woo, Zilin Wang, Byeonghyun Pak, Sangwoo Mo, Stella X. Yu arxiv

Vision-language models such as CLIP often struggle to faithfully understand long, detail-rich captions, relying on dominant scene cues while overlooking fine-grained visual evidence. We propose a hierarchical vision-language learning principle for understanding scenes as part-to-whole compositions: before forming a whole-scene representation, a model should uncover what semantic parts appear where in the image. To this end, we propose CAFT (Cross-domain Alignment of Forests and Trees), a vision-language model that jointly learns local text-region alignment at intermediate representations and global image-text alignment at the final representation. Exploiting the organization of long captions, where local descriptions often correspond to scene parts, CAFT employs a fine-to-coarse image encoder and a part-whole text encoder to discover localized part semantics and progressively compose them into a global image-text representation. Trained on 30M image-text pairs, CAFT achieves state-of-the-art performance on six long-text retrieval benchmarks and exhibits strong scaling behavior. Experiments show that CAFT learns fine-grained representations that localize textual semantics in image regions without explicit region-level supervision.

📄 PDF Abstract BibTeX arXiv:2602.02977

Code (0)

등록된 구현이 없습니다.

Tasks

Text Retrieval

Similar Papers 제목 키워드 기반

Forest-Chat: Adapting Vision-Language Agents for Interactive Forest Change Analysis

2026-01-21 · James Brock, Ce Zhang, Nantheera Anantrasirichai arxiv

The increasing availability of high-resolution satellite imagery, together with advances in deep learning, creates new opportunities for forest monitoring workflows. Two central challenges in this domain are pixel-level …

Change DetectionObject Counting

MATE: Meet At The Embedding -- Connecting Images with Long Texts

2024-06-26 · Young Kyun Jang, Junmo Kang, Yong Jae Lee, Donghyun Kim

While advancements in Vision Language Models (VLMs) have significantly improved the alignment of visual and textual data, these models primarily focus on aligning images with short descriptive captions. This focus limits…

Cross-Modal RetrievalDescriptive

Fréchet random forests for metric space valued regression with non euclidean predictors

2019-06-04 · Louis Capitaine, Jérémie Bigot, Rodolphe Thiébaut, Robin Genuer

Random forests are a statistical learning method widely used in many areas of scientific research because of its ability to learn complex relationships between input and output variables and also its capacity to handle h…

regression

Convolutional Model Trees

2025-11-16 · William Ward Armstrong, Hongyi Li, Jun Xu arxiv

A method for creating a forest of model trees to fit samples of a function defined on images is described in several steps: down-sampling the images, determining a tree's hyperplanes, applying convolutions to the hyperpl…

Deep Learning based Automated Forest Health Diagnosis from Aerial Images

2020-10-16 · Chia-Yen Chiang, Chloe Barnes, Plamen Angelov, Richard Jiang

Global climate change has had a drastic impact on our environment. Previous study showed that pest disaster occured from global climate change may cause a tremendous number of trees died and they inevitably became a fact…

Deep LearningTransfer Learning