paper-with-me

홈 › Papers

Let's Roll a BiFTA: Bi-refinement for Fine-grained Text-visual Alignment in Vision-Language Models

2026-01-28 · Yuhao Sun, Chengyi Cai, Jiacheng Zhang, Zesheng Ye, Xingliang Yuan, Feng Liu arxiv

Recent research has shown that aligning fine-grained text descriptions with localized image patches can significantly improve the zero-shot performance of pre-trained vision-language models (e.g., CLIP). However, we find that both fine-grained text descriptions and localized image patches often contain redundant information, making text-visual alignment less effective. In this paper, we tackle this issue from two perspectives: \emph{View Refinement} and \emph{Description refinement}, termed as \textit{\textbf{Bi}-refinement for \textbf{F}ine-grained \textbf{T}ext-visual \textbf{A}lignment} (BiFTA). \emph{View refinement} removes redundant image patches with high \emph{Intersection over Union} (IoU) ratios, resulting in more distinctive visual samples. \emph{Description refinement} removes redundant text descriptions with high pairwise cosine similarity, ensuring greater diversity in the remaining descriptions. BiFTA achieves superior zero-shot performance on 6 benchmark datasets for both ViT-based and ResNet-based CLIP, justifying the necessity to remove redundant information in visual-text alignment.

📄 PDF Abstract BibTeX arXiv:2601.20419

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation

2025-10-30 · Woojin Kim, Jaeyoung Do arxiv

While diffusion language models (DLMs) enable fine-grained refinement, their practical controllability remains fragile. We identify and formally characterize a central failure mode called update forgetting, in which unif…

Text Generation

VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning

2026-03-30 · Li-Heng Chen, Ke Cheng, Yahui Liu, Lei Shi 외 arxiv

Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse driving videos, while preserving the spatio…

Video Generation

KnobGen: Controlling the Sophistication of Artwork in Sketch-Based Diffusion Models

2024-10-02 · Pouyan Navard, Amin Karimi Monsefi, Mengxi Zhou, Wei-Lun Chao 외

Recent advances in diffusion models have significantly improved text-to-image (T2I) generation, but they often struggle to balance fine-grained precision with high-level control. Methods like ControlNet and T2I-Adapter e…

Image Generation

Fine-grained Cross-modal Fusion based Refinement for Text-to-Image Synthesis

2023-02-17 · Haoran Sun, Yang Wang, Haipeng Liu, Biao Qian

Text-to-image synthesis refers to generating visual-realistic and semantically consistent images from given textual descriptions. Previous approaches generate an initial low-resolution image and then refine it to be high…

Image Generation

VideoRepair: Improving Text-to-Video Generation via Misalignment Evaluation and Localized Refinement

2024-11-22 · Daeun Lee, Jaehong Yoon, Jaemin Cho, Mohit Bansal

Recent text-to-video (T2V) diffusion models have demonstrated impressive generation capabilities across various domains. However, these models often generate videos that have misalignments with text prompts, especially w…

Text-to-Video GenerationVideo AlignmentVideo Generation