paper-with-me

홈 › Papers

Learning from Imperfect Text Guidance: Robust Long-Tail Visual Recognition with High-Noise Label

2026-04-25 · Mengke Li, Haiquan Ling, Yiqun Zhang, Yang Lu, Hui Huang arxiv

Real-world data often exhibit long-tailed distributions with numerous noisy labels, substantially degrading the performance of deep models. While prior research has made progress in addressing this combined challenge, it overlooks the severe label-image mismatch inherent to high-noise settings, thereby limiting their effectiveness. Given that observed labels, though mismatched with images, still retain category information, we propose employing auxiliary text information from labels to address label-image inconsistencies in long-tailed noisy data. Specifically, we leverage the intrinsic cross-modal alignment in pre-trained visual-language models to correct the label-image inconsistencies. This supervisory signal, referred to as Weak Teacher Supervision (WTS), is unaffected by label noise and data distribution biases, albeit exhibits limited accuracy. Therefore, the activation of WTS is determined by evaluating the discrepancy between text-predicted labels and observed labels. Extensive experiments demonstrate the superior performance of WTS across synthetic and real-world datasets, particularly under high-noise conditions. The source code is available at https://anonymous.4open.science/r/WTS-0F3C.

📄 PDF Abstract BibTeX arXiv:2604.23125

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-Level Conditioning by Pairing Localized Text and Sketch for Fashion Image Generation

2026-02-20 · Ziyue Liu, Davide Talon, Federico Girella, Zanxi Ruan 외 arxiv

Sketches offer designers a concise yet expressive medium for early-stage fashion ideation by specifying structure, silhouette, and spatial relationships, while textual descriptions complement sketches to convey material,…

Image Generation

Semantic-guided Fine-tuning of Foundation Model for Long-tailed Visual Recognition

2025-07-17 · Yufei Peng, Yonggang Zhang, Yiu-ming Cheung

The variance in class-wise sample sizes within long-tailed scenarios often results in degraded performance in less frequent classes. Fortunately, foundation models, pre-trained on vast open-world datasets, demonstrate st…

Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions

2026-07-30 · Junrui Zhang, Jiaqi Li, Yiran Wang, Liao Shen 외 arxiv

Monocular depth estimation (MDE) faces challenges with non-Lambertian surfaces and adverse weather conditions due to the visual ambiguities inherent in single-image limited information. Existing works address them in iso…

Monocular Depth EstimationImage Inpainting

Yuan: Yielding Unblemished Aesthetics Through A Unified Network for Visual Imperfections Removal in Generated Images

2025-01-15 · Zhenyu Yu, Chee Seng Chan

Generative AI presents transformative potential across various domains, from creative arts to scientific visualization. However, the utility of AI-generated imagery is often compromised by visual flaws, including anatomi…

Image Generation

HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models

2025-05-10 · Shuhan Zhuang, Mengqi Huang, Fengyi Fu, Nan Chen 외

Visual text rendering, which aims to accurately integrate specified textual content within generated images, is critical for various applications such as commercial design. Despite recent advances, current methods strugg…

Text Generation