paper-with-me

Papers

LEMoN: Label Error Detection using Multimodal Neighbors

2024-07-10 · Haoran Zhang, Aparna Balagopalan, Nassim Oufattole, Hyewon Jeong, Yan Wu, Jiacheng Zhu, Marzyeh Ghassemi

Large repositories of image-caption pairs are essential for the development of vision-language models. However, these datasets are often extracted from noisy data scraped from the web, and contain many mislabeled instances. In order to improve the reliability of downstream models, it is important to identify and filter images with incorrect captions. However, beyond filtering based on image-caption embedding similarity, no prior works have proposed other methods to filter noisy multimodal data, or concretely assessed the impact of noisy captioning data on downstream training. In this work, we propose, theoretically justify, and empirically validate LEMoN, a method to identify label errors in image-caption datasets. Our method leverages the multimodal neighborhood of image-caption pairs in the latent space of contrastively pretrained multimodal models to automatically identify label errors. Through empirical evaluations across eight datasets and twelve baselines, we find that LEMoN outperforms the baselines by over 3% in label error detection, and that training on datasets filtered using our method improves downstream captioning performance by more than 2 BLEU points over noisy training.

📄 PDF Abstract BibTeX arXiv:2407.18941

Code (0)

등록된 구현이 없습니다.

Tasks

Label Error Detection

Similar Papers 제목 키워드 기반

Lemon and Orange Disease Classification using CNN-Extracted Features and Machine Learning Classifier

2024-08-26 · Khandoker Nosiba Arifin, Sayma Akter Rupa, Md Musfique Anwar, Israt Jahan

Lemons and oranges, both are the most economically significant citrus fruits globally. The production of lemons and oranges is severely affected due to diseases in its growth stages. Fruit quality has degraded due to the…

LEMON: Local Explanations via Modality-aware OptimizatioN

2026-02-02 · Yu Qin, Phillip Sloan, Raul Santos-Rodriguez, Majid Mirmehdi 외 arxiv

Multimodal models are ubiquitous, yet existing explainability methods are often single-modal, architecture-dependent, or too computationally expensive to run at scale. We introduce LEMON (Local Explanations via Modality-…

Question Answering

LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?

2026-01-27 · Zhuang Yu, Lei Shen, Jing Zhao, Shiliang Sun arxiv

Recent multimodal large language models (MLLMs) have shown remarkable progress across vision, audio, and language tasks, yet their performance on long-form, knowledge-intensive, and temporally structured educational cont…

Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding

2025-12-14 · Yongyuan Liang, Xiyao Wang, Yuanchen Ju, Jianwei Yang 외 arxiv

Scaling large multimodal models (LMMs) to 3D understanding poses unique challenges: point cloud data is sparse and irregular, existing models rely on fragmented architectures with modality-specific encoders, and training…

Object RecognitionSpatial Reasoning

Towards a new Ontology for Sign Languages

2022-06-01 · LREC 2022 6 · Thierry Declerck

We present the current status of a new ontology for representing constitutive elements of Sign Languages (SL). This development emerged from investigations on how to represent multimodal lexical data in the OntoLex-Lemon…