paper-with-me

홈 › Papers

XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags

2024-06-06 · Faisal Tareque Shohan, Mir Tafseer Nayeem, Samsul Islam, Abu Ubaida Akash, Shafiq Joty

Millions of news articles published online daily can overwhelm readers. Headlines and entity (topic) tags are essential for guiding readers to decide if the content is worth their time. While headline generation has been extensively studied, tag generation remains largely unexplored, yet it offers readers better access to topics of interest. The need for conciseness in capturing readers' attention necessitates improved content selection strategies for identifying salient and relevant segments within lengthy articles, thereby guiding language models effectively. To address this, we propose to leverage auxiliary information such as images and captions embedded in the articles to retrieve relevant sentences and utilize instruction tuning with variations to generate both headlines and tags for news articles in a multilingual context. To make use of the auxiliary information, we have compiled a dataset named XL-HeadTags, which includes 20 languages across 6 diverse language families. Through extensive evaluation, we demonstrate the effectiveness of our plug-and-play multimodal-multilingual retrievers for both tasks. Additionally, we have developed a suite of tools for processing and evaluating multilingual texts, significantly contributing to the research community by enabling more accurate and efficient analysis across languages.

📄 PDF Abstract BibTeX arXiv:2406.03776

Code (1)

faisaltareque/XL-HeadTags 공식 구현

Tasks

ArticlesHeadline GenerationRetrievalTAG

Similar Papers 제목 키워드 기반

Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval

2024-10-02 · Kyle Buettner, Adriana Kovashka

There is a scarcity of multilingual vision-language models that properly account for the perceptual differences that are reflected in image captions across languages and cultures. In this work, through a multimodal, mult…

Image CaptioningRetrieval

CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models

2024-10-17 · Shangda Wu, Yashan Wang, Ruibin Yuan, Zhancheng Guo 외

Challenges in managing linguistic diversity and integrating various musical modalities are faced by current music information retrieval systems. These limitations reduce their effectiveness in a global, multimodal music …

Contrastive LearningDiversityInformation RetrievalMusic Classification+2

CAR-MFL: Cross-Modal Augmentation by Retrieval for Multimodal Federated Learning with Missing Modalities

2024-07-11 · Pranav Poudel, Prashant Shrestha, Sanskar Amgain, Yash Raj Shrestha 외

Multimodal AI has demonstrated superior performance over unimodal approaches by leveraging diverse data sources for more comprehensive analysis. However, applying this effectiveness in healthcare is challenging due to th…

Data AugmentationFederated LearningRetrieval

Multilingual-To-Multimodal (M2M): Unlocking New Languages with Monolingual Text

2026-01-15 · Piyush Singh Pasi arxiv

Multimodal models excel in English, supported by abundant image-text and audio-text data, but performance drops sharply for other languages due to limited multilingual multimodal resources. Existing solutions rely on mac…

Text-to-Image GenerationMachine TranslationImage RetrievalText Retrieval

A Multimodal Recaptioning Framework to Account for Perceptual Diversity in Multilingual Vision-Language Modeling

2025-04-19 · Kyle Buettner, Jacob Emmerson, Adriana Kovashka

There are many ways to describe, name, and group objects when captioning an image. Differences are evident when speakers come from diverse cultures due to the unique experiences that shape perception. Machine translation…

DiversityImage RetrievalLanguage ModelingLanguage Modelling+2