paper-with-me

홈 › Papers

QuTI! Quantifying Text-Image Consistency in Multimodal Documents

2021-04-28 · Matthias Springstein, Eric Müller-Budack, Ralph Ewerth

The World Wide Web and social media platforms have become popular sources for news and information. Typically, multimodal information, e.g., image and text is used to convey information more effectively and to attract attention. While in most cases image content is decorative or depicts additional information, it has also been leveraged to spread misinformation and rumors in recent years. In this paper, we present a Web-based demo application that automatically quantifies the cross-modal relations of entities (persons, locations, and events) in image and text. The applications are manifold. For example, the system can help users to explore multimodal articles more efficiently, or can assist human assessors and fact-checking efforts in the verification of the credibility of news stories, tweets, or other multimodal documents.

📄 PDF Abstract BibTeX arXiv:2104.13748

Code (1)

TIBHannover/cross-modal_entity_consistency 공식 구현 tf

Tasks

ArticlesFact CheckingMisinformation

Similar Papers 제목 키워드 기반

QuTIE: Quantum optimization for Target Identification by Enzymes

2023-03-13 · Hoang M. Ngo, My T. Thai, Tamer Kahveci

Target Identification by Enzymes (TIE) problem aims to identify the set of enzymes in a given metabolic network, such that their inhibition eliminates a given set of target compounds associated with a disease while incur…

Image Realness Assessment and Localization with Multimodal Features

2025-09-16 · Lovish Kaushik, Agnij Biswas, Somdyuti Paul arxiv

A reliable method of quantifying the perceptual realness of AI-generated images and identifying visually inconsistent regions is crucial for practical use of AI-generated images and for improving photorealism of generati…

Multimodal ML: Quantifying the Improvement of Calorie Estimation Through Image-Text Pairs

2025-11-12 · Arya Narang arxiv

This paper determines the extent to which short textual inputs (in this case, names of dishes) can improve calorie estimation compared to an image-only baseline model and whether any improvements are statistically signif…

A Single Image and Multimodality Is All You Need for Novel View Synthesis

2026-02-20 · Amirhosein Javadi, Chi-Shiang Gau, Konstantinos D. Polyzos, Tara Javidi arxiv

Diffusion-based approaches have recently demonstrated strong performance for single-image novel view synthesis by conditioning generative models on geometry inferred from monocular depth estimation. However, in practice,…

Monocular Depth EstimationNovel View SynthesisVideo Generation

BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation

2024-08-12 · Hee Suk Yoon, Eunseop Yoon, Joshua Tian Jin Tee, Kang Zhang 외

Multimodal Dialogue Response Generation (MDRG) is a recently proposed task where the model needs to generate responses in texts, images, or a blend of both based on the dialogue context. Due to the lack of a large-scale …

Response Generation