QuTI! Quantifying Text-Image Consistency in Multimodal Documents
The World Wide Web and social media platforms have become popular sources for news and information. Typically, multimodal information, e.g., image and text is used to convey information more effectively and to attract attention. While in most cases image content is decorative or depicts additional information, it has also been leveraged to spread misinformation and rumors in recent years. In this paper, we present a Web-based demo application that automatically quantifies the cross-modal relations of entities (persons, locations, and events) in image and text. The applications are manifold. For example, the system can help users to explore multimodal articles more efficiently, or can assist human assessors and fact-checking efforts in the verification of the credibility of news stories, tweets, or other multimodal documents.
Code (1)
Tasks
ArticlesFact CheckingMisinformationSimilar Papers 제목 키워드 기반
QuTIE: Quantum optimization for Target Identification by Enzymes
Target Identification by Enzymes (TIE) problem aims to identify the set of enzymes in a given metabolic network, such that their inhibition eliminates a given set of target compounds associated with a disease while incur…
Image Realness Assessment and Localization with Multimodal Features
A reliable method of quantifying the perceptual realness of AI-generated images and identifying visually inconsistent regions is crucial for practical use of AI-generated images and for improving photorealism of generati…
Multimodal ML: Quantifying the Improvement of Calorie Estimation Through Image-Text Pairs
This paper determines the extent to which short textual inputs (in this case, names of dishes) can improve calorie estimation compared to an image-only baseline model and whether any improvements are statistically signif…
A Single Image and Multimodality Is All You Need for Novel View Synthesis
Diffusion-based approaches have recently demonstrated strong performance for single-image novel view synthesis by conditioning generative models on geometry inferred from monocular depth estimation. However, in practice,…
Monocular Depth EstimationNovel View SynthesisVideo GenerationBI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation
Multimodal Dialogue Response Generation (MDRG) is a recently proposed task where the model needs to generate responses in texts, images, or a blend of both based on the dialogue context. Due to the lack of a large-scale …
Response Generation