paper-with-me

홈 › Papers

Image understanding and the web

2020-04-29 · Fariza Fauzi, Mohammed Belkhatir

The contextual information of Web images is investigated to address the issue of characterizing their content with semantic descriptors and therefore bridge the semantic gap, i.e. the gap between their automated low-level representation in terms of colors, textures, shapes. . . and their semantic interpretation. Such characterization allows for understanding the image content and is crucial in important Web-based tasks such as image indexing and retrieval. Although we are highly motivated by the availability of rich knowledge on the Web and the relative success achieved by commercial search engines in automatically characterizing the image content using contextual information in Web pages, we are aware that the unpredictable quality of the contextual information is a major limiting factor. Among the reasons explaining the difficulty to leverage on the image contextual information, some problems are related to the characterization and extraction of this information. Indeed, the first issue is the lack of large-scale studies to highlight what is considered the relevant contextual information of an image, where it is located in a Web page and whether it is consistent across Web pages of different types, content layouts and domains. Also, the matter related to the extraction of this contextual information is topical as state-of-the-art automated extraction tools are unable to handle the heterogeneous Web. As far as the processing of the contextual information is concerned, problems linked to the syntactic and semantic characterizations of the textual components are important to address in order to tackle the semantic gap. Furthermore, questions pertaining to the organization of these textual components into coherent structures that are usable in image indexing and retrieval frameworks shall arise.

📄 PDF Abstract BibTeX arXiv:2005.02127

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval

Similar Papers 제목 키워드 기반

Understanding Non-optical Remote-sensed Images: Needs, Challenges and Ways Forward

2016-12-23 · Amit Kumar Mishra

Non-optical remote-sensed images are going to be used more often in man- aging disaster, crime and precision agriculture. With more small satellites and unmanned air vehicles planning to carry radar and hyperspectral ima…

UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

2025-12-25 · Shuo Cao, Jiayang Li, Xiaohui Li, Yuandong Pu 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image f…

Visual Question AnsweringText-to-Image GenerationVisual Grounding

Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation

2025-09-23 · Yuanhuiyi Lyu, Chi Kit Wong, Chenfei Liao, Lutao Jiang 외 arxiv

Recent works have made notable advancements in enhancing unified models for text-to-image generation through the Chain-of-Thought (CoT). However, these reasoning methods separate the processes of understanding and genera…

Text-to-Image GenerationImage Editing

Multi-modal Visual Understanding with Prompts for Semantic Information Disentanglement of Image

2023-05-16 · Yuzhou Peng

Multi-modal visual understanding of images with prompts involves using various visual and textual cues to enhance the semantic understanding of images. This approach combines both vision and language processing to genera…

Disentanglement

Interpretable Visual Understanding with Cognitive Attention Network

2021-08-06 · Xuejiao Tang, Wenbin Zhang, Yi Yu, Kea Turner 외

While image understanding on recognition-level has achieved remarkable advancements, reliable visual scene understanding requires comprehensive image understanding on recognition-level but also cognition-level, which cal…

Scene UnderstandingVisual Commonsense Reasoning