paper-with-me

홈 › Papers

Collaborative Image Understanding

2022-10-21 · Koby Bibas, Oren Sar Shalom, Dietmar Jannach

Automatically understanding the contents of an image is a highly relevant problem in practice. In e-commerce and social media settings, for example, a common problem is to automatically categorize user-provided pictures. Nowadays, a standard approach is to fine-tune pre-trained image models with application-specific data. Besides images, organizations however often also collect collaborative signals in the context of their application, in particular how users interacted with the provided online content, e.g., in forms of viewing, rating, or tagging. Such signals are commonly used for item recommendation, typically by deriving latent user and item representations from the data. In this work, we show that such collaborative information can be leveraged to improve the classification process of new images. Specifically, we propose a multitask learning framework, where the auxiliary task is to reconstruct collaborative latent item representations. A series of experiments on datasets from e-commerce and social media demonstrates that considering collaborative signals helps to significantly improve the performance of the main task of image classification by up to 9.1%.

📄 PDF Abstract BibTeX arXiv:2210.11907

Code (0)

등록된 구현이 없습니다.

Tasks

image-classificationImage Classification

Similar Papers 제목 키워드 기반

Face Completion with Semantic Knowledge and Collaborative Adversarial Learning

2018-12-08 · Haofu Liao, Gareth Funka-Lea, Yefeng Zheng, Jiebo Luo 외

Unlike a conventional background inpainting approach that infers a missing area from image patches similar to the background, face completion requires semantic knowledge about the target object for realistic outputs. Cur…

Facial InpaintingImage InpaintingSemantic Segmentation

AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning

2025-11-14 · Jirong Zha, Yuxuan Fan, Tianyu Zhang, Geng Chen 외 arxiv

Multimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap is critical, as multi-drone systems pro…

Scene Understanding

Collaboratively Weighting Deep and Classic Representation via L2 Regularization for Image Classification

2018-02-21 · Shaoning Zeng, Bob Zhang, Yanghao Zhang, Jianping Gou

Deep convolutional neural networks provide a powerful feature learning capability for image classification. The deep image features can be utilized to deal with many image understanding tasks like image classification an…

ClassificationGeneral Classificationimage-classificationImage Classification+2

CoVis: A Collaborative Framework for Fine-grained Graphic Visual Understanding

2024-11-27 · Xiaoyu Deng, Zhengjian Kang, Xintao Li, Yongzhe Zhang 외

Graphic visual content helps in promoting information communication and inspiration divergence. However, the interpretation of visual content currently relies mainly on humans' personal knowledge background, thereby affe…

Language ModelingLanguage ModellingLarge Language Model

Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely

2026-05-29 · Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen arxiv

Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected to communicate this understanding through language. Vision-language …

Instruction FollowingQuestion AnsweringSpatial Reasoning