paper-with-me

Papers

Task-Guided Multi-Annotation Triplet Learning for Remote Sensing Representations

2026-04-04 · Meilun Zhou, Alina Zare arxiv

Prior multi-task triplet loss methods relied on static weights to balance supervision between various types of annotation. However, static weighting requires tuning and does not account for how tasks interact when shaping a shared representation. To address this, the proposed task-guided multi-annotation triplet loss removes this dependency by selecting triplets through a mutual-information criteria that identifies triplets most informative across tasks. This strategy modifies which samples influence the representation rather than adjusting loss magnitudes. Experiments on an aerial wildlife dataset compare the proposed task-guided selection against several triplet loss setups for shaping a representation in an effective multi-task manner. The results show improved classification and regression performance and demonstrate that task-aware triplet selection produces a more effective shared representation for downstream tasks.

📄 PDF Abstract BibTeX arXiv:2604.03837

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-Task Learning with Multi-Annotation Triplet Loss for Improved Object Detection

2025-04-10 · Meilun Zhou, Aditya Dutt, Alina Zare

Triplet loss traditionally relies only on class labels and does not use all available information in multi-task scenarios where multiple types of annotations are available. This paper introduces a Multi-Annotation Triple…

Multi-Task Learningobject-detectionObject DetectionTriplet

ReCon1M:A Large-scale Benchmark Dataset for Relation Comprehension in Remote Sensing Imagery

2024-06-10 · Xian Sun, Qiwei Yan, Chubo Deng, Chenglong Liu 외

Scene Graph Generation (SGG) is a high-level visual understanding and reasoning task aimed at extracting entities (such as objects) and their interrelationships from images. Significant progress has been made in the stud…

Graph Generationobject-detectionObject DetectionRelation+1

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

2025-07-16 · Yuchi Ishikawa, Shota Nakada, Hokuto Munakata, Kazuhiro Saito 외

In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to improve audio-visual representation learning. LG-CAV-MAE integrates a pretrained text encoder into contrastive audio-…

Image CaptioningRepresentation LearningRetrieval

Geospatial-Temporal Sensemaking of Remote Sensing Activity Detections with Multimodal Large Language Model

2026-05-11 · David F. Ramirez, Tim Overman, Kristen Jaskie, Andreas Spanias arxiv

We introduce SMART-HC-VQA, a Sentinel-2-based visual question answering dataset derived from the IARPA SMART Heavy Construction dataset, designed for spatiotemporal analysis of human activity. The dataset transforms cons…

Visual Question Answering

AU-Guided Unsupervised Domain Adaptive Facial Expression Recognition

2020-12-18 · Kai Wang, Yuxin Gu, Xiaojiang Peng, Panpan Zhang 외

The domain diversities including inconsistent annotation and varied image collection conditions inevitably exist among different facial expression recognition (FER) datasets, which pose an evident challenge for adapting …

Facial Expression RecognitionFacial Expression Recognition (FER)Triplet