paper-with-me

홈 › Papers

Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images

2026-07-06 · Weikang Wang, Tobias Weißberg, Florian Bernard arxiv

While various works address reflective symmetry understanding in 3D data and images, pixel-level semantic left-right prediction of in-the-wild images remains challenging, due to certain difficulties including the lack of 3D information, occlusion, object pose variation, partiality, etc. In this work, we propose an unsupervised learning framework to tackle this challenge. Leveraging recent advances in vertex-wise semantic left-right understanding of 3D data, our unsupervised learning method jointly utilises 3D shape and image datasets to infer pixel-wise semantic left-right predictions in single-view images. In particular, we show that a medium-scale 3D shape dataset comprising mainly of human- and quadruped animal-like shapes, combined with diverse in-the-wild image data, are sufficient to achieve high-quality semantic left-right prediction in images, even for entirely unseen 3D object categories, such as cars or trains. Overall, our approach achieves superior performance in dense pixel-wise semantic left-right predictions on both rendered and in-the-wild image datasets when compared to existing state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2607.05006

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Semantic-Aware Generative Adversarial Nets for Unsupervised Domain Adaptation in Chest X-ray Segmentation

2018-06-02 · Cheng Chen, Qi Dou, Hao Chen, Pheng-Ann Heng

In spite of the compelling achievements that deep neural networks (DNNs) have made in medical image computing, these deep models often suffer from degraded performance when being applied to new test datasets with domain …

Domain AdaptationSegmentationTransfer LearningUnsupervised Domain Adaptation

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation

2026-06-01 · Muyi Bao, Yuxin Cai, Hang Xu, Zongtai Li 외 arxiv

Vision-language models (VLMs) have become a common foundation for vision-and-language navigation in continuous environments (VLN-CE). Yet most VLM-based methods cast navigation as low-level action prediction, an interfac…

Vision-Language Navigation

DIffSteISR: Harnessing Diffusion Prior for Superior Real-world Stereo Image Super-Resolution

2024-08-14 · Yuanbo Zhou, Xinlin Zhang, Wei Deng, Tao Wang 외

We introduce DiffSteISR, a pioneering framework for reconstructing real-world stereo images. DiffSteISR utilizes the powerful prior knowledge embedded in pre-trained text-to-image model to efficiently recover the lost te…

Image Super-ResolutionStereo Image Super-ResolutionSuper-ResolutionTAG

Integrating Categorical Semantics into Unsupervised Domain Translation

2020-10-03 · ICLR 2021 1 · Samuel Lavoie, Faruk Ahmed, Aaron Courville

While unsupervised domain translation (UDT) has seen a lot of success recently, we argue that mediating its translation via categorical semantic features could broaden its applicability. In particular, we demonstrate tha…

ObjectTranslation

Unsupervised Cross-Lingual Scaling of Political Texts

2017-04-01 · EACL 2017 4 · Goran Glava{\v{s}}, Federico Nanni, Simone Paolo Ponzetto

Political text scaling aims to linearly order parties and politicians across political dimensions (e.g., left-to-right ideology) based on textual content (e.g., politician speeches or party manifestos). Existing models s…