paper-with-me

Papers Scene Recognition

“Scene Recognition” 태그가 달린 논문 222편 · 필터 해제

A Comparative Study of Label-free Representation Quality Metrics in Deep Learning

2026-08-24 · Daniel Richards Arputharaj, Daniel Jönsson, Gabriel Eilertsen arxiv

We present a comparative study of label-free metrics for assessing the quality of representations in deep neural networks to understand their reliability under a wide variety of configurations. We group existing label-fr…

Scene Recognition

Rethinking Text-to-Image as Semantic-Aware Data Augmentation for Indoor Scene Recognition

2026-06-17 · Trong-Vu Hoang, Quang-Binh Nguyen, Dinh-Khoi Vo, Hoai-Danh Vo 외 arxiv

In the realm of computer vision, indoor image recognition presents challenges due to the intricate interplay of lighting conditions, occlusions, and diverse object arrangements within confined spaces. To address the lack…

Data AugmentationScene Recognition

PairWise Image Finder: An Open-source Tool for Finding Visually Aligned Street-Level Image Pairs for Urban Perception Studies

2026-06-07 · Jussi Torkko arxiv

Change detection and scene recognition techniques have been widely applied to Street View Imagery (SVI) to understand changes in scenes across the years. However, metadata alone is often insufficient to reliably find vis…

Semantic SegmentationScene RecognitionChange Detection

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison

2026-05-19 · Tianle Li, Xuyang Shen, Yan Ma, Rongxin Guo 외 arxiv

Long-form image captioning exposes a reward granularity problem in RL: captions are judged as whole sequences, while the important errors occur at the level of individual visual claims. A good dense caption should be bot…

Reinforcement LearningScene RecognitionImage CaptioningObject Counting

Beyond Logit Adjustment: A Residual Decomposition Framework for Long-Tailed Reranking

2026-04-02 · Zhanliang Wang, Hongzhuo Chen, Quan Minh Nguyen, Mian Umair Ahsan 외 arxiv

Long-tailed classification, where a small number of frequent classes dominate many rare ones, remains challenging because models systematically favor frequent classes at inference time. Existing post-hoc methods such as …

Image ClassificationScene Recognition

Dynamic Graph Neural Network with Adaptive Features Selection for RGB-D Based Indoor Scene Recognition

2026-04-01 · Qiong Liu, Ruofei Xiong, Xingzhen Chen, Muyao Peng 외 arxiv

Multi-modality of color and depth, i.e., RGB-D, is of great importance in recent research of indoor scene recognition. In this kind of data representation, depth map is able to describe the 3D structure of scenes and geo…

Graph Neural NetworkScene Recognition

How Class Ontology and Data Scale Affect Audio Transfer Learning

2026-03-26 · Manuel Milling, Andreas Triantafyllopoulos, Alexander Gebhard, Simon Rampp 외 arxiv

Transfer learning is a crucial concept within deep learning that allows artificial neural networks to benefit from a large pre-training data basis when confronted with a task of limited data. Despite its ubiquitous use a…

Activity RecognitionTransfer LearningScene Recognition

Real Eyes Realize Faster: Gaze Stability and Pupil Novelty for Efficient Egocentric Learning

2026-03-04 · Ajan Subramanian, Sumukh Bettadapura, Rohan Sathish arxiv

Always-on egocentric cameras are increasingly used as demonstrations for embodied robotics, imitation learning, and assistive AR, but the resulting video streams are dominated by redundant and low-quality frames. Under t…

Activity RecognitionScene Recognition

A Case Study on Concept Induction for Neuron-Level Interpretability in CNN

2026-02-27 · Moumita Sen Sarma, Samatha Ereshi Akkamahadevi, Pascal Hitzler arxiv

Deep Neural Networks (DNNs) have advanced applications in domains such as healthcare, autonomous systems, and scene understanding, yet the internal semantics of their hidden neurons remain poorly understood. Prior work i…

Scene UnderstandingScene Recognition

LLM-Driven Scenario-Aware Planning for Autonomous Driving

2026-01-29 · He Li, Zhaowei Chen, Rui Gao, Guoliang Li 외 arxiv

Hybrid planner switching framework (HPSF) for autonomous driving needs to reconcile high-speed driving efficiency with safe maneuvering in dense traffic. Existing HPSF methods often fail to make reliable mode transitions…

Scene UnderstandingAutonomous DrivingScene RecognitionMotion Planning

TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D Alignment

2026-01-27 · Jiarun Liu, Qifeng Chen, Yiru Zhao, Minghua Liu 외 arxiv

While visual-language models have profoundly linked features between texts and images, the incorporation of 3D modality data, such as point clouds and 3D Gaussians, further enables pretraining for 3D-related tasks, e.g.,…

Cross-Modal RetrievalScene RecognitionPoint Clouds

Hierarchical Fusion of Local and Global Visual Features with Mixture-of-Experts for Remote Sensing Image Scene Classification

2025-10-31 · Yuanhao Tang, Xuechao Zou, Zhengpei Hu, Junliang Xing 외 arxiv

Remote sensing image scene classification remains a challenging task, primarily due to the complex spatial structures and multi-scale characteristics of ground objects. Although CNN-based methods excel at extracting loca…

Scene ClassificationScene Recognition

A Framework for Low-Effort Training Data Generation for Urban Semantic Segmentation

2025-10-13 · Denis Zavadski, Damjan Kalšan, Tim Küchler, Haebom Lee 외 arxiv

Synthetic datasets are widely used for training urban scene recognition models, but even highly realistic renderings show a noticeable gap to real imagery. This gap is particularly pronounced when adapting to a specific …

Semantic SegmentationScene UnderstandingScene Recognition

Enhancing Self-Driving Segmentation in Adverse Weather Conditions: A Dual Uncertainty-Aware Training Approach to SAM Optimization

2025-09-05 · Dharsan Ravindran, Kevin Wang, Zhuoyuan Cao, Saleh Abdelrahman 외 arxiv

Recent advances in vision foundation models, such as the Segment Anything Model (SAM) and its successor SAM2, have achieved state-of-the-art performance on general image segmentation benchmarks. However, these models str…

Medical Image SegmentationAutonomous DrivingScene Recognition

LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving

2025-08-17 · Nan Song, Bozhou Zhang, Xiatian Zhu, Jiankang Deng 외 arxiv

Large vision-language models (VLMs) have shown promising capabilities in scene understanding, enhancing the explainability of driving behaviors and interactivity with users. Existing methods primarily fine-tune VLMs on o…

Scene UnderstandingAutonomous DrivingScene Recognition

Open-Vocabulary Semantic Segmentation with Uncertainty Alignment for Robotic Scene Understanding in Indoor Building Environments

2025-03-29 · Yifan Xu, Vineet Kamat, Carol Menassa

The global rise in the number of people with physical disabilities, in part due to improvements in post-trauma survivorship and longevity, has amplified the demand for advanced assistive technologies to improve mobility …

NavigateOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationScene Classification+4

Lightweight Multimodal Artificial Intelligence Framework for Maritime Multi-Scene Recognition

2025-03-10 · Xinyu Xi, Hua Yang, Shentai Zhang, Yijie Liu 외

Maritime Multi-Scene Recognition is crucial for enhancing the capabilities of intelligent marine robotics, particularly in applications such as marine conservation, environmental monitoring, and disaster response. Howeve…

Disaster ResponseLarge Language ModelMultimodal Large Language ModelQuantization+1

Contrastive Visual Data Augmentation

2025-02-24 · Yu Zhou, Bingxuan Li, Mohan Tang, Xiaomeng Jin 외

Large multimodal models (LMMs) often struggle to recognize novel concepts, as they rely on pre-trained knowledge and have limited ability to capture subtle visual details. Domain-specific knowledge gaps in training also …

Data AugmentationNovel ConceptsScene Recognition

Advancing ALS Applications with Large-Scale Pre-training: Dataset Development and Downstream Assessment

2025-01-09 · Haoyi Xiu, Xin Liu, TaeHoon Kim, Kyoung-Sook Kim

The pre-training and fine-tuning paradigm has revolutionized satellite remote sensing applications. However, this approach remains largely underexplored for airborne laser scanning (ALS), an important technology for appl…

Scene RecognitionSelf-Supervised LearningSemantic Segmentation

Seeing with Partial Certainty: Conformal Prediction for Robotic Scene Recognition in Built Environments

2025-01-09 · Yifan Xu, Vineet Kamat, Carol Menassa

In assistive robotics serving people with disabilities (PWD), accurate place recognition in built environments is crucial to ensure that robots navigate and interact safely within diverse indoor spaces. Language interfac…

Conformal PredictionHallucinationNavigateScene Recognition
1–20 / 222 다음 →