paper-with-me

홈 › Papers

Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection

2025-11-07 · Xian-Hong Huang, Hui-Kai Su, Chi-Chia Sun, Jun-Wei Hsieh arxiv

This paper introduces a cutting-edge approach to cross-modal interaction for tiny object detection by combining semantic-guided natural language processing with advanced visual recognition backbones. The proposed method integrates the BERT language model with the CNN-based Parallel Residual Bi-Fusion Feature Pyramid Network (PRB-FPN-Net), incorporating innovative backbone architectures such as ELAN, MSP, and CSP to optimize feature extraction and fusion. By employing lemmatization and fine-tuning techniques, the system aligns semantic cues from textual inputs with visual features, enhancing detection precision for small and complex objects. Experimental validation using the COCO and Objects365 datasets demonstrates that the model achieves superior performance. On the COCO2017 validation set, it attains a 52.6% average precision (AP), outperforming YOLO-World significantly while maintaining half the parameter consumption of Transformer-based models like GLIP. Several test on different of backbones such ELAN, MSP, and CSP further enable efficient handling of multi-scale objects, ensuring scalability and robustness in resource-constrained environments. This study underscores the potential of integrating natural language understanding with advanced backbone architectures, setting new benchmarks in object detection accuracy, efficiency, and adaptability to real-world challenges.

📄 PDF Abstract BibTeX arXiv:2511.05474

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language UnderstandingObject Detection

Similar Papers 제목 키워드 기반

Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation

2025-12-24 · Zebin Jiang, Tianle Jin, Xiangtong Yao, Alois Knoll 외 arxiv

Grasping is one of the most fundamental challenging capabilities in robotic manipulation, especially in unstructured, cluttered, and semantically diverse environments. Recent researches have increasingly explored languag…

Resolving Word Vagueness with Scenario-guided Adapter for Natural Language Inference

2024-05-21 · Yonghao Liu, Mengyu Li, Di Liang, Ximing Li 외

Natural Language Inference (NLI) is a crucial task in natural language processing that involves determining the relationship between two sentences, typically referred to as the premise and the hypothesis. However, tradit…

Natural Language InferenceSentenceSentence Fusion

Hierarchical Collaborative Fusion for 3D Instance-aware Referring Expression Segmentation

2026-03-06 · Keshen Zhou, Runnan Chen, Mingming Gong, Tongliang Liu arxiv

Generalised 3D Referring Expression Segmentation (3D-GRES) localizes objects in 3D scenes based on natural language, even when descriptions match multiple or zero targets. Existing methods rely solely on sparse point clo…

Referring Expression SegmentationPoint Clouds

Fuse & Calibrate: A bi-directional Vision-Language Guided Framework for Referring Image Segmentation

2024-05-18 · Yichen Yan, Xingjian He, Sihan Chen, Shichen Lu 외

Referring Image Segmentation (RIS) aims to segment an object described in natural language from an image, with the main challenge being a text-to-pixel correlation. Previous methods typically rely on single-modality feat…

DecoderImage SegmentationSemantic SegmentationSentence

Language Guided Networks for Cross-modal Moment Retrieval

2020-06-18 · Kun Liu, Huadong Ma, Chuang Gan

We address the challenging task of cross-modal moment retrieval, which aims to localize a temporal segment from an untrimmed video described by a natural language query. It poses great challenges over the proper semantic…

Moment RetrievalRetrievalSentenceSentence Embedding+1