paper-with-me

Papers

Progressive Multi-stage Interactive Training in Mobile Network for Fine-grained Recognition

2021-12-08 · Zhenxin Wu, Qingliang Chen, Yifeng Liu, Yinqi Zhang, Chengkai Zhu, Yang Yu

Fine-grained Visual Classification (FGVC) aims to identify objects from subcategories. It is a very challenging task because of the subtle inter-class differences. Existing research applies large-scale convolutional neural networks or visual transformers as the feature extractor, which is extremely computationally expensive. In fact, real-world scenarios of fine-grained recognition often require a more lightweight mobile network that can be utilized offline. However, the fundamental mobile network feature extraction capability is weaker than large-scale models. In this paper, based on the lightweight MobilenetV2, we propose a Progressive Multi-Stage Interactive training method with a Recursive Mosaic Generator (RMG-PMSI). First, we propose a Recursive Mosaic Generator (RMG) that generates images with different granularities in different phases. Then, the features of different stages pass through a Multi-Stage Interaction (MSI) module, which strengthens and complements the corresponding features of different stages. Finally, using the progressive training (P), the features extracted by the model in different stages can be fully utilized and fused with each other. Experiments on three prestigious fine-grained benchmarks show that RMG-PMSI can significantly improve the performance with good robustness and transferability.

📄 PDF Abstract BibTeX arXiv:2112.04223

Code (0)

등록된 구현이 없습니다.

Tasks

Fine-Grained Image Classification

Similar Papers 제목 키워드 기반

Mobile-Aptus: Confidence-Driven Proactive and Robust Interaction in MLLM-based Mobile-Using Agents

2026-05-27 · Zheng Wu, Pengzhou Cheng, Zongru Wu, Yuan Guo 외 arxiv

Recent advancements in multimodal large language models (MLLMs) have shown exceptional potential in enabling mobile-using agents to autonomously execute human instructions. However, fully automated agents often try to ex…

Semantic Similarity

Mobile-R1: Towards Interactive Reinforcement Learning for VLM-Based Mobile Agent via Task-Level Rewards

2025-06-25 · Jihao Gu, Qihang Ai, Yingyao Wang, Pi Bu 외

Vision-language model-based mobile agents have gained the ability to not only understand complex instructions and mobile screenshots, but also optimize their action outputs via thinking and reasoning, benefiting from rei…

reinforcement-learningReinforcement Learning

HumP-KD: A Hybrid Uncertainty-Aware Multi-Stage Progressive Knowledge Distillation Framework for Efficient Fire Classification

2026-06-12 · Mohammed Arif Mainuddin, Najifa Tabassum, Omar Ibne Shahid, Riasat Khan arxiv

Real-time fire classification systems require models that are simultaneously accurate, computationally efficient, and deployable on resource-constrained hardware. This work proposes \textbf{HumP-KD}, a Hybrid Uncertainty…

Knowledge Distillation

HCF: Hierarchical Cascade Framework for Distributed Multi-Stage Image Compression

2025-08-04 · Junhao Cai, Taegun An, Chengjun Jin, Sung Il Choi 외 arxiv

Distributed multi-stage image compression -- where visual content traverses multiple processing nodes under varying quality requirements -- poses challenges. Progressive methods enable bitstream truncation but underutili…

Computational EfficiencyImage Compression

CoME: Empowering Channel-of-Mobile-Experts with Informative Hybrid-Capabilities Reasoning

2026-02-27 · Yuxuan Liu, Weikai Xu, Kun Huang, Changyu Chen 외 arxiv

Mobile Agents can autonomously execute user instructions, which requires hybrid-capabilities reasoning, including screen summary, subtask planning, action decision and action function. However, existing agents struggle t…