paper-with-me

홈 › Papers

A Comprehensive Survey on Architectural Advances in Deep CNNs: Challenges, Applications, and Emerging Research Directions

2025-03-19 · Saddam Hussain Khan, Rashid Iqbal

Deep Convolutional Neural Networks (CNNs) have significantly advanced deep learning, driving breakthroughs in computer vision, natural language processing, medical diagnosis, object detection, and speech recognition. Architectural innovations including 1D, 2D, and 3D convolutional models, dilated and grouped convolutions, depthwise separable convolutions, and attention mechanisms address domain-specific challenges and enhance feature representation and computational efficiency. Structural refinements such as spatial-channel exploitation, multi-path design, and feature-map enhancement contribute to robust hierarchical feature extraction and improved generalization, particularly through transfer learning. Efficient preprocessing strategies, including Fourier transforms, structured transforms, low-precision computation, and weight compression, optimize inference speed and facilitate deployment in resource-constrained environments. This survey presents a unified taxonomy that classifies CNN architectures based on spatial exploitation, multi-path structures, depth, width, dimensionality expansion, channel boosting, and attention mechanisms. It systematically reviews CNN applications in face recognition, pose estimation, action recognition, text classification, statistical language modeling, disease diagnosis, radiological analysis, cryptocurrency sentiment prediction, 1D data processing, video analysis, and speech recognition. In addition to consolidating architectural advancements, the review highlights emerging learning paradigms such as few-shot, zero-shot, weakly supervised, federated learning frameworks and future research directions include hybrid CNN-transformer models, vision-language integration, generative learning, etc. This review provides a comprehensive perspective on CNN's evolution from 2015 to 2025, outlining key innovations, challenges, and opportunities.

📄 PDF Abstract BibTeX arXiv:2503.16546

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionComputational EfficiencyFace RecognitionFederated LearningLanguage ModelingLanguage ModellingMedical Diagnosisobject-detectionObject DetectionPose Estimationspeech-recognitionSpeech Recognitiontext-classificationText ClassificationTransfer Learning

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Transformers in Medical Imaging: A Survey

2022-01-24 · Fahad Shamshad, Salman Khan, Syed Waqas Zamir, Muhammad Haris Khan 외

Following unprecedented success on the natural language tasks, Transformers have been successfully applied to several computer vision problems, achieving state-of-the-art results and prompting researchers to reconsider t…

Image ClassificationImage SegmentationMedical Image DenoisingMedical Image Registration+5

A Comprehensive Survey of Convolutions in Deep Learning: Applications, Challenges, and Future Trends

2024-02-23 · Abolfazl Younesi, Mohsen Ansari, Mohammadamin Fazli, Alireza Ejlali 외

In today's digital age, Convolutional Neural Networks (CNNs), a subset of Deep Learning (DL), are widely used for various computer vision tasks such as image classification, object detection, and image segmentation. Ther…

6D Visionimage-classificationImage ClassificationImage Segmentation+6

Mamba in Vision: A Comprehensive Survey of Techniques and Applications

2024-10-04 · Md Maklachur Rahman, Abdullah Aman Tutul, Ankur Nath, Lamyanba Laishram 외

Mamba is emerging as a novel approach to overcome the challenges faced by Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) in computer vision. While CNNs excel at extracting local features, they often …

MambaState Space ModelsSurvey

A Comprehensive Survey of Time Series Forecasting: Architectural Diversity and Open Challenges

2024-10-24 · Jongseon Kim, Hyungjoon Kim, HyunGi Kim, Dongjun Lee 외

Time series forecasting is a critical task that provides key information for decision-making across various fields. Recently, various fundamental deep learning architectures such as MLPs, CNNs, RNNs, and GNNs have been d…

DiversityMambaTime SeriesTime Series Forecasting

A Survey on Deep Stereo Matching in the Twenties

2024-07-10 · Fabio Tosi, Luca Bartolomei, Matteo Poggi

Stereo matching is close to hitting a half-century of history, yet witnessed a rapid evolution in the last decade thanks to deep learning. While previous surveys in the late 2010s covered the first stage of this revoluti…

Stereo MatchingSurvey