Demystifying Deep Learning: Understanding the Inner Workings of Neural Network
DOI:
https://doi.org/10.60087/jklst.vol1.n1.p129关键词:
Deep Learning, Neural Networks, Artificial Neurons, Activation Functions摘要
Deep learning has emerged as a powerful tool in various domains, revolutionizing fields such as image recognition, natural language processing, and autonomous driving. Despite its widespread applications, the inner workings of neural networks often remain opaque to many practitioners and enthusiasts. This paper aims to demystify deep learning by providing a comprehensive overview of the underlying principles and mechanisms. Beginning with the fundamental building blocks of artificial neurons and activation functions, we delve into the architecture of deep neural networks, elucidating concepts such as feedforward and backpropagation. Additionally, we explore advanced topics including convolutional neural networks (CNNs) and recurrent neural networks (RNNs), shedding light on their applications and intricacies. By elucidating the core concepts and methodologies, this paper empowers readers to develop a deeper understanding of how neural networks operate, paving the way for more informed utilization and innovation in the realm of deep learning.
下载次数
参考
J. Bruna, S. Mallat, "Invariant scattering convolution networks," IEEE Trans. Pattern Anal. Mach. Intell. 35, 1872–1886 (2013).
A. B. Patel, T. Nguyen, R. Baraniuk, "A Probabilistic Framework for Deep Learning," NeurIPS, 2016.
N. Tishby, N. Zaslavsky, "Deep Learning and the Information Bottleneck Principle," IEEE-ITW, 2015.
V. Papyan, Y. Romano, M. Elad, "Convolutional neural networks analyzed via convolutional sparse coding," JMLR 18, 2887–2938 (2017).
J. Sulam, A. Aberdam, A. Beck, M. Elad, "On multi-layer basis pursuit, efficient algorithms and convolutional neural networks," IEEE Trans. Pattern Anal. Mach. Intell. 42, 1968–1980 (2020).
B. D. Haeffele, R. Vidal, "Global Optimality in Neural Network Training," CVPR, 2017.
P. Chaudhari, S. Soatto, "Stochastic Gradient Descent Performs Variational Inference, Converges to Limit Cycles for Deep Networks," IEEE-ITA, 2018.
H. Bölcskei, P. Grohs, G. Kutyniok, P. Petersen, "Optimal approximation with sparsely connected deep neural networks," SIMODS 1, 8–45 (2019).
R. Giryes, G. Sapiro, A. M. Bronstein, "Deep neural networks with random Gaussian weights: A universal classification strategy?" IEEE-TSP 64, 3444–3457 (2016).
V. Papyan, X. Y. Han, D. L. Donoho, "Prevalence of neural collapse during the terminal phase of deep learning training," Proc. Natl. Acad. Sci. U.S.A. 117, 24652–24663 (2020).
M. Belkin, D. Hsu, S. Ma, S. Mandal, "Reconciling modern machine-learning practice and the classical bias-variance trade-off," Proc. Natl. Acad. Sci. U.S.A. 116, 15849–15854 (2019).
S. Ma, R. Bassily, M. Belkin, "The Power of Interpolation: Understanding the Effectiveness of SGD in Modern Over-parametrized Learning," ICML, 2018.
P. Nakkiran et al., "Deep Double Descent: Where Bigger Models and More Data Hurt," ICLR, 2019.
M. Li, M. Soltanolkotabi, S. Oymak, "Gradient Descent with Early Stopping Is Provably Robust to Label Noise for Overparameterized Neural Networks," ICAIS, 2020.
C. Guo, G. Pleiss, Y. Sun, K. Q. Weinberger, "On Calibration of Modern Neural Networks," ICML, 2017.
##submission.downloads##
已出版
##submission.license##
##submission.copyrightStatement##
##submission.license.cc.by4.footer##©2024 All rights reserved by the respective authors and JKLST.



