Next Article in Journal
Some New Refinements of Trapezium-Type Integral Inequalities in Connection with Generalized Fractional Integrals
Next Article in Special Issue
Some New Integral Inequalities Involving Fractional Operator with Applications to Probability Density Functions and Special Means
Previous Article in Journal
Homothetic Symmetries of Static Cylindrically Symmetric Spacetimes—A Rif Tree Approach
Previous Article in Special Issue
Almost Periodic Solution for Forced Perturbed Non-Instantaneous Impulsive Model
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Training Neural Networks by Time-Fractional Gradient Descent

School of Mathematics and Statistics, Guizhou University, Guiyang 550025, China
*
Author to whom correspondence should be addressed.
Axioms 2022, 11(10), 507; https://doi.org/10.3390/axioms11100507
Submission received: 1 September 2022 / Revised: 20 September 2022 / Accepted: 22 September 2022 / Published: 26 September 2022
(This article belongs to the Special Issue Impulsive, Delay and Fractional Order Systems)

Abstract

Motivated by the weighted averaging method for training neural networks, we study the time-fractional gradient descent (TFGD) method based on the time-fractional gradient flow and explore the influence of memory dependence on neural network training. The TFGD algorithm in this paper is studied via theoretical derivations and neural network training experiments. Compared with the common gradient descent (GD) algorithm, the optimization effect of the time-fractional gradient descent algorithm is significant when the value of fractional α is close to 1, under the condition of appropriate learning rate η. The comparison is extended to experiments on the MNIST dataset with various learning rates. It is verified that the TFGD has potential advantages when the fractional α nears 0.95∼0.99. This suggests that the memory dependence can improve training performance of neural networks.
Keywords: time-fractional gradient descent; training neural networks; weighted averaging; memory dependence time-fractional gradient descent; training neural networks; weighted averaging; memory dependence

Share and Cite

MDPI and ACS Style

Xie, J.; Li, S. Training Neural Networks by Time-Fractional Gradient Descent. Axioms 2022, 11, 507. https://doi.org/10.3390/axioms11100507

AMA Style

Xie J, Li S. Training Neural Networks by Time-Fractional Gradient Descent. Axioms. 2022; 11(10):507. https://doi.org/10.3390/axioms11100507

Chicago/Turabian Style

Xie, Jingyi, and Sirui Li. 2022. "Training Neural Networks by Time-Fractional Gradient Descent" Axioms 11, no. 10: 507. https://doi.org/10.3390/axioms11100507

APA Style

Xie, J., & Li, S. (2022). Training Neural Networks by Time-Fractional Gradient Descent. Axioms, 11(10), 507. https://doi.org/10.3390/axioms11100507

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop