Gradient Descent is an optimization technique used in Machine Learning frameworks to train different models. The training process consists of an objective function (or the error function), which determines the error a Machine Learning model has on a given dataset.
While training, the parameters of this algorithm are initialized to random values. As the algorithm iterates, the parameters are updated such that we reach closer and closer to the optimal value of the function.
However, Adaptive Optimization Algorithms are gaining popularity due to their ability to converge swiftly. All these algorithms, in contrast to the conventional Gradient Descent, use statistics from the previous iterations to robustify the process of convergence.
An Adaptive Optimization Algorithm which uses exponentially weighted averages of gradients over previous iterations to stabilize the convergence, resulting in quicker optimization. For example, in most real-world applications of Deep Neural Networks, the training is carried out on noisy data. It is, therefore, necessary to reduce the effect of noise when the data are fed in batches during Optimization. This problem can be tackled using Exponentially Weighted Averages (or Exponentially Weighted Moving Averages).
Implementing Exponentially Weighted Averages:
In order to approximate the trends in a noisy dataset of size N:
, we maintain a set of parameters . As we iterate through all the values in the dataset, we calculate the parameters as below:
On iteration t: Get next
This algorithm averages the value of over its values from previous iterations. This averaging ensures that only the trend is retained and the noise is averaged out. This method is used as a strategy in momentum based gradient descent to make it robust against noise in data samples, resulting in faster training.
As an example, if you were to optimize a function on the parameter , the following pseudo code illustrates the algorithm:
On iteration t: On the current batch, compute
The HyperParameters for this Optimization Algorithm are , called the Learning Rate and, , similar to acceleration in mechanics.
Following is an implementation of Momentum-based Gradient Descent on a function :
- K means Clustering - Introduction
- Introduction To Machine Learning using Python
- Introduction to Dimensionality Reduction
- Artificial Intelligence | An Introduction
- An introduction to Machine Learning
- Introduction to Hill Climbing | Artificial Intelligence
- Decision Tree Introduction with example
- Introduction to Artificial Neutral Networks | Set 1
- Introduction to Artificial Neural Network | Set 2
- Pattern Recognition | Introduction
- ML | Introduction to Data in Machine Learning
- Data Cleansing | Introduction
- ML | Stochastic Gradient Descent (SGD)
- Introduction to Deep Learning
- Introduction to Stemming
If you like GeeksforGeeks and would like to contribute, you can also write an article using contribute.geeksforgeeks.org or mail your article to firstname.lastname@example.org. See your article appearing on the GeeksforGeeks main page and help other Geeks.
Please Improve this article if you find anything incorrect by clicking on the "Improve Article" button below.