> For the complete documentation index, see [llms.txt](https://jgoodman8.gitbook.io/iron-data-science-notebook/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://jgoodman8.gitbook.io/iron-data-science-notebook/ml-datascience/frequent-questions/ridge-vs-lasso.md).

# Ridge vs Lasso

{% hint style="info" %}
*Credit:*

* [*Regularization in Machine Learning (Prashant Gupta)*](https://towardsdatascience.com/regularization-in-machine-learning-76441ddcf99a)
* [*L1 and L2 Regularization Methods* *(Anuja Nagpal)*](https://towardsdatascience.com/l1-and-l2-regularization-methods-ce25e7fc831c)
  {% endhint %}

## Overview

|              Lasso              |   Ridge  |    |                                               |
| :-----------------------------: | :------: | -- | --------------------------------------------- |
|                L1               |    L2    |    |                                               |
| $$loss+ \lambda \sum\_{j=1}^{p} | \beta\_j | $$ | $$loss + \lambda \sum\_{j=1}^{p} \beta\_j^2$$ |

{% content-ref url="/pages/-LpEnkXP2VyMzC73UEvL" %}
[Regularization](/iron-data-science-notebook/ml-datascience/ml-techniques/regularization.md)
{% endcontent-ref %}

## About the equations

Both regressions can be thought of as solving an equation, where the summation of the regularized coefficients is less or equal to *s*. Where *s* is a constant that exists for each value of shrinkage factor *λ.*

* **Ridge**: $$\beta\_1^2 + \beta\_2^2 \le s$$.  This implies that coefficients have the smallest RSS for all points that lie within the **circle** given by the inequation.
* **Lasso**: $$|\beta\_1| + |\beta\_2| \le s$$. This implies that lasso coefficients have the smallest RSS for all points that lie within the **diamond** given by the inequation.

![Source: An Introduction to Statistical Learning by Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani](https://569842953-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LmjpNbCRLUyGiAxD8kn%2F-LpOQEwmJCpJk3RPM3ae%2F-LpOfBv-RoeMTZQi7AmN%2Fimage.png?alt=media\&token=3000b2ae-ea3f-4fc9-999e-b60948844c4f)

## Conclusions

&#x20;The **key difference** between these techniques is that **Lasso shrinks the less important feature’s coefficient to zero** thus, removing some features altogether. So, this works well for **feature selection** in case we have a **huge number of features**.

&#x20;This sheds light on the obvious **disadvantage of ridge regression**, which is **model interpretability.** Given, it will shrink **the coefficients for least important predictors, very close to zero**. But it will never make them exactly zero. In other words, the final model will include all predictors. However, in the case of the lasso, the **L1 penalty** has the eﬀect of forcing some of the **coeﬃcient estimates to be exactly equal to zero** when the tuning parameter λ is suﬃciently large. Therefore, the lasso method also performs variable selection and is said to yield sparse models.
