> For the complete documentation index, see [llms.txt](https://jgoodman8.gitbook.io/iron-data-science-notebook/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://jgoodman8.gitbook.io/iron-data-science-notebook/ml-datascience/machine-learning-algorithms/supervised-learning/classification-algorithms/random-forest.md).

# Random Forest

{% hint style="info" %}
*Sources:*

* [*Ensemble methods: bagging, boosting and stacking (Joseph Rocca)*](https://towardsdatascience.com/ensemble-methods-bagging-boosting-and-stacking-c9214a10a205)
  {% endhint %}

## Overview

The **random forest** approach is a [**bagging**](/iron-data-science-notebook/ml-datascience/ml-techniques/ensemble-methods.md#bagging) **method** where **deep trees**, fitted on **bootstrap samples**, are combined to produce an **output with lower variance**.

Additionally, RF uses another trick to make the multiple fitted trees a bit less correlated with each other: when growing each tree, instead of only sampling over the observations in the dataset to generate a bootstrap sample, we also **sample over features** and keep only a **random subset** of them to build the tree.

{% hint style="success" %}
**Bagging + Feature sampling = Lower Variance Error**
{% endhint %}

This way, all trees do not look at the exact same information to make their decisions and it **reduces the correlation** between the different returned outputs and generates a model more robust to missing data.

![](https://569842953-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LmjpNbCRLUyGiAxD8kn%2F-MggjnVZmNFIrt8ruCOw%2F-Mggk1PwmEvw4wzvJJ7K%2Fimage.png?alt=media\&token=cb6a27d9-de66-4f75-9d55-0a90626ccbd8)
