> For the complete documentation index, see [llms.txt](https://jgoodman8.gitbook.io/iron-data-science-notebook/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://jgoodman8.gitbook.io/iron-data-science-notebook/ml-datascience/frequent-questions/covariance-vs-correlation-matrix.md).

# Covariance vs Correlation Matrix

{% hint style="info" %}
*Sources:*

* [*Baffled by Covariance and Correlation??? Get the Math and the Application in Analytics for both the terms...*](https://towardsdatascience.com/let-us-understand-the-correlation-matrix-and-covariance-matrix-d42e6b643c22) [*(Srishti Saha)*](https://towardsdatascience.com/@srishtisaha?source=post_page-----d42e6b643c22----------------------)
* [*Understanding the Covariance Matrix (Data Science Plus)*](https://datascienceplus.com/understanding-the-covariance-matrix/)
  {% endhint %}

## Overview

* ***Covariance*** $$\Rightarrow$$ **direction** of the linear relationship between variables.
* ***Correlation*** $$\Rightarrow$$ measure of the **strength and direction** of a linear relationship.

> **Correlation values are standardized** whereas, covariance values are not.

## Covariance Matrix

{% hint style="info" %}
Check [covariance definition](/iron-data-science-notebook/ml-datascience/statistics/the-basics.md#covariance).
{% endhint %}

Focusing on the two-dimensional case, the covariance matrix for two dimensions (or $$x$$ and $$y$$variables) is given by:

$$
C =
\begin{pmatrix}
\sigma(x,x)  & \sigma(x,y) \\
\sigma(y,x) & \sigma(y,y)
\end{pmatrix}
$$

{% tabs %}
{% tab title="Sample with Numpy" %}

```python
import numpy as np
import matplotlib.pyplot as plt

plt.style.use('ggplot')
plt.rcParams['figure.figsize'] = (12, 8)

mean = 0
std = 1
num_samples = 500

x = np.random.normal(mean, std, num_samples)
y = np.random.normal(mean, std, num_samples)
X = np.vstack((x, y)).T  # Join both arrays and transpose
# X = np.stack(arrays=[x, y], axis=1) # Equivalent transformation

plt.scatter(X[:, 0], X[:, 1])
plt.title('Generated Data')
plt.axis('equal');
```

![](https://569842953-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LmjpNbCRLUyGiAxD8kn%2F-LoeotXmD7Dz43kwAtVZ%2F-Lof-sStkQvRI4Z076XQ%2Fimage.png?alt=media\&token=3ae55426-520a-4f28-862d-a2aaf3d2c80b)
{% endtab %}

{% tab title="Sample with Tensorflow Probability" %}

```python
import tensorflow_probability as tfp
import matplotlib.pyplot as plt

tfd = tfp.distributions
data = tfd.MultivariateNormalFullCovariance(
      loc = [0., 5], # Mean for each variable ==> mean(a) = 0, mean(b) = 5
      covariance_matrix = [[1., .7], [.7, 1.]] # Covariance matrix
).sample(1000)

plt.scatter(data[:, 0], data[:, 1], color='blue', alpha=0.4)
plt.axis([-5, 5, 0, 10])
plt.title('Data set')
plt.show();
```

![](https://569842953-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LmjpNbCRLUyGiAxD8kn%2F-LoeotXmD7Dz43kwAtVZ%2F-Lof-ZnNtbJb-uv-Bo94%2Fimage.png?alt=media\&token=5390670a-7db0-4671-942e-76e9948fcaea)
{% endtab %}
{% endtabs %}

## Correlation Matrix

Unlike covariance, the correlation has an upper and lower cap on a range $$\[-1, 1]$$.

The correlation coefficient of two variables could be get by dividing the covariance of these variables by the product of the standard deviations of the same values.

$$
\rho\_{x,y} = corr(x,y) = \frac{\sigma\_{x,y}}{\sigma\_{x}^2\sigma\_{y}^2}
$$

```python
import pandas as pd

data = np.random.RandomState(seed=0)
correlation = pd.DataFrame(data.rand(10, 10)).corr()

correlation.style.background_gradient(cmap='coolwarm')
```

![](https://569842953-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LmjpNbCRLUyGiAxD8kn%2F-Lof2tcQlqyuUlMOt4RP%2F-Lof6JssJTUsSvTObypC%2Fimage.png?alt=media\&token=8cca8488-967f-4e0e-b210-9879ff359839)
