> For the complete documentation index, see [llms.txt](https://jgoodman8.gitbook.io/iron-data-science-notebook/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://jgoodman8.gitbook.io/iron-data-science-notebook/ml-datascience/computer-vision/object-detection/two-stage-detectors/faster-r-cnn.md).

# Faster R-CNN

Main contributions with respect to [Fast R-CNN](/iron-data-science-notebook/ml-datascience/computer-vision/object-detection/two-stage-detectors/fast-r-cnn.md),

1. **RPN or Region Proposal Network**: NN to generate proposals for different anchor boxes.
2. Uses **anchor boxes** (combination of scale and aspect ratio)
3. CNN weights are shared by the RPN and the Fast R-CNN.

The Faster R-CNN works as follows:

* The RPN generates region proposals. These regions are generated using a trainable network, that could be customized for each detection scenario. And uses the same convolutional layers used in the Fast R-CNN network.
* For all region proposals in the image, a fixed-length feature vector is extracted from each region using the [ROI Pooling](/iron-data-science-notebook/ml-datascience/computer-vision/object-detection/techniques/roi-pooling.md)layer.
* The extracted feature vectors are then classified using the Fast R-CNN.
* The class scores of the detected objects in addition to their bounding-boxes are returned.

<figure><img src="https://569842953-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-LmjpNbCRLUyGiAxD8kn%2Fuploads%2FquA3RkDwYn0SGWh9KAX6%2Fimage.png?alt=media&amp;token=aea89d4b-914f-4bc5-9ef2-31b36d9365d6" alt=""><figcaption></figcaption></figure>

### How are anchor boxes created?

The RPN works on the output feature map returned from the last convolutional layer shared with the Fast R-CNN. A sliding window of size *NxN* passes through the feature map. For each window, several candidate region proposals are generated. These proposals are not the final proposals as they will be filtered based on their “objectness score”.

Normally, they are defined by two parameters. Combining these parameters we can obtain *K* number of anchor boxes:

* Scale
* Aspect Ratio

For each window, *K* proposals are generated and a feature vector of equal size is extracted. Then, it's fed into two FC layers:

* The *cls* layer generates the **objectness score** for every region proposal as a binary classifier (background vs object).
* The *reg* layer returns a 4-D vector with the bbox of the region.

Given the [Objectness Score](/iron-data-science-notebook/ml-datascience/computer-vision/object-detection/metrics/objectness-score.md), Faster R-CNN computes these two classes based on it:

<figure><img src="https://569842953-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-LmjpNbCRLUyGiAxD8kn%2Fuploads%2F1tsvzplurV1YLsSPeTWm%2Fimage.png?alt=media&amp;token=d5416503-d905-40cc-a64d-fbf91de36db7" alt=""><figcaption></figcaption></figure>
