> For the complete documentation index, see [llms.txt](https://fgheorghe.gitbook.io/machine-learning/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://fgheorghe.gitbook.io/machine-learning/vehicle-detection.md).

# Vehicle Detection

## **Challenge**

Write a software pipeline to detect vehicles in a traffic video taken while driving on a motorway:

![](https://1635893444-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LEUEBQWKhOEp5gdtdP5%2F-LEkXK1CZPohEh5Dt-6P%2F-LEkXNRKpzDOVXX7nU9L%2Fproject_video.gif?alt=media\&token=be4b43c3-925b-41ac-b8b5-64b28a1a53a3)

## **Actions**

* Histogram of Oriented Gradients (HOG)
  * we have been provided with a set of images: 8968 various traffic images, 8792 images with cars;images are color, 64x64 pixels
  * using HOG and other techniques, extract features of these images
  * separate the images in train/test and train a SVM classifier
* Sliding Window Search
  * implemented a sliding window search and classify each window as vehicle or non-vehicle
  * run the function first on test images and afterwards against project video
* Video Implementation
  * output a video with the detected vehicles positions drawn as bounding boxes
  * implement a robust method to avoid false positives (could be a heat map showing the location of repeated detection)

### Tools

This project used a combination of Python, numpy, matplotlib, openCV, scikit-learn and moviepy; this is by definition a computer vision project. These tools are installed in a anaconda environment and ran in a Jupyter notebook.

The complete project implementation is available here: <https://github.com/FlorinGh/SelfDrivingCar-ND-pr4-Advanced-Lane-Lines>.

### Histogram of Oriented Gradients (HOG)

The implementation starts with importing the relevant modules for this project; all images were place in the same directory on a local drive; non-car images have been renamed starting with 'traffic'; this helped separate them in car and non-car lists; the data set has 8792 car images and 8968 non-car images; the data set is well balanced and it doesn't need augmentation.

Here is an example of one of each of the `vehicle` and `non-vehicle` classes:

![](https://1635893444-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LEUEBQWKhOEp5gdtdP5%2F-LEkx6-hv5Qm2HGoC-o4%2F-LEkxfGAHK95kTrU10Eo%2Fcar_not_car.png?alt=media\&token=fb8eb374-4f05-4efd-b887-a15e83a6fe60)

I then explored different colour spaces and different `skimage.hog()` parameters (`orientations`, `pixels_per_cell`, and `cells_per_block`). I grabbed random images from each of the two classes and displayed them to get a feel for what the `skimage.hog()` output looks like.

Here is an example using the `YCrCb` color space and HOG parameters of `orientations=8`, `pixels_per_cell=(8, 8)` and `cells_per_block=(2, 2)`:

![](https://1635893444-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LEUEBQWKhOEp5gdtdP5%2F-LEkx6-hv5Qm2HGoC-o4%2F-LEky9hB87LMLj7M_zkr%2FHOG_example.jpg?alt=media\&token=2dc35a76-f1d6-4e18-b58d-a3d7f5f722be)

In order to choose the HOG parameters, I tested different combinations changing only one parameter at a time, on a range of values; the best accuracy would indicate which value to keep and went ahead to test another parameter; below you can see the effect of each parameter, and in bold the one kept for final run.

![](https://1635893444-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LEUEBQWKhOEp5gdtdP5%2F-LEkx6-hv5Qm2HGoC-o4%2F-LEkygdNe4c1JuUbihqV%2Fparameters_search.png?alt=media\&token=bfd1ce84-897c-4e21-b662-844d5f22a335)

After that the features of each of the car and non-car image are extracted; these are used to train the linear SVM model; before running the training algorithm the features are normalised using the StandardScaler; then data set is initiated and trained; the resulting test accuracy was 98.51%.

### Sliding Window Search

The next section is the main part of the project: using small windows, each image (video frame) is searched; data from each window is than tested against the trained model and a decision is made if it contains a car or not; using several windows we can detect a car in more than one window; this is particularly helpful to reduce false positive cases; using a heat filter we will extract the locations where a car was detected more than 5 times, ignoring all other windows.

I searched with 3 window sizes: 90px, 96px and 112px; these were selected after a few trial and error tests:

![](https://1635893444-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LEUEBQWKhOEp5gdtdP5%2F-LEl0G7ELUPjN7UfJp6y%2F-LEl0nwW5mJToxeIOoIp%2Fsliding_windows.png?alt=media\&token=18f64d9a-ad3c-4ff2-9d49-bbb8afce2c0e)

Ultimately I searched on two scales using YCrCb 3-channel HOG features plus spatially binned color and histograms of color in the feature vector, which provided a nice result. Here are some example images:

![](https://1635893444-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LEUEBQWKhOEp5gdtdP5%2F-LEl0G7ELUPjN7UfJp6y%2F-LEl16dRi9V07Q4gCqSv%2Fsliding_window.jpg?alt=media\&token=4c19a785-0f56-485e-9be5-4256076ad2a9)

### Video Implementation

I recorded the positions of positive detections in each frame of the video. From the positive detections I created a heatmap and then thresholded that map to identify vehicle positions. I then used `scipy.ndimage.measurements.label()` to identify individual blobs in the heatmap. I then assumed each blob corresponded to a vehicle. I constructed bounding boxes to cover the area of each blob detected.

Here's an example result showing the heat map from a series of frames of video, the result of `scipy.ndimage.measurements.label()` and the bounding boxes then overlaid on the last frame of video:

![](https://1635893444-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LEUEBQWKhOEp5gdtdP5%2F-LEl0G7ELUPjN7UfJp6y%2F-LEl1ihGGKVeg-T7avlX%2Fbboxes_and_heat.png?alt=media\&token=058ebe42-6fc7-43be-ab49-5936825174d2)

Here is the output of `scipy.ndimage.measurements.label()` on the integrated heatmap from all six frames:

![](https://1635893444-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LEUEBQWKhOEp5gdtdP5%2F-LEl0G7ELUPjN7UfJp6y%2F-LEl1pv2zPzl_4P20N5_%2Flabels_map.png?alt=media\&token=9c5177b7-1ce2-4d64-a7a1-512296b66a1c)

And here the resulting bounding boxes are drawn onto the last frame in the series:

![](https://1635893444-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LEUEBQWKhOEp5gdtdP5%2F-LEl0G7ELUPjN7UfJp6y%2F-LEl2-26MVI5WoCbyund%2Foutput_bboxes.png?alt=media\&token=6e6eea50-d2ef-474d-bc6c-4cd21be1a8ef)

### Discussion

The most difficult part of the project was eliminating the false positives; even we had good accuracy on the test, in the project this didn't seem good enough; this is a sign the training model had over-fit the training data; One way to improve the project would be to try another training model, maybe a neural network.

## **Results**

All steps described above were captured in a pipeline; applying it over the frames of the traffic video renders the following result:

![](https://1635893444-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LEUEBQWKhOEp5gdtdP5%2F-LEkXK1CZPohEh5Dt-6P%2F-LEkXQV9QBtywT0AvQwf%2Fproject_video_output.gif?alt=media\&token=6a21ddca-a363-4e9f-bae3-b2cda1ee818f)

For more details on this project visit the following github repository: <https://github.com/FlorinGh/SelfDrivingCar-ND-pr5-Vehicle-Detection>.
