You think you know a Neural Net until you have to build it in C...
Have you ever wondered why GPUs are extremely powerful when it comes to Deep Learning?
If you wrote it out by hand (which you basically have to do for a C implementation anyways..,) you
would see that passage through the neural network is just a series of matrix-matrix multiplications.
And who loooooves matrix-matrix multiplications?? GPUs!! This is due to the fact that when you
multiply together matrices, you end up recycling a lot of the computations, meaning a parallel system
will eat it up no problem!
This project compares the performance of various strategies used for matrix-matrix multiplication:
the naive approach (manual calculation), the cBLAS library for the CPU, and the CuBLAS library for a GPU implementation.
For comparing performances accurately, the set of parameters remained constant with a learning rate of 0.1,
a batch size of 200, and a total of 50 epochs.
The data was split into 50,000 images for training and 10,000 images for validation.
If you couldn't guess already, the use of these libraries significantly improved the runtime of the program :D
But you can still read about it if you want.
READ PAPER
CODE REPOSITORY