Tiled MatMul

Load Tiles
__syncthreads()
Dot Product
__syncthreads()
Write C Tile
On-Chip Shared Memory
As[4][4]
Bs[4][4]
Matrix B K×N · global
Matrix A M×K · global
Matrix C M×N · accumulator
Kernel Source tiled_matmul.cu

  
Press RUN to launch the kernel
waiting for kernel launch
CLK