Tiled Conv2D

Load Halo Tile
__syncthreads()
Convolve + Store
Input N 8×8 · global
On-Chip Memory
N_s[6][6] · shared
F_c[3][3] · constant
Output P 8×8 · global
Kernel Source conv2d_tiled.cu

  
Press RUN to launch the kernel
waiting for kernel launch
CLK