Block Reduce

warp shuffle → shared memory → warp shuffle
Warp Shuffle
Lane 0 Writes
__syncthreads()
Warp 0 Loads
Warp Shuffle
Return
Registers · one val per thread
blockDim.x = 64 · 4 warps × 16 lanes
On-Chip Shared Memory
__shared__ float t[32]
Warp 0 · second pass
val = (threadIdx.x < numWarps) ? t[lane] : 0.0f
WARP 0
thread 0 returns····
Kernel Sourceblock_reduce.cu

  
Press RUN to launch the kernel
waiting for kernel launch
CLK