Visualising how model weights W[H×H] reside in GPU VRAM for each execution strategy.
Colour intensity = "hotness" (currently being read by the kernel).
Sequential
One weight matrix loaded at a time.
Kernel cycles through fold₀→fold₅ serially.
N×H² total parameter loads, one after another.
Model: –
Kernel calls: 0
Memory re-use: none
Parallel (Streams)
Up to 8 weight matrices resident simultaneously (one per CUDA stream).
All fold kernels launch concurrently; GPU arbitrates scheduling.
Active streams: 0
Kernel calls: 0
Concurrent: yes
Vectorized (vmap)
Parameters stacked into W_batched[N, H, H].
Single batched forward pass reads all N weight matrices simultaneously.
Batch: N=6
Kernel calls: 0
Memory layout: contiguous