NEEDLE – GPU Weight-Tensor Memory Layout

Visualising how model weights W[H×H] reside in GPU VRAM for each execution strategy.
Colour intensity = "hotness" (currently being read by the kernel).

Sequential

One weight matrix loaded at a time. Kernel cycles through fold₀→fold₅ serially.
N×H² total parameter loads, one after another.
Model: – Kernel calls: 0 Memory re-use: none

Parallel (Streams)

Up to 8 weight matrices resident simultaneously (one per CUDA stream). All fold kernels launch concurrently; GPU arbitrates scheduling.
Active streams: 0 Kernel calls: 0 Concurrent: yes

Vectorized (vmap)

Parameters stacked into W_batched[N, H, H]. Single batched forward pass reads all N weight matrices simultaneously.
Batch: N=6 Kernel calls: 0 Memory layout: contiguous
Idle / not loaded Resident in VRAM Active kernel read Peak activity (hot)