Preallocate declared CUDA workspace
Add an opt-in sequential static-workspace pilot that reuses one aligned buffer per device and routes MatMulNBits slot zero through it with dynamic fallback.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: cdd38fe6-fcf3-45c2-acea-b8e6206a7839