Replace gguf-py NumPy dequant with pure-PyTorch port (3-8x faster)
Port city96 / ComfyUI-GGUF's torch dequant kernels (also used by
`diffusers`) into a new `integrations/gguf_dequant.py`, and route
`spawn_gguf_materialize` through it instead of `gguf.dequantize`.
Covers Q4_0/Q4_1/Q5_0/Q5_1/Q8_0, Q2_K..Q6_K, IQ4_NL, IQ4_XS, BF16, F16,
F32. Bit-exact match with gguf-py on synthetic random tensors and 3-8x
faster on a 4096x4096 weight tile (CPU, single-thread). Falls back to
`gguf.dequantize` for any quant type without a torch kernel.