Fix MPS dequant + Qwen biases; further consolidate T5 FFN
* `dequantize_gguf_tensor` now accepts torch.Tensor input directly (no
numpy conversion), so the dequant op works for tensors that the loader
has already moved to MPS / CUDA.
* Add `\.attn_(q|k|v)\.bias → .self_attn.\1_proj.bias` to the shared
Llama-family rename list — Qwen2/Qwen3/StableLM/Starcoder2 etc. ship
attn biases in GGUF; the rule is a no-op for arches that don't.
* Consolidate T5 FFN renames: redirect `encoder.block.N.ffn_` →
`encoder.block.N.layer.1.ffn_` (and same for decoder → layer.2), then
share the per-tensor ffn_norm/gate/up/down renames between encoder and
decoder. 8 rules → 6.