whisper.cpp
6fb7f1af - SYCL: add BF16 to DMMV kernel path (~4x tg speedup on Intel Arc) (llama/21580)

Commit

3 days ago

SYCL: add BF16 to DMMV kernel path (~4x tg speedup on Intel Arc) (llama/21580) * SYCL: add BF16 to DMMV kernel path for ~4x token generation speedup BF16 models had no dedicated token generation kernel — they fell through to the generic full-GEMM path, resulting in ~14% memory bandwidth utilization on Intel Arc GPUs. This adds BF16 support to the DMMV (dequantize mul-mat-vec) path, matching the existing F16 implementation. Fixes #20478 * SYCL: fix BF16 DMMV out-of-bounds when ncols % 64 != 0 The qk=1 kernel (used for F16 and BF16) iterates with stride 2*GGML_SYCL_DMMV_X (= 64 on Intel targets where WARP_SIZE=16). When ncols is a multiple of DMMV_X (32) but not of 2*DMMV_X (64), the last warp iteration accesses elements at col >= ncols, producing NaN for the final row and wrong values for interior rows. Fix: tighten can_use_dequantize_mul_mat_vec to require ne[0] % (2*DMMV_X) == 0 for F16/BF16 types, and update the ASSERT in the BF16 launcher to match. Quantized types use block-structured kernels with different access patterns and keep the existing DMMV_X check. Verified: test-backend-ops MUL_MAT passes 913/913 on Intel Arc Pro B70. Previously failing: m=128/129 n=1 k=1056 cases (NaN and ERR > 0.0005). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

References

#3824 - sync : ggml

Author

PMZFX

Committer

ggerganov

Parents

2d629533

whisper.cpp 6fb7f1af - SYCL: add BF16 to DMMV kernel path (~4x tg speedup on Intel Arc) (llama/21580)

whisper.cpp
6fb7f1af - SYCL: add BF16 to DMMV kernel path (~4x tg speedup on Intel Arc) (llama/21580)