llama.cpp
0cae4306 - vulkan: support type-aligned GET_ROWS (#28253)

Commit
11 days ago
vulkan: support type-aligned GET_ROWS (#28253) * vulkan: fall back to CPU for GET_ROWS with misaligned offsets The Vulkan GET_ROWS shader asserts when a tensor's backing-buffer offset plus view_offs is misaligned w.r.t. minStorageBufferOffsetAlignment (see init_pushconst_tensor_offsets). Previously this caused a hard crash on models using ggml_view + ggml_get_rows (e.g. Qwen3-TTS, Qwen3-VL). Return false from supports_op() in the misaligned case so the scheduler falls back to CPU, matching the existing pattern for PAD_REFLECT_1D and other unsupported op/shape combinations. Repro: llama-tts -m Qwen3-TTS-*.gguf -mm mmproj-*.gguf -ngl 99 Crash: GGML_ASSERT(dst->op != GGML_OP_GET_ROWS || (a_offset == 0 && ...)) failed * vulkan: trim comment for GET_ROWS misalign fallback * vulkan: fix file corruption in gated_linear_attn struct * vulkan: properly handle misaligned offsets in GET_ROWS quantized path - get_rows_quant.comp was missing get_aoffset()/get_boffset()/get_doffset() calls that are already present in get_rows.comp, causing GGML_ASSERT crashes when GET_ROWS operates on views with non-zero view_offs, as produced by KV cache slices in Qwen3-TTS and Qwen3-VL. - Remove the defensive misalignment GGML_ASSERT in init_pushconst_tensor_offsets for the binary push-constants specialization, since both get_rows.comp and get_rows_quant.comp now correctly apply per-tensor base offsets. - Remove the workaround CPU fallback in supports_op() for GET_ROWS, since the Vulkan backend now handles misaligned offsets natively (no more bailout). - Add backend test coverage with view_src0=true (ggml_view_4d into a padded tensor) for F32, F16, Q4_0, Q4_K, Q8_0, and I32 types, exercising both the non-quantized (get_rows.comp) and quantized (get_rows_quant.comp) paths with non-zero view_offs that reproduce the original Qwen3-TTS crash. * tests: trim redundant comments in test_get_rows vs0 region * tests: trim redundant comments in test_get_rows vs0 region (follow-up) * vulkan: bind tensor base for binary ops, pass full view_offs via push constants For ops using vk_op_binary_push_constants (GET_ROWS, ADD, SUB, MUL, etc.), bind the view_src base and pass the full view_offs divided by type_size via push constant misalign_offsets. This avoids truncation when misalign_bytes is not a multiple of quantized block size. ggml_vk_tensor_subbuffer gains a use_view_offs parameter. When false, the binding points to vk_tensor_offset (base) and size includes view_offs. init_pushconst_tensor_offsets<binary> computes a/b/d_offset directly from tensor->view_offs, which is always row-aligned and therefore exact. Added non-zero view offset (offset_rows=3) backend tests for GET_ROWS across all_types with be1={1,7}, v={false,true}, skipping gradient setup for view tensors (GGML_OP_VIEW fails ggml_set_param). All 223 GET_ROWS tests pass on Vulkan (NVIDIA RTX 5060 Ti). * vulkan: bind aligned offset for binary ops, pass adjusted misalign via push constants For ops using vk_op_binary_push_constants (GET_ROWS, ADD, SUB, etc.), bind the buffer to an aligned position near the view offset (not the tensor base) and pass the adjusted misalignment via push constants. ggml_vk_get_adjusted_misalign finds the smallest misalign that is both a multiple of minStorageBufferOffsetAlignment and type_size, ensuring misalign/type_size is exact (no truncation for quantized block types). ggml_vk_tensor_subbuffer gains use_view_offs parameter. When false, binds to (target - adjusted_misalign) instead of the view_src base, keeping the offset small enough for 16-bit/8-bit push constant fields. Added non-zero view offset (offset_rows=3) backend tests for GET_ROWS across all_types with be1={1,7}, v={false,true}, skipping gradient setup for view tensors (GGML_OP_VIEW fails ggml_set_param). All 223 GET_ROWS tests pass on Vulkan (NVIDIA RTX 5060 Ti). * vulkan: bind aligned offset for binary ops, fix UMA offset mismatch For ops using vk_op_binary_push_constants (GET_ROWS, ADD, SUB, etc.), bind the buffer to an aligned position near the view offset (not the tensor base) and pass the adjusted misalignment via push constants. Added ggml_vk_tensor_physical_offset to unify physical offset lookup across UMA and non-UMA devices. On UMA, resolves via ggml_vk_host_get(tensor->data); otherwise uses vk_tensor_offset(t) + t->view_offs. Both get_misalign_bytes and the new ggml_vk_get_adjusted_misalign helper build on top of this function, so buffer bindings and push constant offsets are always consistent regardless of device memory model. ggml_vk_get_adjusted_misalign finds the smallest misalign that is both a multiple of minStorageBufferOffsetAlignment and type_size, ensuring misalign/type_size is exact (no truncation for quantized block types) while remaining small enough for 16-bit/8-bit push constant fields (adjusted_misalign < lcm(align, type_size)). ggml_vk_tensor_subbuffer gains use_view_offs parameter. When false, binds to (physical_offset - adjusted_misalign) on both UMA and discrete GPUs, fixing a bug where the UMA host_get path previously skipped the adjusted misalign binding and returned the target offset directly. Added non-zero view offset (offset_rows=3) backend tests for GET_ROWS across all_types with be1={1,7}, v={false,true}, skipping gradient setup for view tensors (GGML_OP_VIEW fails ggml_set_param). All 223 GET_ROWS tests pass on Vulkan (NVIDIA GeForce RTX 5060 Ti). * finish misalignment fix * supports_op changes for openvino/webgpu --------- Co-authored-by: AiChiTuDouPian <15327701848@qq.com>
Author
Parents
Loading