onnxruntime
86916eeb - Add Intel WebGPU subgroup-matrix MatMul implementation (#29592)

Commit
21 days ago
Add Intel WebGPU subgroup-matrix MatMul implementation (#29592) - Add an Intel-specific F16 MatMul path built on the 8x16x16 subgroup-matrix config (Xe2/Xe3). A and B are loaded directly from global memory with no prepacking, and results are written through workgroup scratch with bounds-checked stores, so M and N may be any size while only K must be a multiple of 16. The impl is created in MatMul::PrePackInternal and dispatched from Compute when the device reports the required config. - The tile shape and split-K factor are chosen per call rather than fixed: TileM in {8,16,32,64}, TileN in {16,32,64}, and split-K to top up occupancy, sized to the device's resident subgroups. The WGSL kernel is parameterized by sg_mat_count_m/n and split_k, emitting the needed subgroup-matrix accumulators via generation-time guards. - An optional offline-tuned config table (generated from an autotuner sweep) can override the heuristic; selection precedence is tuned table > heuristic. - Adds MatMul subgroup-matrix unit tests covering tile selections and partial/large non-aligned dims.
Author
Parents
Loading