llama.cpp
6a2743f0 - CUDA: bitonic argsort handles rows wider than one block (#28957)

Commit
6 days ago
CUDA: bitonic argsort handles rows wider than one block (#28957) Without CUB (HIP, MUSA) argsort ran the bitonic kernel with one thread per padded column, so any row above 1024 entries launched an invalid block configuration. Each thread now owns several columns, every stage of the network runs all owned columns before the barrier, and the block is capped at 1024 threads. Shared memory becomes the only bound, which supports_op checks against the device instead of a fixed 1024. Rows up to 1024 run the same work as before. Bit-exact with the CUB path on rows of 2048.
Parents
Loading