onnxruntime
8cfbab39 - LightningIndexer: bound the fused scorer by rows, and default to cuBLAS

Commit
61 days ago
LightningIndexer: bound the fused scorer by rows, and default to cuBLAS The fused scorer was chosen on `seq_len <= 32` alone, but its cost is linear in the cache rows it scores while the GEMM's is not. At a 1K context it is 28.7 us and at 256K it is 933 us, against a GEMM that is 84.8 us over the whole 65,664-row capacity -- so the switch picked the wrong kernel exactly where it was most expensive, and a node trace put it at 18.1 ms of a 47.8 ms decode step, 44% of the busy time. Adding `score_capacity` to the condition is necessary but not sufficient: under graph capture that is the replay-invariant capacity unless the engine promises a smaller extent through ORT_LIGHTNING_INDEXER_CAPTURE_MAX_PAST, so with no promise it reads the same at 1K as at 256K. Measuring both arms end to end settled it -- the fused path is not worth discriminating for. It only ever existed to work around a bound the host could not see under capture, and once a bound exists the GEMM is faster everywhere: batch-1 step ms 1K 16K 64K 256K fused (before) 18.03 19.94 24.87 47.71 cuBLAS, no tier 20.06 21.01 22.34 30.21 cuBLAS + tier 17.79 18.85 20.72 30.17 So the default is 0, always cuBLAS, which is ahead of RC2 at every context with no configuration at all and -36.8% at 256K against the fused path. The knob keeps the fused kernel reachable for a capture whose promised extent is small. MMLU-Pro 800 is 529/800 with 0 of 800 answers discordant against the fused path, and acceptance is identical at 1K and 16K, so this is not a numerics change at the contexts a short-prompt eval reaches.
Author
Parents
Loading