LightningIndexer: bound the fused scorer by rows, and default to cuBLAS
The fused scorer was chosen on `seq_len <= 32` alone, but its cost is linear in the
cache rows it scores while the GEMM's is not. At a 1K context it is 28.7 us and at
256K it is 933 us, against a GEMM that is 84.8 us over the whole 65,664-row capacity
-- so the switch picked the wrong kernel exactly where it was most expensive, and a
node trace put it at 18.1 ms of a 47.8 ms decode step, 44% of the busy time.
Adding `score_capacity` to the condition is necessary but not sufficient: under graph
capture that is the replay-invariant capacity unless the engine promises a smaller
extent through ORT_LIGHTNING_INDEXER_CAPTURE_MAX_PAST, so with no promise it reads the
same at 1K as at 256K. Measuring both arms end to end settled it -- the fused path is
not worth discriminating for. It only ever existed to work around a bound the host
could not see under capture, and once a bound exists the GEMM is faster everywhere:
batch-1 step ms 1K 16K 64K 256K
fused (before) 18.03 19.94 24.87 47.71
cuBLAS, no tier 20.06 21.01 22.34 30.21
cuBLAS + tier 17.79 18.85 20.72 30.17
So the default is 0, always cuBLAS, which is ahead of RC2 at every context with no
configuration at all and -36.8% at 256K against the fused path. The knob keeps the
fused kernel reachable for a capture whose promised extent is small.
MMLU-Pro 800 is 529/800 with 0 of 800 answers discordant against the fused path, and
acceptance is identical at 1K and 16K, so this is not a numerics change at the
contexts a short-prompt eval reaches.