Keep allowing bias when key and value are 4-D BNSH tensors
MultiHeadAttention has always accepted a bias together with 4-D BNSH key and
value: both the CPU kernel and the CUDA PrepareQkv_MHA_Cross path apply only
the query part and assume the key/value parts are zero. Rejecting that
combination broke MultiHeadAttentionTest.CrossAttention_DiffSequenceLengths_
UsingDMMHAInsideMHA and test_mha.py TestMultiHeadAttention.test_all, which
covers the same case via the no_bias_k_v reference.
Drop the rejection, document the zero-bias assumption in the schema, and turn
the negative test into one that feeds a non-zero key/value bias and asserts the
output is unchanged.