Fix attention bias broadcast (#24017)
### Description
* Fix broadcast on attention bias dim 1.
* Increase test cases in test_mha.py in pipeline to cover the testing.
### Motivation and Context
This feature was added in
https://github.com/microsoft/onnxruntime/pull/21710.
There was bug when computing the offset when attention bias broadcast on
dim 1 only in both CUDA and CPU kernel.
It can be triggered when attention bias shape is like [batch_size, 1,
sequence_length, total_sequence_length] and batch_size > 1 when unfused
kernel is selected. Note that cudnn flash attention and cutlass fused
attention also supports attention bias, so the bug in unfused kernel was
not discovered previously.