llama.cpp
112c7815 - ggml-cuda: Add NVFP4 dp4a kernel (#20644)

Commit
2 days ago
ggml-cuda: Add NVFP4 dp4a kernel (#20644) Added check for dst_t to cuda_cast template for float Restored ggml_cuda_ue4m3_to_fp32, changed vecdot ints to int32ts Added CUDART/HIP Check and HIP/fp8 include Added NVFP4 to Test-backend-ops Added hip_fp8_e4m3 to __nv_fp8_e4m3 typedef --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Author
Parents
Loading