llama.cpp
CUDA: Add OCP FP8 support
#28413
Open

CUDA: Add OCP FP8 support #28413

ORippler wants to merge 14 commits into osimons/ocp_fp8_kv from osimons/ocp_fp8_cuda
ORippler
ORippler ORippler changed the title Initial CUDA WIP version CUDA: Add OCP FP8 support 35 days ago
github-actions github-actions added ggml
github-actions github-actions added CUDA
0cc4m
0cc4m commented on 2026-09-18
ORippler ORippler force-pushed the osimons/ocp_fp8_kv branch from b6ccff9b to e8b60dda 10 days ago
ORippler ORippler force pushed from 4bcf38c6 to d4b30c45 10 days ago
ORippler ORippler force pushed from d4b30c45 to 902cdea8 10 days ago
ORippler ORippler force-pushed the osimons/ocp_fp8_kv branch from 55a7e532 to b99e4ce0 10 days ago
ORippler ORippler force pushed from 902cdea8 to ac9720e7 10 days ago
ORippler ORippler force-pushed the osimons/ocp_fp8_kv branch from b99e4ce0 to f711e8c8 10 days ago
ORippler ORippler force pushed from ac9720e7 to 2de9bb7a 10 days ago
ORippler ORippler force pushed from 2de9bb7a to e50eed67 10 days ago
ORippler ORippler force-pushed the osimons/ocp_fp8_kv branch from f711e8c8 to d6c8682f 10 days ago
ORippler ORippler force pushed from e50eed67 to 0bc3b72a 9 days ago
ORippler ORippler force-pushed the osimons/ocp_fp8_kv branch from d6c8682f to 51f6a0f2 9 days ago
ORippler ORippler force-pushed the osimons/ocp_fp8_kv branch from 51f6a0f2 to b5f3dfd6 9 days ago
ORippler ORippler force pushed from 0bc3b72a to 93f06440 9 days ago
ORippler ORippler force pushed from 93f06440 to b7d3ef0d 9 days ago
ORippler ORippler force pushed from f8acf520 to 98d0007c 9 days ago
ORippler ORippler force-pushed the osimons/ocp_fp8_kv branch from c8023359 to 4602e642 9 days ago
github-actions github-actions added testing
ORippler Initial CUDA WIP version
983021be
ORippler Use fp8 intrinsics
a1b617dc
ORippler Do W8A16 for FP8 in GEMV case
869abaa3
ORippler Optimize GEMV kernel for FP8
d6c8c36a
ORippler Guard __nv_cvt_fp8x2_to_bf162raw on CTK 13.2
f7c5a98f
ORippler Remove trailing whitespace
1d52850a
ORippler Add supported for strided copies to/from FP8 in CUDA
b841db1b
ORippler Prepare Q_reg for FP8 FlashAttention in fattn-vec.cuh
d2bd12ec
ORippler Fix odd-rows for convert/getrows, possible with FP8
abc23201
ORippler Fix FP8 fallback in cuBlas
6ee4762f
ORippler Gate mmvf_f8x2_e4m3_to_bf162 on FP8 availability, i.e. CTK >= 11.8
8a3884fe
ORippler Restrict (fused) MMVF to 2*ts alignment, and cublas FP8 fo 16-byte
f771a191
ORippler Add todo about accumulator/precision honoring
d34e22aa
ORippler Avoid races in fp8_fallback
b757f0d5
ORippler ORippler force pushed from 021df05a to b757f0d5 21 hours ago
ORippler ORippler force-pushed the osimons/ocp_fp8_kv branch from 4602e642 to 10e80e62 21 hours ago

Login to write a write a comment.

Login via GitHub

Reviewers
Assignees
No one assigned
Labels
Milestone