Go
Home
Pricing
FAQ
Install
Home
Pricing
FAQ
Install
Login
via GitHub
microsoft/onnxruntime
Pull Requests
Commits
Open
Closed
[CUDA] MatMulBlockQuantizedFp8Weight: fold the W8A8 activation QDQ into the decode GEMV
#31481 opened 2026-08-02 07:17 by
tianleiwu
[CUDA] GroupQueryAttention: dequantize the K and V caches in one launch
#31480 opened 2026-08-02 07:14 by
tianleiwu
[CUDA] Skip FP4 QMoE fc1 activation expansion
#31479 opened 2026-08-02 02:31 by
tianleiwu
[CUDA] MatMul: add an opt-in split-K GEMV for small-N fp16 shapes
#31478 opened 2026-08-01 19:54 by
tianleiwu
[CUDA] Fix Abs signed zero handling
#31477 opened 2026-08-01 14:04 by
MohamedElashri
[CUDA] Speed up 8-bit MatMulNBits dequantization with byte permutes
#31350 opened 2026-08-01 00:52 by
tianleiwu
[CUDA] Coalesce QMoE MXFP4/NVFP4 weight dequantization
#31349 opened 2026-07-31 23:55 by
tianleiwu
[Compile API] Plan for adding support for writing external initializers to buffer
#31347 opened 2026-07-31 22:22 by
skottmckay
Ensure ONNX and Protobuf use compatible package configurations
#31346 opened 2026-07-31 22:15 by
Copilot
Add AzureContainerRegistry to network isolation policy for additional pipelines.
#31344 opened 2026-07-31 17:36 by
edgchen1
[Shape Inference] Add shape inference for MatMulIntegerToFloat op
#31199 opened 2026-07-31 15:58 by
roberto-laudani
[WebNN EP] Support LpNormalization op
#31197 opened 2026-07-31 06:55 by
Honry
Make the tellg failure check explicit and document the EPContext write gate error handling
#31196 opened 2026-07-31 03:04 by
GopalakrishnanN
[CUDA] Port ShardedMoE to the new MoE GEMM backend
#31191 opened 2026-07-30 19:09 by
tianleiwu
Fix shape inference failure with file-backed external data
#31169 opened 2026-07-30 13:00 by
yoonseok-kim
[CUDA] Speed up the NVFP4 QMoE decode GEMV and enable it for MTP verify
#31159 opened 2026-07-30 09:36 by
tianleiwu
[CUDA] Add LinearAttentionGate and GatedRMSNorm contrib operators
#31158 opened 2026-07-30 09:36 by
tianleiwu
[CUDA] Add state_window to LinearAttention and CausalConvWithState to support MTP
#31157 opened 2026-07-30 09:36 by
tianleiwu
[WebNN EP] Reuse shared WASM loader for Blob-backed external data
#31151 opened 2026-07-30 05:56 by
Honry
Skip redundant present_key/value copy when aliased to external KV cache (follow-up to #29715)
ep:CUDA
#31150 opened 2026-07-30 04:57 by
titaiwangms
Throw error on negative Split tensor axis [CPU]
#31149 opened 2026-07-30 00:03 by
nenad1002
Fix LayerNorm fusion selecting the Div output as scale or bias
#31148 opened 2026-07-29 23:57 by
Sammy-Dabbas
[ARM] MLAS: SVE i8mm (svmmla) int8 QGEMM kernels, portable machine code
#31146 opened 2026-07-29 21:54 by
mirounga
[ARM] MLAS: portable machine-code SVE elementwise kernels, FEXPA exp
#31145 opened 2026-07-29 21:53 by
mirounga
Do not fuse LayerNorm when ReduceMean has keepdims=0
#31144 opened 2026-07-29 20:01 by
guptaishaan
[ARM] MLAS: route fp32 SGEMM to the KleidiAI SVE ukernel on SVE-only
#31143 opened 2026-07-29 18:10 by
mirounga
Fix segfault in MatMulNBits prepacking when weights/scales are parent-graph initializers in subgraph
#31141 opened 2026-07-29 15:51 by
Copilot
[OpenVINO EP] Fix get/set_float_initializer_data for raw_data backed float initializers
#31138 opened 2026-07-29 11:46 by
wangw-1991
Auto-generated baselines by 1ES Pipeline Templates
#31136 opened 2026-07-29 11:03 by
microsoft-github-policy-service[bot]
[WebGPU EP] Support int64 for Tile and Concat
ep:WebGPU
#31049 opened 2026-07-29 08:50 by
miaobin
Older