llama.cpp
Add OpenVINO backend
#15307
Merged
Go
Login via GitHub
Home
Pricing
FAQ
Install
Login
via GitHub
Overview
Commits
320
Changes
View On
GitHub
Add OpenVINO backend
#15307
ggerganov
merged 320 commits into
ggml-org:master
from
ravi9:dev_backend_openvino
wine99
requested a review
from
ngxson
1 year ago
wine99
requested a review
from
ggerganov
1 year ago
wine99
marked this pull request as draft
1 year ago
github-actions
added
documentation
github-actions
added
testing
github-actions
added
devops
github-actions
added
ggml
wine99
force pushed
1 year ago
wine99
force pushed
to
80f09692
1 year ago
wine99
force pushed
to
76ab76ed
1 year ago
wine99
force pushed
from
76ab76ed
to
2e1dd8df
1 year ago
wine99
force pushed
to
66e503bd
1 year ago
wine99
marked this pull request as ready for review
362 days ago
wine99
requested a review
from
CISC
362 days ago
wine99
requested a review
from
slaren
362 days ago
CISC
commented on 2025-10-14
wine99
force pushed
361 days ago
wine99
force pushed
to
f89292d2
360 days ago
CISC
commented on 2025-10-15
CISC
commented on 2025-10-15
danbev
commented on 2025-10-16
wine99
force pushed
to
d5038aae
340 days ago
CISC
commented on 2025-11-04
yangsu2022
commented on 2026-01-06
cavusmustafa
requested a review
from
allozaur
269 days ago
cavusmustafa
requested a review
from
rgerganov
269 days ago
cavusmustafa
requested a review
from
JohannesGaessler
269 days ago
cavusmustafa
requested a review
from
reeselevine
269 days ago
cavusmustafa
requested a review
from
0cc4m
269 days ago
cavusmustafa
requested a review
from
max-krasnyansky
269 days ago
cavusmustafa
requested a review
from
lhez
269 days ago
cavusmustafa
requested a review
from
am17an
269 days ago
cavusmustafa
requested a review
from
aldehir
269 days ago
cavusmustafa
requested a review
from
pwilkin
269 days ago
github-actions
added
model
github-actions
added
build
github-actions
added
script
github-actions
added
android
github-actions
added
Nvidia GPU
github-actions
added
nix
github-actions
added
Vulkan
github-actions
added
examples
github-actions
added
python
github-actions
added
server
github-actions
added
SYCL
github-actions
added
Apple Metal
github-actions
added
Ascend NPU
github-actions
added
OpenCL
CISC
removed review request
from
rgerganov
268 days ago
CISC
removed review request
from
aldehir
268 days ago
CISC
removed review request
from
max-krasnyansky
268 days ago
CISC
removed review request
from
slaren
268 days ago
CISC
removed review request
from
am17an
268 days ago
CISC
removed review request
from
pwilkin
268 days ago
CISC
removed review request
from
reeselevine
268 days ago
CISC
removed review request
from
0cc4m
268 days ago
CISC
removed review request
from
JohannesGaessler
268 days ago
CISC
removed review request
from
allozaur
268 days ago
CISC
removed review request
from
lhez
268 days ago
CISC
removed
model
CISC
removed
build
CISC
removed
script
CISC
removed
android
CISC
removed
Nvidia GPU
CISC
removed
nix
CISC
removed
Vulkan
CISC
removed
examples
CISC
removed
python
CISC
removed
server
CISC
removed
SYCL
CISC
removed
Apple Metal
CISC
removed
Ascend NPU
CISC
removed
OpenCL
cavusmustafa
force pushed
from
8f75420b
to
ac1c0cc3
268 days ago
Update build doc
fd324366
Add cgraph tensor output name to OV op name
8ce5cc59
Update openvino build instructions
3051d5ae
Add initial NPU support
7fec2233
draft NPU support version 2: prefill + kvcache
34531abc
NPU support version 2: prefill + kvcache
d9ca8f5d
Change due to ggml cgraph changes, not correct yet
f7ad7793
Change due to ggml cgraph changes, llama-3.2 CPU work
592d7f8b
Add AMD64 to CMakeLists
e27738a9
Change due to ggml cgraph changes, all device work
42d42409
Refactor: clean, fix warning
593484ce
Update clang-format
8afee795
Statful transformation for CPU GPU
4c582ac7
Add SwiGLU
73ee84ff
Fuse to SDPA
ebc4fc9f
Update build doc
fd324366
Change due to ggml cgraph changes, not correct yet
f7ad7793
Reduce memory: free ov weights node after graph conversion
f3c05190
Fix CPY due to cgraph change
d61f83c9
Added OpenVINO CI/CD. Updated docs
ea75772e
Fix llama-cli
1ed49bbf
Fix Phi3 ROPE; Add test-backend-ops
44f4cf34
Fix NPU
6dc4b906
Fix llama-bench; Clang-format
75eec626
Fix llama-perplexity
4e7f04a3
temp. changes for mark decomp
9cf56d68
matmul in fp32
01cdf4a9
mulmat input conversion fix
e2fdc1b9
mulmat type conversion update
93b2d09a
add mark decomp pass
1a19566b
Revert changes in fuse_to_sdpa
43489bbf
Update build.md
2f99135c
Fix test-backend-ops
fc865340
Skip test-thread-safety; Run ctest only in ci/run.sh
11413503
Use CiD for NPU
37ff226b
Optimize tensor conversion, improve TTFT
9a91ca6e
Support op SET_ROWS
63d000ba
Fix NPU
7bda5021
Remove CPY
839f8c66
Fix test-backend-ops
f4123be9
Minor updates for raising PR
a7b611bc
Perf: RMS fused to OV internal RMS op
14c8a85c
Fix after rebasing
65e1b1af
Change openvino device_type to GPU; Enable flash_attn
56d59677
Update supports_buft and supports_op for quantized models
3e897df5
Add quant weight conversion functions from genai gguf reader
d4ca760d
Quant models run with accuracy issue
663a0b8c
Fix accuracy: disable cpu_repack
6ab76ed1
Fix CI; Disable test-backend-ops
dd80b042
Fix Q4_1
a1ce4280
Fix test-backend-ops: Treat quantized tensors as weights
9900245e
Add NPU Q4_0 support
9ca53c79
NPU perf: eliminate zp
82c98335
Dequantize q4_1 q4_k q6_k for NPU
b593428e
Add custom quant type: q8_1_c, q4_0_128
6926655f
Set m_is_static=false as default in decoder
c5231a24
Simpilfy translation of get_rows
810eb480
Fix after rebasing
0f7b253c
Improve debug util; Eliminate nop ReshapeReshape
2ad1147b
STYLE: make get_types_to_requant a function
dc77cbb3
Support BF16 model
bcc343af
Fix NPU compile
434059ae
WA for npu 1st token acc issue
da2cc993
Apply EliminateZP only for npu
be07073e
Add GeGLU
59756124
Fix Hunyuan
7d81861a
Support iSWA
9de874cb
Fix NPU accuracy
602f9ca4
Fix ROPE accuracy when freq_scale != 1
1a38339c
Minor: not add attention_size_swa for non-swa model
67e178a2
Minor refactor
2f1d50fb
Add Q5_K to support phi-3-q4_k_m
e4bfe5a2
Requantize Q6_K (gs16) to gs32 on GPU
f3afa7b9
Fix after rebasing
fdadca1e
Always apply Eliminate_ZP to fix GPU compile issue on some platforms
973a80fd
kvcachefusion support
c112bc4e
env variable GGML_OPENVINO_DISABLE_SDPA_OPTIMIZATION added
e7252920
Fix for Phi3
05d7abae
Fix llama-cli (need to run with --no-warmup)
a9371ea6
Fix add_sliced_mask; Revert mulmat, softmax; Remove input attention_s…
8b82d115
fix after rebasing
299f4923
Fix llama-3-8b and phi3-mini q4_0 NPU
2d2f00a4
Update to OV-2025.3 and CMakeLists.txt
841d673b
Add OV CI cache
4c8406eb
Apply CISC review and update CI to OV2025.3
38e8a19f
Update CI to run OV dep install before build
45af912b
Update OV dockerfile to use OV2025.3 and update build docs
3a1129e0
Style: use switch in supports_ops
bd3093f9
Style: middle ptr and ref align, omit optional struct keyword
eba8113d
NPU Unify PD (#14)
b8690bc0
Clean placeholders in ggml-openvino.cpp
303923ab
NPU unify PD (handled internally)
ea2c99be
change graph to 4d, support multi sequences
072dde0b
Fix llama-bench
ae404f7c
Fix NPU
531941b3
Update ggml-decoder.cpp
047bfb5c
Update ggml-decoder.cpp
11b4cc5a
Update ggml-decoder.cpp
bed49522
Update ggml-decoder.cpp
4a57b37d
Update ggml-decoder.cpp
98396b27
Update ggml-decoder.cpp
4400b5cb
Remove the second decoder for node. Moving the function into the mode…
ae936519
Fix error for naive
992dea73
NPU prefill chunking
38254cf5
NPU fix llama-bench
59e7e7c4
fallback naive run with accuracy issue
65348b5d
NPU support llma-perplexity -b 512 --no-warmup
808619e2
Refactor: split ov_graph_compute for dynamic and static
2a9d4ca8
remove unused API GgmlOvDecoder::get_output_stride(const std::string …
0ea8238a
minor update due to ov 2025.4
8f4ee4ee
remove unused API GgmlOvDecoder::get_output_names()
497964af
remove unused API get_output_shape(const std::string & name)
f516db1d
Modified API GgmlOvDecoder::get_output_type(const std::string & name)
6d7a0d60
Removed API GgmlOvDecoder::get_output_op_params(const std::string & n…
ba852f2a
Removed API get_output_ggml_tensor(const std::string & name)
111c96c2
Fix error for naive
992dea73
NPU prefill chunking
38254cf5
NPU fix llama-bench
59e7e7c4
fallback naive run with accuracy issue
65348b5d
NPU support llma-perplexity -b 512 --no-warmup
808619e2
Refactor: split ov_graph_compute for dynamic and static
2a9d4ca8
minor update due to ov 2025.4
8f4ee4ee
remove unused API GgmlOvDecoder::get_output_names()
497964af
remove unused API get_output_shape(const std::string & name)
f516db1d
Modified API GgmlOvDecoder::get_output_type(const std::string & name)
6d7a0d60
Removed API GgmlOvDecoder::get_output_op_params(const std::string & n…
ba852f2a
Removed API m_outputs
8ff73e5d
Removed m_output_names
197ed992
Removed API GgmlOvDecoder::get_input_names()
95c30719
Removed API GgmlOvDecoder::get_input_stride(const std::string& name)
cd611782
Removed API get_input_type
891a3beb
Removed API get_input_type
42ca27f7
Removed API GgmlOvDecoder::get_input_shape(const std::string & name)
acb8a01d
Removed API GgmlOvDecoder::get_input_op_params(const std::string & name)
47c91db3
Fix error for decoder cache
91a1b20c
Reuse cached decoder
28da9a9a
GPU remove Q6_K requantization
469325c6
NPU fix wrong model output shape
ae01322d
NPU fix q4 perf regression
c9234b44
Remove unused variable nodes
9e3163e8
Update build.md for Windows
ae533638
backend buffer: allocate on host
22d9c17a
Use shared_buffer for GPU NPU; Refactor
72bba828
Add ov_backend_host_buffer; Use cached remote context
3fdcb6ab
Put kvcache on GPU
d7578497
Use ggml_aligned_malloc
8273a7c2
only use remote tensor for kvcache
88d1d17e
only use remote tensor for kvcache for GPU
a356b444
FIX: use remote tensor from singleton
cfc47135
Update build.md to include OpenCL
52a44012
NPU always requant to q4_0_128
c1142ddb
Optimize symmetric quant weight extraction: use single zp
67c9720e
Use Q8_0_C in token embd, lm_head, and for 5 and 6 bits quant
4e451778
Update build.md
f5c71e3c
Initial stateful graph support
5f30eacd
Update ggml/src/ggml-openvino/ggml-decoder.cpp
d2fc1522
code cleanup
981ec657
npu perf fix
a40a5dfc
requant to f16 for Q6 embed on NPU
a81b202f
Update ggml/src/ggml-openvino/ggml-decoder.cpp
a92eceec
Update ggml/src/ggml-openvino/ggml-openvino-extra.cpp
599335c6
Create OPENVINO.md in llama.cpp backend docs
416556a8
Update OPENVINO.md
25e65256
Update OPENVINO.md
9ba32472
Update OPENVINO.md
61552e44
Update build.md
63eed0d9
Update OPENVINO.md
f44c60e9
Update OPENVINO.md
e9ed5c4c
Update OPENVINO.md
d3649c11
kq_mask naming fix
d7dccf88
cavusmustafa
force pushed
from
ac1c0cc3
to
d7dccf88
268 days ago
ggerganov
commented on 2026-01-16
Syntax correction for workflows build file
aa4bc900
Change ov backend buffer is_host to false
9a15c8b0
wine99
force pushed
to
4a8fd24e
257 days ago
Fix llama-bench -p -n where p<=256
8fb20b28
Fix --direct-io 0
1c0a47a4
Don't put kvcache on GPU in stateful mode
c8402102
Remove hardcode names
d398214e
Fix stateful shapes
26328fe1
Simplification for stateful and update output shape processing
32599213
Remove hardcode names
18ab0f56
Avoid re-compilation in llama-bench
b6c0697d
Extract zp directly instead of bias
0ee7e054
Refactor weight tensor processing
900dd76c
wine99
force pushed
from
6d71ded5
to
900dd76c
242 days ago
Merge branch 'master' into dev_backend_openvino
7b3b65b0
create_weight_node accept non-ov backend buffer
1d4ec1b2
remove changes in llama-graph.cpp
e0590152
stateful masking fix (#38)
0d74aba2
Fix test-backend-ops crash glu, get_rows, scale, rms_norm, add
d5d673cd
hardcoded name handling for rope_freqs.weight
59e7d730
Suppress logging and add error handling to allow test-backend-ops to …
1a54965c
wine99
force pushed
to
1a54965c
240 days ago
Fix MUL_MAT with broadcast; Add unsupported MUL_MAT FLASH_ATTN cases
2a6a95eb
Use bias instead of zp in test-backend-ops
5525bac0
wine99
force pushed
to
5525bac0
239 days ago
Merge pull request #43 from cavusmustafa/additional_fixes_after_rebase
76775a5b
ggerganov
commented on 2026-02-17
Update OV in CI, Add OV CI Tests in GH Actions
4c1fdd38
CISC
commented on 2026-02-18
Temp fix for multithreading bug
ae8a140c
Update OV CI, fix review suggestions.
20ecf4b2
Merge pull request #45 from cavusmustafa/tmp_fix_multithread
cb92f777
fix editorconfig-checker, update docs
a8e894d8
Fix tabs to spaces for editorconfig-checker
19e4f31c
fix editorconfig-checker
c6ee7c5f
Update docs
7d4d3113
updated model link to be GGUF model links
ed91be28
Remove GGML_CPU_REPACK=OFF
016aa26b
Merge branch 'master' into dev_backend_openvino
21b796b0
Skip permuted ADD and MUL
214838e8
Removed static variables from utils.cpp
40d2bb23
Removed initializing non-existing variable
41179c09
Remove unused structs
252ef84b
Merge pull request #1 from wine99/remove_static_variables
56e89f8c
Removed static variables from utils.cpp
240692b6
Fix test-backend-ops for OV GPU
046669e7
unify api calling
18f0ad76
Update utils.cpp
2e025bf0
When the dim is dynamic, throw an error, need to is stastic forst
8fae1b9d
Add interface compute_model_outputs(), which get the model output thr…
fe6a7eda
No need to return
603e6d2d
Merge branch 'master' into dev_backend_openvino
43ca96a9
Fix test-backend-ops for OV GPU LNL
d25211e3
Fix test-thread-safety
5f0a68c2
use the shape from infer request of output tensor create to avoid issue
0b32b7fa
fix dynamic output shape issue
183f36fc
Merge pull request #49 from zhaixuejun1993/xuejun/unify-api-get_ov_ou…
bc879022
fix issue for the unused node in tests
198e932f
Rewrite the logistic about model outputs computer, add new API comput…
f274f635
Remove unused lock
e8dc98b9
Merge branch 'dev_backend_openvino' into fix-thread-safety
6e9dc504
Fix test-thread-safety
03d03b88
Add comment
cfb395d4
Fix test-backend-ops for OV GPU LNL
415c9b3e
Merge pull request #46 from cavusmustafa/fix-readme-model-links
ba614249
Update openvino docs
aef9e623
update to OV release version 2026.0
9d3d2c4b
add ci ov-gpu self hosted runner
36ce9143
fix editorconfig
091c58f6
Fix perplexity
a613e8bf
Rewrite the model inputs finding mechanism (#54)
82051a9a
Put the iteration logistic in func
db976265
ggerganov
commented on 2026-03-06
ggerganov
commented on 2026-03-06
Added ggml-ci-intel-openvino-gpu and doc update
42a1cb58
.hpp files converted to .h
6d1f94d8
Merge pull request #57 from cavusmustafa/hpp_to_h
c29ccc45
fix ggml-ci-x64-intel-openvino-gpu
f71cc593
Fix for stateful execution bug in llama-bench
eae534e7
Minor updates after stateful llama-bench fix
c2c4211f
Update ggml/src/ggml-openvino/utils.cpp
7b93c50f
Remove multiple get_shape calls
29c217a8
Bring back mutex into compute
8616f120
Fix VIEW op, which slice the input node
0480c2c6
Added token_len_per_seq existence check before slicing masks and move…
f5304c69
Merge pull request #60 from zhaixuejun1993/xuejun/hot-fix-llama-embed…
481d938b
cavusmustafa
force pushed
to
481d938b
214 days ago
Temp. fix for test requant errors
e646c85d
Merge pull request #62 from zhaixuejun1993/xuejun/fix_issue_key_miss
1cf07165
Merge pull request #58 from cavusmustafa/fix_stateful_state_sync
409cc8eb
Update to OV ggml-ci to low-perf
bb40ee89
ci : temporary disable "test-llama-archs"
0aaf8ab0
CISC
commented on 2026-03-13
ggerganov
force pushed
211 days ago
ci : cache v4 -> v5, checkout v4 -> v6, fix runner tag
e73b4d4a
ggerganov
force pushed
to
e73b4d4a
211 days ago
ggerganov
commented on 2026-03-13
docs : update url
5237965b
ggerganov
approved these changes on 2026-03-13
CISC
approved these changes on 2026-03-13
Fix OV link in docker and Update docs
996b739e
ggerganov
merged
9789c4ec
into master
211 days ago
Login to write a write a comment.
Login via GitHub
Reviewers
CISC
ggerganov
ravi9
danbev
cavusmustafa
yangsu2022
ngxson
Assignees
No one assigned
Labels
documentation
testing
devops
ggml
Milestone
No milestone
Login to write a write a comment.
Login via GitHub