OpenVINO HF Exporter (#47003)
* add openvino exporter
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* style
* more OV end ET fixes
* style
* more fixes and targeted skips
* fixes, skips and reverts
* more reverts
* reverts
* claude review
* Add Kimi-K2.5 exportability and standardize packed-vision attention
Exporter support for Kimi-K2.5 (vision/audio attention registered with the
reshaped vision-attention patch) plus:
- Kimi vision rotary now emits the standard packed (seq, head_dim) cos/sin
instead of the batched (1, seq, head_dim) LLM form, matching every other
vision encoder (no exporter-side normalization needed).
- ExecuTorch: materialise non-contiguous reshapes via a backend-local
_patch_reshape rather than an unconditional clone in the shared vision
patch, so ONNX/OpenVINO/dynamo graphs don't carry the copy.
- is_multimodal short-circuits to False for non-PreTrainedModel inputs.
- vision_utils.get_vision_cu_seqlens gains a merge_temporal option.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* more passing models
* more fixes
* more fixes
* more fixes
* revert ET changes
* revert
* fix
* staging model artifacts
* style
* skip kimi
* multi token export
* fix multi token ssm slow forward
* fully fixed multi token decode
* style
* fixes
* remove
* update
* cleanup
* fix
* fix
* fixes
* fix sdpa
* dont compare padding
* fix
* fix
* revert recurrent gemma
* fix
* Fix five OpenVINO export divergences and drop their skips
- sdpa: zero rows that mask every key. OV returns the uniform average the mask
literally describes, which is what torch's own CPU and SDPBackend.MATH paths
return, but its fused CUDA kernel writes zeros — fixes timesfm, where the
first 12 of 16 patches are entirely padding.
- interpolate: build the antialias weights from tensor ops so the patch covers
extents read out of a tensor. siglip2 takes its target size from
`spatial_shapes`, an unbacked symbol, which the numpy weights could not serve.
- cumsum: promote bool and narrow ints to int64 the way torch does. OV's CumSum
keeps a bool input boolean, so its running sum saturates and OPT's
cumsum-derived `position_ids` collapsed to 0 for every token.
- blt: evaluate the byte-group hash in base-256 limbs. OV's CPU plugin executes
i64 nodes in i32, so `1000000007 ** 2` saturated at INT32_MAX and the hash
read the wrong embedding rows.
- drop `_patch_sliding_window_layer`. The stateful conversion handles an
evicting sliding layer on its own (checked over multi-step decode against
eager for gemma2/gemma3/mistral/cohere2/ministral/phi3, ~1e-7), while the
patch made a prefill-only export hand back an untrimmed cache — fixes
shieldgemma2.
- vit_mae: pin `noise` in the tester so eager and the exported graph mask the
same patches, as the integration tests already do.
Also allocate `StaticIndexedLayer`'s indexer counter at lazy initialization,
which is the first point that knows a device, and drop the hy_v4
`_check_outputs_close` override now that the shared helper ignores padded
positions.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Address review feedback
- docs: order the exporter table the way the examples are ordered.
- exporters: define the cache rule once as `is_cache_class`, so `is_cache_object` and
`register_cache_pytrees_for_model` share it instead of each spelling out the `*Cache`
convention.
- onnx: check `input.size(dim)` in `_patch_chunk` before deriving a chunk size from it.
- openvino: split `export()` into `_fix_exported_program` and `_convert_to_openvino` so it
reads as named stages rather than a run of calls; fold the duplicated port lookup into
`_lookup_by_port` plus a `_port_named` helper; and write the complex RoPE rewrite as
`pairs * cos + rotate(pairs) * sin`, the shape `apply_rotary_pos_emb` has, with a pair-wise
rotation standing in for `rotate_half` since these models interleave rather than split halves.
- fsmt: keep `triu_onnx` importable, now deprecated and delegating to `torch.triu`, which is
what it emulated.
- oneformer: append `spatial_shapes_list` instead of inserting it, so every existing call binds
as it did, and deprecate the `spatial_shapes` kwarg on `get_reference_points` rather than
renaming it outright.
- hunyuan_vl: use `torch_compilable_check` rather than a bare `torch._check`.
- funnel, cache_utils: trim comments down to what isn't already obvious from the code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Unskip most of the OpenVINO ledger
Three exporter fixes, and a way to say "run this, just don't compare it".
- cumsum: pass the axis straight through. The patch intercepted `dim` explicitly, so longt5's
`torch.cumsum(mask, axis=1)` — the numpy-style alias torch accepts — landed in `**kwargs` and
`dim=None` reached the original. Only `dtype` needs intercepting.
- scatter_reduce: handle `amin`, and `sum`/`mean` with `include_self=False`. The last two scatter
onto zeros and scatter ones alongside, so the count both divides the mean and says which
positions were touched at all — a position nothing scatters to keeps `self`. Unskips tapas.
- embedding_bag: decompose the 2-D form into an embedding lookup and a reduction along the bag
axis. OpenVINO's CPU plugin cannot compile `aten._embedding_bag` at all — the graph converts and
then `compile_model` raises `to_shape was called on a dynamic shape`. Unskips kimi_k25, whose
vision stack interpolates its position embeddings that way.
The skip lists grow an `.exactness` suffix: those entries still export and run, and only the
comparison against eager is dropped. That covers models whose outputs are unstable for reasons that
say nothing about the export — vibevoice_asr samples a VAE inside the forward, and the detection
and QA models pick among tied `topk` scores — so a real export break still fails rather than
hiding behind a skip. 14 of those were full skips until now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix
* fix slow ci models
* add ort exatness checks too
* fixes
* skip
* Address review round two
Review fixes: drop the limitations bullet from the docs, note on `stateful` that
the NPU plugin takes no stateful model, keep `_OV_NAME_OK` local to the one
function that uses it, `decomp_table.pop(op, None)` instead of try/except, and
shorten the fully-masked-row comment. `_rotate_pairs` becomes
`_apply_rotary_pos_emb_pairs`, which is what it has done since it was rewritten
around `rotate_half`. In `cache_utils`, `indexer_cumulative_length` starts as
`None` beside `indexer_keys` and the `zeros_like` comment goes back to what it
said before.
Five patches were the same fix written twice: `histc`, `searchsorted`,
`bucketize`, the vmap mask expansion and the non-contiguous `reshape`. They move
to `utils.py` next to the other cross-backend patches, one body each with
stacked `@register_patch` decorators, and every backend ends up registering
exactly what it registered before.
That audit also turned up `_patch_chunk` defined twice in the ONNX exporter, both
against `torch.chunk`, with the same dispatch. The second shadowed the first and
was installed on top of it, so the first only ever ran as the `original` of a
call the second had already declined. Removed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Fold the patches that were the same patch twice
`fft` and `ifft` were the same twiddle matmul with a sign flipped and a divide,
so they share `_dft` now. `bernoulli`'s body was already what the `randn` patch
does, so it registers against that one instead. And both symbolic divisions
promoted their integral operands to f32 the same way, which is `_float_operands`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>