transformers
fd290dc1 - OpenVINO HF Exporter (#47003)

Commit
17 days ago
OpenVINO HF Exporter (#47003) * add openvino exporter Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * style * more OV end ET fixes * style * more fixes and targeted skips * fixes, skips and reverts * more reverts * reverts * claude review * Add Kimi-K2.5 exportability and standardize packed-vision attention Exporter support for Kimi-K2.5 (vision/audio attention registered with the reshaped vision-attention patch) plus: - Kimi vision rotary now emits the standard packed (seq, head_dim) cos/sin instead of the batched (1, seq, head_dim) LLM form, matching every other vision encoder (no exporter-side normalization needed). - ExecuTorch: materialise non-contiguous reshapes via a backend-local _patch_reshape rather than an unconditional clone in the shared vision patch, so ONNX/OpenVINO/dynamo graphs don't carry the copy. - is_multimodal short-circuits to False for non-PreTrainedModel inputs. - vision_utils.get_vision_cu_seqlens gains a merge_temporal option. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * more passing models * more fixes * more fixes * more fixes * revert ET changes * revert * fix * staging model artifacts * style * skip kimi * multi token export * fix multi token ssm slow forward * fully fixed multi token decode * style * fixes * remove * update * cleanup * fix * fix * fixes * fix sdpa * dont compare padding * fix * fix * revert recurrent gemma * fix * Fix five OpenVINO export divergences and drop their skips - sdpa: zero rows that mask every key. OV returns the uniform average the mask literally describes, which is what torch's own CPU and SDPBackend.MATH paths return, but its fused CUDA kernel writes zeros — fixes timesfm, where the first 12 of 16 patches are entirely padding. - interpolate: build the antialias weights from tensor ops so the patch covers extents read out of a tensor. siglip2 takes its target size from `spatial_shapes`, an unbacked symbol, which the numpy weights could not serve. - cumsum: promote bool and narrow ints to int64 the way torch does. OV's CumSum keeps a bool input boolean, so its running sum saturates and OPT's cumsum-derived `position_ids` collapsed to 0 for every token. - blt: evaluate the byte-group hash in base-256 limbs. OV's CPU plugin executes i64 nodes in i32, so `1000000007 ** 2` saturated at INT32_MAX and the hash read the wrong embedding rows. - drop `_patch_sliding_window_layer`. The stateful conversion handles an evicting sliding layer on its own (checked over multi-step decode against eager for gemma2/gemma3/mistral/cohere2/ministral/phi3, ~1e-7), while the patch made a prefill-only export hand back an untrimmed cache — fixes shieldgemma2. - vit_mae: pin `noise` in the tester so eager and the exported graph mask the same patches, as the integration tests already do. Also allocate `StaticIndexedLayer`'s indexer counter at lazy initialization, which is the first point that knows a device, and drop the hy_v4 `_check_outputs_close` override now that the shared helper ignores padded positions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Address review feedback - docs: order the exporter table the way the examples are ordered. - exporters: define the cache rule once as `is_cache_class`, so `is_cache_object` and `register_cache_pytrees_for_model` share it instead of each spelling out the `*Cache` convention. - onnx: check `input.size(dim)` in `_patch_chunk` before deriving a chunk size from it. - openvino: split `export()` into `_fix_exported_program` and `_convert_to_openvino` so it reads as named stages rather than a run of calls; fold the duplicated port lookup into `_lookup_by_port` plus a `_port_named` helper; and write the complex RoPE rewrite as `pairs * cos + rotate(pairs) * sin`, the shape `apply_rotary_pos_emb` has, with a pair-wise rotation standing in for `rotate_half` since these models interleave rather than split halves. - fsmt: keep `triu_onnx` importable, now deprecated and delegating to `torch.triu`, which is what it emulated. - oneformer: append `spatial_shapes_list` instead of inserting it, so every existing call binds as it did, and deprecate the `spatial_shapes` kwarg on `get_reference_points` rather than renaming it outright. - hunyuan_vl: use `torch_compilable_check` rather than a bare `torch._check`. - funnel, cache_utils: trim comments down to what isn't already obvious from the code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Unskip most of the OpenVINO ledger Three exporter fixes, and a way to say "run this, just don't compare it". - cumsum: pass the axis straight through. The patch intercepted `dim` explicitly, so longt5's `torch.cumsum(mask, axis=1)` — the numpy-style alias torch accepts — landed in `**kwargs` and `dim=None` reached the original. Only `dtype` needs intercepting. - scatter_reduce: handle `amin`, and `sum`/`mean` with `include_self=False`. The last two scatter onto zeros and scatter ones alongside, so the count both divides the mean and says which positions were touched at all — a position nothing scatters to keeps `self`. Unskips tapas. - embedding_bag: decompose the 2-D form into an embedding lookup and a reduction along the bag axis. OpenVINO's CPU plugin cannot compile `aten._embedding_bag` at all — the graph converts and then `compile_model` raises `to_shape was called on a dynamic shape`. Unskips kimi_k25, whose vision stack interpolates its position embeddings that way. The skip lists grow an `.exactness` suffix: those entries still export and run, and only the comparison against eager is dropped. That covers models whose outputs are unstable for reasons that say nothing about the export — vibevoice_asr samples a VAE inside the forward, and the detection and QA models pick among tied `topk` scores — so a real export break still fails rather than hiding behind a skip. 14 of those were full skips until now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix * fix slow ci models * add ort exatness checks too * fixes * skip * Address review round two Review fixes: drop the limitations bullet from the docs, note on `stateful` that the NPU plugin takes no stateful model, keep `_OV_NAME_OK` local to the one function that uses it, `decomp_table.pop(op, None)` instead of try/except, and shorten the fully-masked-row comment. `_rotate_pairs` becomes `_apply_rotary_pos_emb_pairs`, which is what it has done since it was rewritten around `rotate_half`. In `cache_utils`, `indexer_cumulative_length` starts as `None` beside `indexer_keys` and the `zeros_like` comment goes back to what it said before. Five patches were the same fix written twice: `histc`, `searchsorted`, `bucketize`, the vmap mask expansion and the non-contiguous `reshape`. They move to `utils.py` next to the other cross-backend patches, one body each with stacked `@register_patch` decorators, and every backend ends up registering exactly what it registered before. That audit also turned up `_patch_chunk` defined twice in the ONNX exporter, both against `torch.chunk`, with the same dispatch. The second shadowed the first and was installed on top of it, so the first only ever ran as the `original` of a call the second had already declined. Removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Fold the patches that were the same patch twice `fft` and `ifft` were the same twiddle matmul with a sign flipped and a divide, so they share `_dft` now. `bernoulli`'s body was already what the `randn` patch does, so it registers against that one instead. And both symbolic divisions promoted their integral operands to f32 the same way, which is `_float_operands`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Parents
Loading