-
Notifications
You must be signed in to change notification settings - Fork 5
Add ONNX export and quantization skill #238
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
2 commits
Select commit
Hold shift + click to select a range
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,396 @@ | ||
| --- | ||
| name: onnx-export-quantization | ||
| description: > | ||
| Use this skill when exporting ONNX models with mobius and quantizing | ||
| them with Olive for deployment. Covers the mobius CLI, EP options, | ||
| INT4 quantization (Q4_K_M and NF4), HuggingFace upload structure, | ||
| GPU-accelerated quantization, common issues, and testing quantized | ||
| models. | ||
| --- | ||
|
|
||
| # Skill: ONNX Export and Quantization | ||
|
|
||
| ## When to use | ||
|
|
||
| Use this skill when: | ||
| - Exporting a model from HuggingFace to ONNX format using `mobius build` | ||
| - Quantizing an ONNX model to INT4 (Q4_K_M or NF4) with Olive | ||
| - Uploading ONNX models to HuggingFace Hub in the standard directory layout | ||
| - Debugging export or quantization failures | ||
| - Choosing between execution provider (EP) variants | ||
|
|
||
| ## Exporting models with `mobius build` | ||
|
|
||
| ### Basic command | ||
|
|
||
| ```bash | ||
| mobius build \ | ||
| --model <hf-model-id> \ | ||
| --dtype <f16|bf16> \ | ||
| --ep <default|cuda|onnx-standard> \ | ||
| --runtime ort-genai \ | ||
| --external-data safetensors \ | ||
| --max-shard-size 5GB \ | ||
| <output-directory>/ | ||
| ``` | ||
|
|
||
| ### Flag reference | ||
|
|
||
| | Flag | Description | | ||
| |------|-------------| | ||
| | `--model <id>` | HuggingFace model ID (e.g. `google/gemma-4-27b-it`) | | ||
| | `--dtype <f16\|bf16>` | Model precision — `f16` (float16) or `bf16` (bfloat16) | | ||
| | `--optimize [RULES]` | Apply mobius rewrite rules after building (e.g. `group_query_attention`, `packed_attention`, `skip_norm`). Use without value for all rules, or specify comma-separated names. Not needed for basic exports. | | ||
| | `--ep <variant>` | Execution provider variant (see below) | | ||
| | `--runtime ort-genai` | Generate `genai_config.json` and copy tokenizer files for ORT GenAI runtime | | ||
| | `--external-data safetensors` | Store weights externally in safetensors format | | ||
| | `--max-shard-size 5GB` | Split external data into shards ≤ 5GB | | ||
|
|
||
| ### Execution provider (EP) variants | ||
|
|
||
| Build separate ONNX models per EP because each applies different graph | ||
| rewrites and fused ops: | ||
|
|
||
| | EP | Flag | When to use | | ||
| |----|------|-------------| | ||
| | `default` | `--ep default` | Portable ONNX — no vendor-specific fusions. Compatible with all execution providers and runtimes. This is the default if `--ep` is omitted. | | ||
| | `cuda` | `--ep cuda` | NVIDIA GPU inference. Emits `com.microsoft` fused ops (GroupQueryAttention, MoE, etc.) for maximum CUDA performance. | | ||
| | `onnx-standard` | `--ep onnx-standard` | Strict ONNX-only — inlines all custom-domain functions into standard ONNX ops. Use when targeting runtimes that don't support `com.microsoft` ops. | | ||
|
|
||
| Other EPs are available (`cpu`, `dml`, `webgpu`, `trt-rtx`). Run | ||
| `mobius list eps` to see all options. | ||
|
|
||
| **Typical export matrix:** Build each dtype × EP combination: | ||
|
|
||
| ```bash | ||
| for dtype in f16 bf16; do | ||
| for ep in default cuda onnx-standard; do | ||
| mobius build --model google/gemma-4-12b-it \ | ||
| --dtype $dtype --ep $ep \ | ||
| --runtime ort-genai \ | ||
| --external-data safetensors --max-shard-size 5GB \ | ||
| output/${dtype}/${ep}/ | ||
| done | ||
| done | ||
| ``` | ||
|
|
||
| ### Multi-model outputs | ||
|
|
||
| For multimodal models (VLMs, audio-language), `mobius build` produces | ||
| multiple sub-models: | ||
|
|
||
| ``` | ||
| output/ | ||
| ├── decoder/ # Text decoder | ||
| │ ├── model.onnx | ||
| │ └── model.onnx.data.safetensors | ||
| ├── embedding/ # Embedding model | ||
| │ ├── model.onnx | ||
| │ └── model.onnx.data.safetensors | ||
| ├── vision_encoder/ # Vision encoder (VLMs) | ||
| │ ├── model.onnx | ||
| │ └── model.onnx.data.safetensors | ||
| ├── audio_encoder/ # Audio encoder (ALMs) | ||
| │ ├── model.onnx | ||
| │ └── model.onnx.data.safetensors | ||
| └── genai_config.json | ||
| ``` | ||
|
|
||
| ## Quantization with Olive | ||
|
|
||
| ### Installation | ||
|
|
||
| Olive with ONNX quantization support (install from PR if needed for | ||
| latest features): | ||
|
|
||
| ```bash | ||
| pip install olive-ai | ||
| # Or from a specific PR for bleeding-edge features: | ||
| pip install git+https://github.com/microsoft/Olive.git@refs/pull/2406/head | ||
| ``` | ||
|
|
||
| For GPU-accelerated quantization (highly recommended for large models): | ||
|
|
||
| ```bash | ||
| pip install cupy-cuda12x | ||
| ``` | ||
|
|
||
| ### Q4_K_M quantization (k-quant) | ||
|
|
||
| K-quant quantization uses mixed block sizes with importance-based bit | ||
| allocation. Q4_K_M is a good balance of quality and size. | ||
|
|
||
| The repo uses Olive's config-driven `olive.run()` pattern (see | ||
| `examples/olive/` for working examples). A typical Olive config for | ||
| k-quant quantization: | ||
|
|
||
| ```json | ||
| { | ||
| "input_model": { "type": "OnnxModel", "model_path": "decoder/model.onnx" }, | ||
| "passes": { | ||
| "kquant": { | ||
| "type": "OnnxKQuantQuantization", | ||
| "bits": 4, | ||
| "block_size": 32 | ||
| } | ||
| }, | ||
| "output_dir": "output/Q4_K_M/default/decoder" | ||
| } | ||
| ``` | ||
|
|
||
| ```bash | ||
| olive run --config kquant_config.json | ||
| ``` | ||
|
|
||
| ### NF4 quantization (4-bit NormalFloat) | ||
|
|
||
| NF4 uses a normal-distribution-optimized 4-bit format. Fast native C++ | ||
| implementation — no GPU needed. | ||
|
|
||
| ```json | ||
| { | ||
| "input_model": { "type": "OnnxModel", "model_path": "decoder/model.onnx" }, | ||
| "passes": { | ||
| "nf4": { | ||
| "type": "OnnxBnb4Quantization", | ||
| "precision": "nf4" | ||
| } | ||
| }, | ||
| "output_dir": "output/NF4/default/decoder" | ||
| } | ||
| ``` | ||
|
|
||
| > See `examples/olive/ministral-3-3b-vlm/` for a complete working | ||
| > example that combines mobius export with Olive quantization. | ||
|
|
||
| ### GPU acceleration with cupy | ||
|
|
||
| Installing `cupy-cuda12x` gives a **19–51x speedup** for k-quant | ||
| quantization: | ||
|
|
||
| | Method | CPU time per matrix | GPU time per matrix | Speedup | | ||
| |--------|--------------------|--------------------|---------| | ||
| | K-quant (Q4_K_M) | 3–27s | 0.17–0.52s | 19–51x | | ||
| | NF4 | 42ms for 67M params | N/A (C++ native) | Already fast | | ||
|
|
||
| ```bash | ||
| # Install cupy for CUDA 12.x | ||
| pip install cupy-cuda12x | ||
|
|
||
| # Olive auto-detects cupy and uses GPU when available | ||
| ``` | ||
|
|
||
| ### Quantizing multi-model exports | ||
|
|
||
| Quantize each sub-model independently. Typically only the decoder is | ||
| quantized (it has the most parameters). Copy all other files needed | ||
| for a complete ORT GenAI package: | ||
|
|
||
| ```bash | ||
| # Quantize decoder only (largest model) | ||
| olive run --config kquant_decoder.json | ||
|
|
||
| # Copy other sub-models as-is (already small) | ||
| cp -r output/f16/default/embedding/ output/Q4_K_M/default/embedding/ | ||
| cp -r output/f16/default/vision_encoder/ output/Q4_K_M/default/vision_encoder/ | ||
|
|
||
| # IMPORTANT: Copy config, tokenizer, and processor files too | ||
| cp output/f16/default/genai_config.json output/Q4_K_M/default/ | ||
| cp output/f16/default/tokenizer* output/Q4_K_M/default/ | ||
| cp output/f16/default/image_processor.json output/Q4_K_M/default/ 2>/dev/null | ||
| cp output/f16/default/audio_processor.json output/Q4_K_M/default/ 2>/dev/null | ||
| ``` | ||
|
|
||
| Without the tokenizer and processor config files, ORT GenAI will fail | ||
| to load the model. | ||
|
|
||
| ## HuggingFace upload structure | ||
|
|
||
| ### Standard directory layout | ||
|
|
||
| ``` | ||
| <org>/<model>-onnx/ | ||
| ├── f16/ | ||
| │ ├── default/ # Portable ONNX (no vendor fusions) | ||
| │ ├── cuda/ # CUDA EP (fused ops) | ||
| │ └── onnx-standard/ # Strict ONNX-only (inlined functions) | ||
| ├── bf16/ | ||
| │ ├── default/ | ||
| │ ├── cuda/ | ||
| │ └── onnx-standard/ | ||
| ├── Q4_K_M/ | ||
| │ └── default/ # Quantized models typically CPU-only | ||
| └── NF4/ | ||
| └── default/ | ||
| ``` | ||
|
|
||
| Each EP directory contains the full model structure (decoder/, | ||
| embedding/, vision_encoder/, audio_encoder/ as applicable) plus | ||
| `genai_config.json`. | ||
|
|
||
| ### Upload with huggingface_hub | ||
|
|
||
| ```python | ||
| from huggingface_hub import HfApi | ||
|
|
||
| api = HfApi() | ||
| api.upload_folder( | ||
| folder_path="output/f16/default", | ||
| path_in_repo="f16/default", | ||
| repo_id="org/model-onnx", | ||
| repo_type="model", | ||
| ) | ||
| ``` | ||
|
|
||
| ### Verify uploads | ||
|
|
||
| After uploading, verify all shards are present. Incomplete uploads are | ||
| a common issue with large models: | ||
|
|
||
| ```python | ||
| from huggingface_hub import HfApi | ||
|
|
||
| api = HfApi() | ||
| files = api.list_repo_files("org/model-onnx") | ||
| # Check that all expected .safetensors shards exist | ||
| for variant in ["f16/default", "f16/cuda", "bf16/default"]: | ||
| shards = [f for f in files if f.startswith(variant) and f.endswith(".safetensors")] | ||
| print(f"{variant}: {len(shards)} shards") | ||
| ``` | ||
|
|
||
| ## Common issues and fixes | ||
|
|
||
| ### 1. MoE expert weight mapping | ||
|
|
||
| Models with Mixture-of-Experts (e.g. Gemma4 26b-a4b) may need expert | ||
| weight remapping in `preprocess_weights()`. HuggingFace stores experts | ||
| as 3D tensors (`experts.gate_up_proj [E, 2*inter, H]`) that must be | ||
| mapped to the fused MoE op's parameter names (`fc1_experts_weights`, | ||
| `fc2_experts_weights`). | ||
|
|
||
| **Symptom:** Weight loading errors or incorrect MoE outputs. | ||
|
|
||
| **Fix:** Check the model's `preprocess_weights()` maps HF expert weight | ||
| names to the ONNX parameter names. See the `moe-models` skill for | ||
| the pattern. | ||
|
|
||
| ### 2. Hybrid attention v_proj shape mismatches | ||
|
|
||
| Models with hybrid attention (e.g. Gemma4 31b with different `head_dim` | ||
| for local vs global attention layers) may have shape mismatches in | ||
| value projections. | ||
|
|
||
| **Symptom:** Shape errors during weight loading or forward pass. | ||
|
|
||
| **Fix:** Ensure `v_proj` dimensions account for per-layer head | ||
| configurations. Check `num_global_key_value_heads` vs | ||
| `num_key_value_heads` in the config. | ||
|
|
||
| ### 3. CUDA GQA head_dim limitations | ||
|
|
||
| Older versions of ORT had a limitation where `head_dim > 256` would fail | ||
| with the CUDA GroupQueryAttention kernel. | ||
|
|
||
| **Symptom:** CUDA runtime error during inference with large head | ||
| dimensions. | ||
|
|
||
| **Status:** This limitation has been removed in recent ORT versions. | ||
| If using an older ORT build, fall back to `--ep default` or | ||
| `--ep onnx-standard`. | ||
|
|
||
| ### 4. Incomplete uploads | ||
|
|
||
| Large models with many shards can have incomplete uploads to HuggingFace | ||
| Hub, especially on unstable connections. | ||
|
|
||
| **Symptom:** Model fails to load with file-not-found errors for | ||
| specific shard files. | ||
|
|
||
| **Fix:** Verify all shards are present after upload (see the verify | ||
| script above). Re-upload missing shards with `api.upload_file()`. | ||
|
|
||
| ### 5. BF16 type mismatches | ||
|
|
||
| Some components may produce FP32 outputs when the model is built in | ||
| BF16, causing type mismatch errors in ORT. | ||
|
|
||
| **Symptom:** `Type Error: Type parameter (T) bound to different types | ||
| (tensor(bfloat16) and tensor(float))`. | ||
|
|
||
| **Fix:** Check for constants, initializers, or norm layers that stay | ||
| FP32 when the model is BF16. Add `op.CastLike(result, input)` to | ||
| ensure dtype consistency. See the `reusable-components` skill's | ||
| section on precision behaviour. | ||
|
|
||
| ## Testing quantized models | ||
|
|
||
| ### L4: Golden data generation | ||
|
|
||
| Generate reference outputs from the full-precision HuggingFace model | ||
| using the golden data generation script: | ||
|
|
||
| ```bash | ||
| # Generate golden files for all test cases | ||
| python scripts/generate_golden.py | ||
|
|
||
| # Generate for a specific task type | ||
| python scripts/generate_golden.py --task-type causal-lm | ||
|
|
||
| # Generate for a single test case | ||
| python scripts/generate_golden.py --case testdata/cases/causal-lm/gpt2.yaml | ||
|
|
||
| # Use GPU for large models | ||
| python scripts/generate_golden.py --device cuda | ||
| ``` | ||
|
|
||
| Golden reference files are stored in `testdata/golden/` as JSON. Use | ||
| `compare_golden()` from `mobius._testing.parity` to compare model | ||
| outputs against the reference: | ||
|
|
||
| ```python | ||
| from mobius._testing.parity import compare_golden | ||
|
|
||
| compare_golden( | ||
| model_output=output_logits, | ||
| golden_path="testdata/golden/causal-lm/my_model.json", | ||
| ) | ||
|
justinchuby marked this conversation as resolved.
|
||
| ``` | ||
|
|
||
| ### L5: End-to-end smoke test | ||
|
|
||
| Run inference with the quantized model through ORT GenAI: | ||
|
|
||
| ```python | ||
| import onnxruntime_genai as og | ||
|
|
||
| model = og.Model("output/Q4_K_M/default/") | ||
| tokenizer = og.Tokenizer(model) | ||
| params = og.GeneratorParams(model) | ||
| params.set_search_options(max_length=50, do_sample=False) | ||
| params.input_ids = tokenizer.encode("Hello, world!") | ||
|
|
||
| output_ids = model.generate(params) | ||
| print(tokenizer.decode(output_ids[0])) | ||
| ``` | ||
|
|
||
| ### Numerical parity verification | ||
|
|
||
| Quantized models will have some numerical divergence from the | ||
| full-precision model. Expected tolerances: | ||
|
|
||
| | Quantization | Typical divergence | Notes | | ||
| |-------------|-------------------|-------| | ||
| | Q4_K_M | Moderate | Top-1 token agreement ~95%+ for coherent text | | ||
| | NF4 | Moderate | Similar to Q4_K_M | | ||
| | F16 (no quant) | Minimal | Should match BF16 closely | | ||
|
|
||
| Verify that generated text is coherent and semantically correct rather | ||
| than requiring exact numerical matches. | ||
|
|
||
| ## Cross-references | ||
|
|
||
| - **Adding models:** `.agents/skills/adding-a-new-model/SKILL.md` | ||
| - **MoE weights:** `.agents/skills/moe-models/SKILL.md` | ||
| - **Component precision:** `.agents/skills/reusable-components/SKILL.md` | ||
| - **ORT GenAI config:** `.agents/skills/ort-genai-config/SKILL.md` | ||
| - **Quality checklist:** `.agents/skills/quality-checklist/SKILL.md` | ||
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.