Skill
MLX lm fuse and GGUF export
Fuse MLX LoRA adapters into base model and export to GGUF for Ollama import; covers --de-quantize gotcha and chat-template matching requirement
Primitives inside (7)
finetune-chat-template-exact-matchgotcha-fixA fine-tuned GGUF that loads fine but generates garbage/repetition almost always has a chat-template mismatch — the Ollama Modelfile TEMPLATE must be character-for-character the template the training data used; prevent the whole class by using tokenizer.apply_chat_template() at data-prep time.
When: A fine-tuned model imported into Ollama produces repetition, wrong format, or refusals even though the GGUF loads without error.
mlx-adapter-format-preflightfailsafeBefore fusing, check the adapter dir contains adapter_config.json + adapters.safetensors — mlx-lm 0.31.x does not read legacy .npz adapters from older mlx_lm.lora runs.
When: Pre-flight before mlx_lm.fuse on an adapters/ directory that may come from an older mlx_lm.lora run.
mlx-export-gguf-f16-only-quantize-at-importcalibrationThe mlx_lm --export-gguf path emits f16 ONLY (its GGMLFileType has a single member) — to serve q4_K_M, quantize at Ollama import time (ollama create -q q4_K_M, verified on ollama 0.23.2) or with llama.cpp llama-quantize; ~50% disk/memory saved vs f16 with minimal quality loss.
When: The exported GGUF is huge, or inference after ollama create is very slow / memory-heavy.
mlx-export-gguf-family-allowlistgotcha-fixmlx_lm.fuse --export-gguf works ONLY for model_type llama/mixtral/mistral — any other family (Qwen, Phi, Gemma, ...) hard-raises 'ValueError: Model type X not supported for GGUF conversion'; route those through fuse-to-HF + llama.cpp convert_hf_to_gguf.py.
When: Choosing the GGUF export path after MLX LoRA fine-tuning.
mlx-fuse-dequantize-flag-spellinggotcha-fixThe mlx_lm.fuse de-quantization flag is spelled --dequantize (no inner hyphen) — --de-quantize is rejected with 'error: unrecognized arguments' on mlx-lm 0.31.3.
When: Running mlx_lm.fuse over a quantized base and adding the de-quantization flag.
mlx-fuse-gguf-output-inside-savepathgotcha-fixmlx_lm.fuse --export-gguf writes the GGUF INSIDE the save-path — default fused_model/ggml-model-f16.gguf, not the cwd; 'GGUF not found' after a successful run usually means looking in the wrong place.
When: Locating and verifying the export artifact after an apparently successful fuse run.
qlora-base-fuse-requires-dequantizegotcha-fixIf the base model was 4-bit quantized (QLoRA path via mlx_lm.convert -q), mlx_lm.fuse --export-gguf fails with 'NotImplementedError: Conversion of quantized models is not yet supported' — pass --dequantize so the merge restores full precision and strips the quantization config entry.
When: Exporting GGUF (or fusing to HF format for llama.cpp conversion) after a QLoRA run whose base was 4-bit quantized.
Get the whole skill
All 7 primitives of this skill as one package, with the order to apply them.
Buy only the primitives you need
Each primitive is 1 credit (≈ €0.10). Pick them from the list above — the button is next to each one.
Upgrade your own skill
Paste your skill; we pick the 5 primitives from the shelf that fit it best, as one bundle for 5 credits (≈ €0.50).
Upgrade my skillNeighbour skills
Skills whose primitives are closest to this one (bge-m3 similarity):