Skill
Ollama structured output pipeline step
Drop-in upgrade for Ollama pipeline steps: pass Pydantic schema as format= to engage XGrammar constrained decoding — 6.4x faster, ~100% parse success, no brace-scanning
Primitives inside (5)
flat-schema-for-small-local-modelscalibrationShape JSON schemas to the local model's size: flatten to <=2 nesting levels (at 3+ levels small models return empty arrays at depth), wrap nested lists in a BaseModel instead of List[List[str]], split schema dicts over ~2KB into sequential extraction calls, and reserve deep/complex schemas for >=12B models (flat 2-level verified fine on 9.4B glm4).
When: Designing the Pydantic/JSON schema for a structured-output step targeting a local model, especially one under 12B parameters.
ollama-format-schema-parse-not-speedcalibrationThe guaranteed win of Ollama format=schema constrained decoding is parse reliability, NOT speed: live verification (Ollama 0.23.2, glm4 9.4B, 2026-07-12) reproduced ~100% parse success but the published 6.4x speedup did not — a warm unconstrained call (0.8s) beat constrained calls (3.3-4.6s).
When: Deciding whether to adopt format=schema on a local Ollama pipeline step, or setting performance expectations for it.
pydantic-v2-validationerror-not-constructiblegotcha-fixNever construct pydantic ValidationError(f"...") yourself — in Pydantic v2 it is not constructible from a string and crashes with TypeError (verified on 2.12.5); raise your own exception class with the caught ValidationError chained via 'from last_err' instead.
When: Writing error handling around Pydantic validation, e.g. an exhausted-retries path that must surface the last validation failure.
thinking-models-under-format-slow-but-cleancalibrationThinking models (qwen3, gpt-oss) work correctly under Ollama format=schema — the reasoning lands in message.thinking while message.content stays clean schema-valid JSON (verified live) — but at ~85s on qwen3:32b; route latency-sensitive structured steps to a small non-thinking model.
When: Choosing the model for a schema-constrained Ollama step when candidates include thinking models.
token0-constraint-suppresses-cotcalibrationConstraining output from token 0 makes the model collapse into the schema early and suppresses deliberative CoT — route ambiguous-evidence verdicts to deferred constraint, keep simple extractions on full constraint for speed.
When: Deciding between whole-output constrained decoding and reason-then-commit for a structured-output task.
Get the whole skill
All 5 primitives of this skill as one package, with the order to apply them.
Buy only the primitives you need
Each primitive is 1 credit (≈ €0.10). Pick them from the list above — the button is next to each one.
Upgrade your own skill
Paste your skill; we pick the 5 primitives from the shelf that fit it best, as one bundle for 5 credits (≈ €0.50).
Upgrade my skillNeighbour skills
Skills whose primitives are closest to this one (bge-m3 similarity):