Skill
GBNF lazy grammar CoT constrained
GBNF lazy-activation grammar in llama.cpp: model reasons freely in CoT until trigger token, then JSON schema enforced — reason-then-commit pattern for evidence assessment
Primitives inside (11)
brew-bottle-missing-validator-samplescalibrationThe Homebrew llama.cpp bottle (b9860) ships NEITHER llama-gbnf-validator NOR the grammars/ sample dir the README references - validate grammars via a live request (or external tool) and INLINE grammars in docs instead of pointing at sample files.
When: Writing docs/automation that assume upstream repo artifacts exist in the installed bottle.
gbnf-performance-rulesgotcha-fixIn GBNF, chained optionals (x? x? x?) explode the FSM — use bounded repetition x{0,N} (max 2000 repetitions per rule in llama.cpp), prefer negated character ranges over enumerations, and keep grammars shallow.
When: Hand-writing or debugging a GBNF grammar for llama.cpp constrained decoding.
grammar-self-consistency-static-checkfailsafeBefore shipping a grammar recipe, statically verify every referenced rule is DEFINED and the root accepts the trigger literal (strip string literals and char classes, diff referenced-vs-defined rule names) - catching the two crash classes without starting a server.
When: Grammars embedded in docs/selftests that must not rot as they are edited.
json-schema-param-incompatible-with-lazygotcha-fixThe json_schema convenience param auto-converts to a grammar that starts at the JSON object and cannot accept a word trigger - combining it with grammar_lazy + word triggers 500s; hand-write the trigger-accepting GBNF, or keep json_schema for the constrained call of a two-call flow.
When: Wanting typed-schema convenience together with lazy activation.
lazy-grammar-reason-then-committool-sequenceKeep the grammar INACTIVE while the model reasons in free prose, then activate JSON-schema enforcement only after a trigger token (<answer>) — deliberation before lock-in with a guaranteed-valid structured verdict in one generation.
When: A verdict over ambiguous or conflicting evidence where forcing the schema from token 0 would suppress the deliberative reasoning.
pattern-triggers-unreliable-b9860calibrationPattern triggers (types 2/3) are unreliable for reason-then-commit in b9860: a bare literal pattern succeeds but the tail is NOT strictly enforced (schema-inexact keys emitted), and a full-match capture pattern 500s - stick to type-1 word triggers.
When: Choosing a trigger type for lazy grammars on the installed build.
silently-ignored-request-param-smoke-testgotcha-fixAn unrecognized request parameter can be SILENTLY dropped by a server - grammar_trigger_tokens returned HTTP 200 with fully UNCONSTRAINED output (no error, schema violated) - so never trust a constraint by request success: smoke-test that the output tail actually parses/validates against the schema.
When: Any API where a misspelled/renamed param means a guarantee you rely on silently does not engage.
token0-constraint-suppresses-cotcalibrationConstraining output from token 0 makes the model collapse into the schema early and suppresses deliberative CoT — route ambiguous-evidence verdicts to deferred constraint, keep simple extractions on full constraint for speed.
When: Deciding between whole-output constrained decoding and reason-then-commit for a structured-output task.
trigger-text-fed-to-grammar-rootgotcha-fixWhen a lazy word/pattern trigger fires, the trigger text itself is FED TO THE GRAMMAR - a root starting at the JSON object 500s with 'Unexpected empty grammar stack after accepting piece'; the root must accept the trigger first: root ::= '<answer>' ws object.
When: Writing the GBNF for any lazy-triggered grammar.
two-call-cot-then-constrained-fallbacktool-sequenceWhen lazy grammar is not exposed, split into two calls: call 1 unconstrained with stop=[trigger] to capture CoT, call 2 re-feeding prompt+CoT+trigger under full grammar constraint — functionally equivalent at the cost of a second short pass.
When: Reason-then-commit is needed on a stack lacking grammar_lazy (llama-cpp-python as of 2026-07, Ollama).
word-trigger-multi-token-okcalibrationSingle-token status is only REQUIRED for type-0 (token-id) triggers - a type-1 WORD trigger fires correctly even when the tokenizer splits it (verified on gemma-2b where <answer> is multi-token); single-token remains preferable (preserved-token dylib warning), so check per model.
When: Choosing trigger text across models with different tokenizers.
Get the whole skill
All 11 primitives of this skill as one package, with the order to apply them.
Buy only the primitives you need
Each primitive is 1 credit (≈ €0.10). Pick them from the list above — the button is next to each one.
Upgrade your own skill
Paste your skill; we pick the 5 primitives from the shelf that fit it best, as one bundle for 5 credits (≈ €0.50).
Upgrade my skillNeighbour skills
Skills whose primitives are closest to this one (bge-m3 similarity):