Skill
Dia dialogue TTS
Generate natural two-speaker dialogue audio from [S1]/[S2]-tagged transcript using Dia-1.6B (Apache 2.0), with non-verbal cues (laughter/sighs) in one local model pass — drop-in upgrade for a two-host podcast-avatar pipeline voice stage.
Primitives inside (9)
debunked-flag-doc-tripwirefailsafeAfter debunking flags, add a selftest that greps the doc's ACTUAL invocation lines (not prose) for the debunked flags — prose may still name them to warn readers, but no runnable example may use them.
When: Any skill doc that once contained wrong command syntax and must not regress.
dia-5-20s-length-windowcalibrationDia inputs under ~5s of audio-equivalent sound unnatural and over ~20s become unnaturally fast (1s of audio is roughly 86 tokens) — chunk long dialogues and estimate duration at ~2.5 words/sec to stay in the window.
When: Sizing transcript chunks before generation.
dia-parenthesized-nonverbal-whitelistgotcha-fixDia non-verbal cues are PARENTHESIZED — (laughs), never [laughs] — and only the ~21 README-listed tags are safe; bracketed or unlisted tags are read as text or cause audio artifacts; use sparingly.
When: Adding laughter/sighs/coughs etc. to a Dia transcript.
dia-speaker-tag-grammargotcha-fixDia transcripts must BEGIN with [S1] and strictly alternate [S1]/[S2]; consecutive same-speaker turns are merged into one tag, and if the first turn is speaker B you swap identities rather than start with [S2].
When: Preparing any transcript for Dia-1.6B dialogue generation (podcast pipelines, two-host scripts).
dia-voice-clone-transcript-prependtool-sequenceDia voice cloning: ONE --audio-prompt clip (5-10s, may contain BOTH voices) plus PREPENDING the clip's transcript to the generation text — the model returns audio only for the new text; two-speaker cloning is one clip whose transcript carries both [S1]/[S2].
When: Getting consistent, specific voices from Dia across a generation.
isolated-venv-per-ml-stackdisciplineGive every heavy ML toolchain — each PyTorch-based tool, each research-grade repo with its own pins (torch builds, numpy<2, patched repos) — its OWN isolated venv (e.g. .venv_tts / .venv_avatar / .venv_sadtalker), created by an idempotent install script and invoked through that venv's interpreter path, so mutually conflicting pins never disturb the system Python/torch or another tool's environment; never pip-install a tool into another tool's or the system's environment.
When: Adding a torch-dependent tool or a research-grade ML repo to a machine that already hosts other torch-dependent tools, pipelines or projects sharing the system interpreter.
silence-detect-speaker-resplittool-sequenceWhen a single-pass multi-speaker generator (dialogue TTS such as Dia) emits ONE mixed track but downstream stages (avatar / lip-sync) consume per-speaker tracks, the re-split is an integration SEAM: declare it BEFORE integrating and test the chosen strategy first — the generator's own timestamps if it exposes any (Dia has no timestamps flag), else `ffmpeg -af silencedetect=noise=-40dB:d=0.3` to find turn boundaries and cut, else skip splitting when the consumer accepts the full mixed track — and log the first real run's outcome; until a strategy has run end-to-end the seam stays tagged UNKNOWN in the claims inventory.
When: A single mixed-speaker audio file must feed per-speaker consumers (lip-sync avatar models, per-host processing); swapping a per-speaker TTS for a single-pass dialogue model in such a pipeline.
upstream-example-params-over-cli-defaultscalibrationA tool's CLI defaults and its authors' example parameters can disagree — Dia's cli.py defaults temperature 1.3/top_p 0.95 but upstream examples prefer 1.8/0.90 (+cfg_scale 4.0 for cloning); surface both and prefer the examples for quality.
When: Choosing generation parameters for a model whose repo ships both a CLI and worked examples.
verify-cli-flags-against-source-not-patternsdisciplineNever document CLI flags inferred from 'README patterns' or plausibility — fetch the actual argparse source (or --help output) and transcribe from it; three inferred Dia flags turned out not to exist at all.
When: Authoring or lifting any skill that documents a third-party tool's command surface.
Get the whole skill
All 9 primitives of this skill as one package, with the order to apply them.
Buy only the primitives you need
Each primitive is 1 credit (≈ €0.10). Pick them from the list above — the button is next to each one.
Upgrade your own skill
Paste your skill; we pick the 5 primitives from the shelf that fit it best, as one bundle for 5 credits (≈ €0.50).
Upgrade my skillNeighbour skills
Skills whose primitives are closest to this one (bge-m3 similarity):