Skip to content

Skills ​

videoclaw ships a curated library of skills — reusable, agent-invokable workflows that either produce a video (the video category) or orchestrate the work around it (the workflow category). This doc is the comprehensive per-skill reference. For the machine-readable index see skills/catalog.json; for the full how-to of any individual skill, follow the linked SKILL.md for that skill.


Ecosystem map ​

Skills ecosystem map showing video-framework and brand-presenter as canonical entry points with specialist children, plus a grid of workflow skills

How skills relate ​

The library is not a flat bag of equally-preferred entry points. It uses a small hierarchy:

RoleExamplesWhen you reach for it
Canonical entryvideo-framework, brand-presenterGeneric or unspecified video request — the entry skill routes into a specialist.
Specialistvideo-storyboard, video-clone-ad, movie-director, video-post, ...The mode is clearly known up front (e.g. "clone this ad", "storyboard these 6 scenes").
Compatibility aliasdavendra-presenter, david-sales-presenter, nex-presenter, buntyPersonal/brand presets that exist for discoverability — they all delegate into brand-presenter.
Workflowdeepsearch, ai-slop-cleaner, web-clone, ...Exploration, code-cleanup, and web-clone utilities — independent of any one production mode.

Rule of thumb: start at a canonical entry, specialize only when the mode is clearly known, and treat aliases as discovery handles rather than first-choice workflows.


Skill index ​

GroupSkillStatusOne-liner
🎯 Canonicalvideo-frameworkimportedvclaw-native front door that routes across copy/create/narrated/presentation/long-form/film/UGC.
🎯 Canonicalbrand-presenternative genericGeneric narrated presenter-video workflow over a branded host profile.
🎬 Videogarden-daysnative genericBaby Ram & Sita three-lane series factory: writers → GB identity images → Runway/Higgsfield/Veo queues → automated finish.
🎬 Videodance-breakdownnative genericReference dance video → measured per-shot choreography breakdown (moves sentence for sung windows, timecoded blocks for no-voice windows).
🎬 Videomatch-highlights-localnative genericCricket highlights read off the broadcast scoreboard overlay — one ffmpeg pass, template OCR, a monotonic ball-count prior; no model, no network, no spend.
🎬 Videovox-explainernative genericVox-inspired collage explainers from an approved narration and real archival photos — freely licensed only, object-capable cutouts, a contact sheet approved before anything renders; no generation, no spend.
🎬 Videoad-intelnative genericCompetitor ads torn down with our own tools: the public Meta Ad Library via the Higgsfield browser session (or one public video URL) → whisper → Gemini → per-ad artifacts, a brand teardown and a review page; no third-party API.
🎬 Videovideo-storyboardnative clean-roomBrief or clone plan → scene-by-scene storyboard artifact.
🎬 Videovideo-analyze-templatenative clean-roomReference video → reusable template packet.
🎬 Videovideo-clone-adnative clean-roomSaved template → new product/brand using clone-execute.
🎬 Videovideo-thumbnail-labnative clean-roomFinal render → thumbnail + platform variants.
🎬 Videomovie-directorimportedShort-film production across 12 genres with interview/auto/hybrid entry modes.
🎬 Videovideo-replicatorimported (deep-surface)7-mode professional pipeline: COPY/CREATE/NARRATED/PRESENTATION/LONG-FORM/FILM/UGC.
🎬 Videovideo-postimportedPost-render verify, social variants, thumbnails, archival.
🎬 Videohiggsfield-generateexternal bridgeHiggsfield CLI bridge for Marketing Studio, product photoshoots, Soul IDs, and virality scoring.
🎭 Castcharacter-creatorimportedCreate Go Bananas characters with profile + multi-view reference sheets.
🎭 Castcharacter-libraryimportedAudit, list, patch, and delete entries in the shared Go Bananas library.
🎞️ Promptsseedance-promptsimportedBrowse and apply the clean-room Seedance prompt reference library.
🎞️ Promptsmulti-shot-promptnative clean-roomReference image or storyboard scene → timed multi-shot cinematic prompt sequence, validated against provider-aware videoclaw presets.
🎞️ Promptsmographnative clean-roomStyle-locked motion graphics: one motion-sheet style authority + action-only motion-pack blocks, rendered through the batch queue with zero style drift.
📺 Audioyoutube-audioimportedDownload audio (MP3) or video (MP4) from YouTube using yt-dlp + FFmpeg.
📣 UGCugcimportedBelief-driven UGC campaign generator (E5 method) with multi-video output.
🎤 Aliasesdavendra-presenter · david-sales-presenter · nex-presenter · buntyaliasesAll delegate into brand-presenter with a personal/brand profile.
⚙️ WorkflowdeepsearchimportedThorough codebase search.
⚙️ Workflowai-slop-cleaner · web-cloneimportedOperational utilities (anti-slop cleanup, website cloning).
🎬 Video3d-animation-shortnative genericA one-line story idea → a finished stylized 3D animated short: story spine, labelled cast, no-people plates, a per-second shot table gated by six automated checks before anything renders.
🎞️ Promptsai-directornative genericAI Animation Director: locks a project-level visual blueprint (identity, colour system, lighting grammar, camera language) via director-blueprint + cinema-profile.
🎞️ Promptsai-film-directornative genericPlatform-aware scene director: a minimal idea (plus character / first-frame / last-frame references) → a complete production prompt for Seedance or Google Veo, with continuous multi-scene handoff.
🎞️ Promptsai-filmmakingnative genericProduction-grade prompt chain the videoclaw way: character sheet → storyboard grid → Seedance packets over filmmaking-prompts, storyboard-grid, prompt-lint.
🎬 Videobrand-agencynative genericFull creative-agency brand launch: brand-extract + brand-definition seed several films and a brand kit from a client website.
🎬 Videobrand-explainernativeNarrated 60–70 s brand explainer films in the house motion-graphics style: one shared engine + a .brand data pack per film.
🎬 Videocharacter-adnative genericA short character-led ad the native Google Veo way: a registered Flow character shows off a product via flow-r2v + stitch-ad.
🎯 CanonicalvideoclawaliasPersona alias of concierge — the same guided front door; /videoclaw.
🎯 Canonicalconciergenative genericThe front door for humans, speaking as VideoClaw: a guided menu from idea to finished video, plan → preview → spend, never spending without a go-ahead.
⚙️ WorkflowgraphifyimportedTurns code, docs and media into a persistent knowledge graph (graphify-out/) you query instead of grepping.
⚙️ Workflowimprovement-runnative genericWhole-codebase improvement pass: parallel audit → adversarial verification → plan gate → batched PRs with fresh verifiers.
🎬 Videomotion-reelnative clean-roomTurns an existing talking-head video into a reel with speech-synced motion-graphics overlays via motion-overlay.
🎬 Videorap-avatar-mvnative genericA portrait + a supplied or generated track → rap MV variants: measured vocal windows, image and voice references, clip reuse, timing checks and verified delivery.
🎬 Videorhyme-factorynative genericCharacter pack + nursery-rhyme idea → a YouTube-ready SUNG music video (song, sing-along captions, polish, thumbnail, upload pack) on top of garden-days.
🎬 Videorunwaynative clean-roomDraft renders on Runway through its free queue; paid credits stay opt-in.
⚙️ Workflowskills-auditorported-from-videoclawAudits the skill surface for wrong commands, missing resources, routing overlap and catalogue drift.
🎞️ Promptsstory-bible-buildernative genericInterview-driven canon builder: a story world → one dense story bible that ai-director and the storyboard consume.
⚙️ Workflowui-ux-pro-maximported-from-video-creation-projectsUI/UX design intelligence for admin panels, CMS pages and dashboards (styles, palettes, font pairings, chart and stack guidance).
🎬 Videovideo-portfolio-opsnative clean-roomPortfolio operations: many projects at once — metrics, report, trends, next-actions, doctor-*, CSV/Obsidian exports.
🎬 Videovideo-production-handoffnative clean-roomMoves a project from create/storyboard/assets into review: status, readiness, next-actions, Review UI handoff.
🎬 Videovideo-release-readinessnative clean-roomRelease pre-flight for the repo itself: source checkout, package dry-run, npm handoff and the verification report.
🎬 Videovideo-review-ui-qanative clean-roomQA for the local Review UI: storyboard handoff, review-autopilot, browser behaviour and the decision copy-blocks.

Generic orchestration skills that duplicated the global plugin set (autopilot, ralph/ralph-init/ralplan, team, cancel, trace, hud, git-master, code-review, security-review) and the removed-tooling omx-setup were culled from the repo — use the global versions.


🎯 Canonical entries ​

video-framework ​

Role: OMX-native front door for any "make a video" request. What it does: Routes the request across the seven established workflows — COPY, CREATE, NARRATED, PRESENTATION, LONG-FORM, FILM, UGC — by classifying the intent and reusing proven legacy engines behind clean adapter boundaries. Picks the right specialist instead of forcing the user to. Key features:

  • Single intake surface for both clone-style and from-scratch video requests
  • Adapter pattern preserves legacy engine quality without inheriting legacy mess
  • Hands off to a specialist (storyboard, clone-ad, movie-director, replicator, Higgsfield bridge, ugc, ...) once the mode is decided
  • Useful as the default first-touch when the user's intent is ambiguous

When to reach for it: Any open-ended video request — "I want to make a video", "can you do a video for X?" — where the production mode hasn't been picked yet.

Full guide: skills/video-framework/SKILL.md


brand-presenter ​

Role: Canonical (generic) presenter-video workflow. What it does: Turns a slide deck or structured topic into an intro/slides/outro narrated presentation using a branded host profile (avatar + voice + intro/outro framing). Personal/brand presenter skills (davendra-presenter, david-sales-presenter, nex-presenter, bunty) all delegate here with a different host profile. Key features:

  • One generic workflow with swappable brand profiles (no copy-paste forks)
  • Slide-deck-aware framing (cover slide → body → call-to-action)
  • Lip-synced intro/outro plus TTS narration over body slides
  • Works for product explainers, internal updates, social-first brand cuts

When to reach for it: Anything narrated and host-led — explainers, demos, brand intros, presentation videos. Pick a personal alias instead if a specific host identity is required.

Full guide: skills/brand-presenter/SKILL.md


🎬 Video specialists ​

garden-days ​

Role: Series-factory skill for the "Garden Days" (Baby Ram & Sita) autonomous series factory — status, new episodes, driver relaunch, troubleshooting, QC, and franchise cloning across three render routes (Runway free / Higgsfield free / Veo omni-flash credits). State lives in the vclaw workspace (BRS-SERIES.md, SERIES-PLAYBOOK.md, lane queues/drivers); identity lives in Go Bananas characters 347/348 and Higgsfield reference elements. Triggers: /garden-days, "garden days", "baby ram and sita", "new episode", "series status", "lane A/B/C". References: references/writer-bible.md, references/lanes.md, references/failure-modes.md, references/franchise-cloning.md.

dance-breakdown ​

Role: Video specialist (native generic). A reference dance video → a measured per-shot choreography breakdown: scene detection (or a supplied cut list), each shot's own footage sent inline to a multimodal model, one record per shot with a single moves sentence (for a window that sings — no seconds, the audio is the clock), timecoded blocks (for a window with no voice), framing, camera move, subject count, body-in-frame and hands-at-face facts. Names, titles and "60fps / 4K / hyper-real" lines are stripped; moderation nouns the free route rejects are flagged for you.

Triggers: /dance-breakdown, "break down the dance", "choreography from this video", "dance moves from the video".

References: skills/dance-breakdown/SKILL.md, skills/dance-breakdown/scripts/breakdown.py; consumed by skills/rap-avatar-mv (shot-refs stage).

skills/rap-avatar-mv/scripts/song_breakdown.py reuses this skill's Gemini plumbing (api_keys, call_model, the cooldown) for a second unit of analysis — the sung PHRASE rather than the shot: section, singer and a choreography cue per lyric line, advisory, with whisper keeping the clock (SONG_BREAKDOWN=1 in a rap pack). Both scripts take an opt-in --processing agentic (Gemini's agentic video understanding through the Interactions API — call_interactions, upload_file, gemini-3.8-flash; the song breakdown reads the whole film by reference in one call, the dance breakdown one shot per call); the default static path is unchanged.

match-highlights-local ​

Role: Video specialist (native generic). A cricket match → its boundaries and wickets, read off the broadcast scoreboard overlay instead of out of the picture. The overs counter ticks once per legal ball and the total says what that ball scored, so one ffmpeg pass sampling the score cell at 1 fps recovers every delivery: template OCR with a grammar and a monotonic ball-count prior reads the cell either side of each change, the delta types the ball (four, six, wicket, wide, dot — a no-ball four is typed by the batter's runs so it still reaches the reel), the clip is anchored on the overlay change and made long enough to hold the ball whatever the scorer's lag was, and the proven cutter renders an all-events reel and a highlights reel with an EDL and a filmstrip beside each. The cell is FOUND from the blue block in the lower band, not assumed. An innings change and a scorer's correction are held and confirmed over two reads rather than trusted or dropped. Nothing in it can reach a network — a test asserts that — and providerCalls is 0 in every file it writes. The zero-cost sibling of vclaw video match-highlights, measured against that lane's own verified reels over the same four-part match: 443 of the 448 balls the overlay itself recorded (98.9%), and 57 of the 61 boundaries and wickets whose rendered clip shows the run-up AND the shot (93.4%), at a mean clip of about 31 s. Clips are long on purpose: the overlay is exact on which ball and what happened and only approximate on when — the scorer's keystroke lands a median 10.5 s after the run-up and a p90 of 21 s, with wickets the slowest (median 16.5 s, up to 29.4 s) and one four in this match entered 53 s late — so the window is anchored on the change and made wide enough to hold the ball. Anchoring on the audio instead was measured and is worse to watch. For tight cuts, hand these events to vclaw video match-highlights. At $0.

Triggers: /match-highlights-local, "cricket highlights without the API", "scoreboard highlights", "highlights from the scoreboard", "free match highlights".

References: skills/match-highlights-local/SKILL.md, skills/match-highlights-local/scripts/overlay_read.py (the pure reader), skills/match-highlights-local/scripts/scoreboard_highlights.py (the ffmpeg runner), skills/match-highlights-local/references/frogbox-templates.json (one broadcaster's glyph set — --build-templates makes another).

Templates are per-overlay-font, not per-sport: the committed set is FrogBox at 1280×720, and a different broadcaster needs its own built from labelled crops of that overlay. --eval scores a run against a verified delivery list and splits its timing statistics by anchor, which is how the fallback lead was tuned.

vox-explainer ​

Role: Video specialist (native generic). An approved narration becomes a collage explainer in a Vox-inspired design vocabulary (no affiliation with Vox), adapted from the Apache-2.0 HyperFrames vox-explainer skill. A small per-film pack (skills/vox-explainer/packs/<slug>.vox.json) splits the narration into beats. Each beat carries archival search queries and one of seven frame recipes (evidence stack, zoom-isolation, two-panel compare, specimen grid, lower-third, newsprint layering, circle reveal). The pack is refused when its beats, joined, differ from the approved narration text, so a pack cannot add a claim the narration never made. Phase 1 searches Wikimedia Commons and the Library of Congress and keeps only public domain, CC0, CC BY and CC BY-SA files; everything else, anything without machine-readable licence metadata, and anything categorised as AI-generated is rejected with a named reason. It records source URL, author, licence and licence URL per file in credits.json, and lifts subjects (objects as well as people) off their plates with a local ONNX model, refusing a cutout that keeps under 2% of the frame. It then writes an HTML contact sheet that shows every candidate, the chosen one, and any subject with no free real photograph. Nothing renders until a person approves that sheet; a swap or a changed file after approval makes the approval stale. Phase 1 makes no generations, calls no paid API and spends nothing. Phase 2 transcribes the approved voice-over for word times and generates one HyperFrames composition from the pack (recipes keyed to the words, 12 fps element motion over a smooth camera, continuation cuts, an inverse-zoom arrival on the payoff, and an end card crediting every photo's author and licence). It refuses a scene with a large empty area, a stretch over 3 s without an event, and a failed hyperframes check. It then renders, ducks the music under the voice at -16 LUFS, and sweeps the render for still stretches and empty-ground holes. Generated illustrations are optional; each is sha-checked and labelled on frame.

Triggers: "vox explainer", "vox-style history of", "collage explainer", "archival photo explainer".

References: skills/vox-explainer/SKILL.md, skills/vox-explainer/references/rulebook.md, skills/vox-explainer/scripts/vox.py (the CLI), skills/vox-explainer/references/THIRD_PARTY/NOTICE.md (Apache-2.0 credit).

ad-intel ​

Role: Video specialist (native generic). A brand name, a keyword or one Library ID → the advertiser's ACTIVE ads from the public Meta Ad Library, fetched with the Higgsfield browser session the seedance-direct engine ships (its own profile, never the logged-in Higgsfield one; no login, no cookies, no third-party API) → each card's text, start date, version count, landing link and the ad video → whisper primed with the card's own words → Gemini reads the video with its sound under references/teardown-framework.md → one artifact per ad (verbatim hook, structure, claims, offer, CTA, proof, objections, angles by the skills/ugc belief vocabulary, creative pattern, a UGC script in our voice), a brand teardown (angle clusters, hooks swipe file, longest-running and most-repeated, what they appear to be testing, tests for us) and a self-contained page. It invents no spend or performance numbers; "active" is not "winning", repetition and age are the signals. Single public TikTok / Instagram / Facebook / YouTube video URLs and local files work too.

Triggers: /ad-intel, "competitor ads", "ad library teardown", "what are their ads saying", "tear down this ad", "hooks from their ads".

References: skills/ad-intel/SKILL.md, skills/ad-intel/scripts/adlib_fetch.py, skills/ad-intel/scripts/ad_intel.py, skills/ad-intel/scripts/answer_plan.py, skills/ad-intel/scripts/assemble_answer.py, skills/ad-intel/scripts/before_after.py, skills/ad-intel/scripts/lib/adlib_extract.js, skills/ad-intel/references/teardown-framework.md; schemas ad-intel, ad-intel-teardown; consumed by skills/brand-agency (Phase 2 competitor angles), skills/video-replicator/references/finding-ads.md (Methods D/E) and skills/ugc research.

video-storyboard ​

Role: Brief or clone plan → explicit scene-by-scene storyboard artifact. What it does: Generates a storyboard.json artifact with optional character-to-scene bindings. Scenes can come from raw --scene strings or from a registered storyboard template; characters can be bound per-scene with --scene-character <sceneIndex:name>. Key features:

  • Mode-aware (storyboard vs director) so the right pipeline manifest applies
  • Storyboard-template aware — supports parameterised templates (environment, character A/B)
  • Per-scene character binding flows into character-consistency enforcement
  • Output is canonical JSON and validates against schemas/video/

When to reach for it: "storyboard this brief", "turn this plan into scenes", "assign characters to scenes".

Full guide: skills/video-storyboard/SKILL.md


video-analyze-template ​

Role: Reference video → reusable template packet. What it does: Analyzes a source video (path or URL) and writes a normalized analyze-output.json that can be saved as a reusable template via template-save. With --auto, drives the analysis through the Gemini key pool to fill pacing, beats, keep/change guidance, and reusable variables automatically. Key features:

  • Manual mode (hand-driven beats/keeps/changes) and auto mode (Gemini-backed)
  • Round-robin Gemini key rotation with per-key cooldown for resilient analysis
  • Endpoint override via VCLAW_GEMINI_API_ENDPOINT for local Gemini-compatible targets
  • Output composes directly into template-save → clone-plan → storyboard-from-clone

When to reach for it: "analyze this video style", "break this ad into reusable structure", "turn this reference into a template".

Full guide: skills/video-analyze-template/SKILL.md


video-clone-ad ​

Role: Saved template → new product/brand via the canonical clone-execute flow. What it does: Adapts a known template to a new intent while preserving execution structure (scene count, motion, pacing). Drives the clone-plan → storyboard-from-clone → execution-seed → execute chain in one logical workflow. Key features:

  • Template-driven so structural quality is reused, not re-derived per project
  • Mode-aware (storyboard for fast iteration; director for full approval-gated runs)
  • Execution-profile carrier (aspect-ratio, quality, resolution, audio, outputs) flows into the brief
  • --dry-run lets you validate the payload shape before any provider submission

When to reach for it: "clone this ad", "adapt this launch ad to a new product", "reuse this template for a new campaign".

Full guide: skills/video-clone-ad/SKILL.md


video-thumbnail-lab ​

Role: Final render → click-driving still + platform packaging pass. What it does: Generates thumbnails for a finished render. Social video variants route to video-post, which owns shot-planned reframing and strict QC. Key features:

  • Project-aware (works against the canonical asset trail) and file-mode (works against any local mp4)
  • Optional --text <title> for a simple overlay thumbnail
  • Explicit handoff to video-post for vertical/square/loop delivery
  • Output naming and locations follow the canonical project layout

When to reach for it: "generate a thumbnail for this render", "make square and vertical promo cuts", "package this final video for YouTube/Shorts/social".

Full guide: skills/video-thumbnail-lab/SKILL.md


movie-director ​

Role: Short-film production across 12 genres with structured entry modes. What it does: End-to-end movie production via VideoClaw Director mode. Supports interview-driven, auto-mode, or CLI-hybrid entry. Covers action-thriller, storybook, documentary, UGC-ad, music-video, romance, horror, sci-fi, fantasy, western, short-film, and custom. Bundles cast building, style/color presets, Seedance-safe prompt engineering with content-filter auto-fix, multi-key Gemini rotation, and the storyboard-review gate. Key features:

  • 10 style presets × 9 color gradings × 12 genres
  • Cast building via Go Bananas library lookup or auto-creation from a JSON seed
  • Content-filter auto-fix for Seedance-safe prompts
  • Bundled scripts: verification, interview, auto-mode, cost estimation, iteration, narrated re-mux

When to reach for it: Cinematic, narrative, or multi-genre film work where the bundled genre material and entry-mode structure pays off.

Full guide: skills/movie-director/SKILL.md


video-replicator ​

Role: Deep-reference 7-mode professional video production pipeline. Status: Deprecated for user-facing use — superseded by video-framework (catalog status deprecated-reference). Reach for it only when the canonical entry routes you here. What it does: The legacy comprehensive pipeline: COPY (replicate/clone with subject swap), CREATE (original from scratch), COPY NARRATED (replicate with continuous voiceover), PRESENTATION (slides to animated video), LONG-FORM (10+ minute, 20+ scene batches), FILM (full cinematic with screenplay), or UGC CAMPAIGN. Kept as a deep-surface reference behind video-framework. Key features:

  • 7 distinct production modes covering most real-world video asks
  • SEALCAM+ video analysis for COPY workflows
  • Long-form batch generation across 20+ scenes
  • Image-to-video and text-to-video both supported through the same surface

When to reach for it: When the canonical entry has routed you here, or when an existing legacy workflow needs the deeper reference. Otherwise prefer video-framework first.

NOT for: single image generation, FFmpeg-only scripts without the pipeline, video-player debugging, or static slide deck creation without video output.

Full guide: skills/video-replicator/SKILL.md


video-post ​

Role: Post-render verification, variants, thumbnails, and archival for clean-room outputs. What it does: Closes the loop after render. Verifies final outputs (codec/resolution/duration/audio, first/mid/last frames, contact sheet and media QC), creates shot-planned social variants, and archives finished projects into a tarball with optional cleanup. Key features:

  • verify-final probes structural correctness of the render
  • make-vertical --write-plan-template ... to detect shot boundaries and seed explicitly unverified subject anchors for review
  • make-vertical --reframe-plan ... --strict for verified, subject-centred, caption-safe portrait variants
  • explicit --quick-center-crop compatibility mode for accepted blind crops
  • thumbnail with optional --text overlay
  • archive-project packages a finished project as archives/<slug>-<timestamp>.tar.gz

When to reach for it: Anything that happens after render — verification, packaging, distribution prep, archival.

Full guide: skills/video-post/SKILL.md


character-creator ​

Role: New Go Bananas character creation with reference sheets. What it does: Creates Go Bananas characters with profile images and multi-view reference sheets so the same character can be regenerated consistently across scenes. Inputs feed the character-consistency subsystem and the per-project characters/characters.json store. Key features:

  • Profile image plus multi-view (front / 3/4 / side / back) reference sheet generation
  • Output binds into project character profiles for downstream scene-character mapping
  • Companion to character-library (creator owns creation; library owns audit/patch/delete)
  • Triggers on natural-language asks like "create a character", "new character with reference sheet"

When to reach for it: "create a character", "design a character", "build a character reference", "set up characters", "new character with reference sheet".

Full guide: skills/character-creator/SKILL.md


character-library ​

Role: Audit and hygiene for the shared Go Bananas character library. What it does: Browses the library, flags polluted entries, patches base prompts in place, and deletes bad anchors — without leaving the repo-local skill surface. Companion to character-creator. Key features:

  • library find for exact-name discovery from intent text
  • library clean with dry-run candidate discovery (by ids, name regex, or bloated prompt size)
  • In-place prompt patching for a single character without recreating it
  • Drives the vclaw video library CLI surface

When to reach for it: "list my characters", "audit the character library", "patch this drifting character", "delete polluted characters", "fix library hygiene before a director run".

Full guide: skills/character-library/SKILL.md


seedance-prompts ​

Role: Reference library and prompt-quality assistant for Seedance-targeted scene writing. What it does: Browses the clean-room Seedance prompt reference library and applies current provider guidance to Seedance prompt writing. Built on the actual prompt-lib-list / prompt-lib-show surface. Key features:

  • Searchable Seedance formulas, examples, and prompt-structure guidance
  • Backed by prompt-library references that actually exist in this repo (no hallucinated examples)
  • Triggers on "seedance prompt", "expand prompt for seedance", "prompt quality"
  • Output composes into storyboard and execute flows

When to reach for it: When you need Seedance-specific prompt help, formulas, or examples.

Full guide: skills/seedance-prompts/SKILL.md


multi-shot-prompt ​

Role: Reference image → timed multi-shot cinematic prompt sequence. What it does: Generates structured multi-shot video prompts from a reference image, output as a timed shot sequence validated against the videoclaw cinematic-15s preset. Drives the real CLI: vclaw video multi-shot --plan to scaffold the shot structure, author cinematic prose per shot, then vclaw video multi-shot --validate to enforce the hard rules. An --auto --image <path> path runs the full sequence without manual authoring steps. Existing projects can use --from-storyboard --project &lt;slug&gt; --scene &lt;sceneIndex&gt; to hydrate action, characters, location defaults, and source metadata from the storyboard artifact. Key features:

  • Timed shot sequences anchored to the cinematic-15s preset with hard-rule validation
  • Image-grounded prompting — reference image drives visual continuity across shots
  • Manual (scaffold → author → validate) and automated (--auto --image) entry paths
  • Machine-readable preset discovery, provider-shaped defaults, issue explanations, and parsed shots[] artifacts for downstream agents
  • Prompt-library reference via vclaw video prompt-lib-show --name multi-shot-framework

When to reach for it: "multi-shot prompt", "shot sequence", "cinematic prompt", "video prompt from this image", "shot breakdown".

Full guide: skills/multi-shot-prompt/SKILL.md


mograph ​

Role: Style-locked motion graphics from a script or transcript. What it does: Locks a project's motion-graphics look ONCE into a motion sheet (a master style-reference image + a ≤120-word style lock persisted as artifacts/motion-sheet.json), then turns the voice-over into a motion pack (artifacts/motion-pack.json) of time-coded, action-only clip blocks tagged P1/P2/P3. The engine assembles style lock + SHOT + AVOID per block and compiles batch-eligible blocks into a manifest. batch-submit saves durable Cinema tasks without submitting; use the quote/authorisation/worker flow to render. Shared style text reduces instruction drift, but the actual clips still need visual review. Key features:

  • Built-in style-family register (mograph-sheet --families) + deterministic master-board prompt composer
  • Anti-drift lint gate (mograph-pack --check): style words/hex codes banned from actions, quoted on-screen text budgets, SFX-never-music, coverage-map integrity
  • Plan-only render (mograph-render) — exact prompts + refs reviewed before any submission; v2v blocks routed to omni-flash or prompt-only, never dropped silently
  • Real brand marks via mograph-logos (keyless Clearbit/SimpleIcons/favicon ladder)

When to reach for it: "motion graphics", "explainer B-roll", "kinetic typography", "style lock", "turn this transcript into clips".

Full guide: skills/mograph/SKILL.md · CLI contract: docs/MOGRAPH.md


higgsfield-generate ​

Role: External-provider bridge for Higgsfield AI generation. What it does: Reuses the public MIT-licensed higgsfield-ai/skills command intelligence as a thin videoclaw skill. It routes agents to the official higgsfield CLI for Marketing Studio ad videos, product photoshoots, marketplace cards, Soul Character identity training, generic image/video generation, and finished-video virality scoring. Key features:

  • Keeps higgsfield optional rather than making it a videoclaw dependency
  • Uses Higgsfield's dedicated product-photoshoot and Marketing Studio commands instead of generic prompt guessing
  • Gives agents a clean hand-off rule: standalone Higgsfield URLs directly, project-tracked outputs through existing videoclaw artifacts
  • Captures the upstream source commit and MIT reuse boundary in the skill itself

When to reach for it: "use Higgsfield", "make a UGC ad in Marketing Studio", "product photoshoot", "train Soul ID", "score this video's virality".

Full guide: skills/higgsfield-generate/SKILL.md


youtube-audio ​

Role: YouTube → MP3 audio or MP4 video using yt-dlp + FFmpeg. What it does: Downloads audio or video from YouTube videos and playlists. Supports trimming, resolution selection, and audio quality settings. Key features:

  • Single videos, playlists, or batch URLs
  • Audio-only, video-only, or both in one pass
  • Trim to clip ranges
  • Requires yt-dlp and ffmpeg

When to reach for it: "download audio from this YouTube video", "grab the MP4", "extract music from a playlist".

Full guide: skills/youtube-audio/SKILL.md


ugc ​

Role: Belief-driven UGC campaign generator using the E5 method. What it does: Generates a multi-video belief-targeted UGC campaign (30–60s each) from a product URL and intent. Produces N videos with subtitles plus a campaign report. Key features:

  • Belief-journey decomposition (E5 method: Examine, Educate, Emote, Evidence, Empower)
  • Multi-video campaign output rather than single-clip
  • Per-video subtitles and aggregate campaign report
  • Triggers on "UGC campaign", "belief-driven ads", "E5 method"

When to reach for it: When the goal is a marketing campaign rather than a single creative video.

Full guide: skills/ugc/SKILL.md


🎤 Video aliases ​

These exist for discoverability and personal/brand handoff — they all delegate into brand-presenter with a different host profile. Treat them as compatibility surfaces; the canonical workflow lives in brand-presenter.

AliasProfileTrigger
davendra-presenterDavendra (asset/voice)"davendra video", "davendra presenter"
nex-presenterNex (asset/voice)"nex video", "nex presenter"
david-sales-presenterDavid Sales (Essex comedian, cricket)"david sales", "cricket comedy video"
buntyBunty — cricket commentator (orange blazer)"bunty thing", "match day analysis", "cricket scorecard video"

⚙️ Workflow skills ​

Workflow skills are independent of any one production mode. They orchestrate, debug, review, or operate on top of the production layer.

Diagnostics & exploration ​

deepsearch ​

Thorough codebase search. Use when grep/glob isn't enough and you need an agent to actually understand the codebase as it searches. Full guide

Operational utilities ​

ai-slop-cleaner ​

Run an anti-slop cleanup/refactor/deslop workflow. Removes the kinds of half-baked artifacts that unsupervised agents leave behind. Full guide

web-clone ​

URL-driven website cloning with visual + functional verification. Full guide


Adding or modifying a skill ​

The skill set is curated, not auto-generated. To add or modify a skill:

  1. Add or edit skills/<name>/SKILL.md — the canonical guide for the skill
  2. Update skills/catalog.json with the skill's id, category, status, and any specializes / aliasOf / specializations relationships
  3. Add a section here in docs/SKILLS.md and link it from the index table
  4. Run npm run check:skill-frontdoor to verify the repo-local skill front door stays consistent
  5. Run npm run check:cleanroom-docs to verify clean-room-facing docs and skills don't reference stale paths

See also ​

Built to be driven by agent hosts like Claude Code, Claude Desktop, or Codex · Source-available, commercial use requires a paid license.