Skip to content

Motion-overlay — vclaw video motion-overlay ​

Turn an existing talking-head video into a polished reel with motion-graphics overlays synced to the speech, driven by Google Flow's Omni Flash V2V transport (which paints kinetic typography / icon / metaphor animations on top of footage while preserving the original voice).

The prompt IP — retention principles, a reel-anatomy load map, a concept→animation metaphor map, and two hard-won guardrails (no metadata leaking as on-screen text; no human portraits or the policy filter trips) — lives in deterministic, offline-testable code under src/video/motion-overlay/, not in markdown.

Plan/dry by default. No provider spend until you pass --execute --confirm-spend. The plan renders into the preview portal for human review first.

At a glance ​

vclaw video motion-overlay --input <video-path> (--project <slug> | --output-dir <path>)
  [--layout split|overlay|motion-only|avatar-host]   # default: split
  [--style apple-clean|editorial-dark|knowledge-tool] # default: apple-clean
  [--accent <hex>]                                    # configurable accent (replaces a hardcoded brand colour)
  [--lang <code>]                                     # STT + on-screen-text language (default: auto)
  [--max-take-seconds 10]                             # Omni hard per-clip limit
  [--transcript <path>]                               # optional: bring-your-own transcript JSON
  [--gb-character <Name:ID>]                          # avatar-host layout ONLY: the locked host character
  [--preview]                                         # also render review/review.html
  [--root <path>]
  [--execute --confirm-spend]                         # render (spend-gated)

What happens ​

Default (dry): ingest → STT → slice → compose → write work folder + manifest

  • (optionally) preview HTML. No provider is called.

--execute (requires --confirm-spend, else exit-3 spend_confirmation_required): each take runs through omni-flash V2V, the original take audio is restored, and the takes are restitched into the final reel.

Work folder layout ​

<output>/                            # projects/<slug>/artifacts/motion-overlay/ OR --output-dir
├── motion-overlay-plan.json         # the manifest (schemaVersion 1)
├── README.md
├── source/  original.mp4 · audio.aac · transcript.json
├── takes/   take-01_0s-10s.mp4 ...  # frame-accurate cuts
├── prompts/ take-01_0s-10s.md ...   # one ready-to-send prompt per take
├── frames/  frame_01.jpg ...
├── host/    take-01_0s-10s.mp4 ...  # avatar-host ONLY: generated speaking host clips
├── v2v/     take-01_0s-10s.mp4 ...  # omni-flash overlay output (execute)
├── restored/ take-01_0s-10s.mp4 ... # audio-restored takes (execute)
├── review/  review.html             # preview-portal approval surface (--preview)
└── motion-overlay-reel.mp4          # the final stitched reel (execute)

Layouts ​

  • split (default) — motion graphics in the top half, speaker untouched in the bottom half.
  • overlay — motion graphics composited over the speaker, safe areas around face/hands.
  • motion-only — speaker removed; full-frame motion graphics narrated by their voice (keynote style).
  • avatar-host (Layout D) — like motion-only, but the frame is filled by an identity-locked character host delivering the lines, with synced overlays on top. Requires --gb-character <Name:ID>. This is the heaviest path; see below.

Avatar-host (Layout D) ​

The avatar-host layout replaces the speaker with an identity-locked character. There are two host engines (--host-engine), trading identity-lock against native voice — pick per project:

--host-engineIdentityVoiceBest for
omni-r2v (default)loose (mitigated, see below)native + lip-synceda presenter who actually speaks
veo-i2vpixel-tight (same still every take)silent (captioned / add VO)brand-critical identity, talking-head B-roll

There is no scriptable path that gives both at once (Flow's @Character system locks identity+voice but is web-UI only), so this is a genuine fork.

omni-r2v (default) — the talking avatar ​

character reference image  +  the take's line  +  a voice preset
   → omni-flash R2V (ingredients:) generates the audio AND lip-synced video together
   → host/take-NN.mp4   (the avatar SPEAKING that line, genuinely lip-synced)
   → stitch the host clips in order        →  avatar-reel.mp4
   → re-STT the avatar's own speech + local lower-thirds  →  motion-overlay-avatar.mp4

Because Omni produces the speech and the mouth movement jointly, the avatar is truly lip-synced — no separate TTS, no muxing, no lip-sync model. R2V clears Flow's safety filter far more often than the V2V "add-overlay" edit, but a benign draw is still rejected now and then, per draw: the transport sorts each failed run by what the sidecar printed (classifyHostAttempt) — a safety verdict (PUBLIC_ERROR_*, *_BLOCKED, UNSAFE_GENERATION) or a clean exit with no clip is redrawn at once, a throttle burst waits out Flow's cooldown, a transport blip backs off briefly, and anything a retry cannot change (a bad flag, a dead session) fails immediately with the cause. --host-retries <n> (default 3) is the ceiling. The advisory an account on Google's retiring "labs" backend prints on every render is not read as a dead session, so such an account keeps the throttle cooldown and the no-clip redraw (#710). A take whose job completed is never redrawn unless the job's own verdict was a safety refusal: that includes a clip that came back with no download link and no media id, which the sidecar no longer words as "All operations failed" when Google gave no reason (#710). A create that useapi answered with an error or without a job id is a lost answer too: the take is not drawn again, labs account or not (#710).

veo-i2v — the tight-identity (silent) avatar ​

the SAME character still  (image: literal first frame)  →  Veo 3.1 I2V, per take
   → host/take-NN.mp4   (frame 0 of EVERY take is pixel-identical → identity locked)
   → stitch  →  avatar-reel.mp4
   → caption from the planned script, timed to each take's snapped duration
                                          →  motion-overlay-avatar.mp4 (silent)

Because every take starts from the identical still, identity/wardrobe/backdrop are locked frame-to-frame and take-to-take — far tighter than R2V — but the clip carries no speech (add a voiceover or soundtrack separately, or let the captions carry it). Veo I2V is landscape-only server-side. It shares the same retry + resume + scratch plumbing as omni-r2v, time limit included (600 s; VCLAW_VEO_COMMAND_TIMEOUT_MS changes it, as for veo-useapi). A take stopped before it reports back (that limit, or a kill signal) after its Flow job was submitted is not drawn again, because the job may still be running and billing: the error names the job, so look it up with GET /google-flow/jobs/<id> and save its clip as the take's file before re-running (a re-run reuses that file, as it reuses every completed take). The stopped run leaves its batch running in the take's veo-cli.db, and a re-run of that take stops on it; to draw the take again once the job is looked up, clear it with the command the error gives: vclaw veo cancel <N> for the batch the sidecar named (Created new batch #N or Resuming batch #N on its stdout), with VEO_CLI_DB_PATH set to that database. That only marks the local batch cancelled; it does not stop the Flow job or its billing. Only when the sidecar printed no batch does the error fall back to vclaw veo unwedge --confirm --stale-minutes 0, which clears every running batch in that database, so run it only while no other render uses it (#710).

The sidecar is told when that limit will kill it (VCLAW_VEO_KILL_AT_MS), and a clip download must finish 15 s before then (or a quarter of the limit, when that is shorter), as well as within the sidecar's own download time limit: a slow but steady download now fails as a download ("Failed to download video …", the take is not drawn again and the error says how to fetch the clip) instead of being killed with no error to act on (#710).

Flags: --gb-character Name:ID (label + go-bananas id, split on the final :), --gb-character-image <path> (the reference still — required at --execute; it is the R2V ingredient for omni-r2v and the literal first frame for veo-i2v), --host-engine omni-r2v|veo-i2v (default omni-r2v), --gb-voice <preset> (a Flow voice preset, default Puck; omni-r2v only), --host-retries <n> (per-take generation attempts, default 3; the split/overlay V2V lane has its own --v2v-retries <n>, default 10), --host-look <text> (a stable appearance + setting descriptor — pinned on every take by both engines), --no-host-chain (disable cross-take chaining; omni-r2v only). Captions are timed to the avatar's actual speech (it is re-transcribed because Omni's pacing differs from the source); the caption pass is best-effort — the lip-synced reel stands on its own if it is skipped.

Character consistency (the hard part). omni-flash R2V treats the reference image as a loose influence, so each independently-generated take re-invents the face, wardrobe, and backdrop — the avatar drifts across a multi-take reel. There is no API-scriptable mechanism today that locks identity and generates native voice together (Google Flow's @Character system does, but it is web-UI only). Two levers reduce the drift on the native-voice path:

  • --host-look <text> pins a stable appearance + fixed setting descriptor that is repeated on every take's prompt (e.g. the character's go-bananas base prompt + one locked studio). This is what stops the backdrop/wardrobe from jumping take-to-take. Without it the prompt uses a neutral default. (The old build hardcoded a Bunty-flavoured "friendly cartoon presenter in a sleek modern studio", which actively caused drift for any other character — that is now gone.)
  • Chaining (on by default, --no-host-chain to disable) seeds each take's R2V with the previous take's last frame as an extra ingredients: reference, so the face/wardrobe carry forward. The last frame is extracted with ffmpeg to a hidden host/.chain-NN.png and chained into take N+1. Chaining survives resume (a resumed take re-seeds from its existing clip) and is best-effort (a failed frame extract just drops that one seed). For tight pixel-level identity instead, use the silent go-bananas-character → Veo-I2V path (no native voice).

Robust across multiple takes. The Flow safety filter rejects a benign R2V generation probabilistically, so on a multi-take reel one unlucky take would otherwise fail-fast the whole run. Each take is therefore retried up to --host-retries (a plain retry of the same line usually clears), and generation is resumable — a host/take-NN.mp4 already on disk from a prior run is reused, never regenerated, so re-running after a mid-reel failure never re-spends on the takes that already succeeded.

The generation runs only under --execute --confirm-spend; in plan/dry mode the work folder is laid out and the host clips are merely enumerated.

Design sheet (approve before render) ​

review/review.html opens with a DESIGN SHEET section above the takes: one HTML specimen per named overlay element (ELEMENT_VOCABULARY in motion-style.ts — headline type, outline type, pill, strike, brackets, badge, underline) rendered in the resolved style's real CSS font stacks and the plan's accent, plus an approval row. The vocabulary sentence rides in every style's visualBlock (prose only; the px/ms/hex/font-name invariant still holds), so the prompt, the sheet and the local renderer share one element set. Local primitives behind it: revealLetterCount (per-letter cascade beside the word reveal, animate.ts), strikeSvg / badgeSvg (render-local.ts), and statBarSvg (motifs.ts) — a static percentage bar proven by geometry: 100% track drawn first, fill exactly gauge of it, label printed complete; the count-up ring gauge is unchanged.

On-screen text guardrails (CRITICAL RULES) ​

Every composed prompt emits three standing rules:

  1. AUDIO — pass the original voiceover through; the visual layer only.
  2. NO METADATA ON SCREEN — never render px/ms/hex/font-name/easing/stroke as visible text; only quoted text renders.
  3. TEXT ONLY, NO PORTRAITS — no drawn human silhouettes/portrait outlines/avatar icons (policy-filter guardrail). The avatar-host host base layer is exempt because it is generated separately and never described to Omni as a portrait.

Transcription ​

Gemini STT (reuses the Gemini key pool) by default, or bring your own with --transcript <path>:

json
{ "language": "pt", "segments": [ { "start": 0.0, "end": 3.2, "text": "..." } ] }

Module map (src/video/motion-overlay/) ​

ModulePurposePurity
types.tsshared contract for the whole familytypes
slice.tssegments + maxTakeSeconds → sentence-boundary ≤10s splitspure
analyze-reel.tsper-take reel anatomy + anchor words + pausespure
motion-style.tsstyle registry + concept→animation metaphor map; accent/brand resolutionpure
compose-prompt.tsthe IP — per-take prompt (CRITICAL RULES + layout block + style + shots)pure
plan.tsassemble the MotionOverlayPlan manifestpure
transcribe.tsGemini STT (injectable transport)side-effect, thin
ingest.tsffmpeg probe / extract audio / cut takes / extract framesside-effect, thin
write.tsemit folder + README + manifest + per-take promptsside-effect
preview.tsrender the plan into the preview-portal review surfaceside-effect
execute.tsper-take V2V → audio-restore → stitch (injectable MotionOverlayStepRunner)side-effect, tested via injection
avatar-host.tsavatar-host: per-take speaking-host orchestration (injectable MotionOverlayHostGenerator)side-effect, tested via injection
avatar-host-transport.tsreal omni-flash R2V+voice host generator (shells flow.ts); pure prompt/arg builders tested, spawn injectableside-effect, pure builders tested
v2v-transport.tsreal Google Flow omni-flash V2V transport (upload → submit → poll → download); pure request helpers tested, network path injectableside-effect, tested via injected fetch
render-local.tsnative local overlay renderer — segment → lower-third SVG → sharp PNG → ffmpeg overlay; SVG/filtergraph builders pure, sharp+ffmpeg injectable (--render local)side-effect, pure builders tested
run.tsplan orchestrator — planMotionOverlay, the dry/plan front door wired into vclaw video motion-overlayorchestrator
motifs.tsconcept → visual motif layer — surfaces a matching icon/stat motif at the moment a concept is spokenpure
animate.tsanimation timing math (clamp + easing curves) for the per-frame rendererpure
animate-render.tsper-frame animated overlay renderer (--animate) — animated SVG frames instead of the static fade path; builders pure, sharp+ffmpeg injectableside-effect, pure builders tested
flow-pack.tsFlow hand-off pack + restitch — the productised "manual Omni" loop (write prompts → upload each take → paste back → restitch)side-effect

Every side effect — STT, ffmpeg, V2V, and the omni-flash R2V host generation — is behind an injectable interface, so the whole pipeline is unit- and e2e-tested offline with no network and no spend.

Schema ​

schemas/video/artifacts/motion-overlay-plan.schema.json is the canonical MotionOverlayPlan contract (allowlisted in the artifact-schema-coverage check).

Live execution & moderation (--execute) ​

--execute --confirm-spend runs each take through the Google Flow omni-flash V2V transport (v2v-transport.ts). Requires USEAPI_API_TOKEN and USEAPI_ACCOUNT_EMAIL in the environment (export them from .env); missing creds fail fast as provider_unreachable before any network call.

Proven request recipe (validated live 2026-06-05 against useapi.net):

  1. Upload the take — POST {base}/google-flow/assets/{email} with Content-Type: video/mp4 and the raw bytes → mediaGenerationId.mediaGenerationId (this becomes referenceVideo_1).
  2. Submit — POST {base}/google-flow/videos with { model: "omni-flash", prompt, referenceVideo_1, startFrameIndex_1: 0, endFrameIndex_1: &lt;frameCount&gt;, aspectRatio, email, async: true }. The response field is jobid (lowercase i), not jobId.
  3. Poll — GET {base}/google-flow/jobs/{jobId} with the job id passed RAW / unencoded → status: created → completed | failed.
  4. Result — response.media[].videoUrl; download to the take's output path.

Then execute.ts restores the take's original audio (ffmpeg -map 0:v -map 1:a) and clip-stitches the takes into the final reel.

Moderation gotcha. The Flow safety filter rejects person footage for editing with PUBLIC_ERROR_VIDEO_EDIT / VIDEO_EDIT_BLOCKED (older logs: FINISH_REASON_INPUT_VIDEO_EDIT) — and the verdict is per draw, not per clip: 24 identical submissions of one talking-head take (same body, same reference) cleared 3 times and were blocked 21 times (2026-09-02). Nothing in the prompt or the reference shape predicted which draws cleared. A blocked draw costs 0 credits and returns in ~20 s; a clear costs the normal omni-flash price. So the transport keeps drawing: --v2v-retries <n> (default 10, which finishes a take about seven runs in ten at that rate) and then errors with an actionable message. Realistic-face footage is the usual trigger; stylized / non-photoreal source clips clear far more often. The sync bun run flow.ts --ref-video path returns the generic "All operations failed" and omits the frame-window params — prefer the async path above.

Delivery targets (--delivery) ​

--delivery is the high-level "where the animation comes from" knob (it sets the render mode + flow-pack defaults; --render/--emit-flow-pack stay as the low-level controls):

--deliveryWhat it doesNotes
localrender the finished reel locallyreliable, free, on your exact footage (default product path)
flow-webrender local and emit a Flow hand-off packthe local reel is the instant fallback; the pack is for the real Omni animation, pasted into Flow's web editor by hand
flow-apiomni-flash V2V via the APIgated — needs --execute --confirm-spend; Flow's safety filter blocks ~7 draws in 8 on person footage, so each take draws up to --v2v-retries (blocked draws are free)

The Flow hand-off loop (--emit-flow-pack / --restitch) ​

The source skill's real pipeline is write prompts → upload each take + paste its prompt into Flow → download → stitch. The Omni step was never automated and Flow's API clears the V2V edit only ~1 draw in 8, so the hand-off stays a first-class artifact instead of ad-hoc scripts (flow-pack.ts):

  • --emit-flow-pack writes <work>/flow-pack/: each take's clip + a clean, ready-to-paste prompt .txt (the prompt body extracted from prompts/take-NN.md)
    • a README of the workflow.
  • --restitch <flow-outputs-dir> takes the operator's downloaded animated takes, re-muxes each with its original audio (CRITICAL RULE 1), and concatenates the finished reel (<work>/motion-overlay-reel.mp4). Terminal — no render/spend.
# 1. produce a local reel now + the Flow pack to upgrade it
vclaw video motion-overlay --input clip.mp4 --output-dir out/ --delivery flow-web --style editorial-dark
# 2. (you upload each out/flow-pack/take-NN.mp4 + paste take-NN.txt in Flow, download to out/done/)
# 3. stitch the animated takes back, audio restored
vclaw video motion-overlay --input clip.mp4 --output-dir out/ --restitch out/done/

The pure pieces (prompt extraction, README, output pairing, concat/restore args) are unit-tested; fs + ffmpeg are injectable seams.

--render local (no spend, no moderation) — the reliable path ​

Because the V2V "add-overlay" edit clears moderation only ~1 draw in 8 on person footage, the local renderer is the recommended way to actually produce a reel:

vclaw video motion-overlay --input clip.mp4 --output-dir out/ --render local [--kicker "BRAND"]

--render local (vs the default --render v2v) runs renderLocalOverlay (render-local.ts): each transcript segment becomes a broadcast lower-third — an SVG rasterised to PNG via sharp (resvg text shaping; NO system freetype / ffmpeg drawtext) — composited onto the source with one ffmpeg overlay filtergraph (each card fades + slides up on its window). Accent words reuse the prompt composer's anchor-word detector; --kicker sets the small brand label. It is free (no provider, no --confirm-spend gate) and the original audio is preserved. The pure SVG/filtergraph builders are unit-tested; sharp + ffmpeg are injectable seams.

Card styles (--style) and giant headlines (--headlines) ​

--render local honours --style, rendering three structurally distinct card looks faithful to the source skill's catalog (not just a colour swap — different panel geometry, label treatment and typography so they read as different over busy footage). The same --style/--accent/--kicker/--headlines also drive the avatar-host caption pass.

--styleCard look
apple-clean (default)floating inset rounded frosted card · thin accent left bar · accent-dot + uppercase kicker · clean sans-serif sentence case
editorial-darkfull-bleed edge-to-edge near-black poster bar · thick accent top rule · solid accent tag block label · bold condensed UPPERCASE
knowledge-toolleft-anchored narrower cool-dark study card · thick lavender left rule · [[ MONO ]] bracket label · editorial serif body

Accent words render in the accent colour and get the skill's signature accent * beat-marker after them.

--headlines adds the skill's kinetic-typography beat: each detected anchor word also explodes to a frame-filling centred headline with the accent * beat-marker, synced to the moment the word is spoken, composited above the lower-third. Headline moments are capped (≤6, one per caption) so the reel doesn't strobe. Pure lowerThirdSvg/anchorHeadlineSvg builders are unit-tested.

--icons adds the skill's "visualize what's being said" layer (motifs.ts): the transcript is scanned and, at the moment a concept is spoken, a matching hairline icon (money → $, team → people, launch → rocket, AI → node graph, time → clock, …) or an infographic stat callout (a number, or a percentage with a ring gauge) pulses in the upper-right corner like a broadcast stat bug — in the style's accent, capped to ≤5 per reel. Icons are vector geometry (no portraits, no bitmaps), so they render reliably. Both --headlines and --icons work on --render local and the avatar-host caption pass. detectMotifs + the iconMotifSvg/statMotifSvg builders are pure and unit-tested.

--animate — real motion graphics (per-frame) ​

By default the local overlays fade and slide in, then hold. --animate swaps in a per-frame motion-graphics layer (animate.ts + animate-render.ts): every overlay is evaluated as a function of time t and one transparent frame is rasterised per output frame (30fps), then ffmpeg overlays the sequence on the footage. The motions:

  • lower-third — fade + slide-up entrance, then a word-by-word reveal (words appear across the first ~60% of the take; the panel holds full size so it never reflows — later words render at fill-opacity 0);
  • stat callout — the number counts up 0→N and the % gauge fills (eased);
  • icons / headlines — pop in with a small overshoot (spring).

All the timing is pure math (easeOutCubic/easeOutBack, revealWordCount, countUpValue, gaugeProgress) and unit-tested; the sharp rasterise + ffmpeg run are injectable. It costs more (N frames rasterised — a 16s clip ≈ 480 frames, ~20s on a laptop) but stays fully local/free. Works with any --style and composes with --icons/--headlines.

vclaw video motion-overlay --input clip.mp4 --output-dir out/ \
  --render local --style apple-clean --icons --animate --kicker "BRAND"
vclaw video motion-overlay --input clip.mp4 --output-dir out/ \
  --render local --style editorial-dark --headlines --kicker "BRAND"

Built to be driven by agent hosts like Claude Code, Claude Desktop, or Codex · Source-available, commercial use requires a paid license.