Patch release. v0.20.0 shipped two entry points that aborted with a NameError before doing any work: - `migrate_to_codex.py` / `migrate_to_kiro.py --force` — undefined `write_text`, dropped by the _migrate_common extraction (#85, thanks @Anai-Guo) - `tools/dewatermark.py --setup` — undefined `get_runpod_config` (#86) Also consolidates the duplicated migrate helpers and adds a Lint Python CI gate for undefined names, which had no coverage before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
40 KiB
Toolkit Changelog
All notable changes to claude-code-video-toolkit.
Releases are automated via GitHub Actions. See
.github/workflows/release.yml.
Unreleased
2026-09-08 (v0.20.1)
Patch release. v0.20.0 shipped with two tools that aborted with a NameError
before doing any work; both are fixed here.
Fixed
- Codex and Kiro migrations (
scripts/migrate_to_codex.py,scripts/migrate_to_kiro.py) —write_text()was dropped by the_migrate_commonextraction without being moved to the shared module, so--forceaborted withNameError: name 'write_text' is not definedon the first real write.--dry-runreturned before that path, which is why it went unnoticed. Restored in_migrate_common, and the two orphaned function bodies the same refactor left stranded are gone. Thanks @Anai-Guo (#85). tools/dewatermark.py --setup—setup_runpod()calledget_runpod_config(), which is defined nowhere in the repo, so the documented one-time setup crashed immediately. Now uses theload_dotenv()+os.getenv("RUNPOD_API_KEY")patterntools/upscale.pyalready uses (#86)._migrate_common.load_mappingreturned raw JSON instead of the normalized shape, so any caller wouldKeyErroronmapping["skip_commands"]. It was unreachable only because both scripts shadowed it with a local copy. The real version is now shared and the duplicates (plus a duplicatefind_repo_rootand an unusedyaml_quoteimport) are gone — verified byte-identical migration output before and after (#86).
Added
Lint PythonCI workflow — gates on undefined names (ruff --select F821,F811,F822,F823) acrossscripts/andtools/. Deliberately narrow: only rules where a hit is a real defect, not style. Run against the v0.20.0 tree it reports 26 errors, including bothwrite_textsites. There was previously no Python lint in CI (#86).
2026-08-31 (v0.20.0)
Added
- SoulX-FlashHead talking head (
tools/soulx.py+docker/modal-soulx/) — the toolkit's default talking head generator (Soul AI Lab, Apache 2.0, 1.3B), Modal-only. Preserves the input aspect ratio, so 16:9 presenter images come back 16:9 with no--preprocessworkaround, and--sizesnaps to the model's latent grid while keeping the aspect. It is the default because identity holds over a long take: measured at 97% of frame-zero sharpness at 70s, flat across 72 segments, so per-scene generation is a choice rather than a workaround. ~$0.0024 per second of output against SadTalker's ~$0.0014. Decision table in CLAUDE.md, full detail indocs/soulx.md. (#81)
Changed
docker/modal-soulx/keeps its weights in a Modal Volume, unlike the other Modal apps which bake them into the image. Measured: redeploy after a code change is ~2.3s against 79-385s baked, while cold start and generation speed are unchanged. Needs a one-offmodal run …::populate_weights(14.7 GB). The settled apps stay baked.- Upstream repo ref and both model revisions in
modal-soulxpinned by SHA — weights in a Volume aren't tied to the image, so nothing else prevents drift. (same lesson as #71/#74) NarratorPiPhonours itsobjectPositionprop, which was declared and documented but silently ignored by a hardcoded value, and gains anobjectFitprop (defaultcontain, unchanged behaviour).tools/cloud_gpu.pylearnedsoulx, including its GPU tier, so its jobs report a cost estimate.
Removed
- EchoMimicV3 (
tools/echomimic3.py,docker/modal-echomimic3/,docs/echomimic3.md), added earlier in this unreleased cycle and never part of a tagged release. Identity drift made it unusable for real narration: on a controlled A/B — same photo, same 80s audio, same 544x736 — it fell to 50% of frame-zero sharpness by 70s, collapsing to a featureless smear, where SoulX-FlashHead held 97%. The failure is absorbing rather than gradual, since each segment re-anchors on the previous one's output, so it could not be tuned out; it implied a ~30s render ceiling that shaped the surrounding design. SoulX supersedes it on every axis that mattered — drift, cost (~3.7x cheaper), wall clock (~3-6x faster) — while keeping aspect-ratio preservation. (#81, closes #80)
2026-08-27 (v0.19.0)
Added
- uv is the install path —
pyproject.toml+uv.lock;uv sync(extras:whisper,modal,youtube) and every doc/tool now saysuv run tools/….tools/requirements.txtstays as the pip fallback. Contributed by @AsharibAli (#33, rebased as #59). - Kiro CLI support —
scripts/migrate_to_kiro.pyinstalls the toolkit's skills/commands for Kiro with steering-file generation; seedocs/kiro.md. Contributed by @homeo26 (#55). Follow-up refactor tracked in #58. - Version maintenance (closes #23) —
scripts/check_versions.py(toolkit-wide Remotion/Python/toolkit staleness, wired into/versions);tests/render-baseline/+scripts/render-baseline.mjs+render-baseline.yml(same-runner before/after render diff that gates Remotion bumps in CI);scripts/bump_remotion.pyas/versions --bump-toolkit; grouped monthly Dependabot for theremotiongroup. (#63, #64, #65) sync_timing.pygeneric mode — pairs any*AudioFilewith*DurationSecondsso ticket-driven and future config layouts work without tool changes. (#62)- Community add-ons README section for integrations that live in their own repos (first entry: VOICEPEAK offline Japanese TTS, #42), plus CONTRIBUTING guidance for single-vendor integrations and automated/AI-generated PRs. (#60, #61)
Changed
- Remotion 4.0.425 → 4.0.518 across all templates, examples, showcases and the baseline harness — first bump gated by the render-baseline check (all frames identical). (#67)
- All Remotion projects now pin the same exact version (four stragglers were on 4.0.381/4.0.382/
^4.0.0).
Housekeeping
- External PR queue triaged: MiniMax video/music and Atlas Cloud bot PRs closed under the new policy; Edge-TTS (#46) has review feedback pending.
- Stale branches removed; unmerged experiments preserved as
archive/*tags.
2026-08-26 (v0.18.0)
Added
- YouTube publishing —
tools/youtube_upload.py+/publishcommand. OAuth login (--auth), private-by-default upload, scheduled go-live (--publish-at), thumbnails,--dry-runvalidation, 403s classified by API reason. Setup guide indocs/youtube-upload.md; google-* deps are opt-in intools/requirements.txt. (#29, #37, #38, #39, #40) - 60db TTS provider — third
voiceover.py --provider 60dbbackend plus standalonetools/sixtydb_tts.py(synthesize / stream / websocket transports,--list-voices). Unified 0-1 voice settings auto-convert to 60db's 0-100 scale;redub.py --tts-provider 60db. Brands carry asixtydbblock invoice.json. (#27) _internal/narrative-patterns.md— Discovery Story structure for sprint reviews (reorder by narrative arc, turning-point slides, honest bridge framing).docs/ltx2.mdWindows gotchas + LTX-2 Modal config wired intoverify_setup.py. (#51, #53)
Changed
- Remotion skills sync now tracks upstream's new per-topic layout via
skills/remotion-best-practices/(router skill embedding markup/maps/captions/render/saas/studio/docs/upgrade). Workflow fails loudly on the next re-layout and carries the upstream skill version in PR titles. First sync from the new layout: 39 → 137 files. (#56) - Modal image-edit requests A100-80GB (OOMed on 40GB).
.gitignorecoversscratch/, Playwright run artifacts, recursiveassets/voices; duplicate project copies underprojects/untracked.
Fixed
- Weekly Remotion skills sync had failed every run since 2026-07-13 (upstream directory moved).
verify_setup.pyaccepts bothmodalCLI JSON key casings when detecting deployed apps.
Registry
- Added missing entries:
align_captionstool,quick-spot/data-viz-chart/ds-crt-stinger/sky-blue-shortexamples.the-space-betweenmarked showcase-only (no source directory).
2026-06-10 (v0.17.0)
Added
- concept-explainer-short template — 9:16 vertical TikTok/Reels/Shorts explainers. First pure Python/moviepy template:
scenes.json→gen_vo.py(per-scene TTS) →gen_captions.py(whisper forced-align) →build.py(Ken Burns stills, boomerang clips, karaoke caption pills, ducked music). (#31) examples/sky-blue-short— 52s showcase of the template with all assets committed.- TTS pacing QC —
tools/pacing.py;voiceover.pyandqwen3_tts.pyreport per-takewpm+pacinglabel and accept--max-wpmfor a pitch-preserving atempo clamp. Fixes cloned voices inheriting a rushed pace from their reference. (#30)
Docs
- README restructured to lead with finished videos; "Using with Codex" moved to
docs/codex.md.
2026-06-08 (v0.16.0)
Added
- Ideogram 4 text-to-image —
tools/ideogram4.py+ideogram4skill. Structured JSON captions for legible in-image text and exact brand colours; hosted paid API. FLUX-vs-Ideogram decision rule added to CLAUDE.md. (#28) - Qwen3-TTS v0.2 — VoiceDesign mode (
--design-instruct), shared-prompt cloning for consistent multi-scene narration,clone.designbrand schema with self-caching reference;tools/align_captions.py(Scribe-based caption timing alignment). - Codex migration flow for toolkit skills/commands (#16, thanks @kimhoontae-gogo); thin
AGENTS.md.
Changed
- Remotion pinned to exact versions in lab examples; CLAUDE.md clarifies the
<OffthreadVideo>choice vs upstream@remotion/media. - Official Remotion skills synced (upstream 277510e). (#17)
- GitHub Actions bumped to Node 24-compatible versions; docker-build workflow handles multi-commit pushes and skips
modal-*dirs.
Fixed
- Python 3.9 compatibility restored across
tools/viafrom __future__ import annotations. (#22) modal-upscalebasicsr build conflict. (#19, thanks @ATECHPCS)- Codex migration name read from registry; no-op skip entry dropped. (closes #20)
2026-04-21 (v0.15.0)
Added
- LTX-2 style LoRA support with
crt-terminalpreset. examples/ds-crt-stinger— 6s CRT LoRA + grunged-logo brand stinger.- Animated synthwave README banner (
showcase/banner).
Changed
- Official Remotion skills synced (upstream 41a3b06); "Stsrt" → "Start" typo fixed. (#14)
2026-04-09 (v0.14.2)
Added
- moviepy skill (
.claude/skills/moviepy/) — Python video composition for overlaying deterministic text on AI-generated video (LTX-2, SadTalker). Leads with the "trustworthy text" framing: news, trailers, lower thirds, and social captions all need deterministic overlay because AI models can't guarantee readable text. Covers the PIL workaround for moviepy 2.x's TextClip ascender-clipping bug, audio-anchored timelines, common recipes (labels on LTX-2 clips, lower thirds on SadTalker heads, tinted overlays for contrast, split-screen composites, music + VO mixing), 2.x API gotchas, and a genres-where-this-shines table. examples/quick-spot— runnable 15-second ad-style moviepy example. Audio-anchored timeline, PIL text rendering with cross-platform font loading, optional per-scene VO + ducked music, solid-colour backgrounds. Renders withpython3 build.pyand zero external assets.examples/data-viz-chart— runnable animated time-series chart (real GitHub star history) demonstrating the "matplotlib for data + moviepy for trustworthy text" split that mirrors real news-graphics pipelines. Cache-aware matplotlib step.examples/hello-world— minimal Remotion sprint-review example finally committed. The root README's quick-start callout has been pointing atcd examples/hello-world && npm install && npm run rendersince Feb 23, but the example had zero git history on any branch. Fixes the broken link for anyone cloning the repo.- LTX-2 skill: Stylized Character Cameo — new use case in
.claude/skills/ltx2/SKILL.mddocumenting LTX-2 image-to-video as a SadTalker alternative for non-photoreal faces (fantasy characters, heavy beards, masks, helmets) where lip-sync precision matters less than motion + atmosphere. - CLAUDE.md: Audio-Anchored Timelines — new subsection in the Video Timing guide complementing the existing reactive
sync_timing.pyflow with a prevention-first pattern (generate audio first, anchor visuals to absolute timestamps). Especially useful for single-file moviepybuild.pyprojects.
Changed
tools/requirements.txtaddsPillow>=10.0,moviepy>=2.0,matplotlib>=3.7. One install command (python3 -m pip install -r tools/requirements.txt) now covers every Python feature in the toolkit: AI voiceover, image gen, music gen, and the new moviepy examples.- Root README quick-start switches to
python3 -m pip install— the PyPA-recommended invocation that works identically on macOS, Linux, and Windows and guaranteespipandpython3resolve to the same environment. docs/sadtalker.mdadds a "When NOT to Use SadTalker" section pointing at LTX-2 image-to-video for stylized characters, heavy facial hair, masks, and non-frontal angles.examples/README.mdsplits "Using Examples" into Remotion (npm installflow) and moviepy (python3 build.pyflow) sections and adds both new examples to the listing table.
Fixed
- Hardcoded macOS font paths in
examples/quick-spot/build.pyandexamples/data-viz-chart/build.py— both examples used/System/Library/Fonts/Supplemental/Arial Bold.ttfdirectly, which would crash withFileNotFoundErroron Linux and Windows. Replaced with aplatform.system()-keyed fallback chain that tries OS-appropriate sans-serif fonts (Arial on macOS and Windows, DejaVu Sans on Linux, with multiple candidate paths each) and falls back toImageFont.load_default()if nothing matches, so the examples never crash on unusual setups. Both examples now genuinely run cross-platform as their READMEs promise. - Latent Pillow gap in
tools/flux2.pyandtools/image_edit.py— both tools imported PIL inside atry/except ImportErrorguard with a friendly error message, but Pillow was never actually declared intools/requirements.txt. Users had to install Pillow manually after hitting the runtime error. Free-rider fix via the moviepy declaration. - Guarded imports added to both moviepy example
build.pyfiles — matches the existingflux2.py/image_edit.pyfriendly-error pattern. Prevents bareModuleNotFoundErrortracebacks for users who run the examples before installing dependencies.
Docs
- README examples table adds
q2-townhall-longarm-ad(Super Bowl-style launch ad with LTX-2 animated Lugh cameo) andq2-townhall-stars(GitHub star history time-lapse) as finished demo entries dated 2026-04-08. - moviepy skill added to the README's skills table and
_internal/toolkit-registry.json.
2026-04-04 (v0.14.1)
Added
tools/chain_video.py— Chain LTX-2 clips with visual continuity (each scene uses previous clip's last frame as input). Supports per-scene prompts via JSON, auto-resume of interrupted runs, and structured progress reporting.- Structured progress reporting (
--progress json) for all cloud GPU tools - OpenClaw skill: yieldMs polling — agent stays in poll loop for live progress instead of going silent with
background:true - OpenClaw skill: style drift prevention — rules to prevent anime/Asian style drift in chained LTX-2 sequences (per-scene prompts + negative prompts)
Changed
- acemusic cloud API is now default music generation provider (official API, XL Turbo 4B model with 5Hz LM thinking)
- SadTalker timeouts increased: 20x multiplier + 300s buffer for size=512+gfpgan on A10G
- Synced official Remotion skills (upstream d5d3955)
Fixed
- chain_video: glob matching for manually-renamed clips (e.g.
chain-05-brigid.mp4) - chain_video: 20-min subprocess timeout, stderr in error messages, JSON validation, temp file cleanup
- chain_video: registered in toolkit-registry.json
- R2 recovery note for SadTalker client timeouts
2026-03-26 (v0.13.2)
Changed
- Default cloud provider switched to Modal — All 6 cloud GPU tools (flux2, image_edit, music_gen, qwen3_tts, sadtalker, voiceover) now default to
--cloud modalinstead of--cloud runpod. RunPod remains available via--cloud runpod.
Fixed
- Remove unused imports in
tools/ltx2.py(json,requests) — flagged by code quality bot
2026-03-25 (v0.13.1)
Added
tools/ltx2.py— AI video generation using LTX-2.3 22B DiT model- Text-to-video generation from prompts (~5s clips at up to 1024x1536)
- Image-to-video — animate still images with motion prompts
- Joint audio+video generation (synchronized ambient audio)
- Quality presets:
standard(30 steps) andfast(15 steps) - Valid frame counts: 25-193 frames ((n-1)%8==0), 24fps default
- Modal deployment on A100-80GB with baked weights (~55GB)
- Estimated cost: ~$0.23 per 5-second clip
- LTX-2 skill (
.claude/skills/ltx2/) — prompting guide, parameters, video production use cases - Openclaw skill updated — LTX-2 added as section 4d (video clips), setup instructions, cost table
Technical Notes
torch.inference_mode()required around pipeline call — without it, Gemma text encoder retains ~37GB of autograd activations causing OOM (upstream bug: Lightricks/LTX-2#152)- Weights hosted on
Lightricks/LTX-2.3HuggingFace repo (separate from the code repoLightricks/LTX-2) transformerspinned to 4.57.6 (5.x breaksGemma3TextConfig.rope_local_base_freq)- HuggingFace token required for both LTX-2 (rate limiting) and Gemma 3 (gated model)
2026-03-22 (v0.12.0)
Added
tools/music_gen.py— AI music generation using ACE-Step 1.5- Text-to-music with precise BPM, key, time signature, and duration control
- Vocal music with lyrics in 50+ languages
- Cover / style transfer from reference audio (
--cover) - Stem extraction — isolate vocals, drums, bass, etc. (
--extract) - 8 scene presets for video production:
corporate-bg,upbeat-tech,ambient,dramatic,tension,hopeful,cta,lofi - Brand-aware generation —
--brandloads style hints from brand.json --setupcreates RunPod template + endpoint via GraphQL- Docker image:
ghcr.io/conalmullan/video-toolkit-acestep:latest(CUDA 12.8, baked model weights) - MIT licensed model — free alternative to ElevenLabs Music with more control
- ~2-3s inference on GPU (turbo mode, 8 steps)
- ACE-Step skill (
.claude/skills/acestep/) — prompt engineering patterns, lyrics formatting, scene preset guide, video production integration
2026-03-22 (v0.11.1)
Fixed
- Queue timeout safeguard — All 6 RunPod tools now auto-cancel jobs stuck in queue for >5 min, preventing runaway billing from GPU unavailability
- Cancel on exit — Jobs are cancelled on RunPod when client polling times out (previously left orphaned jobs in queue indefinitely)
- R2 upload support — flux2.py and image_edit.py now pass R2 config to handlers (was already supported server-side but never triggered)
Changed
- flux2 default GPU — Changed from ADA_24 (RTX 4090, frequently throttled) to AMPERE_24,ADA_24 (fallback chain)
- flux2 setup —
--setup-gpunow accepts comma-separated GPU types for fallback - RunPod skill — Added API reference (GraphQL queries/mutations, REST endpoints, GPU IDs, R2 CLI patterns)
- R2 docs — Added operations guide (listing, cleanup,
--region autogotcha) to runpod-setup.md
2026-03-15 (v0.11.0)
Added
tools/flux2.py— AI image generation and editing using FLUX.2 Klein 4B- Text-to-image generation from prompts (~2.5s fast mode, ~8s quality mode)
- Image editing with reference images (
--input photo.jpg --prompt "...") - Multi-image compositing (up to 3 reference images)
- 8 scene presets for video production:
title-bg,problem,solution,demo-bg,stats-bg,cta,thumbnail,portrait-bg - Brand-aware generation —
--brand digital-sambareads brand.json and injects color palette into prompts - Preset + prompt layering — preset provides style/mood,
--promptadds subject context --setupcreates RunPod template + endpoint via GraphQL--list-presetsshows all available scene presets- Docker image:
ghcr.io/conalmullan/video-toolkit-flux2:latest(baked model weights, ~15GB) - Apache 2.0 licensed model (commercial OK)
- "The Space Between" — showcase video demonstrating end-to-end AI video creation
- Avatar generated with flux2, voiced with Qwen3-TTS, animated with SadTalker, composed in Remotion
- Concept, script, voice, visuals — all AI-generated
- New video format: video essay (beyond sprint reviews and product demos)
Changed
- Baked Docker images rebuilt — both
flux2andqwen3-ttsimages now include model weights baked in, eliminating cold-start model downloads (~30s cold start vs minutes) - Updated
_internal/toolkit-registry.jsonwith flux2 tool entry, cloud endpoint, scene presets, and showcase example - Updated README with flux2 tool, Docker image, demo table entry, and Cloud GPU tools table
2026-02-25 (v0.10.1)
Added
tools/sync_timing.py— Audio-to-config timing sync tool- Measures actual per-scene audio durations via ffprobe
- Auto-detects config file and template type (sprint-review v1/v2, product-demo)
- 3-pass audio-to-scene matching:
audioFilefield → index → name - Comparison table with delta indicators; skips changes < 0.3s
--applyupdatesdurationSecondsin-place (creates.bakbackup)--voiceover-jsonaccepts voiceover.py output directly- Suggests
playbackRateadjustments for demo scenes
Changed
- CLAUDE.md slimmed 44% (861 → 480 lines)
- Removed catalog data duplicated in
toolkit-registry.json(skills, commands, components, transitions, presets, Docker images, duplicate CLI examples) - Added cross-references to registry for structured data
- All workflow guidance, timing knowledge, code patterns, and tool-specific gotchas retained
- Removed catalog data duplicated in
- Integrated
sync_timing.pyinto Video Production Workflow (step 7) and TTS drift feedback loop - Updated toolkit-registry.json with
sync_timingtool entry
2026-02-24 (v0.10.0)
Added
- Qwen3-TTS integration — Self-hosted TTS via RunPod as free alternative to ElevenLabs
tools/qwen3_tts.pystandalone CLI with 9 speakers, tone presets, voice cloningvoiceover.py --provider qwen3for per-scene generation- Docker image:
ghcr.io/conalmullan/video-toolkit-qwen3-tts:latest - Temperature/top_p params for expressiveness control
/voice-clonecommand — Record, test, and save a cloned voice to a brand profilesprint-review-v2template — Composable scene-based architecture for sprint reviewsFilmGraincomponent — SVG noise overlay for cinematic film texturehello-worldexample — Minimal 25s video, zero config, renders in 2 minutesrunpodskill — Cloud GPU setup, Docker images, endpoint management, troubleshooting- Remotion skills sync update — 6 new upstream rules (audio-visualization, ffmpeg, light-leaks, subtitles, transparent-videos, voiceover)
Improved
- Onboarding & developer experience — All API keys now optional, videos render with just Node.js
.env.examplereorganized with sections, all keys commented out by default- README and getting-started.md restructured with "Try It Now" quick start
- All Python tools show actionable guidance when API keys are missing (add key, use alternative, or skip)
voiceover.pycatches missingelevenlabspip package with helpful message pointing to Qwen3-TTS
Changed
- Updated README with prerequisites table, Qwen3-TTS, Docker images, voice-clone command
- Updated CLAUDE.md with FilmGrain component and checkerboard transition
- Updated roadmap metrics and skill status table
- Fixed cho-oyu demo link; added cortina sprint to demos table
2026-02-19 (v0.9.3)
Added
- Official Remotion skills — Synced 33 rule files from remotion-dev/skills into
.claude/skills/remotion-official/ - Weekly sync workflow — GitHub Actions checks upstream every Monday and opens a PR if files changed
- Sync documentation —
docs/remotion-skills-sync.mdexplaining the split and sync process
Changed
- Remotion skill split — Custom
remotionskill trimmed from ~470 to ~160 lines, now covers only toolkit-specific patterns (transitions, components, conventions). Core framework knowledge deferred toremotion-official - Updated CLAUDE.md skills table with both remotion skills
2026-01-25 (v0.9.2)
Added
-
Per-scene voiceover generation - New recommended workflow for audio
voiceover.py --scene-dirprocesses all.txtfiles in a directoryvoiceover.py --concatjoins scene audio for SadTalker narrator- Each scene's
<Audio>starts at frame 0 within itsSeries.Sequence - Regenerate individual scenes without re-doing the whole video
- Scene durations match audio naturally (no manual offset calculations)
-
Sprint-review template per-scene audio support
audioFilefield onDemoConfig,info,overview,summary- SprintReview.tsx renders per-scene
<Audio>elements - Backward compatible: global voiceover track still works
-
Updated
/generate-voiceovercommand- Detects
public/audio/scenes/*.txtand offers per-scene mode - Per-scene is now the default when scene scripts exist
- Concat option for SadTalker integration
- Detects
Changed
- Updated documentation (CLAUDE.md, README.md, getting-started.md)
2025-12-30 (v0.7.0)
Added
-
tools/dewatermark.py- AI-powered watermark removal using ProPainter- Cloud GPU processing via RunPod serverless (~$0.05-0.30/video)
- Local processing for NVIDIA GPU users (8GB+ VRAM)
- Auto resize-ratio: 1.0 for <30s, 0.75 for <1min, 0.5 for longer
- Chunked processing for long videos with overlap stitching
- Cloudflare R2 for reliable file transfer
- Automated
--setupfor one-command RunPod configuration - Presets: notebooklm, tiktok, sora, stock-br, stock-bl, stock-center
-
tools/locate_watermark.py- Helper tool for finding watermark coordinates- Coordinate grid overlay for visual identification
- Region verification across multiple frames
- Preset support for common watermarks
- Requires ImageMagick (
brew install imagemagick)
-
tools/notebooklm_brand.py- Post-processing for NotebookLM videos- Trims NotebookLM visual outro while preserving full audio
- Adds custom branded outro with logo and URL
- Handles freeze-frame bridging when audio exceeds video
-
Docker infrastructure for RunPod serverless (
docker/runpod-propainter/)- Dockerfile, handler.py, README
- Public image available for quick deployment
-
Documentation
docs/optional-components.md- ML component installation guidedocs/runpod-setup.md- Cloud GPU setup guide
Changed
- ElevenLabs TTS: Added
eleven_v3(alpha) model option - Updated
CLAUDE.mdwith dewatermark and locate_watermark documentation - Updated
.env.examplewith RunPod and R2 configuration
2025-12-28 (v0.6.0)
Added
redub.py --sync- Word-level time remapping for TTS synchronization- Uses ElevenLabs
convert_with_timestampsAPI for character-level alignment - Uses Scribe word-level timestamps from original audio
- Builds segment mapping (configurable via
--segment-size, default: 15 words) - Applies variable speed per segment via FFmpeg filtergraph
- Fixes audio drift when TTS voice speaks at different pace than original
- Solves: TTS often starts fast and ends slow, causing 3-4+ second drift
- Uses ElevenLabs
Changed
- Updated
CLAUDE.mdwith Redub Sync Mode documentation - TTS duration now measured via ffprobe (more accurate than timestamp data)
2025-12-28 (v0.5.0)
Added
tools/addmusic.py- Add background music to existing videos- Generate music via ElevenLabs or use existing audio file
- Options:
--music-volume,--fade-in,--fade-out - Works on any video file without requiring a project structure
Changed
- Updated
CLAUDE.mdwith addmusic documentation
2025-12-28 (v0.4.0)
Added
-
/redubcommand - Redub existing videos with a different voice- Guided workflow for voice selection and transcript handling
- Supports transcript review/editing before TTS generation
- Works on any video file without requiring a project structure
-
tools/redub.py- Voice replacement utility tool- Extracts audio from video (FFmpeg)
- Transcribes audio to text (ElevenLabs Scribe STT)
- Generates new audio with target voice (ElevenLabs TTS)
- Replaces audio track in video (FFmpeg)
- Options:
--save-transcript,--transcript,--keep-temp,--speed - Warns about duration mismatches between original and new audio
-
Utility tools concept - Documented distinction between project tools and utility tools
- Project tools (voiceover, music, sfx): Used during video creation workflow
- Utility tools (redub): Quick transformations on existing videos, no project needed
Changed
- Updated
CLAUDE.mdwith utility tools documentation - Updated Python Tools section to include redub
2025-12-20
Added
-
GitHub Actions release automation (
.github/workflows/release.yml)- Tag-triggered releases (
v*tags) - Manual dispatch option with version input
- Auto-generates changelog from commits since last tag
- Reads version from
toolkit-registry.json
- Tag-triggered releases (
-
/versionscommand - Check dependency versions and toolkit updates- Detects Remotion package version mismatches in projects
- Compares local toolkit version against GitHub releases
- Offers to fix version mismatches by pinning and reinstalling
- Documents common issues and prevention strategies
Fixed
- Remotion version mismatch in digital-samba-free-intro project
- Pinned all Remotion packages to 4.0.387 (was mixing 4.0.383 and 4.0.387)
- Removed
^prefix to prevent future drift
2025-12-10
Added
/scene-reviewcommand - Dedicated scene-by-scene review with Remotion Studio- Starts Remotion Studio for visual verification
- Walks through scenes one by one (not summary tables)
- Generic - works with any template's config
/videonow delegates to/scene-reviewwhen phase isreview/generate-voiceoverwarns if review not complete- Fixes: Review kept getting skipped because
/videocommand was too long
Changed
- Consolidated tracking files - Simplified from 4 files to 3:
ROADMAP.md- What we're building (removed duplicate "Next Actions" and "In Progress" sections)BACKLOG.md- What we might build (removed all implemented items, now only future ideas)CHANGELOG.md- What we built (historical record)- Removed
FEEDBACK.md- Evolution principles moved todocs/contributing.md
- Created
docs/contributing.mdwith evolution principles and contribution workflow
Fixed
- Slash commands not loading - Renamed
/skillto/skillsto avoid conflict with built-inSkilltool. The naming collision was silently preventing ALL custom commands from loading. Bug reported to Anthropic.
Removed
/reviewcommand - Clashed with Claude Code's built-in PR review command. Replaced by/scene-review.
Added
- Animation components (
lib/components/)- Envelope - 3D envelope with opening flap animation, configurable message
- PointingHand - Animated hand emoji with directional slide-in and pulse effect
/contributecommand - Guided contribution workflow- Report issues via
gh issue create - Submit PRs for improvements
- Share skills and templates
- Safety checks to exclude private project work
- Report issues via
- Evolution narrative across all commands
- Consistent "## Evolution" section in each command
- Local improvement workflow
- Remote contribution links (GitHub issues + PRs)
- Command history tracking
- Template evolution guidance in
/templatecommand- How to add features to existing templates
- Template maturity indicators
- Pattern extraction to shared lib/
- Product demo template README (
templates/product-demo/README.md)- Quick start guide, configuration, project structure
- All 7 scene types documented with examples
- Narrator PiP and demo chrome options
Changed
projects/now fully gitignored - User video work stays completely private- Only
.gitkeepis tracked to preserve directory structure - Safe to contribute without exposing project content
- Only
- New
examples/directory for shareable showcase projects- Configs, scripts, and docs are tracked
- Large media files (mp4, mp3) are gitignored
- Each example includes
ASSETS-NEEDED.mddocumenting required media
/contributenow supports example projects (Option 5)- Guides copying from projects/ to examples/
- Auto-generates ASSETS-NEEDED.md
- Creates README with quick start instructions
- Contributor recognition with backlinks to website/org
- CONTRIBUTORS.md - Recognition for organizations and individuals who share examples
- Documentation updates
docs/getting-started.md- Updated for/videocommand, added multi-session workflowdocs/creating-brands.md- Updated for/brandcommand integration
- Renamed
skills-registry.json→toolkit-registry.json- Consistent format across all entries (path, description, status, created, updated)
- Added
componentssection for shared lib/components - Synced with actual commands (7), templates (2), components (9)
2025-12-09
Added
-
Multi-session project system (
lib/project/)types.ts- TypeScript schema for project.jsonREADME.md- Documentation for project lifecycle and phases- Projects now track: phase, scenes, assets, audio, session history
- Filesystem reconciliation (compares intent vs reality)
- Auto-generated CLAUDE.md per project for instant context
-
Unified commands - Context-aware entry points that list existing items or create new:
/video- Replaces/new-video. Scans projects, offers resume or new/brand- Replaces/new-brand. Lists brands, edit or create new/template- Lists templates or creates new ones (copy, minimal, from project)/skill- Lists installed skills or creates new ones
Changed
- Command pattern unified - All domain commands now scan first, then offer actions
- Commands integrate with project.json for state tracking
- README.md updated with new commands and multi-session workflow
Removed
/new-video- Replaced by/video/new-brand- Replaced by/brand
Notes
- After creating/modifying commands or skills, restart Claude Code to load changes
2025-12-09
Added
- Shared component library (
lib/)lib/theme/- ThemeProvider, useTheme, createStyles, type definitionslib/components/- Reusable video components:- AnimatedBackground (supports subtle, tech, warm, dark variants)
- SlideTransition (fade, zoom, slide-up, blur-fade)
- Label (floating badge with optional JIRA ref)
- Vignette (cinematic edge darkening)
- LogoWatermark (corner branding)
- SplitScreen (side-by-side video layout)
- NarratorPiP (picture-in-picture presenter - needs refinement)
lib/index.ts- Unified exports for templates
Changed
- Templates now import from shared library
- sprint-review: Uses lib components via re-exports
- product-demo: Uses lib components (keeps local NarratorPiP for different API)
- Both templates use shared ThemeProvider from lib/theme
- Deleted duplicate component files from templates
- Updated ROADMAP.md - Shared component library marked complete
- Updated BACKLOG.md - Added NarratorPiP API refinement task
2025-12-09 (earlier)
Added
- Product demo template (
templates/product-demo/)- Scene-based composition (title, problem, solution, demo, stats, CTA)
- Config-driven content via
demo-config.ts - Dark tech aesthetic with animated background
- Narrator PiP (picture-in-picture presenter)
- Browser/terminal chrome for demo videos
- Stats cards with spring animations
/new-brandcommand - guided brand profile creation- Extract colors from website URL
- Manual color entry with palette generation
- Logo and voice configuration guidance
- Digital Samba brand profile (
brands/digital-samba/) - Template-brand integration
- Brand loader utility (
lib/brand.ts) brand.tsin each template (generated at project creation)project.jsonfor tracking project metadata- Brand generator script (
lib/generate-brand-ts.ts)
- Brand loader utility (
/new-videocommand - unified project creation wizard- Choose template (sprint-review, product-demo)
- Choose brand from available brands
- Scene-centric workflow:
- Content gathering (URLs, notes, paste)
- Claude proposes scene breakdown
- Interactive scene refinement
- Scene types: title, overview, demo, split-demo, stats, credits, problem, solution, feature, cta
- Visual types:
[DEMO],[SCREENSHOT],[EXTERNAL VIDEO],[SLIDE]
- Generates VOICEOVER-SCRIPT.md with narration and asset checklist
- Creates project.json with scene tracking
- Guides through asset creation phase
- Ongoing project support (resume any project)
Changed
- Replaced
/new-sprint-videowith/new-video- single entry point for all templates - Templates now load theme from
brand.tsinstead of hardcoded values - Updated CLAUDE.md with new workflow and commands
- Updated ROADMAP.md - Phase 3 template-brand integration complete
- Added Video Timing section to CLAUDE.md - pacing rules, scene durations, timing calculations
- Removed
/convert-assetfrom backlog - FFmpeg skill handles this conversationally - Removed
/sync-timingfrom backlog - timing knowledge now in CLAUDE.md
2025-12-08
Added
- Open source release - Published to GitHub at digitalsamba/claude-code-video-toolkit
- Brand profiles system (
brands/)brand.jsonfor colors, fonts, typographyvoice.jsonfor ElevenLabs voice settingsassets/for logos and backgrounds- Default brand profile included
- Documentation (
docs/)getting-started.md- First video walkthroughcreating-brands.md- Brand profile guidecreating-templates.md- Template creation guide
- Environment variable support
ELEVENLABS_VOICE_ID- Set voice ID via env var- Falls back to
_internal/skills-registry.jsonif not set
/generate-voiceovercommand - guided ElevenLabs TTS generation/record-democommand - guided Playwright browser recording- Interactive recording stop controls (Escape key, Stop button)
- Window scaling for laptop screens (
--scaleoption, default 0.75) - FFmpeg skill (beta) - common video/audio conversion commands
- Playwright recording skill (beta) - browser demo capture
- Playwright infrastructure (
playwright/) with recording scripts - Python tools:
voiceover.py,music.py,sfx.py - Skills registry for centralized config
- README.md, LICENSE (MIT), CONTRIBUTING.md
.env.exampletemplate
Changed
- Directory restructure for open source:
templates/- Video templates (moved from root)projects/- User video projects (moved from root)brands/- Brand profiles (new)assets/- Shared assets (consolidated)_internal/- Toolkit metadata (renamed from_toolkit/)
- Updated
/new-sprint-videocommand paths tools/config.pyreads from_internal/and supports env vars- Playwright recordings output at 30fps (matches Remotion)
Fixed
- Removed hardcoded voice ID from committed files
- Proper
.gitignorefor secrets and build artifacts - FFmpeg trim command syntax (use
-tonot-tfor end time) - Playwright double navigation issue
- Recording frame rate mismatch (was 25fps, now 30fps)
2025-12-04
Added
- Sprint review template (
templates/sprint-review/)- Theme system with colors, fonts, spacing
- Config-driven content via
sprint-config.ts - Slide components: Title, Overview, Summary, EndCredits
- Demo components: DemoSection, SplitScreen
- NarratorPiP component for picture-in-picture narrator
- Audio integration (voiceover, background music, SFX)
/new-sprint-videoslash command for guided project creation- Initial workspace setup
- Remotion skill documentation
- ElevenLabs skill documentation
- First video project: sprint-review-cho-oyu
- Voice cloning workflow with ElevenLabs