mirror of
https://github.com/digitalsamba/claude-code-video-toolkit.git
synced 2026-09-18 19:41:13 +08:00
REMOVE: EchoMimicV3 — drift made it unusable, SoulX is the default (#81)
EchoMimicV3 was added earlier in this same unreleased cycle and never reached a tagged release, so this removes it rather than deprecating it — no public API changes. Identity drift is why. On a controlled A/B (same photo, same 80s audio, same 544x736) it fell to 50% of frame-zero sharpness by 70s, glasses dissolving around 55s and no recognisable face by 67s, where SoulX-FlashHead held 97%. The failure is absorbing rather than gradual — each segment re-anchors on the previous segment's output, so one bad segment poisons everything after it — and colour correction recovers none of it because the damage is structural. That implied a ~30s render ceiling which shaped the surrounding design. The one capability SoulX lacks is text-prompt and CFG steering, and our own docs already recorded from testing that the prompt is close to inert. Everything else EchoMimicV3 offered, SoulX matches or beats: aspect preservation, motion quality, ~3.7x cheaper, ~3-6x faster. The genuine loss is gesture and upper-body motion, which nothing in the toolkit currently uses. Also in this commit: - Talking head docs, registry entries (tools and modal endpoints), cloud_gpu dispatch, modal-setup, NarratorPiP's comment and both env files repointed to soulx. cloud_gpu also gains a GPU tier for it, which echomimic3 never had -- its jobs silently reported no cost estimate. - The Unreleased Kiro changelog entry is dropped; it shipped in v0.19.0 and was flagged in-file for removal when this section was cut. Findings worth keeping outlived the tool and moved into docs/soulx.md: the warning against scoring talking heads with similarity metrics (two have now misled — one ranked highest a variant with a visible eye defect, the other plateaued straight through a total collapse), the volume-weights rationale, and the image guidelines. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -34,6 +34,9 @@
|
||||
# LTX-2 also needs a Modal secret named `huggingface-token` (gated weights) and an
|
||||
# A100-80GB, which requires a payment method on your Modal account.
|
||||
# MODAL_LTX2_ENDPOINT_URL=https://yourname--video-toolkit-ltx2-....modal.run
|
||||
# SoulX-FlashHead talking head. Needs a one-off weights fetch before first use:
|
||||
# modal run docker/modal-soulx/app.py::populate_weights
|
||||
# MODAL_SOULX_ENDPOINT_URL=https://yourname--video-toolkit-soulx-....modal.run
|
||||
|
||||
|
||||
# --- Cloud GPU: RunPod (alternative) ---
|
||||
|
||||
@@ -130,7 +130,7 @@ This is especially critical for background commands where the working directory
|
||||
|------|-------|-------------|
|
||||
| **Project tools** | voiceover, music, music_gen, sfx, sync_timing | During video creation workflow |
|
||||
| **Utility tools** | redub, addmusic, notebooklm_brand, locate_watermark | Quick transformations on existing videos |
|
||||
| **Cloud GPU** | image_edit, upscale, dewatermark, sadtalker, echomimic3, qwen3_tts, music_gen, flux2 | AI processing via RunPod or Modal (`--cloud runpod\|modal`; echomimic3 is Modal-only) |
|
||||
| **Cloud GPU** | image_edit, upscale, dewatermark, sadtalker, soulx, qwen3_tts, music_gen, flux2 | AI processing via RunPod or Modal (`--cloud runpod\|modal`; soulx is Modal-only) |
|
||||
| **Publishing** | youtube_upload | Upload a finished render to YouTube (use `/publish` for the guided workflow) |
|
||||
|
||||
Utility tools work on any video file without requiring a project structure.
|
||||
@@ -340,47 +340,60 @@ uv run tools/dewatermark.py --setup # One-time setup
|
||||
|
||||
**Local mode** requires NVIDIA GPU (8GB+ VRAM). Mac users should use `--runpod`.
|
||||
|
||||
### Talking Head Generation (EchoMimicV3 vs SadTalker)
|
||||
### Talking Head Generation (SoulX-FlashHead vs SadTalker)
|
||||
|
||||
Two generators, and the deciding factor is **how big the narrator is on screen and how long
|
||||
the viewer looks at it**.
|
||||
Two generators, and the deciding factor is **whether a viewer actually watches the shot**.
|
||||
|
||||
```bash
|
||||
# EchoMimicV3 — diffusion, follows the input aspect ratio, ~6.5x the cost (Modal only)
|
||||
uv run tools/echomimic3.py --image presenter_16x9.png --audio voiceover.mp3 \
|
||||
--steps 5 --size 640 --output narrator.mp4
|
||||
# SoulX-FlashHead — the default. Diffusion, follows the input aspect ratio (Modal only)
|
||||
uv run tools/soulx.py --image presenter_16x9.png --audio voiceover.mp3 \
|
||||
--size 768 --output narrator.mp4
|
||||
|
||||
# SadTalker — warp-based, fast and cheap, square crop unless --preprocess full
|
||||
# SadTalker — warp-based, cheap and near-realtime; square crop unless --preprocess full
|
||||
uv run tools/sadtalker.py --image presenter_16x9.png --audio voiceover.mp3 \
|
||||
--preprocess full --still --expression-scale 0.8 --output narrator.mp4
|
||||
```
|
||||
|
||||
| Need | Use |
|
||||
|------|-----|
|
||||
| Narrator large in frame, or a shot held long enough to watch | **echomimic3** |
|
||||
| Small PiP overlay, drafts, or many takes to choose between | **sadtalker** |
|
||||
| Non-square source image you don't want to fight | **echomimic3** (no `--preprocess` needed) |
|
||||
| Anything a viewer watches — narrator in frame, held shot, finished video | **soulx** |
|
||||
| Throwaway drafts, or many takes to choose between | **sadtalker** |
|
||||
| Non-square source image you don't want to fight | **soulx** (no `--preprocess` needed) |
|
||||
|
||||
**Cost is the trade-off, and wall clock more than money.** EchoMimicV3 is ~$0.009 per second
|
||||
of output against SadTalker's ~$0.0014 — a 3-minute narrator is ~$1.76 vs ~$0.27. But it also
|
||||
runs at 22.8-47.8x realtime, so that same 3 minutes is **1.5-2.4 hours** of generation.
|
||||
Generate per-scene narrator clips ahead of time rather than one long render.
|
||||
**Identity holds over long takes, which is why this is the default.** Segment-chained
|
||||
talking heads re-anchor each segment on the previous segment's output, so the failure is
|
||||
*absorbing*: one bad segment poisons everything after it. SoulX is trained with
|
||||
Oracle-Guided Bidirectional Distillation against exactly that. Measured at **97% of
|
||||
frame-zero sharpness at 70s**, flat across all 72 segments. There is no short-render
|
||||
ceiling to design around, so per-scene generation is a choice rather than a workaround.
|
||||
|
||||
**Key flags for NarratorPiP:**
|
||||
- echomimic3: `--steps 5 --size 640` — the cheap pass; 16:9 in gives 16:9 out, no crop workaround
|
||||
- sadtalker: `--preprocess full` — **Critical!** Preserves input dimensions (default `crop` outputs square)
|
||||
**Cost:** ~$0.0024 per second of output against SadTalker's ~$0.0014 — only ~1.7x, so cost
|
||||
is rarely the deciding factor between them any more.
|
||||
|
||||
**Key flags:**
|
||||
- soulx: `--size 768` — aspect follows the image and snaps to the model's grid
|
||||
- soulx: `--width`/`--height` for exact dimensions; **both must be multiples of 16**, and
|
||||
nothing upstream validates that (an off-grid size floors silently and renders wrong)
|
||||
- sadtalker: `--preprocess full` — **Critical!** Preserves input dimensions (default `crop`
|
||||
outputs square)
|
||||
- sadtalker: `--still` and `--expression-scale 0.8` — calmer, more professional look
|
||||
|
||||
**Image requirements (both):** Face 30-70% of frame, front-facing, 16:9 for NarratorPiP, 512px+.
|
||||
|
||||
**Gotchas that cost hours** (full list in `docs/echomimic3.md`):
|
||||
- EchoMimicV3's `transformers==4.49.0` pin is load-bearing. A newer version silently removes
|
||||
**all lip sync** — no error, just a dead mouth.
|
||||
- Keep `--wav2vec chinese` even for English audio; the `english` encoder under-articulates.
|
||||
- Don't score talking-head quality with a mouth-crop metric. It ranked highest the one
|
||||
variant with a visible eye defect. Whole-face, or human review.
|
||||
**Gotchas that cost hours** (full list in `docs/soulx.md`):
|
||||
- `torch.compile` is on upstream, costs ~600s on a cold container, and **recompiles on every
|
||||
resolution change**. Pick one narrator resolution per project; a batch at one size
|
||||
amortises it, switching size per scene pays it per scene.
|
||||
- Width and height must be multiples of 16. `--size` handles this; exact dimensions do not
|
||||
forgive you.
|
||||
- `--face-crop` is unconditionally square whatever the target size says, so it throws away a
|
||||
16:9 framing. Leave it off.
|
||||
- Don't score talking-head quality with an automated similarity metric. Two have now given
|
||||
confidently wrong answers — one ranked highest a variant with a visible eye defect, the
|
||||
other plateaued straight through a total collapse. Sharpness-over-time plus a contact
|
||||
sheet, or human review.
|
||||
|
||||
See `docs/echomimic3.md` and `docs/sadtalker.md` for detailed options and troubleshooting.
|
||||
See `docs/soulx.md` and `docs/sadtalker.md` for detailed options and troubleshooting.
|
||||
|
||||
### Redub Sync Mode
|
||||
|
||||
|
||||
@@ -206,8 +206,8 @@ uv run tools/image_edit.py --input photo.jpg --prompt "Add sunglasses" --cloud m
|
||||
# AI video generation (LTX-2.3 — text-to-video, image-to-video)
|
||||
uv run tools/ltx2.py --prompt "A sunset over the ocean, cinematic" --cloud modal
|
||||
|
||||
# Talking head from a portrait + audio (SadTalker)
|
||||
uv run tools/sadtalker.py --image portrait.png --audio voiceover.mp3 --output talking.mp4 --cloud modal
|
||||
# Talking head from a portrait + audio (SoulX-FlashHead)
|
||||
uv run tools/soulx.py --image portrait.png --audio voiceover.mp3 --output talking.mp4
|
||||
```
|
||||
|
||||
<details>
|
||||
@@ -254,8 +254,8 @@ uv run tools/dewatermark.py --input video.mp4 --preset sora --output clean.mp4 -
|
||||
# Locate watermark coordinates
|
||||
uv run tools/locate_watermark.py --input video.mp4 --grid --output-dir ./review/
|
||||
|
||||
# Generate talking head video from image + audio (SadTalker)
|
||||
uv run tools/sadtalker.py --image portrait.png --audio voiceover.mp3 --output talking.mp4 --cloud modal
|
||||
# Generate talking head video from image + audio (SoulX-FlashHead)
|
||||
uv run tools/soulx.py --image portrait.png --audio voiceover.mp3 --output talking.mp4
|
||||
|
||||
# AI image generation (FLUX.2 Klein 4B — text-to-image + editing)
|
||||
uv run tools/flux2.py --prompt "A sunset over mountains" --cloud modal
|
||||
@@ -279,7 +279,7 @@ uv run tools/youtube_upload.py --video out/video.mp4 --title "My video" --privac
|
||||
|------|-------|---------|
|
||||
| **Project** | voiceover, music, music_gen, sfx | Used during video creation workflow |
|
||||
| **Utility** | redub, addmusic, notebooklm_brand, locate_watermark | Quick transformations, no project needed |
|
||||
| **Cloud GPU** | image_edit, upscale, dewatermark, sadtalker, qwen3_tts, flux2, music_gen, ltx2 | AI processing via Modal or RunPod |
|
||||
| **Cloud GPU** | image_edit, upscale, dewatermark, sadtalker, soulx, qwen3_tts, flux2, music_gen, ltx2 | AI processing via Modal or RunPod |
|
||||
| **Publishing** | youtube_upload | Upload a finished render to YouTube (or use `/publish`) |
|
||||
|
||||
### Cloud GPU (Modal + RunPod)
|
||||
@@ -293,7 +293,8 @@ uv run tools/youtube_upload.py --video out/video.mp4 --title "My video" --privac
|
||||
| `image_edit` | AI image editing & style transfer | ~$0.03 |
|
||||
| `upscale` | AI image upscaling (2x/4x) | ~$0.01 |
|
||||
| `music_gen` | AI music generation (8 scene presets) | Free (acemusic) / ~$0.05 (self-hosted) |
|
||||
| `sadtalker` | Talking head video from portrait + audio | ~$0.10 |
|
||||
| `soulx` | Talking head video from portrait + audio — holds identity over long takes | ~$0.0024/sec |
|
||||
| `sadtalker` | Talking head video, warp-based — fast, cheap drafts | ~$0.10 |
|
||||
| `ltx2` | AI video generation (text-to-video, image-to-video) | ~$0.23 |
|
||||
| `dewatermark` | Video watermark removal | ~$0.10 |
|
||||
|
||||
|
||||
+29
-19
@@ -9,28 +9,38 @@ All notable changes to claude-code-video-toolkit.
|
||||
## Unreleased
|
||||
|
||||
### Added
|
||||
- **EchoMimicV3 talking head** (`tools/echomimic3.py` + `docker/modal-echomimic3/`) — diffusion-based
|
||||
audio-driven talking head (Ant Group, Apache 2.0), Modal-only. Preserves the input aspect ratio, so
|
||||
16:9 presenter images come back 16:9 with no `--preprocess` workaround. Runs alongside SadTalker
|
||||
rather than replacing it: ~$0.009/second of output against ~$0.0014, and 22.8-47.8x realtime, so
|
||||
SadTalker stays the right call for small overlays and drafts. Decision table in CLAUDE.md, full
|
||||
detail in `docs/echomimic3.md`. (#77)
|
||||
- **`--anchor-retreat`** — fixes a bug where a segment seam landing mid-blink made the next segment
|
||||
start closed-eyed and hold it. Anchor windows are now scored by upper-frame motion and the calmest
|
||||
is chosen, backing off up to N frames. `0` restores the old behaviour. (#77)
|
||||
<!-- NOTE: the Kiro entry below already shipped in v0.19.0; left in place rather than silently
|
||||
dropped, but it should come out when this section is cut into a release. -->
|
||||
- **Kiro CLI support** (`scripts/migrate_to_kiro.py`) — sibling of the Codex migration script. Installs the toolkit skills into `~/.kiro/skills` (Kiro shares Claude Code's `SKILL.md` frontmatter format, so they copy verbatim), generates a wrapper skill per `.claude/commands/*.md` invoked as the same `/video`, `/setup`, … slash commands, and generates `.kiro/steering/video-toolkit.md` from `CLAUDE.md` inside a managed marker block. Wrappers pin the toolkit's absolute path so commands work from any directory (Claude Code parity — Kiro doesn't walk up the directory tree). Supports `--force`, `--dry-run`, `--reset`, `--workspace-skills`, and `kiro/migration_map.json` for skips/renames. See `docs/kiro.md`.
|
||||
- **SoulX-FlashHead talking head** (`tools/soulx.py` + `docker/modal-soulx/`) — the toolkit's
|
||||
default talking head generator (Soul AI Lab, Apache 2.0, 1.3B), Modal-only. Preserves the
|
||||
input aspect ratio, so 16:9 presenter images come back 16:9 with no `--preprocess`
|
||||
workaround, and `--size` snaps to the model's latent grid while keeping the aspect. It is
|
||||
the default because identity holds over a long take: measured at **97% of frame-zero
|
||||
sharpness at 70s, flat across 72 segments**, so per-scene generation is a choice rather
|
||||
than a workaround. ~$0.0024 per second of output against SadTalker's ~$0.0014. Decision
|
||||
table in CLAUDE.md, full detail in `docs/soulx.md`. (#81)
|
||||
|
||||
### Changed
|
||||
- **`docker/modal-echomimic3/` keeps its weights in a Modal Volume**, unlike the other six Modal apps
|
||||
which bake them into the image. Measured: rebuild after a dependency change is 1.8-8.2s against
|
||||
- **`docker/modal-soulx/` keeps its weights in a Modal Volume**, unlike the other Modal apps
|
||||
which bake them into the image. Measured: redeploy after a code change is ~2.3s against
|
||||
79-385s baked, while cold start and generation speed are unchanged. Needs a one-off
|
||||
`modal run …::populate_weights`. The settled apps stay baked. (#76)
|
||||
- Upstream repo ref and all four model revisions in `modal-echomimic3` pinned by SHA — weights in a
|
||||
Volume aren't tied to the image, so nothing else prevents drift. (#76, same lesson as #71/#74)
|
||||
- **`NarratorPiP`** honours its `objectPosition` prop, which was declared and documented but silently
|
||||
ignored by a hardcoded value, and gains an `objectFit` prop (default `contain`, unchanged behaviour).
|
||||
`modal run …::populate_weights` (14.7 GB). The settled apps stay baked.
|
||||
- Upstream repo ref and both model revisions in `modal-soulx` pinned by SHA — weights in a
|
||||
Volume aren't tied to the image, so nothing else prevents drift. (same lesson as #71/#74)
|
||||
- **`NarratorPiP`** honours its `objectPosition` prop, which was declared and documented but
|
||||
silently ignored by a hardcoded value, and gains an `objectFit` prop (default `contain`,
|
||||
unchanged behaviour).
|
||||
- `tools/cloud_gpu.py` learned `soulx`, including its GPU tier, so its jobs report a cost
|
||||
estimate.
|
||||
|
||||
### Removed
|
||||
- **EchoMimicV3** (`tools/echomimic3.py`, `docker/modal-echomimic3/`, `docs/echomimic3.md`),
|
||||
added earlier in this unreleased cycle and never part of a tagged release. Identity drift
|
||||
made it unusable for real narration: on a controlled A/B — same photo, same 80s audio,
|
||||
same 544x736 — it fell to 50% of frame-zero sharpness by 70s, collapsing to a featureless
|
||||
smear, where SoulX-FlashHead held 97%. The failure is absorbing rather than gradual, since
|
||||
each segment re-anchors on the previous one's output, so it could not be tuned out; it
|
||||
implied a ~30s render ceiling that shaped the surrounding design. SoulX supersedes it on
|
||||
every axis that mattered — drift, cost (~3.7x cheaper), wall clock (~3-6x faster) — while
|
||||
keeping aspect-ratio preservation. (#81, closes #80)
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -464,40 +464,29 @@
|
||||
"created": "2026-01-11",
|
||||
"updated": "2026-01-12"
|
||||
},
|
||||
"echomimic3": {
|
||||
"path": "tools/echomimic3.py",
|
||||
"description": "Generate talking head videos from portrait image + audio using EchoMimicV3-Flash — preserves the input aspect ratio (16:9 in, 16:9 out)",
|
||||
"usage": "uv run tools/echomimic3.py --image presenter_16x9.png --audio voiceover.mp3 --steps 5 --size 640 --output narrator.mp4",
|
||||
"soulx": {
|
||||
"path": "tools/soulx.py",
|
||||
"description": "Generate talking head videos from portrait image + audio using SoulX-FlashHead — holds identity across long renders and preserves the input aspect ratio (16:9 in, 16:9 out)",
|
||||
"usage": "uv run tools/soulx.py --image presenter_16x9.png --audio voiceover.mp3 --size 768 --output narrator.mp4",
|
||||
"status": "beta",
|
||||
"category": "video-generation",
|
||||
"backend": "echomimicv3-flash",
|
||||
"backend": "soulx-flashhead-1.3b-pro",
|
||||
"requires": "Modal account",
|
||||
"options": {
|
||||
"size": [
|
||||
512,
|
||||
640,
|
||||
768
|
||||
],
|
||||
"steps": "5 (fast) to 8+ (quality)",
|
||||
"wav2vec": [
|
||||
"chinese",
|
||||
"english"
|
||||
],
|
||||
"videoLength": "frames per segment, default 81",
|
||||
"overlap": "frames cross-faded between segments, default 8",
|
||||
"anchorRetreat": "max frames to back off a blink at a seam, default 6, 0 disables",
|
||||
"guidanceScale": "3-6 text CFG",
|
||||
"audioGuidanceScale": "1.8-3.0 audio CFG",
|
||||
"size": "target long edge, default 768; aspect ratio follows the input image",
|
||||
"width/height": "exact dimensions instead of size; both must be multiples of 16",
|
||||
"seed": true,
|
||||
"faceCrop": "upstream face detect+crop, square-only so it discards a 16:9 framing, default off",
|
||||
"compare": "also write a labelled side-by-side against an existing render",
|
||||
"fps": 25
|
||||
},
|
||||
"envVars": [
|
||||
"MODAL_ECHOMIMIC3_ENDPOINT_URL"
|
||||
"MODAL_SOULX_ENDPOINT_URL"
|
||||
],
|
||||
"estimatedCost": "~$0.009 per second of output (measured 28.9x realtime at --steps 5 --size 640)",
|
||||
"documentation": "docs/echomimic3.md",
|
||||
"created": "2026-08-30",
|
||||
"updated": "2026-08-30"
|
||||
"estimatedCost": "~$0.0024 per second of output (measured 6.4-7.9x realtime on A10G), plus a one-off ~600s torch.compile per container per resolution",
|
||||
"documentation": "docs/soulx.md",
|
||||
"created": "2026-08-31",
|
||||
"updated": "2026-08-31"
|
||||
},
|
||||
"qwen3_tts": {
|
||||
"path": "tools/qwen3_tts.py",
|
||||
@@ -946,15 +935,15 @@
|
||||
"gpu": "A10G",
|
||||
"estimatedCost": "$0.05-0.30 per video"
|
||||
},
|
||||
"echomimic3": {
|
||||
"appFile": "docker/modal-echomimic3/app.py",
|
||||
"envVar": "MODAL_ECHOMIMIC3_ENDPOINT_URL",
|
||||
"soulx": {
|
||||
"appFile": "docker/modal-soulx/app.py",
|
||||
"envVar": "MODAL_SOULX_ENDPOINT_URL",
|
||||
"operations": [
|
||||
"echomimic3"
|
||||
"soulx"
|
||||
],
|
||||
"gpu": "A10G",
|
||||
"weights": "modal-volume",
|
||||
"estimatedCost": "~$0.11 per 12s of output"
|
||||
"estimatedCost": "~$0.0024 per second of output, plus a one-off ~600s torch.compile per container per resolution"
|
||||
},
|
||||
"dewatermark": {
|
||||
"appFile": "docker/modal-propainter/app.py",
|
||||
|
||||
@@ -1,237 +0,0 @@
|
||||
# EchoMimicV3 (Modal)
|
||||
|
||||
Audio-driven talking head generation with [EchoMimicV3](https://github.com/antgroup/echomimic_v3)
|
||||
(Ant Group, Apache 2.0, AAAI 2026). Runs alongside `docker/modal-sadtalker/`, which is built on a
|
||||
model unmaintained since 2023 but remains much cheaper and faster.
|
||||
|
||||
**Status: shipping.** User-facing docs are `docs/echomimic3.md`; this file keeps the build
|
||||
detail, the measurements, and the traps. Still not run at full narration length.
|
||||
|
||||
## Why this model
|
||||
|
||||
| | SadTalker | EchoMimicV3-Flash |
|
||||
|---|---|---|
|
||||
| Released | 2023, unmaintained | Flash variant Jan 2026 |
|
||||
| Licence | Apache 2.0 | Apache 2.0 |
|
||||
| Params | ~0.3B (warp-based) | 1.3B (Wan2.1-Fun diffusion) |
|
||||
| VRAM | ~8GB | 12GB (Flash) / 16GB with offload |
|
||||
| Aspect ratio | square crop unless `--preprocess full` | follows the input image |
|
||||
| Motion | head + light expression | head, upper body, gestures |
|
||||
| Speed | fast (~$0.04 / 30s) | **28.9x realtime measured** (~$0.11 / 12s) |
|
||||
|
||||
The alternatives considered and rejected for this slot: **LongCat-Video-Avatar-1.5** (MIT, newer,
|
||||
probably better output, but an A100/H100 tier); **InfiniteTalk** (Apache 2.0, best for long-form,
|
||||
Wan-14B-based); **OmniHuman-1.5** (best quality, closed weights, API-only); hosted fal.ai endpoints
|
||||
($0.15–0.30 per video-second, i.e. $27+ for a 3-minute narrator).
|
||||
|
||||
## Measured results (2026-08-27)
|
||||
|
||||
12s of narration, `--steps 5 --size 640`, 1024x576 presenter crop, Modal A10 (22.06GB usable):
|
||||
|
||||
| Metric | Value |
|
||||
|---|---|
|
||||
| Output | **848x480 — 16:9 preserved**, vs SadTalker's 512x512 square crop from the same portrait |
|
||||
| Wall clock | 347s for 12.0s of video = **28.9x realtime** |
|
||||
| Segments | 4, at a steady 78-82s each (exactly what the loop dry-run predicted) |
|
||||
| Resident VRAM | 0.0GB idle; fits 24GB only via `enable_model_cpu_offload` |
|
||||
| Cost | ~$0.11 per 12s ≈ **$0.0089 per second of output** |
|
||||
|
||||
Extrapolated to the 199s `pluribus-sprint` narration: **~96 minutes and ~$1.76**, against roughly
|
||||
$0.27 for SadTalker. Call it **~6.5x the cost and far longer wall clock** — real, but not
|
||||
prohibitive for a handful of narrator tracks per video.
|
||||
|
||||
**Quality:** clear viseme articulation and natural head rotation where SadTalker with `--still`
|
||||
is near-frozen. Identity holds against the source portrait (frame 0 is essentially the input) and
|
||||
across all four segments. **No visible seam or identity pop** at the segment joins (frames 73,
|
||||
146, 219) — the overlap cross-fade does its job at these settings.
|
||||
|
||||
## Tuning matrix (2026-08-27) — and the bug it exposed
|
||||
|
||||
Six variants on a 5.8s clip (Rob's SadTalker render: frame 0 as the still, its audio as the driver),
|
||||
one factor changed each from a fixed baseline, seed pinned to 43.
|
||||
|
||||
| Variant | mouth motion | sync_r | Human verdict |
|
||||
|---|---|---|---|
|
||||
| A baseline (audio CFG 3.0, cn, 8 steps) | 3.488 | 0.194 | best-validated config |
|
||||
| B audio CFG 1.8 | 3.326 | 0.291 | **rejected — eye artifact** |
|
||||
| C `--wav2vec english` | 2.264 | 0.157 | clearly worst |
|
||||
| D prompt = `"A person is speaking."` | 3.630 | 0.201 | fine |
|
||||
| E steps 5 | 4.035 | 0.168 | fine, most motion |
|
||||
| F rich descriptive prompt | 3.387 | 0.205 | marginally preferred |
|
||||
|
||||
**Settled:** `--wav2vec english` is worse despite English audio — under-articulates throughout
|
||||
(motion 2.26 vs 3.5) and visibly sits half-open. Keep `chinese` regardless of language; the Flash
|
||||
model was trained with it.
|
||||
|
||||
**Do NOT trust `sync_r` alone.** It scores a *mouth crop only* and is structurally blind to eye,
|
||||
hair and background artifacts. It ranked B highest — and B is the variant with the visible defect.
|
||||
Any future scoring needs a whole-face check, or just human eyes.
|
||||
|
||||
### Fixed: segment re-anchoring could latch a blink
|
||||
|
||||
Variant B holds the eyes closed across frames ~72-80. Segment 2 starts at frame 73. The loop
|
||||
re-anchors each segment on the last `overlap` frames of the previous one, so **if those anchor
|
||||
frames land mid-blink, the next segment starts from closed eyes and holds them.** A blinks at f80
|
||||
and recovers; B latched.
|
||||
|
||||
This is a flaw in the chunking, not an audio-CFG property — low audio CFG probably worsens it
|
||||
(less audio drive, more deference to the anchor pose) but does not cause it.
|
||||
|
||||
**Fixed by `--anchor-retreat` (default 6).** Detecting eyes would need the face-landmark stack
|
||||
this image deliberately omits, so `_pick_anchor_retreat` instead scores candidate anchor windows
|
||||
by how much motion they contain in the **upper half** of the frame — where blinks live and mouth
|
||||
movement does not — and backs off up to N frames to anchor on the calmest one. A blink is the
|
||||
largest short transient up there, so it scores worst and gets skipped. It only pays the cost of
|
||||
regenerating those frames when there is a clear improvement (< 0.8x the score at retreat 0), so
|
||||
clips with clean seams regenerate nothing. `--anchor-retreat 0` restores the old behaviour.
|
||||
|
||||
Anchoring on a settled pose is the better default regardless: a continuation segment has to
|
||||
extrapolate from whatever frames it is handed.
|
||||
|
||||
The selection logic was verified offline against synthetic blinks (detects a blink in the anchor
|
||||
window, leaves clean windows alone, chooses a window that excludes the blink frames, clamps on
|
||||
short clips) — no GPU needed. Whether it *looks* fixed on a real latching clip is still unwatched.
|
||||
|
||||
Note the earlier "no visible seam" finding was about *identity* continuity, which did hold. It did
|
||||
not rule out a pose getting stuck across the join.
|
||||
|
||||
**Retracted from earlier in this session:** the recommendation to default `--audio-guidance-scale`
|
||||
to 2.0. It rested on `sync_r`, which the above invalidates as a selector.
|
||||
|
||||
**Still unverified:** behaviour at full narration length (drift over 60+ segments), and whether the
|
||||
prompt matters (F edged it by eye, but the metric that called it inert is the one that failed).
|
||||
|
||||
## Is the gap visible at NarratorPiP size? (2026-08-30)
|
||||
|
||||
Controlled A/B, same still and same 12s of audio: `conal-narrator.png` (772x440) against the
|
||||
existing SadTalker render made from it. Both downscaled to 240x135 — NarratorPiP `sm`, ~3% of a
|
||||
1080p frame — and measured for how much motion survives the downscale.
|
||||
|
||||
| | whole frame | mouth band | eye band |
|
||||
|---|---|---|---|
|
||||
| SadTalker (`--preprocess full --still`) | 0.325 | 0.130 | 0.359 |
|
||||
| EchoMimicV3 (`--steps 5 --size 640`) | 1.151 | 0.800 | 1.208 |
|
||||
| ratio | **3.5x** | **6.2x** | **3.4x** |
|
||||
|
||||
The difference is not washed out by the downscale. This is a *motion* measure, not a quality
|
||||
one — more motion is not automatically better, and the tuning matrix above is the standing
|
||||
warning against reading it that way — but it settles the narrow question the test was for:
|
||||
whatever gap exists is still there at PiP size rather than being invisible.
|
||||
|
||||
**Human verdict on the same pair: "echo is significantly better but sadtalker isn't terrible
|
||||
either."** So both tools stay. EchoMimicV3 is the better picture at any size; SadTalker remains
|
||||
good enough for a small overlay, and at ~1/6th the cost and a fraction of the wall clock it keeps
|
||||
earning its place for PiP boxes, drafts, and generating several takes to choose between. Note the
|
||||
SadTalker side was rendered `--preprocess full --still`, so it was already 16:9 and deliberately
|
||||
steady — this compared articulation, not framing.
|
||||
|
||||
That run also re-measured throughput at 22.8x realtime (274s of GPU for 12.0s of output),
|
||||
against 28.9x in the original measurement.
|
||||
|
||||
## Deploy
|
||||
|
||||
```bash
|
||||
uv sync --extra modal && uv run modal setup
|
||||
|
||||
# One-off: fill the weights volume (~26GB, ~10 min). Must run before the first deploy.
|
||||
uv run modal run docker/modal-echomimic3/app.py::populate_weights
|
||||
|
||||
uv run modal deploy docker/modal-echomimic3/app.py
|
||||
```
|
||||
|
||||
Then put the printed URL in `.env`:
|
||||
|
||||
```
|
||||
MODAL_ECHOMIMIC3_ENDPOINT_URL=https://....modal.run
|
||||
```
|
||||
|
||||
Weights (Wan2.1-Fun-V1.1-1.3B-InP, the `echomimicv3-flash-pro` transformer, and both
|
||||
wav2vec2 encoders) live in a **Modal Volume**, not in the image — the one place this app
|
||||
diverges from the other six. Rationale and numbers in #76; the short version is that
|
||||
rebuild after a dependency change is 1.8-8.2s instead of 79-385s, while cold start and
|
||||
generation speed are unchanged. Because the weights are no longer pinned by the image,
|
||||
the upstream repo ref and all four model revisions are pinned by SHA in `app.py` — bump
|
||||
them deliberately and re-run `populate_weights`.
|
||||
|
||||
For the baked variant (deploys separately as `video-toolkit-echomimic3-baked`):
|
||||
|
||||
```bash
|
||||
ECHOMIMIC_WEIGHTS=image uv run modal deploy docker/modal-echomimic3/app.py
|
||||
```
|
||||
|
||||
## Use
|
||||
|
||||
```bash
|
||||
# Plain generation
|
||||
uv run tools/echomimic3.py --image portrait.png --audio vo.mp3 --output talking.mp4
|
||||
|
||||
# NarratorPiP settings — 16:9 in, 16:9 out, cheap 5-step pass
|
||||
uv run tools/echomimic3.py \
|
||||
--image presenter_16x9.png --audio scene_01.mp3 \
|
||||
--steps 5 --size 640 --output narrator.mp4
|
||||
|
||||
# Side-by-side against the existing SadTalker clip for the same inputs
|
||||
uv run tools/echomimic3.py \
|
||||
--image presenter_16x9.png --audio scene_01.mp3 \
|
||||
--output narrator_echo.mp4 --compare narrator_sadtalker.mp4
|
||||
```
|
||||
|
||||
## Getting it working — four blockers, for anyone repeating this
|
||||
|
||||
1. **OOM on load.** `pipeline.to("cuda")` consumed 21.98 of 22.06GB before inference allocated
|
||||
anything, and died in `@modal.enter()`. Fixed with `enable_model_cpu_offload()`. Upstream's
|
||||
`infer_flash.py` parses a `GPU_memory_mode` flag but never applies it, so following upstream
|
||||
literally does not work on a 24GB card.
|
||||
2. **`chinese-wav2vec2-base` ships `pytorch_model.bin`**, and transformers >=4.51.3 refuses
|
||||
`torch.load` below torch 2.6 (CVE-2025-32434) — a hard failure, not a warning. Converted to
|
||||
safetensors at build time rather than bumping torch, which would flip `torch.load`'s
|
||||
`weights_only` default and break EchoMimic's own `.pth` loads for VAE/text encoder/CLIP.
|
||||
3. **`hidden_states=None` from wav2vec.** transformers 5.x drives `output_hidden_states` from
|
||||
config, not the kwarg EchoMimic's `Wav2Vec2Model` subclass passes — so the audio embeddings,
|
||||
which *are* the lip sync, came back empty. Setting `config.output_hidden_states` did **not**
|
||||
fix it; pinning `transformers==4.49.0` did.
|
||||
4. **Version floors are a trap here.** The repo's `diffusers>=0.30.1` / `transformers>=4.46.2`
|
||||
resolve to releases its vendored code does not survive. Both are now pinned exactly.
|
||||
|
||||
## What to check next
|
||||
|
||||
1. **Segment seams and drift.** Upstream's `infer_flash.py` generates a single 81-frame (3.2s)
|
||||
clip and silently truncates longer audio; the Flash pipeline does not take the long-video kwargs
|
||||
the preview pipeline does. So the segment loop lives in `app.py` instead — re-anchoring each
|
||||
segment on the last `overlap` frames and cross-fading the seam, mirroring `infer_preview.py`.
|
||||
|
||||
At the defaults each segment advances only `81 - 8 = 73` frames (2.9s), so a 30s narrator is
|
||||
**11 full diffusion passes** and a 3-minute one is 62. That, not the per-step cost, is what
|
||||
makes this expensive — and it is the number to attack first if the realtime factor is bad
|
||||
(raise `--video-length` until VRAM complains). Watch a 30s clip for identity drift and for pops
|
||||
at ~2.9s intervals; a larger `--overlap` softens seams at the cost of more segments.
|
||||
|
||||
The loop's frame arithmetic was dry-run separately across audio durations 0.5s–180s and sweeps
|
||||
of both `--video-length` and `--overlap`: coverage is complete, the silence padding always
|
||||
covers the rounded-up final segment, and there is no runaway. That is bookkeeping only — it
|
||||
says nothing about whether the output *looks* right.
|
||||
`--anchor-retreat` adds a little to this: each retreat discards frames that then have to be
|
||||
regenerated. It only fires on a detected transient, so the common case costs nothing.
|
||||
2. **`--audio-guidance-scale`.** Defaulted to 3.0 to match `run_flash.sh`, but the upstream README
|
||||
recommends 1.8–2.0 for lip sync. Try both — and judge by eye, since the metric that would
|
||||
otherwise decide it is the discredited one.
|
||||
3. **Watch a real latching clip with `--anchor-retreat` on.** The selection logic is verified
|
||||
against synthetic blinks, but variant B (audio CFG 1.8) is the known reproducer and has not
|
||||
been re-run.
|
||||
|
||||
(`--wav2vec chinese` vs `english` was on this list and is now settled — see the tuning matrix.)
|
||||
|
||||
## Known deviations from upstream
|
||||
|
||||
- `_build_inputs` reimplements `src.utils.get_image_to_video_latent2` because that helper calls
|
||||
`.resize()` on its argument before checking whether it is a list, so it raises `AttributeError`
|
||||
on the multi-frame re-anchoring the segment loop needs.
|
||||
- The image omits `tensorflow`, `retina-face`, `gradio`, `decord` and `moviepy` from upstream's
|
||||
`requirements.txt`. None are reachable from the Flash path — tensorflow and retina-face are only
|
||||
used by `src.face_detect` for the preview variant's `ip_mask`. Add them back if the preview
|
||||
variant is ever wired up.
|
||||
- `REPO_REF` is pinned to a commit SHA, and so are all four model revisions. An unpinned ref
|
||||
silently re-resolves on rebuild, which is how #71/#74 happened to flux2. Pinning matters more
|
||||
here than in the baked apps, because weights in a Volume are not tied to the image at all.
|
||||
Fetching by SHA needs `git init` + `fetch --depth 1 <sha>`; `clone --branch` only takes a
|
||||
branch or tag name.
|
||||
@@ -1,838 +0,0 @@
|
||||
"""
|
||||
Modal deployment for EchoMimicV3-Flash talking head generation.
|
||||
|
||||
Deploy:
|
||||
modal deploy docker/modal-echomimic3/app.py
|
||||
|
||||
Candidate replacement for docker/modal-sadtalker/app.py. Generates a talking
|
||||
head video from a portrait image + audio file, using Ant Group's EchoMimicV3
|
||||
(Apache 2.0, 1.3B params, built on Wan2.1-Fun-V1.1-1.3B-InP).
|
||||
|
||||
Why the Flash variant: 5-8 denoise steps instead of 25, and it is the one that
|
||||
fits the toolkit's existing 24GB endpoint tier. `echomimicv3-flash-pro` weights
|
||||
override the base Wan transformer.
|
||||
|
||||
Long audio:
|
||||
Upstream `infer_flash.py` generates ONE segment and silently truncates the
|
||||
audio to `video_length` frames (81 = 3.24s at 25fps), and the Flash pipeline
|
||||
(`pipeline_wan_fun_inpaint_audio_2512`) does not accept the long-video kwargs
|
||||
that the preview pipeline does. So the segment loop lives here instead:
|
||||
re-anchor each segment on the last `overlap` frames of the previous one and
|
||||
cross-fade the seam. This mirrors the loop in upstream `infer_preview.py`.
|
||||
|
||||
Note this is NOT a drop-in for SadTalker's aspect behaviour by accident -- it is
|
||||
better on purpose. Output aspect ratio follows the input image (see _fit_size),
|
||||
so a 16:9 portrait yields a 16:9 clip with no `--preprocess full` workaround.
|
||||
|
||||
Input format (POST JSON to web endpoint):
|
||||
{
|
||||
"image_url" | "image_base64": str,
|
||||
"audio_url" | "audio_base64": str,
|
||||
"prompt": str, # default: "A person is speaking."
|
||||
"steps": int, # default: 8 (5 is enough for talking head)
|
||||
"video_length": int, # frames per segment, default 81
|
||||
"overlap": int, # blended frames between segments, default 8
|
||||
"anchor_retreat": int, # max frames to back off a bad seam, default 6
|
||||
"sample_size": [int, int], # area target, default [768, 768]
|
||||
"guidance_scale": float, # default 6.0
|
||||
"audio_guidance_scale": float, # default 3.0 (1.8-2.0 per upstream README)
|
||||
"audio_scale": float, # default 1.0
|
||||
"shift": float, # default 5.0
|
||||
"seed": int, # default 43
|
||||
"fps": int, # default 25
|
||||
"teacache_threshold": float, # default 0.1, 0 disables
|
||||
"negative_prompt": str,
|
||||
"wav2vec": "chinese" | "english",
|
||||
"r2": dict # optional R2 upload config
|
||||
}
|
||||
"""
|
||||
|
||||
import os
|
||||
|
||||
import modal
|
||||
|
||||
REPO_URL = "https://github.com/antgroup/echomimic_v3.git"
|
||||
# Pinned so a rebuild cannot silently pick up a new upstream main whose src/
|
||||
# module layout no longer matches the loader below. Bump deliberately, then
|
||||
# smoke-test the endpoint (same lesson as flux2/diffusers in #71, #74).
|
||||
REPO_REF = "7e89489ca51c0d008fc1963ec6c03fc5bd0b9397"
|
||||
|
||||
# Weights are pinned by revision for the same reason the repo is. It matters more
|
||||
# here than in the baked apps: with the weights in a Volume they are no longer
|
||||
# part of the image, so nothing but these SHAs stops image and weights drifting
|
||||
# apart (#76). Bump deliberately, then re-run populate_weights.
|
||||
BASE_MODEL = "alibaba-pai/Wan2.1-Fun-V1.1-1.3B-InP"
|
||||
BASE_MODEL_REV = "fc913c34361f4ec879e2f9c78b4f11ae50a937d1"
|
||||
ECHO_MODEL = "BadToBest/EchoMimicV3"
|
||||
ECHO_MODEL_REV = "311e176905a8c4c24b240b530488fe636ce4d249"
|
||||
WAV2VEC_CN = "TencentGameMate/chinese-wav2vec2-base" # what run_flash.sh uses
|
||||
WAV2VEC_CN_REV = "3991242c806928916fff4a8c0e4f76acf661b743"
|
||||
WAV2VEC_EN = "facebook/wav2vec2-base-960h" # preview default
|
||||
WAV2VEC_EN_REV = "22aad52d435eb6dbaf354bdad9b0da84ce7d6156"
|
||||
|
||||
APP_DIR = "/app/echomimic_v3"
|
||||
MODELS_DIR = "/models"
|
||||
|
||||
# Where weights come from. This app keeps them in a Modal Volume, unlike the
|
||||
# other six modal-* apps which bake them into the image. That split is deliberate
|
||||
# and measured (#76): rebuild after a dependency change is 1.8-8.2s on a volume
|
||||
# against 79-385s baked, while cold start and generation speed are the same
|
||||
# either way. Getting this model working took four dependency changes, so the
|
||||
# rebuild cost is the one that bites. The settled apps gain nothing by moving.
|
||||
#
|
||||
# Set ECHOMIMIC_WEIGHTS=image at DEPLOY time for the baked variant, which
|
||||
# deploys as a separate app -- self-contained, but re-downloads 26GB whenever a
|
||||
# layer above the weights is invalidated.
|
||||
WEIGHTS_SOURCE = os.environ.get("ECHOMIMIC_WEIGHTS", "volume")
|
||||
if WEIGHTS_SOURCE not in ("image", "volume"):
|
||||
raise ValueError(f"ECHOMIMIC_WEIGHTS must be 'image' or 'volume', got {WEIGHTS_SOURCE!r}")
|
||||
|
||||
USE_VOLUME = WEIGHTS_SOURCE == "volume"
|
||||
|
||||
app = modal.App(
|
||||
"video-toolkit-echomimic3" if USE_VOLUME else "video-toolkit-echomimic3-baked"
|
||||
)
|
||||
|
||||
# Created in both modes: unused (and empty, so free) in image mode, but having
|
||||
# the handle unconditionally keeps the module importable either way.
|
||||
volume = modal.Volume.from_name("echomimic3-weights", create_if_missing=True)
|
||||
|
||||
_WEIGHT_FETCH = [
|
||||
(BASE_MODEL, BASE_MODEL_REV, f"{MODELS_DIR}/Wan2.1-Fun-V1.1-1.3B-InP", None),
|
||||
# Only the Flash transformer -- skips the larger preview checkpoint.
|
||||
(ECHO_MODEL, ECHO_MODEL_REV, f"{MODELS_DIR}/EchoMimicV3", ["echomimicv3-flash-pro/*"]),
|
||||
(WAV2VEC_CN, WAV2VEC_CN_REV, f"{MODELS_DIR}/chinese-wav2vec2-base", None),
|
||||
(WAV2VEC_EN, WAV2VEC_EN_REV, f"{MODELS_DIR}/wav2vec2-base-960h", None),
|
||||
]
|
||||
|
||||
# chinese-wav2vec2-base ships pytorch_model.bin, and current transformers refuses
|
||||
# torch.load outright below torch 2.6 (CVE-2025-32434) -- it fails the load rather
|
||||
# than warning. Converting to safetensors is the cheap fix: bumping to torch 2.6
|
||||
# would flip torch.load's weights_only default and break EchoMimic's own .pth
|
||||
# loads for the VAE, text encoder and CLIP, trading one breakage for three.
|
||||
def _fetch_weights():
|
||||
"""Download weights into MODELS_DIR. Runs at image build OR into the volume."""
|
||||
import os
|
||||
|
||||
import safetensors.torch
|
||||
import torch
|
||||
from huggingface_hub import snapshot_download
|
||||
|
||||
for repo, revision, dest, patterns in _WEIGHT_FETCH:
|
||||
snapshot_download(repo, revision=revision, local_dir=dest, allow_patterns=patterns)
|
||||
|
||||
for d in (f"{MODELS_DIR}/chinese-wav2vec2-base", f"{MODELS_DIR}/wav2vec2-base-960h"):
|
||||
b, sf = os.path.join(d, "pytorch_model.bin"), os.path.join(d, "model.safetensors")
|
||||
if os.path.exists(b) and not os.path.exists(sf):
|
||||
sd = torch.load(b, map_location="cpu", weights_only=True)
|
||||
safetensors.torch.save_file({k: v.contiguous() for k, v in sd.items()}, sf)
|
||||
os.remove(b)
|
||||
print("converted", d)
|
||||
else:
|
||||
print("already safetensors", d)
|
||||
|
||||
total = sum(
|
||||
os.path.getsize(os.path.join(r, f))
|
||||
for r, _, fs in os.walk(MODELS_DIR) for f in fs
|
||||
)
|
||||
print(f"weights ready: {total / 1e9:.1f} GB")
|
||||
|
||||
|
||||
image = (
|
||||
modal.Image.debian_slim(python_version="3.10")
|
||||
.apt_install("git", "ffmpeg", "libgl1-mesa-glx", "libglib2.0-0")
|
||||
.pip_install("torch==2.5.1", "torchvision==0.20.1", "torchaudio==2.5.1")
|
||||
.pip_install(
|
||||
# Pinned, not floored. EchoMimicV3's requirements.txt says
|
||||
# diffusers>=0.30.1 / transformers>=4.46.2, but those floors resolve to
|
||||
# current releases that the repo's vendored code does not survive:
|
||||
# - transformers 5.x drives output_hidden_states from config rather than
|
||||
# the kwarg EchoMimic's Wav2Vec2Model subclass passes, so the encoder
|
||||
# returns hidden_states=None and there are no audio embeddings at all
|
||||
# (setting config.output_hidden_states does not rescue it).
|
||||
# - transformers >=4.51.3 also refuses torch.load below torch 2.6.
|
||||
# - current diffusers moved load_model_dict_into_meta, which silently
|
||||
# disables low_cpu_mem_usage on the transformer load.
|
||||
# These two are contemporary with the repo and mutually compatible.
|
||||
"diffusers==0.32.2",
|
||||
"transformers==4.49.0",
|
||||
"accelerate>=0.25.0",
|
||||
"safetensors",
|
||||
"omegaconf",
|
||||
"einops",
|
||||
"timm",
|
||||
"tomesd",
|
||||
"torchdiffeq",
|
||||
"torchsde",
|
||||
"imageio[ffmpeg]",
|
||||
"imageio[pyav]",
|
||||
"opencv-python-headless",
|
||||
"scikit-image",
|
||||
"librosa",
|
||||
"pyloudnorm",
|
||||
"SentencePiece",
|
||||
"ftfy",
|
||||
"func_timeout",
|
||||
"Pillow",
|
||||
"numpy<2",
|
||||
"boto3",
|
||||
"requests",
|
||||
"fastapi[standard]",
|
||||
"huggingface_hub>=0.25.0",
|
||||
)
|
||||
# Upstream requirements.txt also pins tensorflow==2.15.0 + retina-face +
|
||||
# gradio + decord + moviepy. None of those are reachable from the Flash
|
||||
# path: tensorflow/retina-face are only for src.face_detect (the preview
|
||||
# ip_mask), and gradio/moviepy only for the demo UIs. Left out to keep the
|
||||
# image small -- add them back if the preview variant is ever wired up.
|
||||
# Fetch-by-SHA rather than `clone --branch`, which only accepts a branch or
|
||||
# tag name. Still a single-commit download.
|
||||
.run_commands(
|
||||
f"git init {APP_DIR}",
|
||||
f"git -C {APP_DIR} remote add origin {REPO_URL}",
|
||||
f"git -C {APP_DIR} fetch --depth 1 origin {REPO_REF}",
|
||||
f"git -C {APP_DIR} checkout FETCH_HEAD",
|
||||
)
|
||||
.env({
|
||||
"PYTHONPATH": APP_DIR,
|
||||
"TOKENIZERS_PARALLELISM": "false",
|
||||
# Modal re-imports this module inside the container, where the local
|
||||
# shell env does NOT exist. Without baking the mode in, USE_VOLUME reads
|
||||
# False in-container regardless of how it was deployed -- which silently
|
||||
# drops the volume mount and hides populate_weights.
|
||||
"ECHOMIMIC_WEIGHTS": WEIGHTS_SOURCE,
|
||||
})
|
||||
)
|
||||
|
||||
if not USE_VOLUME:
|
||||
# Baked variant: weights become image layers. Every dependency change above
|
||||
# this point invalidates them and re-downloads ~26GB.
|
||||
image = image.run_function(_fetch_weights)
|
||||
|
||||
# Volume variant carries code only, so a dependency change rebuilds in seconds
|
||||
# and leaves the weights untouched. Downloading needs no GPU.
|
||||
download_image = (
|
||||
modal.Image.debian_slim(python_version="3.10")
|
||||
# numpy is not optional here: torch.load pulls it in during the
|
||||
# .bin -> safetensors conversion below.
|
||||
.pip_install("huggingface_hub>=0.25.0", "safetensors", "torch==2.5.1", "numpy<2")
|
||||
.env({"ECHOMIMIC_WEIGHTS": WEIGHTS_SOURCE})
|
||||
)
|
||||
|
||||
|
||||
@app.function(image=download_image, volumes={MODELS_DIR: volume}, timeout=3600)
|
||||
def populate_weights():
|
||||
"""One-off, idempotent: fill the volume before first use.
|
||||
|
||||
ECHOMIMIC_WEIGHTS=volume modal run docker/modal-echomimic3/app.py::populate_weights
|
||||
|
||||
Defined in both modes on purpose. Gating it behind USE_VOLUME made it vanish
|
||||
inside the container, where the module is re-imported without the deploy-time
|
||||
shell env.
|
||||
"""
|
||||
_fetch_weights()
|
||||
volume.commit()
|
||||
|
||||
|
||||
@app.cls(
|
||||
image=image,
|
||||
# A10G (24GB, 22.06 usable) matches the tier the toolkit's other endpoints
|
||||
# already run on. It does NOT fit with everything resident -- see
|
||||
# enable_model_cpu_offload in load_pipeline. With offload on it does fit,
|
||||
# at the cost of paging modules over PCIe on every pipeline call.
|
||||
gpu="A10G",
|
||||
volumes=({MODELS_DIR: volume} if USE_VOLUME else {}),
|
||||
timeout=7200,
|
||||
scaledown_window=120,
|
||||
)
|
||||
@modal.concurrent(max_inputs=1)
|
||||
class EchoMimicV3:
|
||||
@modal.enter()
|
||||
def load_pipeline(self):
|
||||
import os
|
||||
import time
|
||||
import torch
|
||||
|
||||
_load_started = time.time()
|
||||
from omegaconf import OmegaConf
|
||||
from transformers import AutoTokenizer, Wav2Vec2FeatureExtractor
|
||||
|
||||
from src.wan_vae import AutoencoderKLWan
|
||||
from src.wan_image_encoder import CLIPModel
|
||||
from src.wan_text_encoder import WanT5EncoderModel
|
||||
from src.wan_transformer3d_audio_2512 import WanTransformerAudioMask3DModel
|
||||
from src.pipeline_wan_fun_inpaint_audio_2512 import WanFunInpaintAudioPipeline
|
||||
from src.fm_solvers_unipc import FlowUniPCMultistepScheduler
|
||||
from src.utils import filter_kwargs
|
||||
from src.wav2vec2 import Wav2Vec2Model
|
||||
|
||||
print(f"PyTorch {torch.__version__}, CUDA available: {torch.cuda.is_available()}")
|
||||
if torch.cuda.is_available():
|
||||
print(f"GPU: {torch.cuda.get_device_name(0)}")
|
||||
|
||||
self.device = "cuda"
|
||||
self.dtype = torch.bfloat16
|
||||
self.model_name = f"{MODELS_DIR}/Wan2.1-Fun-V1.1-1.3B-InP"
|
||||
|
||||
cfg = OmegaConf.load(f"{APP_DIR}/config/config.yaml")
|
||||
self.cfg = cfg
|
||||
|
||||
# config.yaml uses transformer_subpath "./" -- the transformer config
|
||||
# lives at the root of the Wan2.1-Fun repo, and the Flash safetensors
|
||||
# below replace its weights.
|
||||
transformer = WanTransformerAudioMask3DModel.from_pretrained(
|
||||
os.path.join(self.model_name, cfg["transformer_additional_kwargs"].get("transformer_subpath", "./")),
|
||||
transformer_additional_kwargs=OmegaConf.to_container(cfg["transformer_additional_kwargs"]),
|
||||
low_cpu_mem_usage=True,
|
||||
torch_dtype=self.dtype,
|
||||
)
|
||||
|
||||
flash_ckpt = f"{MODELS_DIR}/EchoMimicV3/echomimicv3-flash-pro/diffusion_pytorch_model.safetensors"
|
||||
from safetensors.torch import load_file
|
||||
|
||||
state_dict = load_file(flash_ckpt)
|
||||
state_dict = state_dict.get("state_dict", state_dict)
|
||||
missing, unexpected = transformer.load_state_dict(state_dict, strict=False)
|
||||
print(f"Flash checkpoint: {len(missing)} missing keys, {len(unexpected)} unexpected")
|
||||
# A large `missing` count means the Flash weights did not line up with
|
||||
# the base transformer -- output would be garbage rather than an error.
|
||||
if len(missing) > 50:
|
||||
raise RuntimeError(
|
||||
f"Flash checkpoint mismatch: {len(missing)} missing keys. "
|
||||
"Check that BASE_MODEL and the flash weights are the matching pair."
|
||||
)
|
||||
|
||||
vae = AutoencoderKLWan.from_pretrained(
|
||||
os.path.join(self.model_name, cfg["vae_kwargs"].get("vae_subpath", "Wan2.1_VAE.pth")),
|
||||
additional_kwargs=OmegaConf.to_container(cfg["vae_kwargs"]),
|
||||
).to(self.dtype)
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained(
|
||||
os.path.join(self.model_name, cfg["text_encoder_kwargs"].get("tokenizer_subpath", "google/umt5-xxl")),
|
||||
)
|
||||
text_encoder = WanT5EncoderModel.from_pretrained(
|
||||
os.path.join(self.model_name, cfg["text_encoder_kwargs"].get("text_encoder_subpath")),
|
||||
additional_kwargs=OmegaConf.to_container(cfg["text_encoder_kwargs"]),
|
||||
low_cpu_mem_usage=True,
|
||||
torch_dtype=self.dtype,
|
||||
).eval()
|
||||
clip_image_encoder = CLIPModel.from_pretrained(
|
||||
os.path.join(self.model_name, cfg["image_encoder_kwargs"].get("image_encoder_subpath")),
|
||||
).to(self.dtype).eval()
|
||||
|
||||
# Flow_Unipc is what run_flash.sh uses; it wants shift folded into the
|
||||
# pipeline call rather than the scheduler config.
|
||||
scheduler_cfg = OmegaConf.to_container(cfg["scheduler_kwargs"])
|
||||
scheduler_cfg["shift"] = 1
|
||||
scheduler = FlowUniPCMultistepScheduler(
|
||||
**filter_kwargs(FlowUniPCMultistepScheduler, scheduler_cfg)
|
||||
)
|
||||
|
||||
self.pipeline = WanFunInpaintAudioPipeline(
|
||||
transformer=transformer,
|
||||
vae=vae,
|
||||
tokenizer=tokenizer,
|
||||
text_encoder=text_encoder,
|
||||
scheduler=scheduler,
|
||||
clip_image_encoder=clip_image_encoder,
|
||||
)
|
||||
# Model CPU offload, NOT pipeline.to("cuda"). Measured on an A10 (22.06GB
|
||||
# usable): keeping the umt5-xxl text encoder, CLIP-huge image encoder,
|
||||
# transformer and VAE all resident OOMs during load -- 21.98GB consumed
|
||||
# before inference allocates anything. The pipeline declares
|
||||
# model_cpu_offload_seq, so diffusers can keep only the executing module
|
||||
# on GPU and page the rest back to host RAM.
|
||||
# Do not add a .to(device) call alongside this; diffusers rejects both.
|
||||
self.pipeline.enable_model_cpu_offload(device=self.device)
|
||||
self.vae_ratio = vae.config.temporal_compression_ratio
|
||||
|
||||
# Audio encoders stay on CPU (upstream does the same) -- they are small
|
||||
# and this keeps VRAM for the transformer.
|
||||
self.audio_encoders = {}
|
||||
for key, path in (("chinese", f"{MODELS_DIR}/chinese-wav2vec2-base"),
|
||||
("english", f"{MODELS_DIR}/wav2vec2-base-960h")):
|
||||
enc = Wav2Vec2Model.from_pretrained(path, local_files_only=True).to("cpu")
|
||||
# EchoMimic's Wav2Vec2Model subclass passes output_hidden_states as a
|
||||
# kwarg to the inner encoder, but current transformers drives it from
|
||||
# config instead (the same release that deprecates use_return_dict).
|
||||
# Without this the encoder returns hidden_states=None and the audio
|
||||
# embedding stack fails -- and the audio embeddings ARE the lip sync,
|
||||
# so this has to be right, not merely non-crashing.
|
||||
enc.config.output_hidden_states = True
|
||||
enc.feature_extractor._freeze_parameters()
|
||||
self.audio_encoders[key] = (
|
||||
enc,
|
||||
Wav2Vec2FeatureExtractor.from_pretrained(path, local_files_only=True),
|
||||
)
|
||||
|
||||
if torch.cuda.is_available():
|
||||
print(f"VRAM resident after load: {torch.cuda.memory_allocated() / 1e9:.1f}GB "
|
||||
f"(offloaded; modules page in per call)")
|
||||
self.load_seconds = time.time() - _load_started
|
||||
print(f"Pipeline ready in {self.load_seconds:.1f}s")
|
||||
|
||||
# -- helpers ------------------------------------------------------------
|
||||
|
||||
def _round_frames(self, n: int, up: bool = False) -> int:
|
||||
"""Snap a frame count to the VAE's temporal compression grid (4k+1).
|
||||
|
||||
Segment lengths round UP so the final segment still covers the tail of
|
||||
the audio: rounding down leaves a remainder that the loop can only chew
|
||||
through a frame at a time. The total is rounded DOWN, and the render is
|
||||
trimmed to it, so the extra frames never reach the output.
|
||||
"""
|
||||
if n <= 1:
|
||||
return 1
|
||||
r = self.vae_ratio
|
||||
steps = -(-(n - 1) // r) if up else (n - 1) // r
|
||||
return int(steps * r) + 1
|
||||
|
||||
@staticmethod
|
||||
def _fit_size(img, target):
|
||||
"""Pick an output size that keeps the image's aspect ratio.
|
||||
|
||||
Scales to roughly `target` pixel area, rounded to /16. This is why a
|
||||
16:9 input gives a 16:9 output -- no square crop, unlike SadTalker.
|
||||
"""
|
||||
import math
|
||||
|
||||
w, h = img.size
|
||||
ori_a, tgt_a = w * h, target[0] * target[1]
|
||||
if tgt_a < ori_a:
|
||||
ratio = math.sqrt(ori_a / tgt_a)
|
||||
w, h = w / ratio // 16 * 16, h / ratio // 16 * 16
|
||||
else:
|
||||
w, h = w // 16 * 16, h // 16 * 16
|
||||
return int(h), int(w)
|
||||
|
||||
@staticmethod
|
||||
def _build_inputs(start_images, video_length, height, width):
|
||||
"""Build (input_video, input_video_mask, clip_image) for one segment.
|
||||
|
||||
Reimplemented rather than calling src.utils.get_image_to_video_latent2:
|
||||
that helper calls .resize() on its argument before checking whether it
|
||||
is a list, so the multi-frame re-anchoring this loop needs would raise
|
||||
AttributeError there.
|
||||
"""
|
||||
import numpy as np
|
||||
import torch
|
||||
from PIL import Image
|
||||
|
||||
if not isinstance(start_images, list):
|
||||
start_images = [start_images]
|
||||
imgs = [
|
||||
im.convert("RGB").resize((width, height), resample=Image.Resampling.LANCZOS)
|
||||
for im in start_images
|
||||
]
|
||||
clip_image = imgs[0]
|
||||
|
||||
start = torch.cat(
|
||||
[torch.from_numpy(np.array(im)).permute(2, 0, 1).unsqueeze(1).unsqueeze(0) for im in imgs],
|
||||
dim=2,
|
||||
) # [1, 3, n, H, W]
|
||||
n = min(start.shape[2], video_length)
|
||||
|
||||
video = torch.tile(start[:, :, :1], [1, 1, video_length, 1, 1]).clone()
|
||||
video[:, :, :n] = start[:, :, :n]
|
||||
video = video / 255
|
||||
|
||||
mask = torch.zeros_like(video[:, :1])
|
||||
mask[:, :, n:] = 255
|
||||
return video, mask, clip_image
|
||||
|
||||
@staticmethod
|
||||
def _pick_anchor_retreat(accumulated, overlap, max_retreat):
|
||||
"""How many trailing frames to drop before re-anchoring the next segment.
|
||||
|
||||
Each continuation segment starts from the last `overlap` frames of the
|
||||
previous one, so whatever pose those frames hold becomes the next
|
||||
segment's opening pose. When they land mid-blink the model starts
|
||||
closed-eyed and *holds* it -- observed as a prolonged closure across a
|
||||
segment boundary, and the reason this exists.
|
||||
|
||||
Rather than detect eyes (which needs the face-landmark stack this image
|
||||
deliberately omits), score each candidate window by how much motion it
|
||||
contains in the upper half of the frame -- where blinks live and mouth
|
||||
movement does not -- and anchor on the calmest one. A blink is the
|
||||
largest short transient up there, so it scores worst and gets skipped.
|
||||
Anchoring on a settled pose is the better default regardless, since a
|
||||
continuation has to extrapolate from whatever it is handed.
|
||||
|
||||
Returns frames to discard: 0 keeps the current behaviour, which is also
|
||||
what a clip with no transient near the seam gets, because dropping
|
||||
frames means regenerating them.
|
||||
"""
|
||||
import torch
|
||||
|
||||
total = accumulated.shape[2]
|
||||
# Never retreat past having `overlap` frames left to anchor on.
|
||||
budget = max(0, min(int(max_retreat), total - overlap - 1))
|
||||
if budget <= 0:
|
||||
return 0
|
||||
|
||||
# Upper half only, and greyscale: blink transients are small, and mouth
|
||||
# motion in the lower half would otherwise dominate every score.
|
||||
tail = accumulated[0, :, -(overlap + budget + 1):] # [C, n, H, W]
|
||||
upper = tail[:, :, : max(1, tail.shape[2] // 2)].mean(dim=0) # [n, H/2, W]
|
||||
motion = (upper[1:] - upper[:-1]).abs().mean(dim=(1, 2)) # [n-1] per-frame
|
||||
|
||||
# Window r covers the `overlap` frames ending `r` from the end. Score it
|
||||
# by its worst frame, not its mean: one blink frame in the window is
|
||||
# enough to poison the anchor.
|
||||
n = len(motion)
|
||||
scores = [
|
||||
torch.max(motion[n - overlap - r: n - r]).item()
|
||||
for r in range(budget + 1)
|
||||
]
|
||||
|
||||
best = min(range(len(scores)), key=lambda r: scores[r])
|
||||
# Only pay for a retreat when it is a clear improvement. Without this,
|
||||
# sensor-level noise picks an arbitrary r on every segment and quietly
|
||||
# adds a regenerated frame budget to clips that never needed one.
|
||||
if best == 0 or scores[best] >= 0.8 * scores[0]:
|
||||
return 0
|
||||
return best
|
||||
|
||||
def _audio_embed(self, wav, start_frame, seg_frames, fps, which, sr=16000):
|
||||
"""Wav2Vec embeddings for one segment, windowed the way Flash expects.
|
||||
|
||||
Upstream builds a +/-2 frame window per output frame -> [F, 5, 12, 768].
|
||||
Slicing the waveform per segment (rather than slicing a whole-clip
|
||||
embedding) is what infer_flash.py does, and it keeps the encoder's
|
||||
seq_len argument consistent with the segment length.
|
||||
"""
|
||||
import numpy as np
|
||||
import torch
|
||||
from einops import rearrange
|
||||
|
||||
encoder, extractor = self.audio_encoders[which]
|
||||
|
||||
lo = int(start_frame / fps * sr)
|
||||
hi = int((start_frame + seg_frames) / fps * sr)
|
||||
chunk = wav[lo:hi]
|
||||
|
||||
feature = np.squeeze(extractor(chunk, sampling_rate=sr).input_values)
|
||||
feature = torch.from_numpy(feature).float().unsqueeze(0)
|
||||
with torch.no_grad():
|
||||
out = encoder(feature, seq_len=int(seg_frames), output_hidden_states=True,
|
||||
return_dict=True)
|
||||
|
||||
if getattr(out, "hidden_states", None) is None:
|
||||
raise RuntimeError(
|
||||
"wav2vec returned no hidden_states -- transformers is not honouring "
|
||||
"output_hidden_states for this encoder. Check config.output_hidden_states "
|
||||
"in load_pipeline, or pin transformers to a release contemporary with "
|
||||
"EchoMimicV3 (>=4.46.2, before use_return_dict was deprecated)."
|
||||
)
|
||||
|
||||
emb = torch.stack(out.hidden_states[1:], dim=1).squeeze(0)
|
||||
emb = rearrange(emb, "b s d -> s b d").cpu().detach()
|
||||
|
||||
indices = torch.arange(5) - 2
|
||||
centers = torch.arange(0, seg_frames).unsqueeze(1) + indices.unsqueeze(0)
|
||||
centers = torch.clamp(centers, min=0, max=emb.shape[0] - 1)
|
||||
return emb[centers].unsqueeze(0)
|
||||
|
||||
# -- endpoint -----------------------------------------------------------
|
||||
|
||||
@modal.fastapi_endpoint(method="GET")
|
||||
def health(self) -> dict:
|
||||
"""Return as soon as @modal.enter() has finished.
|
||||
|
||||
Timing a GET against a cold container measures schedule + weight fetch +
|
||||
model load, without spending GPU minutes on a generation that would tell
|
||||
us nothing new (inference speed cannot differ between image and volume).
|
||||
"""
|
||||
return {
|
||||
"ok": True,
|
||||
"weights_source": WEIGHTS_SOURCE,
|
||||
"load_seconds": round(getattr(self, "load_seconds", -1), 1),
|
||||
}
|
||||
|
||||
@modal.fastapi_endpoint(method="POST")
|
||||
def generate(self, request: dict) -> dict:
|
||||
import base64
|
||||
import shutil
|
||||
import subprocess
|
||||
import tempfile
|
||||
import time
|
||||
import uuid
|
||||
from pathlib import Path
|
||||
|
||||
import librosa
|
||||
import numpy as np
|
||||
import pyloudnorm as pyln
|
||||
import requests as req
|
||||
import torch
|
||||
from PIL import Image
|
||||
|
||||
from src.utils import save_videos_grid
|
||||
from src.cache_utils import get_teacache_coefficients
|
||||
|
||||
start_time = time.time()
|
||||
|
||||
image_url = request.get("image_url")
|
||||
image_base64 = request.get("image_base64")
|
||||
audio_url = request.get("audio_url")
|
||||
audio_base64 = request.get("audio_base64")
|
||||
|
||||
if not image_url and not image_base64:
|
||||
return {"error": "Missing image_url or image_base64"}
|
||||
if not audio_url and not audio_base64:
|
||||
return {"error": "Missing audio_url or audio_base64"}
|
||||
|
||||
prompt = request.get("prompt") or "A person is speaking."
|
||||
steps = int(request.get("steps", 8))
|
||||
seg_length = int(request.get("video_length", 81))
|
||||
overlap = int(request.get("overlap", 8))
|
||||
anchor_retreat = max(0, int(request.get("anchor_retreat", 6)))
|
||||
sample_size = request.get("sample_size") or [768, 768]
|
||||
guidance_scale = float(request.get("guidance_scale", 6.0))
|
||||
audio_guidance_scale = float(request.get("audio_guidance_scale", 3.0))
|
||||
audio_scale = float(request.get("audio_scale", 1.0))
|
||||
shift = float(request.get("shift", 5.0))
|
||||
seed = int(request.get("seed", 43))
|
||||
fps = int(request.get("fps", 25))
|
||||
teacache_threshold = float(request.get("teacache_threshold", 0.1))
|
||||
which_wav2vec = request.get("wav2vec", "chinese")
|
||||
negative_prompt = request.get("negative_prompt") or (
|
||||
"Gesture is bad. Gesture is unclear. Strange and twisted hands. Bad hands. "
|
||||
"Bad fingers. Unclear and blurry hands."
|
||||
)
|
||||
r2_config = request.get("r2")
|
||||
|
||||
if which_wav2vec not in self.audio_encoders:
|
||||
return {"error": f"wav2vec must be one of {list(self.audio_encoders)}"}
|
||||
|
||||
work_dir = Path(tempfile.mkdtemp(prefix="modal_echomimic3_"))
|
||||
|
||||
try:
|
||||
image_path = work_dir / "input_image.png"
|
||||
if image_url:
|
||||
resp = req.get(image_url, stream=True, timeout=300)
|
||||
resp.raise_for_status()
|
||||
with open(image_path, "wb") as f:
|
||||
for chunk in resp.iter_content(8192):
|
||||
f.write(chunk)
|
||||
else:
|
||||
data = image_base64.split(",", 1)[-1]
|
||||
image_path.write_bytes(base64.b64decode(data))
|
||||
|
||||
audio_path = work_dir / "input_audio.wav"
|
||||
if audio_url:
|
||||
resp = req.get(audio_url, stream=True, timeout=300)
|
||||
resp.raise_for_status()
|
||||
with open(audio_path, "wb") as f:
|
||||
for chunk in resp.iter_content(8192):
|
||||
f.write(chunk)
|
||||
else:
|
||||
data = audio_base64.split(",", 1)[-1]
|
||||
audio_path.write_bytes(base64.b64decode(data))
|
||||
|
||||
ref_image = Image.open(image_path).convert("RGB")
|
||||
height, width = self._fit_size(ref_image, sample_size)
|
||||
|
||||
wav, sr = librosa.load(str(audio_path), sr=16000)
|
||||
total_duration = len(wav) / sr
|
||||
meter = pyln.Meter(sr)
|
||||
loudness = meter.integrated_loudness(wav)
|
||||
if abs(loudness) <= 100:
|
||||
wav = pyln.normalize.loudness(wav, loudness, -23)
|
||||
|
||||
total_frames = self._round_frames(int(total_duration * fps))
|
||||
seg_length = self._round_frames(seg_length)
|
||||
|
||||
# Checked after rounding, which can pull seg_length down onto overlap.
|
||||
# overlap must be >= 1: continuation segments are anchored on those
|
||||
# frames, and `tensor[:, :, -0:]` selects the whole tensor rather than
|
||||
# nothing, so a zero would silently blend over the entire clip.
|
||||
if not 1 <= overlap < seg_length:
|
||||
return {"error": f"overlap ({overlap}) must be >= 1 and < video_length ({seg_length})"}
|
||||
|
||||
# A rounded-up final segment asks for audio past the end of the clip.
|
||||
# Wav2Vec resamples whatever it is given to seq_len, so a short tail
|
||||
# would be stretched and drift out of sync -- pad with silence instead.
|
||||
padded_frames = total_frames + seg_length
|
||||
wanted_samples = int(padded_frames / fps * sr)
|
||||
if len(wav) < wanted_samples:
|
||||
wav = np.pad(wav, (0, wanted_samples - len(wav)))
|
||||
|
||||
print(
|
||||
f"Audio {total_duration:.1f}s -> {total_frames} frames @ {fps}fps, "
|
||||
f"output {width}x{height}, {steps} steps, segment={seg_length}, overlap={overlap}"
|
||||
)
|
||||
|
||||
if teacache_threshold > 0:
|
||||
coefficients = get_teacache_coefficients(self.model_name)
|
||||
if coefficients is not None:
|
||||
self.pipeline.transformer.enable_teacache(
|
||||
coefficients, steps, teacache_threshold, num_skip_start_steps=5, offload=False
|
||||
)
|
||||
|
||||
generator = torch.Generator(device=self.device).manual_seed(seed)
|
||||
mix_ratio = torch.linspace(0, 1, steps=overlap).view(1, 1, -1, 1, 1)
|
||||
|
||||
accumulated = None
|
||||
starts = ref_image
|
||||
produced = 0
|
||||
segments = 0
|
||||
|
||||
with torch.no_grad():
|
||||
while produced < total_frames:
|
||||
# A continuation segment re-generates the `overlap` frames it
|
||||
# is anchored on, so it has to be that much longer to still
|
||||
# advance by the frames actually remaining.
|
||||
remaining = total_frames - produced
|
||||
want = remaining if accumulated is None else remaining + overlap
|
||||
seg_frames = min(seg_length, self._round_frames(want, up=True))
|
||||
if seg_frames - (0 if accumulated is None else overlap) < 1:
|
||||
break
|
||||
|
||||
seg_start = produced if accumulated is None else produced - overlap
|
||||
audio_embeds = self._audio_embed(
|
||||
wav, seg_start, seg_frames, fps, which_wav2vec
|
||||
).to(device=self.device, dtype=self.dtype)
|
||||
|
||||
input_video, input_video_mask, clip_image = self._build_inputs(
|
||||
starts, seg_frames, height, width
|
||||
)
|
||||
|
||||
seg_started = time.time()
|
||||
sample = self.pipeline(
|
||||
prompt,
|
||||
num_frames=seg_frames,
|
||||
negative_prompt=negative_prompt,
|
||||
audio_embeds=audio_embeds,
|
||||
audio_scale=audio_scale,
|
||||
ip_mask=None,
|
||||
use_un_ip_mask=False,
|
||||
height=height,
|
||||
width=width,
|
||||
generator=generator,
|
||||
neg_scale=1.0,
|
||||
neg_steps=0,
|
||||
use_dynamic_cfg=False,
|
||||
use_dynamic_acfg=False,
|
||||
guidance_scale=guidance_scale,
|
||||
audio_guidance_scale=audio_guidance_scale,
|
||||
num_inference_steps=steps,
|
||||
video=input_video,
|
||||
mask_video=input_video_mask,
|
||||
clip_image=clip_image,
|
||||
cfg_skip_ratio=0.0,
|
||||
shift=shift,
|
||||
).videos
|
||||
|
||||
segments += 1
|
||||
print(
|
||||
f" segment {segments}: frames {seg_start}-{seg_start + seg_frames} "
|
||||
f"in {time.time() - seg_started:.0f}s"
|
||||
)
|
||||
|
||||
if accumulated is None:
|
||||
accumulated = sample
|
||||
else:
|
||||
# Cross-fade the re-generated overlap so the seam between
|
||||
# segments does not pop.
|
||||
accumulated[:, :, -overlap:] = (
|
||||
accumulated[:, :, -overlap:] * (1 - mix_ratio)
|
||||
+ sample[:, :, :overlap] * mix_ratio
|
||||
)
|
||||
accumulated = torch.cat([accumulated, sample[:, :, overlap:]], dim=2)
|
||||
|
||||
produced = accumulated.shape[2]
|
||||
if produced >= total_frames:
|
||||
break
|
||||
|
||||
# Drop a few trailing frames if they hold a blink or another
|
||||
# transient, so the next segment does not start from it and
|
||||
# latch it. Costs the dropped frames, which get regenerated.
|
||||
retreat = self._pick_anchor_retreat(
|
||||
accumulated, overlap, anchor_retreat
|
||||
)
|
||||
if retreat:
|
||||
accumulated = accumulated[:, :, : produced - retreat]
|
||||
produced = accumulated.shape[2]
|
||||
print(f" re-anchored {retreat} frames back (transient at seam)")
|
||||
|
||||
starts = [
|
||||
Image.fromarray(
|
||||
(accumulated[0, :, i].permute(1, 2, 0) * 255)
|
||||
.clamp(0, 255)
|
||||
.numpy()
|
||||
.astype(np.uint8)
|
||||
)
|
||||
for i in range(-overlap, 0)
|
||||
]
|
||||
|
||||
if accumulated is None:
|
||||
return {"error": "No frames generated (audio too short?)"}
|
||||
|
||||
silent = work_dir / "silent.mp4"
|
||||
save_videos_grid(accumulated[:, :, :total_frames], str(silent), fps=fps)
|
||||
|
||||
final_video = work_dir / "final.mp4"
|
||||
subprocess.run(
|
||||
["ffmpeg", "-y", "-i", str(silent), "-i", str(audio_path),
|
||||
"-c:v", "copy", "-c:a", "aac", "-shortest", str(final_video)],
|
||||
capture_output=True, timeout=300, check=True,
|
||||
)
|
||||
|
||||
elapsed = time.time() - start_time
|
||||
realtime_factor = elapsed / total_duration if total_duration else 0
|
||||
print(f"Done: {elapsed:.1f}s for {total_duration:.1f}s of video "
|
||||
f"({realtime_factor:.1f}x realtime), {segments} segments")
|
||||
|
||||
result = {
|
||||
"success": True,
|
||||
"duration_seconds": round(total_duration, 2),
|
||||
"segments": segments,
|
||||
"width": width,
|
||||
"height": height,
|
||||
"steps": steps,
|
||||
"processing_time_seconds": round(elapsed, 2),
|
||||
"realtime_factor": round(realtime_factor, 1),
|
||||
}
|
||||
|
||||
if r2_config:
|
||||
import boto3
|
||||
from botocore.config import Config
|
||||
|
||||
client = boto3.client(
|
||||
"s3",
|
||||
endpoint_url=r2_config["endpoint_url"],
|
||||
aws_access_key_id=r2_config["access_key_id"],
|
||||
aws_secret_access_key=r2_config["secret_access_key"],
|
||||
config=Config(signature_version="s3v4"),
|
||||
)
|
||||
object_key = f"echomimic3/results/{uuid.uuid4().hex[:12]}.mp4"
|
||||
client.upload_file(
|
||||
str(final_video), r2_config["bucket_name"], object_key,
|
||||
ExtraArgs={"ContentType": "video/mp4"},
|
||||
)
|
||||
result["video_url"] = client.generate_presigned_url(
|
||||
"get_object",
|
||||
Params={"Bucket": r2_config["bucket_name"], "Key": object_key},
|
||||
ExpiresIn=7200,
|
||||
)
|
||||
result["r2_key"] = object_key
|
||||
else:
|
||||
result["video_base64"] = base64.b64encode(final_video.read_bytes()).decode("utf-8")
|
||||
print("Warning: Returning video as base64 (use R2 for large files)")
|
||||
|
||||
return result
|
||||
|
||||
except torch.cuda.OutOfMemoryError:
|
||||
return {
|
||||
"error": "CUDA OOM. Lower sample_size (e.g. [576, 576]) or "
|
||||
"video_length (e.g. 65), or redeploy the app on L40S."
|
||||
}
|
||||
except subprocess.CalledProcessError as e:
|
||||
return {"error": f"ffmpeg mux failed: {e.stderr[-300:] if e.stderr else e}"}
|
||||
except Exception as e:
|
||||
import traceback
|
||||
|
||||
print(traceback.format_exc())
|
||||
return {"error": f"Internal error: {e}"}
|
||||
finally:
|
||||
shutil.rmtree(work_dir, ignore_errors=True)
|
||||
@@ -1,242 +0,0 @@
|
||||
# EchoMimicV3 - Talking Head Video Generation
|
||||
|
||||
Audio-driven talking head generation with [EchoMimicV3-Flash](https://github.com/antgroup/echomimic_v3)
|
||||
(Ant Group, Apache 2.0, AAAI 2026). Runs on Modal.
|
||||
|
||||
The toolkit has two talking-head generators. The short version: **EchoMimicV3 for anything
|
||||
the viewer looks at directly, SadTalker for small overlays and quick drafts.** See
|
||||
[Choosing between EchoMimicV3 and SadTalker](#choosing-between-echomimicv3-and-sadtalker).
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# Basic usage
|
||||
uv run tools/echomimic3.py --image portrait.png --audio voiceover.mp3 --output talking.mp4
|
||||
|
||||
# NarratorPiP settings - 16:9 in, 16:9 out, cheap 5-step pass
|
||||
uv run tools/echomimic3.py \
|
||||
--image presenter_16x9.png --audio scene_01.mp3 \
|
||||
--steps 5 --size 640 --output narrator.mp4
|
||||
|
||||
# Side by side against an existing SadTalker render of the same inputs
|
||||
uv run tools/echomimic3.py \
|
||||
--image presenter_16x9.png --audio scene_01.mp3 \
|
||||
--output narrator_echo.mp4 --compare narrator_sadtalker.mp4
|
||||
```
|
||||
|
||||
## Choosing between EchoMimicV3 and SadTalker
|
||||
|
||||
| | SadTalker | EchoMimicV3-Flash |
|
||||
|---|---|---|
|
||||
| Released | 2023, unmaintained | Flash variant Jan 2026 |
|
||||
| Licence | Apache 2.0 | Apache 2.0 |
|
||||
| Params | ~0.3B (warp-based) | 1.3B (Wan2.1-Fun diffusion) |
|
||||
| Aspect ratio | square crop unless `--preprocess full` | follows the input image |
|
||||
| Motion | head + light expression | head, upper body, gestures |
|
||||
| Cost | ~$0.0014 per second of output | ~$0.009 per second (**~6.5x**) |
|
||||
| Speed | faster than realtime | 22.8-47.8x realtime |
|
||||
| Cloud | RunPod or Modal | Modal only |
|
||||
|
||||
**Use EchoMimicV3 when** the narrator is large in frame, the shot is held long enough to
|
||||
watch, or the source image is not square and you do not want to fight the crop.
|
||||
|
||||
**Use SadTalker when** the output is a small overlay, you need a draft in a minute rather
|
||||
than an hour, or you are generating many takes to choose between.
|
||||
|
||||
This was checked rather than assumed. A controlled A/B — same still, same audio, both downscaled
|
||||
to NarratorPiP `sm` (240x135) — found EchoMimicV3 carries 3.5x the whole-frame motion and 6.2x in
|
||||
the mouth band, so the gap is not washed out by the shrink. On review the verdict was that
|
||||
EchoMimicV3 is clearly better while SadTalker is still perfectly usable at that size. Hence two
|
||||
tools rather than a replacement.
|
||||
|
||||
The cost gap is real but bounded: a 3-minute narrator is roughly $1.76 against $0.27.
|
||||
The **wall clock** is the sharper constraint — that same 3 minutes is 1.5-2.4 hours of
|
||||
generation, so per-scene narrator clips generated in advance beat a single long render.
|
||||
|
||||
## Setup
|
||||
|
||||
EchoMimicV3 is Modal-only.
|
||||
|
||||
```bash
|
||||
uv sync --extra modal && uv run modal setup
|
||||
|
||||
# One-off: fill the weights volume (~26GB, ~10 min). Needed before first deploy.
|
||||
uv run modal run docker/modal-echomimic3/app.py::populate_weights
|
||||
|
||||
uv run modal deploy docker/modal-echomimic3/app.py
|
||||
```
|
||||
|
||||
Then add the printed URL to `.env`:
|
||||
|
||||
```
|
||||
MODAL_ECHOMIMIC3_ENDPOINT_URL=https://....modal.run
|
||||
```
|
||||
|
||||
Unlike the other Modal apps, this one keeps its weights in a **Modal Volume** rather than
|
||||
baking them into the image — which is why `populate_weights` exists and why it must run
|
||||
first. See [Weight storage](#weight-storage) for why.
|
||||
|
||||
## Parameters
|
||||
|
||||
### Core settings
|
||||
|
||||
| Flag | Default | Notes |
|
||||
|------|---------|-------|
|
||||
| `--steps` | 8 | 5 is materially cheaper and holds up well; 8+ for hero shots |
|
||||
| `--size` | 768 | Generation resolution; 640 pairs well with `--steps 5` |
|
||||
| `--fps` | 25 | |
|
||||
| `--seed` | 43 | |
|
||||
| `--prompt` | "A person is speaking to the camera." | Effect is weak — see below |
|
||||
| `--wav2vec` | chinese | **Leave it.** See [Settled findings](#settled-findings) |
|
||||
|
||||
### Tuning
|
||||
|
||||
| Flag | Default | Notes |
|
||||
|------|---------|-------|
|
||||
| `--video-length` | 81 | Frames per segment. Raising it cuts the segment count — the single biggest lever on cost — until VRAM complains |
|
||||
| `--overlap` | 8 | Frames cross-faded between segments |
|
||||
| `--anchor-retreat` | 6 | Max frames to back off when a seam lands on a blink; 0 restores the old behaviour |
|
||||
| `--guidance-scale` | 6.0 | Text CFG, 3-6 |
|
||||
| `--audio-guidance-scale` | 3.0 | Audio CFG. Upstream suggests 1.8-2.0 for lip sync, but see below |
|
||||
| `--audio-scale` | 1.0 | Audio conditioning strength |
|
||||
|
||||
## Image guidelines
|
||||
|
||||
Same as SadTalker with one difference that matters: **EchoMimicV3 follows the input
|
||||
aspect ratio**, so a 16:9 presenter image comes back 16:9. There is no `--preprocess`
|
||||
equivalent and none is needed.
|
||||
|
||||
Good source images are front-facing, evenly lit, with the face 30-70% of the frame.
|
||||
Avoid heavy backlighting, extreme angles, and faces that fill the entire frame.
|
||||
|
||||
## Performance and cost
|
||||
|
||||
Measured on Modal A10G (24GB, 22.06 usable):
|
||||
|
||||
| Config | Realtime factor | Cost per second of output |
|
||||
|--------|-----------------|---------------------------|
|
||||
| `--steps 5 --size 640` | 22.8-28.9x | ~$0.009 |
|
||||
| `--steps 8 --size 768` | 47.8x | ~$0.015 |
|
||||
|
||||
The dominant cost is the **number of segments**, not the per-step cost. At the defaults
|
||||
each segment advances only `81 - 8 = 73` frames (2.9s), so a 30s narrator is 11 full
|
||||
diffusion passes and a 3-minute one is 62. Raise `--video-length` before reaching for
|
||||
fewer steps.
|
||||
|
||||
The model needs `enable_model_cpu_offload()` to fit 24GB, which trades PCIe paging for
|
||||
resident VRAM on every pipeline call.
|
||||
|
||||
## How long audio is handled
|
||||
|
||||
Upstream's `infer_flash.py` generates one 81-frame (3.2s) clip and **silently truncates
|
||||
longer audio** — the Flash pipeline does not accept the long-video kwargs the preview
|
||||
pipeline does. The segment loop therefore lives in `docker/modal-echomimic3/app.py`:
|
||||
each segment re-anchors on the last `--overlap` frames of the previous one, and the seam
|
||||
is cross-faded.
|
||||
|
||||
### Blinks at segment seams
|
||||
|
||||
Because each segment starts from the previous segment's final frames, whatever pose those
|
||||
frames hold becomes the next segment's opening pose. When they landed mid-blink the model
|
||||
started closed-eyed and **held it** — a prolonged closure straddling the boundary.
|
||||
|
||||
`--anchor-retreat` fixes this: before re-anchoring, the loop scores candidate anchor
|
||||
windows by how much motion they contain in the **upper half** of the frame (where blinks
|
||||
live and mouth movement does not) and backs off up to N frames to anchor on the calmest
|
||||
one. It only pays that cost when there is a clear improvement to be had, so clips with no
|
||||
transient near a seam regenerate nothing. `--anchor-retreat 0` restores the old behaviour.
|
||||
|
||||
## Settled findings
|
||||
|
||||
Things established by measurement, so they do not get re-litigated:
|
||||
|
||||
- **`--wav2vec english` is worse than `chinese`, even for English audio.** It
|
||||
under-articulates throughout (motion 2.26 vs 3.5) and visibly sits half-open. The Flash
|
||||
model was trained with the Chinese encoder; `run_flash.sh` uses it regardless of
|
||||
language. Keep the default.
|
||||
- **Do not trust a mouth-crop sync metric.** It scores a mouth crop only and is
|
||||
structurally blind to eye, hair and background artifacts. In the tuning matrix it ranked
|
||||
*highest* the one variant with a visible eye defect. Any automated scoring needs
|
||||
whole-face coverage, or use human review.
|
||||
- **The dependency pins are load-bearing.** `diffusers==0.32.2` and
|
||||
`transformers==4.49.0`. Newer transformers drives `output_hidden_states` from config
|
||||
rather than the kwarg EchoMimic's `Wav2Vec2Model` subclass passes, so audio embeddings
|
||||
come back empty and **all lip sync silently disappears**. Setting
|
||||
`config.output_hidden_states` does not fix it.
|
||||
- **The prompt is close to inert.** A rich descriptive prompt was marginally preferred by
|
||||
eye over `"A person is speaking."`, but not decisively. Do not spend effort here.
|
||||
- **`--audio-guidance-scale` is unsettled.** Upstream recommends 1.8-2.0; the default here
|
||||
is 3.0 to match `run_flash.sh`. An earlier recommendation to lower it was retracted
|
||||
because it rested on the mouth-crop metric above.
|
||||
|
||||
## Weight storage
|
||||
|
||||
This app keeps its ~26GB of weights in a Modal Volume. Every other `docker/modal-*` app
|
||||
bakes them into the image. That split is deliberate and measured:
|
||||
|
||||
| | Baked image | Volume |
|
||||
|---|---|---|
|
||||
| Rebuild after a dependency change | 79-385s | 1.8-8.2s |
|
||||
| Cold start | 57-94s | 60-67s |
|
||||
| Generation speed | identical | identical |
|
||||
| Storage cost | £0 | £0 (26GB against a 1TiB/month allowance) |
|
||||
|
||||
Cold start and generation are a wash; the rebuild difference is 20-100x. Getting this
|
||||
model working took four separate dependency changes, each of which re-downloaded 26GB
|
||||
under the baked scheme. The settled apps (`upscale`, `image-edit`) change rarely and gain
|
||||
nothing by moving, so they stay baked.
|
||||
|
||||
**The volume is optional, and it is free at this scale.** Modal charges $0.09/GiB/month
|
||||
for volume storage with **1 TiB/month included free**, so the 26.6GB this app stores is
|
||||
about 2.6% of the free allowance — nothing is being spent to hold it, and nothing is saved
|
||||
by not holding it. That is what makes the choice a pure engineering one rather than a cost
|
||||
trade-off: the only real question is whether faster rebuilds are worth an extra step, and
|
||||
for an app that needed four dependency changes to get working, they are.
|
||||
|
||||
If you would rather have one self-contained artifact — no `populate_weights` step, no
|
||||
ordering requirement, weights pinned to the image — the baked variant is fully supported
|
||||
and deploys as a separate app. Neither path is deprecated.
|
||||
|
||||
The one genuine cost of the volume is reproducibility: weights are no longer pinned to the
|
||||
image, so nothing but explicit revisions stops image and weights drifting apart. That is
|
||||
why the upstream repo ref and all four model revisions are pinned by SHA in `app.py`. Bump
|
||||
them deliberately and re-run `populate_weights`.
|
||||
|
||||
To deploy the baked variant instead (as a separate app, `video-toolkit-echomimic3-baked`):
|
||||
|
||||
```bash
|
||||
ECHOMIMIC_WEIGHTS=image uv run modal deploy docker/modal-echomimic3/app.py
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**Lip sync is completely absent, no error.** The `transformers` pin has drifted. It must
|
||||
be exactly `4.49.0` — see [Settled findings](#settled-findings).
|
||||
|
||||
**OOM during `@modal.enter()`.** `enable_model_cpu_offload()` is not being applied.
|
||||
Upstream's `infer_flash.py` parses a `GPU_memory_mode` flag but never applies it, so
|
||||
following upstream literally does not work on a 24GB card.
|
||||
|
||||
**A `pytorch_model.bin` refuses to load.** `chinese-wav2vec2-base` ships one, and
|
||||
transformers >= 4.51.3 refuses `torch.load` below torch 2.6 (CVE-2025-32434) as a hard
|
||||
failure. The image converts it to safetensors at build time rather than bumping torch,
|
||||
which would flip `torch.load`'s `weights_only` default and break EchoMimic's own `.pth`
|
||||
loads for the VAE, text encoder and CLIP.
|
||||
|
||||
**A prolonged eye closure across a segment boundary.** See
|
||||
[Blinks at segment seams](#blinks-at-segment-seams) — raise `--anchor-retreat`.
|
||||
|
||||
## Known deviations from upstream
|
||||
|
||||
- `_build_inputs` reimplements `src.utils.get_image_to_video_latent2`, which calls
|
||||
`.resize()` on its argument before checking whether it is a list and so raises
|
||||
`AttributeError` on the multi-frame re-anchoring the segment loop needs.
|
||||
- The image omits `tensorflow`, `retina-face`, `gradio`, `decord` and `moviepy` from
|
||||
upstream's `requirements.txt`. None are reachable from the Flash path — tensorflow and
|
||||
retina-face are only used by `src.face_detect` for the preview variant's `ip_mask`. Add
|
||||
them back if the preview variant is ever wired up.
|
||||
|
||||
## Still unverified
|
||||
|
||||
- Behaviour at full narration length — drift across 60+ segments has not been watched.
|
||||
- Whether the prompt matters at all.
|
||||
+14
-14
@@ -75,9 +75,9 @@ uv run modal deploy docker/modal-sadtalker/app.py
|
||||
uv run modal deploy docker/modal-propainter/app.py
|
||||
|
||||
# Talking head, diffusion-based. Weights live in a Volume, so populate it FIRST
|
||||
# (one-off, ~26GB, ~10 min) or the app deploys with nothing to load.
|
||||
uv run modal run docker/modal-echomimic3/app.py::populate_weights
|
||||
uv run modal deploy docker/modal-echomimic3/app.py
|
||||
# (one-off, ~15GB, ~5 min) or the app deploys with nothing to load.
|
||||
uv run modal run docker/modal-soulx/app.py::populate_weights
|
||||
uv run modal deploy docker/modal-soulx/app.py
|
||||
|
||||
# Video generation (see the LTX-2 prerequisites note below)
|
||||
uv run modal deploy docker/modal-ltx2/app.py
|
||||
@@ -111,7 +111,7 @@ MODAL_IMAGE_EDIT_ENDPOINT_URL=https://yourname--video-toolkit-image-edit-...moda
|
||||
MODAL_UPSCALE_ENDPOINT_URL=https://yourname--video-toolkit-upscale-...modal.run
|
||||
MODAL_MUSIC_GEN_ENDPOINT_URL=https://yourname--video-toolkit-music-gen-...modal.run
|
||||
MODAL_SADTALKER_ENDPOINT_URL=https://yourname--video-toolkit-sadtalker-...modal.run
|
||||
MODAL_ECHOMIMIC3_ENDPOINT_URL=https://yourname--video-toolkit-echomimic3-...modal.run
|
||||
MODAL_SOULX_ENDPOINT_URL=https://yourname--video-toolkit-soulx-...modal.run
|
||||
MODAL_DEWATERMARK_ENDPOINT_URL=https://yourname--video-toolkit-dewatermark-...modal.run
|
||||
MODAL_LTX2_ENDPOINT_URL=https://yourname--video-toolkit-ltx2-...modal.run
|
||||
```
|
||||
@@ -166,7 +166,7 @@ uv run tools/music_gen.py --preset corporate-bg --duration 60 --output bg.mp3
|
||||
|
||||
# Talking head from portrait + audio
|
||||
uv run tools/sadtalker.py --image portrait.png --audio voiceover.mp3 --output talking.mp4 --cloud modal
|
||||
uv run tools/echomimic3.py --image portrait.png --audio voiceover.mp3 --output talking.mp4
|
||||
uv run tools/soulx.py --image portrait.png --audio voiceover.mp3 --output talking.mp4
|
||||
|
||||
# Watermark removal
|
||||
uv run tools/dewatermark.py --input video.mp4 --region 1080,660,195,40 --output clean.mp4 --cloud modal
|
||||
@@ -182,29 +182,29 @@ uv run tools/dewatermark.py --input video.mp4 --region 1080,660,195,40 --output
|
||||
| `upscale` | RealESRGAN | AI image upscaling (2x/4x) | ~$0.005-0.02 |
|
||||
| `music_gen` | ACE-Step 1.5 | AI music generation | Free (acemusic) / ~$0.02-0.10 (Modal) |
|
||||
| `sadtalker` | SadTalker | Talking head video | ~$0.05-0.30 |
|
||||
| `echomimic3` | EchoMimicV3-Flash | Talking head video, aspect-preserving | ~$0.009 per second of output |
|
||||
| `soulx` | SoulX-FlashHead 1.3B Pro | Talking head video, aspect-preserving | ~$0.0024 per second of output |
|
||||
| `dewatermark` | ProPainter | AI video inpainting | ~$0.05-0.50 |
|
||||
|
||||
All apps use A10G GPUs (24GB VRAM) except `image_edit` which uses A100 for its 25GB model.
|
||||
|
||||
### Weight storage
|
||||
|
||||
Most apps **bake** their model weights into the image at build time. `echomimic3` is the
|
||||
exception: it keeps its ~26GB in a **Modal Volume**, which is why it needs the one-off
|
||||
Most apps **bake** their model weights into the image at build time. `soulx` is the
|
||||
exception: it keeps its ~15GB in a **Modal Volume**, which is why it needs the one-off
|
||||
`populate_weights` run above before its first deploy.
|
||||
|
||||
The split is measured, not stylistic. Cold start and generation speed are the same either
|
||||
way; what differs is rebuild time after a dependency change — 1.8-8.2s on a volume against
|
||||
79-385s baked, because any invalidated layer re-downloads everything below it. Apps that
|
||||
way; what differs is rebuild time after a dependency change — measured at 1.8-8.2s on a
|
||||
volume against 79-385s baked, because any invalidated layer re-downloads everything below it. Apps that
|
||||
still change often earn a volume; settled ones don't need one.
|
||||
|
||||
**Volumes are optional and free at this scale.** Modal charges $0.09/GiB/month for volume
|
||||
storage with **1 TiB/month included free**, so `echomimic3`'s 26.6GB costs nothing — it is
|
||||
storage with **1 TiB/month included free**, so `soulx`'s 14.7GB costs nothing — it is
|
||||
~2.6% of the free allowance. There is no bill either way, which is precisely why the choice
|
||||
comes down to rebuild speed versus having one self-contained artifact rather than to cost.
|
||||
|
||||
Every app can be built either way. `echomimic3` defaults to a volume and falls back with
|
||||
`ECHOMIMIC_WEIGHTS=image`; the rest bake by default. See `docs/echomimic3.md` for the full
|
||||
Every app can be built either way. `soulx` defaults to a volume; the rest bake by
|
||||
default. See `docs/soulx.md` for the full
|
||||
comparison.
|
||||
|
||||
## Cold Starts
|
||||
@@ -219,7 +219,7 @@ First request after idle triggers a cold start while Modal loads the model:
|
||||
| `upscale` | ~25-30s | ~3-5s |
|
||||
| `music_gen` | ~60-90s | ~10-30s |
|
||||
| `sadtalker` | ~45-60s | ~30-60s |
|
||||
| `echomimic3` | ~60-95s | 22.8-47.8x realtime |
|
||||
| `soulx` | ~15s + ~600s first-call torch.compile | 6.4-7.9x realtime |
|
||||
| `dewatermark` | ~60-70s | varies by video length |
|
||||
|
||||
After 60 seconds of no requests, containers scale back to zero. No charges while idle.
|
||||
|
||||
@@ -50,7 +50,7 @@ export interface NarratorPiPProps {
|
||||
* CSS object-fit for the video (default: 'contain').
|
||||
*
|
||||
* 'contain' is right for both generators. A 16:9 source fills a PiP box
|
||||
* exactly, since every SIZE_PRESET is 16:9 -- that is what tools/echomimic3.py
|
||||
* exactly, since every SIZE_PRESET is 16:9 -- that is what tools/soulx.py
|
||||
* produces, and what tools/sadtalker.py produces with --preprocess full.
|
||||
* SadTalker's default square crop letterboxes instead; use 'cover' to fill
|
||||
* the box at the cost of cropping the sides.
|
||||
|
||||
+2
-1
@@ -45,7 +45,7 @@ _MODAL_ENV_VARS = {
|
||||
"music_gen": "MODAL_MUSIC_GEN_ENDPOINT_URL",
|
||||
"dewatermark": "MODAL_DEWATERMARK_ENDPOINT_URL",
|
||||
"ltx2": "MODAL_LTX2_ENDPOINT_URL",
|
||||
"echomimic3": "MODAL_ECHOMIMIC3_ENDPOINT_URL",
|
||||
"soulx": "MODAL_SOULX_ENDPOINT_URL",
|
||||
}
|
||||
|
||||
|
||||
@@ -188,6 +188,7 @@ _TOOL_GPU = {
|
||||
"music_gen": "A10G",
|
||||
"dewatermark": "A10G",
|
||||
"ltx2": "A100-80GB",
|
||||
"soulx": "A10G",
|
||||
},
|
||||
"runpod": {
|
||||
"qwen3_tts": "ADA_24",
|
||||
|
||||
@@ -1,353 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Generate talking head videos using EchoMimicV3-Flash.
|
||||
|
||||
Candidate replacement for tools/sadtalker.py. EchoMimicV3 (Ant Group, Apache 2.0,
|
||||
1.3B params) is a Wan2.1-Fun-based audio-driven human animation model. Unlike
|
||||
SadTalker it preserves the input image's aspect ratio, so 16:9 presenter images
|
||||
come back 16:9 with no --preprocess workaround.
|
||||
|
||||
Usage:
|
||||
# Basic
|
||||
uv run tools/echomimic3.py --image portrait.png --audio voiceover.mp3 --output talking.mp4
|
||||
|
||||
# NarratorPiP settings (16:9 input, cheaper 5-step pass)
|
||||
uv run tools/echomimic3.py \
|
||||
--image presenter_16x9.png --audio scene_01.mp3 \
|
||||
--steps 5 --size 640 --output narrator.mp4
|
||||
|
||||
# A/B against the current SadTalker narrator for the same inputs
|
||||
uv run tools/echomimic3.py --image p.png --audio vo.mp3 --output new.mp4 --compare old.mp4
|
||||
|
||||
Setup:
|
||||
uv sync --extra modal && uv run modal setup
|
||||
uv run modal deploy docker/modal-echomimic3/app.py
|
||||
# then add the printed URL to .env:
|
||||
MODAL_ECHOMIMIC3_ENDPOINT_URL=https://....modal.run
|
||||
|
||||
Cost:
|
||||
Diffusion video generation, not SadTalker's warp-based animation -- expect
|
||||
roughly an order of magnitude more GPU time per second of output. The tool
|
||||
prints the measured realtime factor so the real number replaces this guess.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import base64
|
||||
import json
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).parent))
|
||||
from file_transfer import (
|
||||
upload_to_storage, download_from_r2, r2_cleanup,
|
||||
download_from_url, get_r2_payload_config,
|
||||
)
|
||||
|
||||
# Wall-clock budget per second of audio. Generous: an A10G cold start has to
|
||||
# page ~20GB of weights into VRAM before the first segment starts.
|
||||
PROCESSING_TIME_MULTIPLIER = 90
|
||||
PROCESSING_TIME_BUFFER = 420
|
||||
|
||||
DEFAULT_PROMPT = "A person is speaking to the camera."
|
||||
|
||||
|
||||
def get_audio_duration(audio_path: str) -> float | None:
|
||||
"""Get audio duration in seconds using ffprobe."""
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["ffprobe", "-v", "quiet", "-show_entries", "format=duration",
|
||||
"-of", "default=noprint_wrappers=1:nokey=1", audio_path],
|
||||
capture_output=True, text=True,
|
||||
)
|
||||
if result.returncode == 0:
|
||||
return float(result.stdout.strip())
|
||||
except Exception:
|
||||
pass
|
||||
return None
|
||||
|
||||
|
||||
def calculate_timeout(audio_duration: float) -> int:
|
||||
return int(audio_duration * PROCESSING_TIME_MULTIPLIER + PROCESSING_TIME_BUFFER)
|
||||
|
||||
|
||||
def build_comparison(new_video: str, old_video: str, output_path: str,
|
||||
verbose: bool = True) -> str | None:
|
||||
"""Stack two talking head renders side by side for eyeballing.
|
||||
|
||||
Labels each half so the pair stays readable once it is out of context.
|
||||
Heights are matched to the taller input; the audio comes from the new clip.
|
||||
"""
|
||||
if verbose:
|
||||
print(f"Building comparison: {old_video} | {new_video}", file=sys.stderr)
|
||||
|
||||
labelled = (
|
||||
"[0:v]scale=-2:720,pad=iw:ih+40:0:40:black,"
|
||||
"drawtext=text='SadTalker':x=10:y=8:fontsize=24:fontcolor=white[a];"
|
||||
"[1:v]scale=-2:720,pad=iw:ih+40:0:40:black,"
|
||||
"drawtext=text='EchoMimicV3':x=10:y=8:fontsize=24:fontcolor=white[b];"
|
||||
"[a][b]hstack=inputs=2[v]"
|
||||
)
|
||||
# drawtext needs a fontconfig default that not every ffmpeg build ships, so
|
||||
# fall back to an unlabelled stack rather than losing the comparison.
|
||||
plain = "[0:v]scale=-2:720[a];[1:v]scale=-2:720[b];[a][b]hstack=inputs=2[v]"
|
||||
|
||||
for filt in (labelled, plain):
|
||||
cmd = [
|
||||
"ffmpeg", "-y", "-i", old_video, "-i", new_video,
|
||||
"-filter_complex", filt, "-map", "[v]", "-map", "1:a?",
|
||||
"-c:v", "libx264", "-crf", "18", "-pix_fmt", "yuv420p",
|
||||
"-c:a", "aac", output_path,
|
||||
]
|
||||
proc = subprocess.run(cmd, capture_output=True, text=True)
|
||||
if proc.returncode == 0:
|
||||
if verbose:
|
||||
note = "" if filt is labelled else " (unlabelled: drawtext unavailable)"
|
||||
print(f" Comparison: {output_path}{note}", file=sys.stderr)
|
||||
return output_path
|
||||
|
||||
print(f"Comparison render failed: {proc.stderr[-400:]}", file=sys.stderr)
|
||||
return None
|
||||
|
||||
|
||||
def process_with_cloud(
|
||||
image_path: str,
|
||||
audio_path: str,
|
||||
output_path: str,
|
||||
prompt: str = DEFAULT_PROMPT,
|
||||
steps: int = 8,
|
||||
size: int = 768,
|
||||
video_length: int = 81,
|
||||
overlap: int = 8,
|
||||
anchor_retreat: int = 6,
|
||||
guidance_scale: float = 6.0,
|
||||
audio_guidance_scale: float = 3.0,
|
||||
audio_scale: float = 1.0,
|
||||
seed: int = 43,
|
||||
fps: int = 25,
|
||||
wav2vec: str = "chinese",
|
||||
timeout: int = 0,
|
||||
verbose: bool = True,
|
||||
cloud: str = "modal",
|
||||
progress=None,
|
||||
) -> dict:
|
||||
"""Generate a talking head via the EchoMimicV3 cloud endpoint."""
|
||||
with r2_cleanup() as r2_keys_to_cleanup:
|
||||
if verbose:
|
||||
print(f"Cloud provider: {cloud}", file=sys.stderr)
|
||||
|
||||
audio_duration = get_audio_duration(audio_path)
|
||||
if timeout <= 0:
|
||||
timeout = calculate_timeout(audio_duration) if audio_duration else 1800
|
||||
|
||||
if verbose and audio_duration:
|
||||
total_frames = int(audio_duration * fps)
|
||||
stride = max(1, video_length - overlap)
|
||||
segments = max(1, -(-total_frames // stride))
|
||||
print(f"Audio: {audio_duration:.1f}s -> ~{total_frames} frames, "
|
||||
f"~{segments} segment{'s' if segments > 1 else ''}, timeout {timeout}s",
|
||||
file=sys.stderr)
|
||||
|
||||
image_url, image_r2_key = upload_to_storage(image_path, "echomimic3/input")
|
||||
if not image_url:
|
||||
return {"error": "Failed to upload image"}
|
||||
if image_r2_key:
|
||||
r2_keys_to_cleanup.append(image_r2_key)
|
||||
|
||||
audio_url, audio_r2_key = upload_to_storage(audio_path, "echomimic3/input")
|
||||
if not audio_url:
|
||||
return {"error": "Failed to upload audio"}
|
||||
if audio_r2_key:
|
||||
r2_keys_to_cleanup.append(audio_r2_key)
|
||||
|
||||
payload = {
|
||||
"input": {
|
||||
"image_url": image_url,
|
||||
"audio_url": audio_url,
|
||||
"prompt": prompt,
|
||||
"steps": steps,
|
||||
"sample_size": [size, size],
|
||||
"video_length": video_length,
|
||||
"overlap": overlap,
|
||||
"anchor_retreat": anchor_retreat,
|
||||
"guidance_scale": guidance_scale,
|
||||
"audio_guidance_scale": audio_guidance_scale,
|
||||
"audio_scale": audio_scale,
|
||||
"seed": seed,
|
||||
"fps": fps,
|
||||
"wav2vec": wav2vec,
|
||||
}
|
||||
}
|
||||
|
||||
r2_payload = get_r2_payload_config()
|
||||
if r2_payload:
|
||||
payload["input"]["r2"] = r2_payload
|
||||
else:
|
||||
print("Warning: R2 not configured. Video will be returned as base64.", file=sys.stderr)
|
||||
|
||||
from cloud_gpu import call_cloud_endpoint
|
||||
|
||||
result, elapsed = call_cloud_endpoint(
|
||||
provider=cloud,
|
||||
payload=payload,
|
||||
tool_name="echomimic3",
|
||||
timeout=timeout,
|
||||
progress_label="Generating talking head",
|
||||
verbose=verbose,
|
||||
progress=progress,
|
||||
)
|
||||
|
||||
if isinstance(result, dict) and result.get("error"):
|
||||
return {"error": result["error"]}
|
||||
|
||||
Path(output_path).parent.mkdir(parents=True, exist_ok=True)
|
||||
downloaded = False
|
||||
|
||||
output_r2_key = result.get("r2_key") if isinstance(result, dict) else None
|
||||
output_url = result.get("video_url") if isinstance(result, dict) else None
|
||||
|
||||
if output_r2_key:
|
||||
downloaded = download_from_r2(output_r2_key, output_path)
|
||||
if downloaded:
|
||||
r2_keys_to_cleanup.append(output_r2_key)
|
||||
|
||||
if not downloaded and output_url:
|
||||
downloaded = download_from_url(output_url, output_path, verbose=verbose)
|
||||
if downloaded and output_r2_key:
|
||||
r2_keys_to_cleanup.append(output_r2_key)
|
||||
|
||||
if not downloaded:
|
||||
video_base64 = result.get("video_base64") if isinstance(result, dict) else None
|
||||
if video_base64:
|
||||
Path(output_path).write_bytes(base64.b64decode(video_base64))
|
||||
downloaded = True
|
||||
|
||||
if not downloaded:
|
||||
return {"error": f"No video in result: {list(result.keys()) if isinstance(result, dict) else result}"}
|
||||
|
||||
if verbose:
|
||||
size_kb = Path(output_path).stat().st_size // 1024
|
||||
rtf = result.get("realtime_factor")
|
||||
print(f" Downloaded: {output_path} ({size_kb}KB, "
|
||||
f"{result.get('width')}x{result.get('height')}"
|
||||
f"{f', {rtf}x realtime' if rtf else ''})", file=sys.stderr)
|
||||
|
||||
return {
|
||||
"success": True,
|
||||
"output": output_path,
|
||||
"processing_time_seconds": round(elapsed, 2),
|
||||
"duration_seconds": result.get("duration_seconds"),
|
||||
"segments": result.get("segments"),
|
||||
"width": result.get("width"),
|
||||
"height": result.get("height"),
|
||||
"realtime_factor": result.get("realtime_factor"),
|
||||
}
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(
|
||||
description="Generate talking head videos with EchoMimicV3-Flash",
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||||
epilog=__doc__,
|
||||
)
|
||||
|
||||
io_group = parser.add_argument_group("Input/output")
|
||||
io_group.add_argument("--image", "-i", required=True, help="Portrait image (16:9 for NarratorPiP)")
|
||||
io_group.add_argument("--audio", "-a", required=True, help="Driving audio file")
|
||||
io_group.add_argument("--output", "-o", default="talking.mp4", help="Output video path")
|
||||
io_group.add_argument("--compare", metavar="OLD_VIDEO",
|
||||
help="Also render a side-by-side against an existing (e.g. SadTalker) clip")
|
||||
|
||||
gen_group = parser.add_argument_group("Generation")
|
||||
gen_group.add_argument("--prompt", "-p", default=DEFAULT_PROMPT,
|
||||
help=f"Motion prompt (default: {DEFAULT_PROMPT!r})")
|
||||
gen_group.add_argument("--steps", type=int, default=8,
|
||||
help="Denoise steps: 5 for talking head, 15-25 for talking body (default: 8)")
|
||||
gen_group.add_argument("--size", type=int, default=768,
|
||||
help="Target size; output keeps the image's aspect ratio (default: 768)")
|
||||
gen_group.add_argument("--seed", type=int, default=43, help="Random seed")
|
||||
gen_group.add_argument("--fps", type=int, default=25, help="Output frame rate (default: 25)")
|
||||
gen_group.add_argument("--wav2vec", choices=["chinese", "english"], default="chinese",
|
||||
help="Audio encoder. 'chinese' is what upstream run_flash.sh uses "
|
||||
"for both languages; 'english' is worth A/B-ing (default: chinese)")
|
||||
|
||||
tune_group = parser.add_argument_group("Tuning")
|
||||
tune_group.add_argument("--video-length", type=int, default=81,
|
||||
help="Frames per segment; lower to cut VRAM (default: 81)")
|
||||
tune_group.add_argument("--overlap", type=int, default=8,
|
||||
help="Frames cross-faded between segments (default: 8)")
|
||||
tune_group.add_argument("--anchor-retreat", type=int, default=6,
|
||||
help="Max frames to back off when the seam lands on a blink "
|
||||
"or other transient; 0 disables (default: 6)")
|
||||
tune_group.add_argument("--guidance-scale", type=float, default=6.0, help="Text CFG, 3-6")
|
||||
tune_group.add_argument("--audio-guidance-scale", type=float, default=3.0,
|
||||
help="Audio CFG; upstream suggests 1.8-2.0 for tightest lip sync")
|
||||
tune_group.add_argument("--audio-scale", type=float, default=1.0, help="Audio conditioning strength")
|
||||
|
||||
out_group = parser.add_argument_group("Output control")
|
||||
out_group.add_argument("--timeout", type=int, default=0, help="Override auto-calculated timeout")
|
||||
out_group.add_argument("--json", action="store_true", help="Output result as JSON")
|
||||
out_group.add_argument("--quiet", "-q", action="store_true", help="Suppress progress output")
|
||||
out_group.add_argument("--cloud", default="modal", choices=["modal"],
|
||||
help="Cloud provider (RunPod not deployed for this tool yet)")
|
||||
|
||||
args = parser.parse_args()
|
||||
verbose = not args.quiet and not args.json
|
||||
|
||||
for label, path in (("image", args.image), ("audio", args.audio)):
|
||||
if not Path(path).exists():
|
||||
msg = f"{label.capitalize()} not found: {path}"
|
||||
print(json.dumps({"error": msg}) if args.json else f"Error: {msg}", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
if args.compare and not Path(args.compare).exists():
|
||||
msg = f"Comparison video not found: {args.compare}"
|
||||
print(json.dumps({"error": msg}) if args.json else f"Error: {msg}", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
result = process_with_cloud(
|
||||
image_path=args.image,
|
||||
audio_path=args.audio,
|
||||
output_path=args.output,
|
||||
prompt=args.prompt,
|
||||
steps=args.steps,
|
||||
size=args.size,
|
||||
video_length=args.video_length,
|
||||
overlap=args.overlap,
|
||||
anchor_retreat=args.anchor_retreat,
|
||||
guidance_scale=args.guidance_scale,
|
||||
audio_guidance_scale=args.audio_guidance_scale,
|
||||
audio_scale=args.audio_scale,
|
||||
seed=args.seed,
|
||||
fps=args.fps,
|
||||
wav2vec=args.wav2vec,
|
||||
timeout=args.timeout,
|
||||
verbose=verbose,
|
||||
)
|
||||
|
||||
if result.get("error"):
|
||||
print(json.dumps(result) if args.json else f"Error: {result['error']}", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
if args.compare:
|
||||
compare_path = str(Path(args.output).with_name(Path(args.output).stem + "_vs_sadtalker.mp4"))
|
||||
built = build_comparison(args.output, args.compare, compare_path, verbose=verbose)
|
||||
if built:
|
||||
result["comparison"] = built
|
||||
|
||||
if args.json:
|
||||
print(json.dumps(result, indent=2))
|
||||
else:
|
||||
print(f"\nGenerated: {result['output']}")
|
||||
if result.get("realtime_factor"):
|
||||
print(f" {result['duration_seconds']}s of video in "
|
||||
f"{result['processing_time_seconds']}s ({result['realtime_factor']}x realtime)")
|
||||
if result.get("comparison"):
|
||||
print(f" Side-by-side: {result['comparison']}")
|
||||
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
Reference in New Issue
Block a user