execute() sends sampleCount=number_of_images and estimate_cost() bills
0.04 * n, but result handling decoded only predictions[0] and wrote it to
a single output_path. Images 2..n were dropped: never decoded, never
written, absent from artifacts. The user paid for n and received one.
The result also misreported the drop rather than failing loudly --
images_generated returned len(predictions) (what the API sent) while
artifacts held a single path, so an agent picking between variants read a
count that did not match the artifact list.
Add _output_paths() and loop over every prediction, mirroring the pattern
already used by openai_image and grok_image: suffix multi-image paths
_1/_2/... so none overwrite each other, keep the exact requested path when
n=1, return all paths in artifacts, and report images_generated as the
count actually written.
Closes#388
Add an Azure AI Speech transcription tool. It is opt-in: when
AZURE_SPEECH_KEY is configured the agent may prefer it for cloud STT,
while the local faster-whisper `transcriber` stays the default offline
path. Shared pipeline manifests are intentionally left unchanged, so no
default provider selection is altered for existing users.
- tools/analysis/azure_stt.py: new `azure_stt` tool (capability=analysis,
provider=azure) calling the Fast Transcription REST API. The local file
is uploaded via multipart and transcribed synchronously with word-level
timestamps and optional diarization — no Blob storage or async polling.
Output schema mirrors `transcriber` exactly, so it is a drop-in for
`subtitle_gen` and other transcript consumers. Follows the existing
provider-tool conventions (env-var status check, `_transcribe` helper,
cost_usd/model on the result, fallback="transcriber").
- Auto-discovered by the registry; no registry or selector changes.
- tests/tools/test_azure_stt.py: contract, discovery, status, response
mapping, execute guardrails, and a mocked-network success path (no live
API calls).
- .agents/skills + .claude/skills: azure-speech-to-text Layer-3 skill.
- docs/PROVIDERS.md: Azure AI Speech setup, API notes, and pricing.
- .env.example, skills/INDEX.md, AGENT_GUIDE.md: document the optional
cloud STT path alongside the default whisper transcriber.
Include every request field that can alter Kling video, image, avatar, or lip-sync media in the public idempotency contract. Add a shared regression matrix that detects future cache-key collisions while excluding transport-only controls.
Isolate Kling contract tests from the singleton registry so discovery state cannot leak into later selector tests. Align lip-sync face, audio, and timing payloads with the current official API and extend the live smoke coverage.
Review findings from PR #333:
P1: _download_via_uri assumed output_video.uri is always files/<id>.
The API can return a full resource URI or a ready-made
.../files/<id>:download?alt=media download URL, which produced an
invalid poll path with a second :download appended. New
_file_id_from_uri() extracts the bare id from every documented shape;
regression tests cover the full-URL form plus a parametrized matrix of
URI shapes.
P2: docs/PROVIDERS.md still described the Google key as TTS + Imagen
only. The shared-key section now covers gemini_omni_video (model id,
~$0.10/sec pricing table, paid-tier-only, edit-turn billing note), and
the env snippet, provider-to-tool mapping, and capability coverage
tables include the new provider.
Add gemini_omni_video, a native Gemini API provider wrapping
gemini-omni-flash-preview via the Interactions API. Text-to-video,
image/reference-to-video with <FIRST_FRAME>/<IMAGE_REF_N> prompt tags,
and stateful edit_video turns via previous_interaction_id — the only
provider in the fleet that can refine a clip without regenerating it.
Reuses the existing GOOGLE_API_KEY / GEMINI_API_KEY, so one Google key
now unlocks images, TTS, and video.
- New Layer 3 skill .agents/skills/gemini-omni (prompting, edit-loop
rules, tag/timecode syntax, preview limits) sourced from official
Google docs; linked via agent_skills and the AGENT_GUIDE Layer 3 map
- ai-video-gen gains the Gemini API gateway row + editing pointer
- veo_video/sora_video fallback lists and video_selector agent_skills
reference the new provider; quality_score 0.85 with rationale
- Contract tests: registry discovery, selector routing, status from
env keys, uri + inline delivery, edit turns, typed image parts,
store=false editability, cost clamp
_validate_artifacts_for_stage looked up CANONICAL_STAGE_ARTIFACTS[stage]
unconditionally, but the valid stage list comes from the pipeline manifest via
get_pipeline_stages(), which declares stages beyond the 9 canonical ones — e.g.
character-animation adds `character_design`/`rig_plan`. Such a stage passes the
`stage in valid_stages` guard, then raised an unhandled KeyError on the
canonical lookup, so those stages could never be checkpointed (the crash hits
write/read_checkpoint and friends, even for in_progress checkpoints).
Look the canonical artifact up defensively with `.get()` and skip the
required-artifact check when there is none. Canonical stages still require their
artifact when completed.
_ts_srt/_ts_vtt computed the seconds and millisecond fields independently:
`ms = int(round((seconds % 1) * 1000))`. When the fractional part is >= 0.9995
that rounds to 1000, emitting a malformed 4-digit `…,1000` value with no carry
into the seconds field (and, at 59.9999/3599.9999, no carry into minutes/hours).
For example 0.9999s became `00:00:00,1000` instead of `00:00:01,000`. ASR word
and segment end-times routinely land on such fractional boundaries, and the
resulting cue is rejected or mistimed by strict SRT/VTT parsers (ffmpeg
subtitles filter, VLC, browser WebVTT).
Decompose from a single rounded total-milliseconds value so the carry
propagates across all fields. Both formatters now share one `_hmsms` helper.
skills/creative/music-gen-usage.md mandates 'Always set
force_instrumental=true for video background', but music_gen.py never
sent the kwarg, so ElevenLabs could return vocal tracks that collide
with narration/dialogue.
- Add force_instrumental to input_schema (default True) so the mandate
holds by default; callers may opt out only with an explicit False.
- Include force_instrumental in the /v1/music payload.
- Add tests pinning: kwarg sent True by default, explicit opt-out
honored, and the schema default.
Refs: docs/REVIEW-image-to-video-voice.md §8 #8
Co-Authored-By: Claude <noreply@anthropic.com>
COGVIDEO_VARIANTS declares cogvideo-2b i2v=False (it is t2v-only), but
cogvideo_video advertised image_to_video + reference_image unconditionally
and the variant flag was never consulted. An image_to_video brief against
the 2B variant reached the diffusion pipeline and failed opaquely.
- Add is_operation_available(operation) that derives capability from the
variant table (the selector calls it without inputs, so it reports the
DEFAULT variant cogvideo-5b: t2v + i2v both True). This replaces an
implicit unconditional-True.
- Add an execute()-time guard that consults the CALLER's chosen variant
and fails fast with a clear error when it lacks the requested mode
(2B + image_to_video), instead of dropping into generate_local_video.
- Add _variant_for(inputs) helper shared by estimate_runtime / the guard.
Tests pin: the 2B premise (i2v=False), default-variant capability
reporting, fast-fail for 2B+i2v (generation never runs), and that 5B+i2v
still routes through to generate_local_video.
Refs: docs/REVIEW-image-to-video-voice.md §8 #4
Co-Authored-By: Claude <noreply@anthropic.com>
Three routing defects in video_selector, none previously covered by
tests (REVIEW §8 #3, #5, #7); plus the routing-test coverage itself (#10).
#3 Seedance dedup race
tool_by_provider keyed by provider STRING, so two tools legitimately
sharing provider="seedance" (seedance_video=fal, seedance_replicate)
collided — only the first-registered was ever selectable; the other
was invisible to the selector regardless of rank. Key selectable tools
by NAME instead; ranking picks the best of the shared-provider backends.
#5 preferred_provider had no score-gap gate
The selector returned the preferred provider on the first ranking match
no matter how far below the top it scored (the comment claimed "unless
drastically worse" but nothing enforced it). Add a configurable
preferred_provider_gap (default 0.15): honor the preference only when
its best ranked tool is within the gap of the overall top, else yield
to the top-ranked provider.
#7 fallback_tools appended image_selector unconditionally
The motion-required prohibition lived only in director skills, so a
direct caller could silently fall back to an image-only tool for an
image_to_video / reference_to_video brief. Add input-aware
fallback_tools_for(inputs) that drops image_selector for
motion-required operations; keep the static fallback_tools property
(with image_selector) for external consumers / contracts.
#10 routing coverage
First routing tests for video_selector: dedup reachability, the gap
gate (honored / ignored / configurable), motion-aware fallback, and
estimate_cost / estimate_runtime delegation. 13 tests, scoring patched
for determinism so they test routing logic, not the scorer.
Full tools + contracts suite green (638 passed, 6 skipped).
Refs: docs/REVIEW-image-to-video-voice.md §8 #3, #5, #7, #10
Co-Authored-By: Claude <noreply@anthropic.com>
audio_mixer hard-coded loudnorm I=-16 (Apple Podcasts) in both _mix
and _full_mix. sound-design.md targets -14 for YouTube/TikTok/IG, and
edit_decisions.metadata.loudnorm_target is the declarative form — but
the mixer never read it, so the executed loudness silently defaulted
to podcast levels regardless of the target platform.
- Add loudnorm_target to input_schema (default -16, clamped to [-40, 0]).
- Extract _loudnorm_filter() helper and use it in _mix and _full_mix so
a director can forward edit_decisions.metadata.loudnorm_target (or a
caller can pass it directly) to hit the right platform target.
- Add tests pinning: default -16, -14 honored, out-of-range clamped,
non-numeric fallback, and the schema default.
Refs: docs/REVIEW-image-to-video-voice.md §8 #1
Co-Authored-By: Claude <noreply@anthropic.com>
Every premium video provider sets quality_score (seedance 0.95, runway /
higgsfield 0.9) so the scorer ranks them above stock/local options.
grok_video had none, so it was scored only on supports/stability flags
despite shipping native synchronized audio (lip-sync + dialogue + SFX
in a single generation pass) — likely under-ranked.
Set quality_score=0.9, on par with the other native-audio premium
providers. Add a regression pinning the field and its get_info() surface.
Refs: docs/REVIEW-image-to-video-voice.md §8 #6
Co-Authored-By: Claude <noreply@anthropic.com>
`_segmented_music` mixed the video's audio with the shaped music via
`amix=inputs=2`, whose default `normalize=1` scales every input by 1/inputs
(x0.5, -6 dB). Unlike `_mix` and `_full_mix`, this path has no `loudnorm` stage
afterward to re-normalize, so the narration was permanently attenuated across
the entire timeline — including the stretches where the music volume expression
evaluates to 0. A one-second music segment quietly dropped the narration by
~6 dB for the whole video.
Add `normalize=0` to the amix: the music is already scaled to `music_volume`
by the `volume` expression, so speech passes at unity. Verified with ffmpeg —
narration in a no-music region tracks the stereo/aac conversion baseline
instead of sitting 6 dB below it.
The tool advertised `multiple_outputs: True`, accepted `n` (1-4) in its schema,
requested `n` images from the API, and scaled `estimate_cost` by `n` — but the
result handling was hardcoded to `response.data[0]`. Images 1..n-1 were decoded
never, written never, and absent from `artifacts`, so a caller who set `n=4`
paid for four images and received one.
Iterate over `response.data`, writing each image to a distinct path (suffixed
`_1`, `_2`, … when several are requested, mirroring `grok_image` /
`dashscope_image`), and return `outputs` / `images_generated` alongside the
full `artifacts` list. A single image keeps its exact requested path.
video_compose.get_info() reported render_engines.ffmpeg as always
available, unlike the real availability checks used for remotion and
hyperframes. On a machine without ffmpeg on PATH, preflight would
falsely report ffmpeg as usable, letting render_runtime="ffmpeg" get
locked at proposal time only to fail at compose.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
From an independent review of the branch:
- Regression: existing non-raster-but-showable visuals (.svg diagrams) were
dropped by the renderable filter — they were served fine before via <img>
(/thumb passes SVG through). Add .svg to MEDIA_IMAGE_EXT.
- Doc accuracy: the board dedupes decisions by (category, subject), not category
alone (category-only would wrongly merge distinct decisions that share a
category, e.g. TTS vs image provider_selection). Correct AGENT_GUIDE to say
the pair is the key and to reuse the same subject when re-logging.
- Coverage: the new visual-selection logic was untested (which let the earlier
missing-file regression through). Add TestStoryboardVisualSelection covering
the .tsx-animation exclusion, snapshot fallback (exact + <id>_* match), SVG
renderability, the preserved missing-file indicator, and takes = renderable.
Not changed (reviewed, deliberate): .narr clamp at --fs-scale 1.16 degrades
gracefully via the fade + click-to-expand modal; decision dedupe stays keyed on
(category, subject) as the more-correct behavior.