mirror of
https://github.com/calesthio/OpenMontage.git
synced 2026-08-05 15:20:40 +08:00
dashscope: add Layer 3 skill documentation
This commit is contained in:
136
.agents/skills/dashscope/SKILL.md
Normal file
136
.agents/skills/dashscope/SKILL.md
Normal file
@@ -0,0 +1,136 @@
|
||||
---
|
||||
name: dashscope
|
||||
description: DashScope (Alibaba Cloud Bailian / 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans). Use when generating images via Qwen-Image, narrating via Qwen-TTS, or transcribing with word-level timestamps via Qwen-ASR.
|
||||
---
|
||||
|
||||
# DashScope
|
||||
|
||||
Requires `DASHSCOPE_API_KEY` in `.env`. Get one at https://dashscope.aliyun.com/.
|
||||
|
||||
## Current API
|
||||
|
||||
**CRITICAL:** DashScope's `/compatible-mode/v1/` only supports `/chat/completions` and `/embeddings`. Image generation, TTS, and ASR all use **DashScope-native endpoints** — not OpenAI-compatible paths.
|
||||
|
||||
All three tools use `Authorization: Bearer $DASHSCOPE_API_KEY`.
|
||||
|
||||
### Image Generation
|
||||
|
||||
```text
|
||||
POST https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
|
||||
```
|
||||
|
||||
- Model: `qwen-image-2.0-pro` (default), `qwen-image-max`, `wan2.7-image`, `z-image-turbo`
|
||||
- Body: `{model, input: {messages: [{role: "user", content: [{text: "prompt"}]}]}, parameters: {size: "W*H", n, prompt_extend, watermark}}`
|
||||
- **Size format uses asterisk:** `"1024*1024"` not `"1024x1024"`
|
||||
- Response: `output.choices[0].message.content[0].image` (URL, valid ~24h) — must download separately
|
||||
|
||||
### Text-to-Speech
|
||||
|
||||
```text
|
||||
POST https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
|
||||
```
|
||||
|
||||
Same endpoint as image gen, different body.
|
||||
|
||||
- Model: `qwen3-tts-flash` (default), `qwen3-tts-instruct-flash`, `qwen-tts-2025-05-22`
|
||||
- Body: `{model, input: {text, voice: "Cherry", language_type: "Auto"}}`
|
||||
- Response: `output.audio.url` (WAV, valid ~24h) — must download separately
|
||||
|
||||
### ASR with Word-Level Timestamps
|
||||
|
||||
```text
|
||||
POST https://dashscope.aliyuncs.com/api/v1/services/audio/asr/transcription
|
||||
Header: X-DashScope-Async: enable
|
||||
```
|
||||
|
||||
- Model: `qwen3-asr-flash-filetrans` (NOT `qwen3-asr-flash` — the sync version has no word timestamps)
|
||||
- Body: `{model, input: {file_url: "https://public-url/audio.mp3"}, parameters: {enable_words: true, language_hints: ["zh","en"]}}`
|
||||
- Returns `task_id` → poll `GET /api/v1/tasks/{task_id}` until `SUCCEEDED` → download `output.result.transcription_url` → JSON with `transcripts[].sentences[].words[]`
|
||||
- Timestamps in `begin_time`/`end_time` are in **milliseconds** — the tool normalizes to seconds
|
||||
|
||||
## OpenMontage Usage
|
||||
|
||||
### Image via selector
|
||||
|
||||
```python
|
||||
from tools.graphics.image_selector import ImageSelector
|
||||
|
||||
result = ImageSelector().execute({
|
||||
"preferred_provider": "dashscope",
|
||||
"prompt": "一只猫坐在沙发上",
|
||||
"output_path": "projects/my-video/assets/images/cat.png",
|
||||
})
|
||||
```
|
||||
|
||||
### TTS via selector
|
||||
|
||||
```python
|
||||
from tools.audio.tts_selector import TTSSelector
|
||||
|
||||
result = TTSSelector().execute({
|
||||
"preferred_provider": "dashscope",
|
||||
"text": "如果 AI 真的会改变未来,普通人到底该怎么参与?",
|
||||
"voice": "Cherry",
|
||||
"output_path": "projects/my-video/assets/audio/narration.wav",
|
||||
})
|
||||
```
|
||||
|
||||
### ASR directly (word timestamps for subtitles)
|
||||
|
||||
```python
|
||||
from tools.analysis.dashscope_asr import DashscopeAsr
|
||||
|
||||
result = DashscopeAsr().execute({
|
||||
"audio_url": "https://example.com/narration.wav",
|
||||
"output_path": "projects/my-video/assets/audio/transcription.json",
|
||||
})
|
||||
|
||||
# result.data["words"] is a flat list of {text, begin_time_seconds, end_time_seconds}
|
||||
```
|
||||
|
||||
## Recommended Workflow
|
||||
|
||||
1. **Image:** Generate a sample first. Check `prompt_extend: true` (default) — DashScope rewrites your prompt for better results. Disable if you need literal prompt adherence.
|
||||
2. **TTS:** Generate a 10-15 second sample before full narration. Approve voice and pacing before committing to full generation.
|
||||
3. **ASR:** Audio must be at a **publicly accessible URL**. Upload to any public host (S3, etc.) first. Local paths are rejected with a clear error.
|
||||
4. **Subtitles:** Build from `result.data["words"]` — each word has `begin_time_seconds` and `end_time_seconds`. Group words into caption phrases by language semantics, not fixed character count.
|
||||
|
||||
## Parameters
|
||||
|
||||
### Image (`dashscope_image`)
|
||||
- `prompt` (required): text prompt
|
||||
- `model`: default `qwen-image-2.0-pro`
|
||||
- `size`: default `"1024*1024"` — **asterisk separator, not "x"**
|
||||
- `n`: 1-6 images
|
||||
- `negative_prompt`: things to avoid (max 500 chars)
|
||||
- `prompt_extend`: default `true` — auto-rewrite prompt for better results
|
||||
- `watermark`: default `false`
|
||||
- `seed`: for reproducibility
|
||||
|
||||
### TTS (`dashscope_tts`)
|
||||
- `text` (required): text to synthesize (max 600 chars for qwen3-tts-flash)
|
||||
- `model`: default `qwen3-tts-flash`
|
||||
- `voice`: default `"Cherry"` — other voices: `"Ethan"`, `"Chelsie"`, etc.
|
||||
- `language_type`: default `"Auto"` — `"Chinese"`, `"English"`, `"Japanese"`, `"Korean"`
|
||||
- `instructions`: natural language delivery instructions (only for `qwen3-tts-instruct-flash`)
|
||||
|
||||
### ASR (`dashscope_asr`)
|
||||
- `audio_url` (required): **must be publicly accessible URL**
|
||||
- `model`: `qwen3-asr-flash-filetrans` (only model that supports word timestamps)
|
||||
- `language_hints`: default `["zh", "en"]`
|
||||
- `enable_words`: default `true` — required for word-level timestamps
|
||||
- `poll_interval_seconds`: default `5.0`
|
||||
- `timeout_seconds`: default `300`
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- **Image size error:** Use `"W*H"` with asterisk, not `"WxH"`. Example: `"2048*2048"`.
|
||||
- **TTS no audio URL:** Check `output.audio.url` — if empty, the model name or voice may be wrong.
|
||||
- **ASR "file not accessible":** `audio_url` must be publicly reachable. DashScope servers fetch the file; local paths and auth-gated URLs don't work.
|
||||
- **ASR poll timeout:** Increase `timeout_seconds` (default 300). Long audio files take longer to transcribe.
|
||||
- **ASR no word timestamps:** Ensure `enable_words: true` and model is `qwen3-asr-flash-filetrans` (not the sync `qwen3-asr-flash`).
|
||||
- **Auth error (401):** Verify `DASHSCOPE_API_KEY` is set. Use `Authorization: Bearer $KEY` header.
|
||||
|
||||
## Safety
|
||||
|
||||
Never print or write the API key to logs, metadata, patches, or project artifacts. `.env.example` should contain only empty variable names. The tool's `_safe_error()` method redacts the key from error messages.
|
||||
Reference in New Issue
Block a user