mirror of
https://github.com/modelstudioai/cli.git
synced 2026-09-14 19:49:23 +08:00
Compare commits
12 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| a755388734 | |||
| 2220f932b5 | |||
| 5c390fad74 | |||
| 6218637f7c | |||
| 0c7ba721a7 | |||
| 0dbf071368 | |||
| 16e30686f0 | |||
| 7803d91e3f | |||
| c2d17707c0 | |||
| 1dfb4400cc | |||
| fd2077b373 | |||
| 3477a08dad |
@@ -6,6 +6,14 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
|
||||
|
||||
[中文版](CHANGELOG.zh.md) · [README](README.md) · [Contributing](CONTRIBUTING.md)
|
||||
|
||||
## [1.23.0] - 2026-09-10
|
||||
|
||||
### Added
|
||||
|
||||
- **Profile-level watermark control** — configure `watermark` with `bl config set --key watermark --value true|false` to control the default watermark behavior for image generation and editing, video generation and editing, and reference-to-video commands.
|
||||
- **ASR accuracy controls** — `bl speech recognize` now supports instant hot words with `--vocabulary`, contextual word enhancement with `--context`, and reusable pre-built vocabularies with `--vocabulary-id` for supported ASR models.
|
||||
- **Speech vocabulary management** — added `bl speech vocabulary create|list|get|update|delete` to manage reusable pre-built hot-word vocabularies.
|
||||
|
||||
## [1.22.0] - 2026-09-08
|
||||
|
||||
### Changed
|
||||
|
||||
@@ -6,6 +6,14 @@
|
||||
|
||||
[English](CHANGELOG.md) · [README](README.zh.md) · [参与贡献](CONTRIBUTING.zh.md)
|
||||
|
||||
## [1.23.0] - 2026-09-10
|
||||
|
||||
### 新增
|
||||
|
||||
- **Profile 级水印控制** —— 可通过 `bl config set --key watermark --value true|false` 设置图片生成与编辑、视频生成与编辑以及参考生视频命令的默认水印行为。
|
||||
- **ASR 准确率增强** —— `bl speech recognize` 现支持通过 `--vocabulary` 传入即时热词、通过 `--context` 增强上下文词表,以及通过 `--vocabulary-id` 使用适用于对应 ASR 模型的预编译热词表。
|
||||
- **语音热词表管理** —— 新增 `bl speech vocabulary create|list|get|update|delete`,用于管理可复用的预编译热词表。
|
||||
|
||||
## [1.22.0] - 2026-09-08
|
||||
|
||||
### 变更
|
||||
|
||||
@@ -115,6 +115,7 @@ Once installed, just describe your task to your AI Agent — no need to assemble
|
||||
| ------------------------ | --------------------------------------------------------------------------------- |
|
||||
| Managed Agent | "Create a Managed Agent that can generate short-film storyboards and videos." |
|
||||
| Image & video generation | "Generate an image of a cat in a spacesuit on Mars, then turn it into a video." |
|
||||
| Speech recognition | "Transcribe this audio; if proper nouns are wrong, add hot words and try again." |
|
||||
| Usage & quota | "Show my recent model usage, free-tier quota, and rate limits." |
|
||||
| Model selection | "Recommend a model for image understanding and customer support." |
|
||||
| About Bailian CLI | "Tell me what Bailian CLI can do for me, and suggest how to use it for my needs." |
|
||||
|
||||
@@ -114,6 +114,7 @@ irm https://bailian.aliyun.com/cli/install.ps1 | iex
|
||||
| ---------------- | ----------------------------------------------------------------------- |
|
||||
| Managed Agent | “帮我创建一个能够生成短片分镜和视频的 Managed Agent。” |
|
||||
| 图片和视频生成 | “生成一张穿着太空服的猫站在火星上的图片,再把它制作成一段视频。” |
|
||||
| 语音识别 | “把这段音频转写成文字,专有名词识别不准的话帮我加上热词再试。” |
|
||||
| 用量与额度 | “查看最近的模型用量、免费额度和限流情况。” |
|
||||
| 模型选型 | “推荐一个适合图片理解和智能客服的模型。” |
|
||||
| 了解 Bailian CLI | “介绍一下 Bailian CLI 能帮我完成哪些任务,并根据我的需求推荐使用方式。” |
|
||||
|
||||
@@ -115,6 +115,7 @@ Once installed, just describe your task to your AI Agent — no need to assemble
|
||||
| ------------------------ | --------------------------------------------------------------------------------- |
|
||||
| Managed Agent | "Create a Managed Agent that can generate short-film storyboards and videos." |
|
||||
| Image & video generation | "Generate an image of a cat in a spacesuit on Mars, then turn it into a video." |
|
||||
| Speech recognition | "Transcribe this audio; if proper nouns are wrong, add hot words and try again." |
|
||||
| Usage & quota | "Show my recent model usage, free-tier quota, and rate limits." |
|
||||
| Model selection | "Recommend a model for image understanding and customer support." |
|
||||
| About Bailian CLI | "Tell me what Bailian CLI can do for me, and suggest how to use it for my needs." |
|
||||
|
||||
@@ -114,6 +114,7 @@ irm https://bailian.aliyun.com/cli/install.ps1 | iex
|
||||
| ---------------- | ----------------------------------------------------------------------- |
|
||||
| Managed Agent | “帮我创建一个能够生成短片分镜和视频的 Managed Agent。” |
|
||||
| 图片和视频生成 | “生成一张穿着太空服的猫站在火星上的图片,再把它制作成一段视频。” |
|
||||
| 语音识别 | “把这段音频转写成文字,专有名词识别不准的话帮我加上热词再试。” |
|
||||
| 用量与额度 | “查看最近的模型用量、免费额度和限流情况。” |
|
||||
| 模型选型 | “推荐一个适合图片理解和智能客服的模型。” |
|
||||
| 了解 Bailian CLI | “介绍一下 Bailian CLI 能帮我完成哪些任务,并根据我的需求推荐使用方式。” |
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "bailian-cli",
|
||||
"version": "1.22.0",
|
||||
"version": "1.23.0",
|
||||
"description": "CLI for Aliyun Model Studio (DashScope) AI Platform.",
|
||||
"keywords": [
|
||||
"agent",
|
||||
|
||||
@@ -70,6 +70,11 @@ import {
|
||||
searchWeb,
|
||||
speechSynthesize,
|
||||
speechRecognize,
|
||||
speechVocabularyCreate,
|
||||
speechVocabularyList,
|
||||
speechVocabularyGet,
|
||||
speechVocabularyUpdate,
|
||||
speechVocabularyDelete,
|
||||
fileUpload,
|
||||
consoleCall,
|
||||
usageFree,
|
||||
@@ -286,6 +291,11 @@ export const commands: Record<string, AnyCommand> = {
|
||||
"search web": searchWeb,
|
||||
"speech synthesize": speechSynthesize,
|
||||
"speech recognize": speechRecognize,
|
||||
"speech vocabulary create": speechVocabularyCreate,
|
||||
"speech vocabulary list": speechVocabularyList,
|
||||
"speech vocabulary get": speechVocabularyGet,
|
||||
"speech vocabulary update": speechVocabularyUpdate,
|
||||
"speech vocabulary delete": speechVocabularyDelete,
|
||||
"file upload": fileUpload,
|
||||
"console call": consoleCall,
|
||||
"usage free": usageFree,
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "bailian-cli-commands",
|
||||
"version": "1.22.0",
|
||||
"version": "1.23.0",
|
||||
"description": "Command library for bailian-cli products (knowledge, memory, media, …). See https://www.npmjs.com/package/bailian-cli for usage.",
|
||||
"homepage": "https://bailian.console.aliyun.com/cli",
|
||||
"bugs": {
|
||||
|
||||
@@ -12,9 +12,9 @@ export default defineCommand({
|
||||
valueHint: "<key>",
|
||||
description: {
|
||||
"en-US":
|
||||
"Config key (language, base_url, output, output_dir, timeout, api_key, api_key_capabilities, access_token, access_key_id, access_key_secret, security_token, default_*_model, workspace_id)",
|
||||
"Config key (language, base_url, output, output_dir, timeout, watermark, api_key, api_key_capabilities, access_token, access_key_id, access_key_secret, security_token, default_*_model, workspace_id)",
|
||||
"zh-CN":
|
||||
"配置项名称(language、base_url、output、output_dir、timeout、api_key、api_key_capabilities、access_token、access_key_id、access_key_secret、security_token、default_*_model、workspace_id)",
|
||||
"配置项名称(language、base_url、output、output_dir、timeout、watermark、api_key、api_key_capabilities、access_token、access_key_id、access_key_secret、security_token、default_*_model、workspace_id)",
|
||||
},
|
||||
required: true,
|
||||
},
|
||||
@@ -29,6 +29,7 @@ export default defineCommand({
|
||||
"--key language --value zh-CN",
|
||||
"--key output --value json",
|
||||
"--key timeout --value 600",
|
||||
"--key watermark --value false",
|
||||
"--key base_url --value https://dashscope.aliyuncs.com",
|
||||
"--config company-plan --key api-key-capabilities --value text.chat,image.generate",
|
||||
],
|
||||
|
||||
@@ -3,6 +3,7 @@ import {
|
||||
ExitCode,
|
||||
isApiKeyCapability,
|
||||
normalizeModelBaseUrl,
|
||||
parseBooleanValue,
|
||||
SUPPORTED_LANGUAGES,
|
||||
} from "bailian-cli-core";
|
||||
|
||||
@@ -13,6 +14,7 @@ export const VALID_KEYS = [
|
||||
"output",
|
||||
"output_dir",
|
||||
"timeout",
|
||||
"watermark",
|
||||
"api_key",
|
||||
"access_token",
|
||||
"access_key_id",
|
||||
@@ -62,7 +64,7 @@ export const UI_ENUM_KEYS: Record<string, string[]> = {
|
||||
};
|
||||
|
||||
// Keys the UI renders as a true/false dropdown and stores as a boolean.
|
||||
export const UI_BOOLEAN_KEYS = new Set<string>(["telemetry"]);
|
||||
export const UI_BOOLEAN_KEYS = new Set<string>(["telemetry", "watermark"]);
|
||||
|
||||
// Default model each `default_*_model` key falls back to when left unset. These
|
||||
// mirror the inline `|| "<model>"` fallbacks in the generation commands
|
||||
@@ -164,7 +166,10 @@ export function resolveKey(key: string): string {
|
||||
* Validate a single config entry and coerce its value to the stored type.
|
||||
* Throws BailianError(USAGE) for unknown keys or invalid values.
|
||||
*/
|
||||
export function validateAndCoerce(key: string, value: string): string | number | string[] {
|
||||
export function validateAndCoerce(
|
||||
key: string,
|
||||
value: string,
|
||||
): string | number | boolean | string[] {
|
||||
const resolvedKey = resolveKey(key);
|
||||
|
||||
if (!(VALID_KEYS as readonly string[]).includes(resolvedKey)) {
|
||||
@@ -201,6 +206,8 @@ export function validateAndCoerce(key: string, value: string): string | number |
|
||||
|
||||
if (resolvedKey === "base_url") return normalizeModelBaseUrl(value);
|
||||
|
||||
if (resolvedKey === "watermark") return parseBooleanValue(value, "watermark");
|
||||
|
||||
if (resolvedKey === "api_key_capabilities") {
|
||||
let rawCapabilities: unknown;
|
||||
if (value.trim().startsWith("[")) {
|
||||
|
||||
@@ -18,6 +18,7 @@ export default defineCommand({
|
||||
base_url: client.baseUrl,
|
||||
output: settings.output,
|
||||
timeout: settings.timeout,
|
||||
watermark: settings.watermark,
|
||||
config: settings.configName ?? "default",
|
||||
config_file: store.path,
|
||||
};
|
||||
|
||||
@@ -207,7 +207,7 @@ export default defineCommand({
|
||||
"prompt-extend",
|
||||
);
|
||||
|
||||
const watermark = resolveWatermark(flags.watermark);
|
||||
const watermark = resolveWatermark(flags.watermark, settings.watermark);
|
||||
|
||||
const parameters: NonNullable<DashScopeImageRequest["parameters"]> = {
|
||||
size: resolveImageSize(flags.size, route.sizeProfile),
|
||||
|
||||
@@ -185,7 +185,7 @@ export default defineCommand({
|
||||
"prompt-extend",
|
||||
);
|
||||
|
||||
const watermark = resolveWatermark(flags.watermark);
|
||||
const watermark = resolveWatermark(flags.watermark, settings.watermark);
|
||||
|
||||
const parameters: NonNullable<DashScopeImageRequest["parameters"]> = {
|
||||
size,
|
||||
|
||||
@@ -14,9 +14,11 @@ import {
|
||||
speechRecognizePath,
|
||||
resolveAsrApi,
|
||||
buildAsrFlashRequest,
|
||||
buildAsrContextMessages,
|
||||
buildAsyncAsrLanguageFields,
|
||||
collectAsrTranscriptionItems,
|
||||
extractAsrFlashText,
|
||||
parseInstantVocabulary,
|
||||
type AsrApiRoute,
|
||||
type AsrFlashFamily,
|
||||
type OutputFormat,
|
||||
@@ -73,8 +75,30 @@ const RECOGNIZE_FLAGS = {
|
||||
type: "string",
|
||||
valueHint: "<id>",
|
||||
description: {
|
||||
"en-US": "Hot-word vocabulary ID for improved accuracy",
|
||||
"zh-CN": "用于提升识别准确率的热词表 ID",
|
||||
"en-US":
|
||||
"Pre-built hot-word vocabulary ID (create it via `speech vocabulary create`). Its target_model must exactly match --model, otherwise it is silently ignored. Wider model support than --vocabulary, including Fun-ASR and Paraformer",
|
||||
"zh-CN":
|
||||
"预编译热词列表 ID(可用 `speech vocabulary create` 创建)。其 target_model 必须与 --model 完全一致,否则静默失效且无报错。支持模型比 --vocabulary 更广,含 Fun-ASR 与 Paraformer 系列",
|
||||
},
|
||||
},
|
||||
vocabulary: {
|
||||
type: "string",
|
||||
valueHint: "<json>",
|
||||
description: {
|
||||
"en-US":
|
||||
"Instant hot words as JSON object of word→weight, e.g. '{\"Fendouzhe\":4}'. Weight 1-5 (4 recommended; higher values can hurt other words), 50 for super hot word. No pre-built vocabulary needed. Takes effect only on Qwen-Audio-3.0-ASR-Flash models",
|
||||
"zh-CN":
|
||||
"即时热词,JSON 对象「热词→权重」,例如 '{\"奋斗者\":4}'。权重 1-5(推荐 4,过高会拖累其他词),50 表示超级热词。无需预先创建热词表。仅 Qwen-Audio-3.0-ASR-Flash 系列模型生效",
|
||||
},
|
||||
},
|
||||
context: {
|
||||
type: "string",
|
||||
valueHint: "<text>",
|
||||
description: {
|
||||
"en-US":
|
||||
"Context enhancement word list to improve accuracy on proper nouns; must contain the target words themselves (a topic description alone has little effect); max 400 chars. Takes effect only on Qwen-Audio-3.0-ASR-Flash and Fun-ASR-Flash models",
|
||||
"zh-CN":
|
||||
"上下文增强词表,提升专有名词准确率;须包含待识别的原词本身(只写主题描述效果有限),最长 400 字符。仅 Qwen-Audio-3.0-ASR-Flash 系列与 Fun-ASR-Flash 模型生效",
|
||||
},
|
||||
},
|
||||
channelId: {
|
||||
@@ -110,9 +134,11 @@ function assertSyncFlashFlagsAllowed(
|
||||
const unsupported: string[] = [];
|
||||
if (flags.diarization === true) unsupported.push("--diarization");
|
||||
if (flags.speakerCount !== undefined) unsupported.push("--speaker-count");
|
||||
// qwen3 sync Flash does not use vocabulary_id; input-audio Flash (fun-asr-flash* / qwen-audio-*-asr-flash) does
|
||||
if (flashFamily === "qwen3" && flags.vocabularyId !== undefined) {
|
||||
unsupported.push("--vocabulary-id");
|
||||
// qwen3 sync Flash has no place for vocabulary_id / vocabulary / context in its body shape
|
||||
if (flashFamily === "qwen3") {
|
||||
if (flags.vocabularyId !== undefined) unsupported.push("--vocabulary-id");
|
||||
if (flags.vocabulary !== undefined) unsupported.push("--vocabulary");
|
||||
if (flags.context !== undefined) unsupported.push("--context");
|
||||
}
|
||||
if (flags.channelId !== undefined) unsupported.push("--channel-id");
|
||||
if (flags.async === true) unsupported.push("--async");
|
||||
@@ -121,12 +147,34 @@ function assertSyncFlashFlagsAllowed(
|
||||
if (unsupported.length > 0) {
|
||||
throw new BailianError(
|
||||
`Model "${model}" uses sync Flash ASR and does not support: ${unsupported.join(", ")}.\n` +
|
||||
`Hint: Use an async filetrans model (e.g. fun-asr, qwen3-asr-flash-filetrans) for those flags.`,
|
||||
syncFlashUnsupportedHint(unsupported),
|
||||
ExitCode.USAGE,
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/** Pick a hint that matches the rejected flags (vocab/context vs diarization/async/…). */
|
||||
function syncFlashUnsupportedHint(unsupported: string[]): string {
|
||||
const vocabularyRelated = new Set(["--vocabulary", "--vocabulary-id", "--context"]);
|
||||
const hasVocabularyRelated = unsupported.some((flag) => vocabularyRelated.has(flag));
|
||||
const hasOtherFlags = unsupported.some((flag) => !vocabularyRelated.has(flag));
|
||||
|
||||
if (hasVocabularyRelated && !hasOtherFlags) {
|
||||
return (
|
||||
"Hint: Use qwen-audio-3.0-asr-flash (or an async filetrans model such as " +
|
||||
"qwen-audio-3.0-asr-flash-filetrans) for vocabulary/context flags."
|
||||
);
|
||||
}
|
||||
if (hasVocabularyRelated && hasOtherFlags) {
|
||||
return (
|
||||
"Hint: For vocabulary/context flags use qwen-audio-3.0-asr-flash or " +
|
||||
"qwen-audio-3.0-asr-flash-filetrans; for the other flags use an async filetrans model " +
|
||||
"(e.g. fun-asr)."
|
||||
);
|
||||
}
|
||||
return "Hint: Use an async filetrans model (e.g. fun-asr, qwen3-asr-flash-filetrans) for those flags.";
|
||||
}
|
||||
|
||||
export default defineCommand({
|
||||
description: {
|
||||
"en-US": "Recognize speech from audio files (FunAudio-ASR / Qwen-ASR Flash)",
|
||||
@@ -141,6 +189,8 @@ export default defineCommand({
|
||||
"--url https://example.com/meeting.wav --diarization --speaker-count 3",
|
||||
"--url https://example.com/audio.mp3 --language zh",
|
||||
"--url https://example.com/audio.mp3 --vocabulary-id vocab-abc123",
|
||||
'--url https://example.com/audio.mp3 --model qwen-audio-3.0-asr-flash-filetrans --vocabulary \'{"奋斗者":4,"鲸落":4}\'',
|
||||
'--url https://example.com/audio.mp3 --model qwen-audio-3.0-asr-flash-filetrans --context "奋斗者号 鲸落 深海勇士"',
|
||||
"--url https://example.com/audio.mp3 --out result.json",
|
||||
"--url https://example.com/audio.mp3 --async --quiet",
|
||||
"--url https://example.com/audio.mp3 --model qwen-audio-3.0-asr-flash --language en",
|
||||
@@ -198,6 +248,9 @@ export default defineCommand({
|
||||
|
||||
const format = detectOutputFormat(settings.output);
|
||||
|
||||
const vocabulary =
|
||||
flags.vocabulary !== undefined ? parseInstantVocabulary(flags.vocabulary) : undefined;
|
||||
|
||||
// Auto-upload local files in parallel
|
||||
const resolvedUrls = await Promise.all(rawUrls.map((url) => ctx.client.uploadFile(url, model)));
|
||||
|
||||
@@ -210,6 +263,7 @@ export default defineCommand({
|
||||
model,
|
||||
route,
|
||||
resolvedUrls[0]!,
|
||||
vocabulary,
|
||||
);
|
||||
return;
|
||||
}
|
||||
@@ -223,16 +277,19 @@ export default defineCommand({
|
||||
|
||||
const body: DashScopeASRRequest = {
|
||||
model,
|
||||
input:
|
||||
route.asyncInputStyle === "file_url"
|
||||
input: {
|
||||
...(route.asyncInputStyle === "file_url"
|
||||
? { file_url: resolvedUrls[0]! }
|
||||
: { file_urls: resolvedUrls },
|
||||
: { file_urls: resolvedUrls }),
|
||||
...(flags.context !== undefined ? { context: buildAsrContextMessages(flags.context) } : {}),
|
||||
},
|
||||
parameters: {
|
||||
channel_id: channelId !== undefined ? [channelId] : [0],
|
||||
...languageFields,
|
||||
diarization_enabled: diarization ? true : undefined,
|
||||
speaker_count: speakerCount,
|
||||
vocabulary_id: vocabularyId,
|
||||
vocabulary,
|
||||
},
|
||||
};
|
||||
|
||||
@@ -260,6 +317,7 @@ async function handleSyncFlashMode(
|
||||
model: string,
|
||||
route: AsrApiRoute,
|
||||
audioUrl: string,
|
||||
vocabulary: Record<string, number> | undefined,
|
||||
): Promise<void> {
|
||||
const flashFamily = route.flashFamily as AsrFlashFamily;
|
||||
const body = buildAsrFlashRequest({
|
||||
@@ -267,6 +325,8 @@ async function handleSyncFlashMode(
|
||||
audioUrl,
|
||||
language: flags.language,
|
||||
vocabularyId: flags.vocabularyId,
|
||||
vocabulary,
|
||||
context: flags.context,
|
||||
flashFamily,
|
||||
});
|
||||
|
||||
|
||||
@@ -0,0 +1,117 @@
|
||||
import {
|
||||
defineCommand,
|
||||
detectOutputFormat,
|
||||
speechVocabularyPath,
|
||||
buildVocabularyRequest,
|
||||
createVocabulary,
|
||||
type FlagsDef,
|
||||
type ParsedFlags,
|
||||
} from "bailian-cli-core";
|
||||
import { emitResult, emitBare } from "bailian-cli-runtime";
|
||||
import {
|
||||
VOCABULARY_BODY_FLAGS,
|
||||
VOCABULARY_LIMIT_NOTES,
|
||||
validateVocabularySource,
|
||||
readVocabularyEntries,
|
||||
} from "./shared.ts";
|
||||
|
||||
const CREATE_FLAGS = {
|
||||
model: {
|
||||
type: "string",
|
||||
valueHint: "<model>",
|
||||
description: {
|
||||
"en-US":
|
||||
"ASR model this vocabulary is built for (required). Must exactly match the --model passed to `speech recognize` later, otherwise the vocabulary is silently ignored",
|
||||
"zh-CN":
|
||||
"该热词表服务的 ASR 模型(必填)。必须与后续 `speech recognize` 的 --model 完全一致,否则热词表静默失效",
|
||||
},
|
||||
required: true,
|
||||
},
|
||||
prefix: {
|
||||
type: "string",
|
||||
valueHint: "<prefix>",
|
||||
description: {
|
||||
"en-US":
|
||||
"Custom vocabulary prefix (required). Digits and lowercase letters only, max 10 chars",
|
||||
"zh-CN": "热词表自定义前缀(必填)。仅允许数字和小写字母,最长 10 个字符",
|
||||
},
|
||||
required: true,
|
||||
},
|
||||
...VOCABULARY_BODY_FLAGS,
|
||||
} satisfies FlagsDef;
|
||||
type CreateFlags = ParsedFlags<typeof CREATE_FLAGS>;
|
||||
|
||||
export default defineCommand({
|
||||
description: {
|
||||
"en-US": "Create a precompiled hot-word vocabulary for ASR",
|
||||
"zh-CN": "创建用于语音识别的预编译热词表",
|
||||
},
|
||||
auth: "apiKey",
|
||||
usageArgs: "--model <model> --prefix <prefix> (--words <json> | --words-file <path>) [flags]",
|
||||
flags: CREATE_FLAGS,
|
||||
notes: [
|
||||
{
|
||||
"en-US":
|
||||
"The --model must exactly match the --model used later with `speech recognize --vocabulary-id`; a mismatch causes silent failure with no error.",
|
||||
"zh-CN":
|
||||
"--model 必须与后续 `speech recognize --vocabulary-id` 使用的 --model 完全一致;不一致时热词表会静默失效且无报错。",
|
||||
},
|
||||
...VOCABULARY_LIMIT_NOTES,
|
||||
],
|
||||
exampleArgs: [
|
||||
{
|
||||
"en-US": '--model fun-asr --prefix demo --words \'{"Fendouzhe":4,"Jingluo":4}\'',
|
||||
"zh-CN": '--model fun-asr --prefix demo --words \'{"奋斗者":4,"鲸落":4}\'',
|
||||
},
|
||||
{
|
||||
"en-US":
|
||||
'--model paraformer-v2 --prefix demo --words \'[{"text":"Fendouzhe","weight":4,"lang":"zh"}]\'',
|
||||
"zh-CN":
|
||||
'--model paraformer-v2 --prefix demo --words \'[{"text":"奋斗者","weight":4,"lang":"zh"}]\'',
|
||||
},
|
||||
{
|
||||
"en-US": "--model fun-asr --prefix demo --words '{\"Fendouzhe\":4}' --lang zh",
|
||||
"zh-CN": "--model fun-asr --prefix demo --words '{\"奋斗者\":4}' --lang zh",
|
||||
},
|
||||
"--model fun-asr --prefix demo --words-file ./hotwords.json",
|
||||
{
|
||||
"en-US": "--model fun-asr --prefix demo --words '{\"Fendouzhe\":4}' --quiet",
|
||||
"zh-CN": "--model fun-asr --prefix demo --words '{\"奋斗者\":4}' --quiet",
|
||||
},
|
||||
],
|
||||
validate: (flags: CreateFlags) => validateVocabularySource(flags),
|
||||
async run(ctx) {
|
||||
const { settings, flags } = ctx;
|
||||
const vocabulary = readVocabularyEntries(flags);
|
||||
const format = detectOutputFormat(settings.output);
|
||||
|
||||
const request = buildVocabularyRequest("create_vocabulary", {
|
||||
target_model: flags.model,
|
||||
prefix: flags.prefix,
|
||||
vocabulary,
|
||||
});
|
||||
|
||||
if (settings.dryRun) {
|
||||
emitResult(
|
||||
{
|
||||
endpoint: ctx.client.url(speechVocabularyPath()),
|
||||
request,
|
||||
},
|
||||
format,
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
const response = await createVocabulary(ctx.client, {
|
||||
targetModel: flags.model,
|
||||
prefix: flags.prefix,
|
||||
vocabulary,
|
||||
});
|
||||
|
||||
if (settings.quiet || format === "text") {
|
||||
emitBare(response.output?.vocabulary_id ?? "");
|
||||
} else {
|
||||
emitResult(response, format);
|
||||
}
|
||||
},
|
||||
});
|
||||
@@ -0,0 +1,60 @@
|
||||
import {
|
||||
defineCommand,
|
||||
detectOutputFormat,
|
||||
speechVocabularyPath,
|
||||
buildVocabularyRequest,
|
||||
deleteVocabulary,
|
||||
type FlagsDef,
|
||||
type ParsedFlags,
|
||||
} from "bailian-cli-core";
|
||||
import { emitResult, emitBare } from "bailian-cli-runtime";
|
||||
import { VOCABULARY_ID_FLAG } from "./shared.ts";
|
||||
|
||||
const DELETE_FLAGS = {
|
||||
...VOCABULARY_ID_FLAG,
|
||||
} satisfies FlagsDef;
|
||||
type DeleteFlags = ParsedFlags<typeof DELETE_FLAGS>;
|
||||
|
||||
export default defineCommand({
|
||||
description: {
|
||||
"en-US": "Delete a precompiled hot-word vocabulary",
|
||||
"zh-CN": "删除预编译热词表",
|
||||
},
|
||||
auth: "apiKey",
|
||||
risk: {
|
||||
level: "high",
|
||||
message: {
|
||||
"en-US": "This permanently deletes the specified hot-word vocabulary and cannot be undone.",
|
||||
"zh-CN": "该操作会永久删除指定的热词表,且无法撤销。",
|
||||
},
|
||||
},
|
||||
usageArgs: "--id <id>",
|
||||
flags: DELETE_FLAGS,
|
||||
exampleArgs: ["--id vocab-demo-xxx --dry-run", "--id vocab-demo-xxx --yes"],
|
||||
async run(ctx) {
|
||||
const { settings, flags } = ctx;
|
||||
const vocabularyId = (flags as DeleteFlags).id;
|
||||
const format = detectOutputFormat(settings.output);
|
||||
|
||||
if (settings.dryRun) {
|
||||
emitResult(
|
||||
{
|
||||
endpoint: ctx.client.url(speechVocabularyPath()),
|
||||
request: buildVocabularyRequest("delete_vocabulary", {
|
||||
vocabulary_id: vocabularyId,
|
||||
}),
|
||||
},
|
||||
format,
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
const response = await deleteVocabulary(ctx.client, vocabularyId);
|
||||
|
||||
if (settings.quiet || format === "text") {
|
||||
emitBare(vocabularyId);
|
||||
} else {
|
||||
emitResult(response, format);
|
||||
}
|
||||
},
|
||||
});
|
||||
@@ -0,0 +1,79 @@
|
||||
import {
|
||||
defineCommand,
|
||||
detectOutputFormat,
|
||||
speechVocabularyPath,
|
||||
buildVocabularyRequest,
|
||||
queryVocabulary,
|
||||
type FlagsDef,
|
||||
type ParsedFlags,
|
||||
} from "bailian-cli-core";
|
||||
import { emitResult, emitBare } from "bailian-cli-runtime";
|
||||
import { VOCABULARY_ID_FLAG } from "./shared.ts";
|
||||
|
||||
const GET_FLAGS = {
|
||||
...VOCABULARY_ID_FLAG,
|
||||
} satisfies FlagsDef;
|
||||
type GetFlags = ParsedFlags<typeof GET_FLAGS>;
|
||||
|
||||
export default defineCommand({
|
||||
description: {
|
||||
"en-US": "Get details of a precompiled hot-word vocabulary",
|
||||
"zh-CN": "查看预编译热词表详情",
|
||||
},
|
||||
auth: "apiKey",
|
||||
usageArgs: "--id <id>",
|
||||
flags: GET_FLAGS,
|
||||
notes: [
|
||||
{
|
||||
"en-US":
|
||||
"Use this command to confirm target_model before calling `speech recognize --vocabulary-id`; a model mismatch causes silent failure.",
|
||||
"zh-CN":
|
||||
"调用 `speech recognize --vocabulary-id` 前请用本命令确认 target_model;模型不一致会导致静默失效。",
|
||||
},
|
||||
],
|
||||
exampleArgs: ["--id vocab-demo-xxx", "--id vocab-demo-xxx --quiet"],
|
||||
async run(ctx) {
|
||||
const { settings, flags } = ctx;
|
||||
const vocabularyId = (flags as GetFlags).id;
|
||||
const format = detectOutputFormat(settings.output);
|
||||
|
||||
if (settings.dryRun) {
|
||||
emitResult(
|
||||
{
|
||||
endpoint: ctx.client.url(speechVocabularyPath()),
|
||||
request: buildVocabularyRequest("query_vocabulary", {
|
||||
vocabulary_id: vocabularyId,
|
||||
}),
|
||||
},
|
||||
format,
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
const response = await queryVocabulary(ctx.client, vocabularyId);
|
||||
const output = response.output;
|
||||
|
||||
if (settings.quiet) {
|
||||
emitBare(output?.target_model ?? "");
|
||||
return;
|
||||
}
|
||||
|
||||
if (format === "text") {
|
||||
emitBare(`vocabulary_id: ${vocabularyId}`);
|
||||
emitBare(`status: ${output?.status ?? ""}`);
|
||||
emitBare(`target_model: ${output?.target_model ?? ""}`);
|
||||
emitBare(`gmt_create: ${output?.gmt_create ?? ""}`);
|
||||
emitBare(`gmt_modified: ${output?.gmt_modified ?? ""}`);
|
||||
const entries = output?.vocabulary ?? [];
|
||||
if (entries.length > 0) {
|
||||
emitBare("vocabulary:");
|
||||
for (const entry of entries) {
|
||||
const langPart = entry.lang ? ` lang=${entry.lang}` : "";
|
||||
emitBare(` ${entry.text} weight=${entry.weight}${langPart}`);
|
||||
}
|
||||
}
|
||||
} else {
|
||||
emitResult(response, format);
|
||||
}
|
||||
},
|
||||
});
|
||||
@@ -0,0 +1,116 @@
|
||||
import {
|
||||
defineCommand,
|
||||
detectOutputFormat,
|
||||
speechVocabularyPath,
|
||||
buildVocabularyRequest,
|
||||
listVocabularies,
|
||||
type FlagsDef,
|
||||
type ParsedFlags,
|
||||
} from "bailian-cli-core";
|
||||
import { emitResult, emitBare } from "bailian-cli-runtime";
|
||||
|
||||
const LIST_FLAGS = {
|
||||
prefix: {
|
||||
type: "string",
|
||||
valueHint: "<prefix>",
|
||||
description: {
|
||||
"en-US": "Filter by vocabulary prefix",
|
||||
"zh-CN": "按热词表前缀过滤",
|
||||
},
|
||||
},
|
||||
page: {
|
||||
type: "number",
|
||||
valueHint: "<n>",
|
||||
description: {
|
||||
"en-US": "Page number, 1-based (default: 1). Mapped to API page_index (0-based) as page - 1",
|
||||
"zh-CN": "页码,从 1 开始(默认:1)。映射为 API 的 page_index(从 0 开始):page - 1",
|
||||
},
|
||||
},
|
||||
pageSize: {
|
||||
type: "number",
|
||||
valueHint: "<n>",
|
||||
description: {
|
||||
"en-US": "Results per page (default: 10)",
|
||||
"zh-CN": "每页结果数(默认:10)",
|
||||
},
|
||||
},
|
||||
} satisfies FlagsDef;
|
||||
type ListFlags = ParsedFlags<typeof LIST_FLAGS>;
|
||||
|
||||
export default defineCommand({
|
||||
description: {
|
||||
"en-US": "List precompiled hot-word vocabularies",
|
||||
"zh-CN": "列出预编译热词表",
|
||||
},
|
||||
auth: "apiKey",
|
||||
usageArgs: "[--prefix <prefix>] [--page <n>] [--page-size <n>]",
|
||||
flags: LIST_FLAGS,
|
||||
notes: [
|
||||
{
|
||||
"en-US":
|
||||
"List responses do not include target_model; use `speech vocabulary get` to inspect the model a vocabulary was built for.",
|
||||
"zh-CN": "list 响应不含 target_model;要对齐模型请使用 `speech vocabulary get`。",
|
||||
},
|
||||
{
|
||||
"en-US": "Vocabularies with status UNDEPLOYED are silently ignored by ASR.",
|
||||
"zh-CN": "status 为 UNDEPLOYED 的热词表会被 ASR 静默忽略。",
|
||||
},
|
||||
],
|
||||
exampleArgs: ["", "--prefix demo", "--page 2 --page-size 20"],
|
||||
validate: (flags: ListFlags) => {
|
||||
if (flags.page !== undefined && flags.page < 1) {
|
||||
return "--page must be >= 1.";
|
||||
}
|
||||
return undefined;
|
||||
},
|
||||
async run(ctx) {
|
||||
const { settings, flags } = ctx;
|
||||
const format = detectOutputFormat(settings.output);
|
||||
|
||||
const pageIndex = flags.page !== undefined ? flags.page - 1 : undefined;
|
||||
const input: Record<string, unknown> = {};
|
||||
if (flags.prefix !== undefined) input.prefix = flags.prefix;
|
||||
if (pageIndex !== undefined) input.page_index = pageIndex;
|
||||
if (flags.pageSize !== undefined) input.page_size = flags.pageSize;
|
||||
|
||||
if (settings.dryRun) {
|
||||
emitResult(
|
||||
{
|
||||
endpoint: ctx.client.url(speechVocabularyPath()),
|
||||
request: buildVocabularyRequest("list_vocabulary", input),
|
||||
},
|
||||
format,
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
const response = await listVocabularies(ctx.client, {
|
||||
prefix: flags.prefix,
|
||||
pageIndex,
|
||||
pageSize: flags.pageSize,
|
||||
});
|
||||
|
||||
if (settings.quiet || format === "text") {
|
||||
const items = response.output?.vocabulary_list ?? [];
|
||||
if (items.length === 0) {
|
||||
emitBare("No vocabularies found.");
|
||||
} else {
|
||||
let hasUndeployed = false;
|
||||
for (const item of items) {
|
||||
const id = item.vocabulary_id ?? "";
|
||||
const status = item.status ?? "";
|
||||
const modified = item.gmt_modified ?? "";
|
||||
if (status === "UNDEPLOYED") hasUndeployed = true;
|
||||
emitBare(`[${id}] ${status} ${modified}`.trimEnd());
|
||||
}
|
||||
if (hasUndeployed) {
|
||||
emitBare(
|
||||
"Note: UNDEPLOYED vocabularies are silently ignored by ASR. Use `speech vocabulary get` to inspect them.",
|
||||
);
|
||||
}
|
||||
}
|
||||
} else {
|
||||
emitResult(response, format);
|
||||
}
|
||||
},
|
||||
});
|
||||
@@ -0,0 +1,103 @@
|
||||
import {
|
||||
readTextFromPathOrStdin,
|
||||
parseVocabularyEntries,
|
||||
UsageError,
|
||||
type FlagsDef,
|
||||
type ParsedFlags,
|
||||
type VocabularyEntry,
|
||||
} from "bailian-cli-core";
|
||||
|
||||
/** Shared --id flag for get / update / delete. */
|
||||
export const VOCABULARY_ID_FLAG = {
|
||||
id: {
|
||||
type: "string",
|
||||
valueHint: "<id>",
|
||||
description: {
|
||||
"en-US": "Hot-word vocabulary ID (required)",
|
||||
"zh-CN": "热词表 ID(必填)",
|
||||
},
|
||||
required: true,
|
||||
},
|
||||
} satisfies FlagsDef;
|
||||
|
||||
/** Shared hot-word body flags for create / update. */
|
||||
export const VOCABULARY_BODY_FLAGS = {
|
||||
words: {
|
||||
type: "string",
|
||||
valueHint: "<json>",
|
||||
description: {
|
||||
"en-US":
|
||||
"Hot words as JSON object of word→weight, e.g. '{\"Fendouzhe\":4}'; or the API entry array for per-entry lang. Weight 1-5 (4 recommended); when the vocabulary target_model is a Qwen-Audio-3.0-ASR-Flash series model, 50 is also allowed as super hot word. Or use --words-file",
|
||||
"zh-CN":
|
||||
"热词,JSON 对象「热词→权重」,例如 '{\"奋斗者\":4}';需要逐条指定语言时可传 API 的条目数组。权重 1-5(推荐 4);当热词表的 target_model 为 Qwen-Audio-3.0-ASR-Flash 系列时还可使用 50(超级热词)。也可使用 --words-file",
|
||||
},
|
||||
},
|
||||
wordsFile: {
|
||||
type: "string",
|
||||
valueHint: "<path>",
|
||||
description: {
|
||||
"en-US": "JSON file with the hot words (use - for stdin)",
|
||||
"zh-CN": "包含热词的 JSON 文件(使用 - 从 stdin 读取)",
|
||||
},
|
||||
},
|
||||
lang: {
|
||||
type: "string",
|
||||
valueHint: "<code>",
|
||||
description: {
|
||||
"en-US":
|
||||
"Language code applied to every hot word when using object form (optional; ignored for array form). Paraformer: zh/en/ja/yue/ko/de/fr/ru; Fun-ASR: zh/en/ja",
|
||||
"zh-CN":
|
||||
"对象形态时应用到所有热词的语言代码(选填;数组形态忽略)。Paraformer 支持 zh/en/ja/yue/ko/de/fr/ru;Fun-ASR 支持 zh/en/ja",
|
||||
},
|
||||
},
|
||||
} satisfies FlagsDef;
|
||||
|
||||
type VocabularySourceFlags = ParsedFlags<typeof VOCABULARY_BODY_FLAGS>;
|
||||
|
||||
/** Cross-flag validation for --words / --words-file. */
|
||||
export function validateVocabularySource(flags: VocabularySourceFlags): string | undefined {
|
||||
if (!flags.words && !flags.wordsFile) {
|
||||
return "Provide --words or --words-file.";
|
||||
}
|
||||
if (flags.words && flags.wordsFile) {
|
||||
return "Use either --words or --words-file, not both.";
|
||||
}
|
||||
if (!flags.words) {
|
||||
return undefined;
|
||||
}
|
||||
try {
|
||||
parseVocabularyEntries(flags.words, flags.lang);
|
||||
} catch (error) {
|
||||
if (error instanceof UsageError) {
|
||||
return error.message;
|
||||
}
|
||||
throw error;
|
||||
}
|
||||
return undefined;
|
||||
}
|
||||
|
||||
/** Read and parse vocabulary entries from flag or file. */
|
||||
export function readVocabularyEntries(flags: VocabularySourceFlags): VocabularyEntry[] {
|
||||
const raw = flags.wordsFile ? readTextFromPathOrStdin(flags.wordsFile) : (flags.words as string);
|
||||
return parseVocabularyEntries(raw, flags.lang);
|
||||
}
|
||||
|
||||
/** Shared notes covering account limits and silent-failure pitfalls. */
|
||||
export const VOCABULARY_LIMIT_NOTES = [
|
||||
{
|
||||
"en-US":
|
||||
"Each account may have at most 10 vocabularies; updates must be at least 5 minutes apart. See improve-asr-accuracy for full limits.",
|
||||
"zh-CN": "每个账号最多 10 个热词表;两次更新间隔至少 5 分钟。完整限制见 improve-asr-accuracy。",
|
||||
},
|
||||
{
|
||||
"en-US":
|
||||
"Hot-word vocabularies are not supported in Singapore sub-workspaces; the server error is passed through as-is.",
|
||||
"zh-CN": "新加坡子业务空间不支持热词表;服务端错误会原样透传。",
|
||||
},
|
||||
{
|
||||
"en-US":
|
||||
"Weight 1-5 (4 recommended); when the vocabulary target_model is a Qwen-Audio-3.0-ASR-Flash series model, 50 is also allowed as super hot word.",
|
||||
"zh-CN":
|
||||
"权重 1-5(推荐 4);当热词表的 target_model 为 Qwen-Audio-3.0-ASR-Flash 系列时还可使用 50(超级热词)。",
|
||||
},
|
||||
] as const;
|
||||
@@ -0,0 +1,90 @@
|
||||
import {
|
||||
defineCommand,
|
||||
detectOutputFormat,
|
||||
speechVocabularyPath,
|
||||
buildVocabularyRequest,
|
||||
updateVocabulary,
|
||||
type FlagsDef,
|
||||
type ParsedFlags,
|
||||
} from "bailian-cli-core";
|
||||
import { emitResult, emitBare } from "bailian-cli-runtime";
|
||||
import {
|
||||
VOCABULARY_ID_FLAG,
|
||||
VOCABULARY_BODY_FLAGS,
|
||||
VOCABULARY_LIMIT_NOTES,
|
||||
validateVocabularySource,
|
||||
readVocabularyEntries,
|
||||
} from "./shared.ts";
|
||||
|
||||
const UPDATE_FLAGS = {
|
||||
...VOCABULARY_ID_FLAG,
|
||||
...VOCABULARY_BODY_FLAGS,
|
||||
} satisfies FlagsDef;
|
||||
type UpdateFlags = ParsedFlags<typeof UPDATE_FLAGS>;
|
||||
|
||||
export default defineCommand({
|
||||
description: {
|
||||
"en-US": "Replace the contents of a precompiled hot-word vocabulary",
|
||||
"zh-CN": "完全替换预编译热词表的内容",
|
||||
},
|
||||
auth: "apiKey",
|
||||
risk: {
|
||||
level: "high",
|
||||
message: {
|
||||
"en-US":
|
||||
"This fully replaces all hot words in the vocabulary. Entries not listed will be discarded and cannot be undone.",
|
||||
"zh-CN": "该操作会完全替换热词表中的全部词条。未列出的词将被丢弃,且无法撤销。",
|
||||
},
|
||||
},
|
||||
usageArgs: "--id <id> (--words <json> | --words-file <path>) [flags]",
|
||||
flags: UPDATE_FLAGS,
|
||||
notes: [
|
||||
{
|
||||
"en-US":
|
||||
"update is a full replace, not an append. Prefer --dry-run first to preview the complete vocabulary that will be written.",
|
||||
"zh-CN": "update 是完全替换,不是增量追加。建议先用 --dry-run 预览将要写入的完整词表。",
|
||||
},
|
||||
...VOCABULARY_LIMIT_NOTES,
|
||||
],
|
||||
exampleArgs: [
|
||||
{
|
||||
"en-US": "--id vocab-demo-xxx --words '{\"Fendouzhe\":4}' --dry-run",
|
||||
"zh-CN": "--id vocab-demo-xxx --words '{\"奋斗者\":4}' --dry-run",
|
||||
},
|
||||
{
|
||||
"en-US": '--id vocab-demo-xxx --words \'{"Fendouzhe":4,"Jingluo":4}\' --yes',
|
||||
"zh-CN": '--id vocab-demo-xxx --words \'{"奋斗者":4,"鲸落":4}\' --yes',
|
||||
},
|
||||
],
|
||||
validate: (flags: UpdateFlags) => validateVocabularySource(flags),
|
||||
async run(ctx) {
|
||||
const { settings, flags } = ctx;
|
||||
const vocabularyId = flags.id;
|
||||
const vocabulary = readVocabularyEntries(flags);
|
||||
const format = detectOutputFormat(settings.output);
|
||||
|
||||
const request = buildVocabularyRequest("update_vocabulary", {
|
||||
vocabulary_id: vocabularyId,
|
||||
vocabulary,
|
||||
});
|
||||
|
||||
if (settings.dryRun) {
|
||||
emitResult(
|
||||
{
|
||||
endpoint: ctx.client.url(speechVocabularyPath()),
|
||||
request,
|
||||
},
|
||||
format,
|
||||
);
|
||||
return;
|
||||
}
|
||||
|
||||
const response = await updateVocabulary(ctx.client, vocabularyId, vocabulary);
|
||||
|
||||
if (settings.quiet || format === "text") {
|
||||
emitBare(vocabularyId);
|
||||
} else {
|
||||
emitResult(response, format);
|
||||
}
|
||||
},
|
||||
});
|
||||
@@ -53,10 +53,15 @@ export default defineCommand({
|
||||
});
|
||||
|
||||
if (taskInfo.output.task_status !== "SUCCEEDED") {
|
||||
const status = taskInfo.output.task_status;
|
||||
const detail = taskInfo.output.message || taskInfo.output.code;
|
||||
const message = detail
|
||||
? `Task is not complete (status: ${status}): ${detail}`
|
||||
: `Task is not complete (status: ${status}).`;
|
||||
throw new BailianError(
|
||||
`Task is not complete (status: ${taskInfo.output.task_status}).`,
|
||||
message,
|
||||
ExitCode.GENERAL,
|
||||
"Wait for the task to complete before downloading.",
|
||||
status === "FAILED" ? undefined : "Wait for the task to complete before downloading.",
|
||||
);
|
||||
}
|
||||
|
||||
|
||||
@@ -197,7 +197,7 @@ export default defineCommand({
|
||||
|
||||
// --- Build request body ---
|
||||
const promptExtend = resolveBooleanFlag(flags.promptExtend, undefined, "prompt-extend");
|
||||
const watermark = resolveWatermark(flags.watermark);
|
||||
const watermark = resolveWatermark(flags.watermark, settings.watermark);
|
||||
|
||||
const body: DashScopeVideoEditRequest = {
|
||||
model,
|
||||
|
||||
@@ -208,7 +208,7 @@ export default defineCommand({
|
||||
resolvedFileUrl = await ctx.client.uploadFile(fileUrl, model);
|
||||
}
|
||||
|
||||
const watermark = resolveWatermark(flags.watermark);
|
||||
const watermark = resolveWatermark(flags.watermark, settings.watermark);
|
||||
const promptExtend = resolveBooleanFlag(flags.promptExtend, undefined, "prompt-extend");
|
||||
|
||||
const body: DashScopeVideoRequest = {
|
||||
|
||||
@@ -249,7 +249,7 @@ export default defineCommand({
|
||||
|
||||
// --- Build request body ---
|
||||
const promptExtend = resolveBooleanFlag(flags.promptExtend, undefined, "prompt-extend");
|
||||
const watermark = resolveWatermark(flags.watermark);
|
||||
const watermark = resolveWatermark(flags.watermark, settings.watermark);
|
||||
|
||||
const body: DashScopeVideoRefRequest = {
|
||||
model,
|
||||
|
||||
@@ -2,6 +2,7 @@ import {
|
||||
defineCommand,
|
||||
taskPath,
|
||||
detectOutputFormat,
|
||||
stripUndefined,
|
||||
type DashScopeTaskResponse,
|
||||
} from "bailian-cli-core";
|
||||
import { emitResult, emitBare } from "bailian-cli-runtime";
|
||||
@@ -42,15 +43,20 @@ export default defineCommand({
|
||||
return;
|
||||
}
|
||||
|
||||
// 透传服务端失败字段:FAILED 时 output.code / output.message 是排查依据;request_id 便于工单溯源。
|
||||
emitResult(
|
||||
{
|
||||
stripUndefined({
|
||||
task_id: response.output.task_id,
|
||||
task_status: response.output.task_status,
|
||||
video_url: response.output.video_url,
|
||||
results: response.output.results,
|
||||
submit_time: response.output.submit_time,
|
||||
scheduled_time: response.output.scheduled_time,
|
||||
end_time: response.output.end_time,
|
||||
},
|
||||
code: response.output.code,
|
||||
message: response.output.message,
|
||||
request_id: response.request_id,
|
||||
}),
|
||||
format,
|
||||
);
|
||||
},
|
||||
|
||||
@@ -73,6 +73,11 @@ export { default as mcpTools } from "./commands/mcp/tools.ts";
|
||||
export { default as searchWeb } from "./commands/search/web.ts";
|
||||
export { default as speechSynthesize } from "./commands/speech/synthesize.ts";
|
||||
export { default as speechRecognize } from "./commands/speech/recognize.ts";
|
||||
export { default as speechVocabularyCreate } from "./commands/speech/vocabulary/create.ts";
|
||||
export { default as speechVocabularyList } from "./commands/speech/vocabulary/list.ts";
|
||||
export { default as speechVocabularyGet } from "./commands/speech/vocabulary/get.ts";
|
||||
export { default as speechVocabularyUpdate } from "./commands/speech/vocabulary/update.ts";
|
||||
export { default as speechVocabularyDelete } from "./commands/speech/vocabulary/delete.ts";
|
||||
export { default as fileUpload } from "./commands/file/upload.ts";
|
||||
export { default as consoleCall } from "./commands/console/call.ts";
|
||||
export { default as usageFree } from "./commands/usage/free.ts";
|
||||
|
||||
@@ -33,3 +33,9 @@ test("default-speech-recognition-model alias accepts an ASR model ID", () => {
|
||||
"qwen-audio-3.0-asr-flash",
|
||||
);
|
||||
});
|
||||
|
||||
test("watermark config accepts only boolean text and stores a boolean", () => {
|
||||
expect(validateAndCoerce("watermark", "false")).toBe(false);
|
||||
expect(validateAndCoerce("watermark", "TRUE")).toBe(true);
|
||||
expect(() => validateAndCoerce("watermark", "yes")).toThrow(/true or false/i);
|
||||
});
|
||||
|
||||
@@ -200,8 +200,10 @@ test("GET /api/config 返回全部 profile、明文密钥与持久化激活项",
|
||||
expect(res.json.keys).toContain("console_site");
|
||||
expect(res.json.keys).toContain("telemetry");
|
||||
expect(res.json.keys).toContain("default_speech_recognition_model");
|
||||
expect(res.json.keys).toContain("watermark");
|
||||
expect(res.json.enums.console_site).toEqual(["domestic", "international"]);
|
||||
expect(res.json.booleanKeys).toContain("telemetry");
|
||||
expect(res.json.booleanKeys).toContain("watermark");
|
||||
// Default field hints are surfaced as prefilled values in the UI.
|
||||
expect(res.json.fieldDefaults.default_image_model).toBe("qwen-image-3.0");
|
||||
expect(res.json.fieldDefaults.default_text_model).toBe("qwen3.8-max");
|
||||
|
||||
@@ -64,10 +64,53 @@ describe("e2e: config", () => {
|
||||
config_file?: string;
|
||||
base_url?: string;
|
||||
timeout?: number;
|
||||
watermark?: boolean;
|
||||
}>(stdout);
|
||||
expect(data.config_file).toBeDefined();
|
||||
expect(data.base_url).toBeDefined();
|
||||
expect(data.timeout).toBeDefined();
|
||||
expect(data.watermark).toBe(true);
|
||||
});
|
||||
|
||||
test("config set 将 watermark 作为 boolean 写入并由 config show 读回", async () => {
|
||||
const configDir = mkdtempSync(join(tmpdir(), "bl-config-watermark-"));
|
||||
try {
|
||||
const env = { BAILIAN_CONFIG_DIR: configDir };
|
||||
const setResult = await runCommandE2e(
|
||||
CONFIG_ROUTES,
|
||||
[
|
||||
"config",
|
||||
"set",
|
||||
"--config",
|
||||
"media",
|
||||
"--key",
|
||||
"watermark",
|
||||
"--value",
|
||||
"false",
|
||||
"--output",
|
||||
"json",
|
||||
],
|
||||
env,
|
||||
);
|
||||
expect(setResult.exitCode, setResult.stderr).toBe(0);
|
||||
expect(parseStdoutJson<{ watermark?: boolean }>(setResult.stdout).watermark).toBe(false);
|
||||
|
||||
const persisted = JSON.parse(readFileSync(join(configDir, "config.json"), "utf8")) as Record<
|
||||
string,
|
||||
Record<string, unknown>
|
||||
>;
|
||||
expect(persisted.media?.watermark).toBe(false);
|
||||
|
||||
const showResult = await runCommandE2e(
|
||||
CONFIG_ROUTES,
|
||||
["config", "show", "--config", "media", "--output", "json"],
|
||||
env,
|
||||
);
|
||||
expect(showResult.exitCode, showResult.stderr).toBe(0);
|
||||
expect(parseStdoutJson<{ watermark?: boolean }>(showResult.stdout).watermark).toBe(false);
|
||||
} finally {
|
||||
rmSync(configDir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test("config show --output text", async () => {
|
||||
|
||||
@@ -240,4 +240,75 @@ describe("e2e: pipeline", () => {
|
||||
expect(stdout).toBe("");
|
||||
expect(stderr).toMatch(/--events must be one of: jsonl/i);
|
||||
});
|
||||
|
||||
test("pipeline image/generate dry-run 继承 Profile watermark=false", async () => {
|
||||
const configDir = await mkdtemp(join(tmpdir(), "bl-pipeline-wm-"));
|
||||
const workflowPath = join(configDir, "image-generate.json");
|
||||
try {
|
||||
await writeFile(
|
||||
join(configDir, "config.json"),
|
||||
JSON.stringify({ api_key: "sk-test-placeholder", watermark: false }, null, 2) + "\n",
|
||||
);
|
||||
await writeFile(
|
||||
workflowPath,
|
||||
JSON.stringify({
|
||||
version: "workflow/v1",
|
||||
steps: [{ id: "gen", type: "image/generate", input: { prompt: "A cat" } }],
|
||||
}),
|
||||
);
|
||||
const { stdout, stderr, exitCode } = await runCommandE2e(
|
||||
PIPELINE_ROUTES,
|
||||
["pipeline", "run", "--file", workflowPath, "--dry-run", "--output", "json"],
|
||||
{ BAILIAN_CONFIG_DIR: configDir, DASHSCOPE_API_KEY: "", DASHSCOPE_BASE_URL: "" },
|
||||
);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
const report = parseStdoutJson<{
|
||||
status?: string;
|
||||
steps?: Array<{ type?: string; input?: { watermark?: boolean; prompt?: string } }>;
|
||||
}>(stdout);
|
||||
expect(report.status).toBe("planned");
|
||||
expect(report.steps?.[0]).toMatchObject({
|
||||
type: "image/generate",
|
||||
input: { prompt: "A cat", watermark: false },
|
||||
});
|
||||
} finally {
|
||||
await rm(configDir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test("pipeline 步骤显式 watermark=true 覆盖 Profile false", async () => {
|
||||
const configDir = await mkdtemp(join(tmpdir(), "bl-pipeline-wm-ov-"));
|
||||
const workflowPath = join(configDir, "image-generate.json");
|
||||
try {
|
||||
await writeFile(
|
||||
join(configDir, "config.json"),
|
||||
JSON.stringify({ api_key: "sk-test-placeholder", watermark: false }, null, 2) + "\n",
|
||||
);
|
||||
await writeFile(
|
||||
workflowPath,
|
||||
JSON.stringify({
|
||||
version: "workflow/v1",
|
||||
steps: [
|
||||
{
|
||||
id: "gen",
|
||||
type: "image/generate",
|
||||
input: { prompt: "A cat", watermark: true },
|
||||
},
|
||||
],
|
||||
}),
|
||||
);
|
||||
const { stdout, stderr, exitCode } = await runCommandE2e(
|
||||
PIPELINE_ROUTES,
|
||||
["pipeline", "run", "--file", workflowPath, "--dry-run", "--output", "json"],
|
||||
{ BAILIAN_CONFIG_DIR: configDir, DASHSCOPE_API_KEY: "", DASHSCOPE_BASE_URL: "" },
|
||||
);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
const report = parseStdoutJson<{
|
||||
steps?: Array<{ input?: { watermark?: boolean } }>;
|
||||
}>(stdout);
|
||||
expect(report.steps?.[0]?.input?.watermark).toBe(true);
|
||||
} finally {
|
||||
await rm(configDir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
@@ -40,11 +40,18 @@ describe("e2e: speech recognize", () => {
|
||||
language_hints?: string[];
|
||||
language?: string;
|
||||
vocabulary_id?: string;
|
||||
vocabulary?: Record<string, number>;
|
||||
};
|
||||
input?: {
|
||||
file_url?: string;
|
||||
file_urls?: string[];
|
||||
messages?: Array<{ content?: Array<{ type?: string }> }>;
|
||||
context?: Array<{
|
||||
role?: string;
|
||||
content?: Array<{ type?: string; text?: string }>;
|
||||
}>;
|
||||
messages?: Array<{
|
||||
content?: Array<{ type?: string; text?: string; input_audio?: { data?: string } }>;
|
||||
}>;
|
||||
};
|
||||
};
|
||||
}>(stdout);
|
||||
@@ -120,6 +127,116 @@ describe("e2e: speech recognize", () => {
|
||||
expect(body.request?.input?.messages?.[0]?.content?.[0]?.type).toBe("input_audio");
|
||||
});
|
||||
|
||||
test("speech recognize async dry-run 注入 input.context 与 parameters.vocabulary", async () => {
|
||||
const body = await runRecognizeDryRun([
|
||||
"--model",
|
||||
"qwen-audio-3.0-asr-flash-filetrans",
|
||||
"--url",
|
||||
"https://example.com/audio.mp3",
|
||||
"--vocabulary",
|
||||
'{"奋斗者":4,"鲸落":4}',
|
||||
"--context",
|
||||
"奋斗者号 鲸落 深海勇士",
|
||||
]);
|
||||
expect(body.mode).toBe("async");
|
||||
expect(body.request?.parameters?.vocabulary).toEqual({ 奋斗者: 4, 鲸落: 4 });
|
||||
expect(body.request?.input?.context).toEqual([
|
||||
{
|
||||
role: "user",
|
||||
content: [{ type: "input_text", text: "奋斗者号 鲸落 深海勇士" }],
|
||||
},
|
||||
]);
|
||||
});
|
||||
|
||||
test("speech recognize sync input-audio dry-run 将 context 前置且 input_audio 在最后", async () => {
|
||||
const body = await runRecognizeDryRun([
|
||||
"--model",
|
||||
"qwen-audio-3.0-asr-flash",
|
||||
"--url",
|
||||
"https://example.com/audio.wav",
|
||||
"--vocabulary",
|
||||
'{"奋斗者":4}',
|
||||
"--context",
|
||||
"奋斗者号",
|
||||
]);
|
||||
expect(body.mode).toBe("sync");
|
||||
expect(body.request?.parameters?.vocabulary).toEqual({ 奋斗者: 4 });
|
||||
const messages = body.request?.input?.messages ?? [];
|
||||
expect(messages).toHaveLength(2);
|
||||
expect(messages[0]?.content?.[0]).toMatchObject({ type: "input_text", text: "奋斗者号" });
|
||||
expect(messages[1]?.content?.[0]?.type).toBe("input_audio");
|
||||
});
|
||||
|
||||
test("speech recognize 非法 --vocabulary JSON 返回用法错误", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"recognize",
|
||||
"--model",
|
||||
"qwen-audio-3.0-asr-flash-filetrans",
|
||||
"--url",
|
||||
"https://example.com/a.wav",
|
||||
"--vocabulary",
|
||||
"{bad json",
|
||||
"--dry-run",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/not valid JSON|--vocabulary/i);
|
||||
});
|
||||
|
||||
test("speech recognize 空 --vocabulary 返回用法错误", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"recognize",
|
||||
"--model",
|
||||
"qwen-audio-3.0-asr-flash-filetrans",
|
||||
"--url",
|
||||
"https://example.com/a.wav",
|
||||
"--vocabulary",
|
||||
"",
|
||||
"--dry-run",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/not valid JSON|--vocabulary/i);
|
||||
});
|
||||
|
||||
test("speech recognize qwen3 sync 拒绝 --vocabulary", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"recognize",
|
||||
"--model",
|
||||
"qwen3-asr-flash",
|
||||
"--url",
|
||||
"https://example.com/a.wav",
|
||||
"--vocabulary",
|
||||
'{"奋斗者":4}',
|
||||
"--dry-run",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/--vocabulary|does not support/i);
|
||||
expect(stderr).toMatch(/qwen-audio-3\.0-asr-flash|vocabulary\/context/i);
|
||||
});
|
||||
|
||||
test("speech recognize sync Flash 拒绝 --diarization 时提示 async filetrans", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"recognize",
|
||||
"--model",
|
||||
"fun-asr-flash",
|
||||
"--url",
|
||||
"https://example.com/a.wav",
|
||||
"--diarization",
|
||||
"--dry-run",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/--diarization|does not support/i);
|
||||
expect(stderr).toMatch(/async filetrans|fun-asr/i);
|
||||
expect(stderr).not.toMatch(/vocabulary\/context/i);
|
||||
});
|
||||
|
||||
test("speech recognize qwen3 filetrans dry-run 使用 file_url 与 language", async () => {
|
||||
const body = await runRecognizeDryRun([
|
||||
"--model",
|
||||
@@ -366,5 +483,88 @@ describe.skipIf(!isBailianE2EMediaEnabled() || !isDashScopeE2EReady())(
|
||||
const raw = readFileSync(asrJson, "utf8");
|
||||
expect(raw.length).toBeGreaterThan(2);
|
||||
}, 300_000);
|
||||
|
||||
test("【qwen-audio】synthesize → recognize 即时热词/上下文", async () => {
|
||||
// 生造专名:无热词时常被听错;带 --vocabulary/--context 后应能正确召回。
|
||||
const hotwordScript =
|
||||
"请把录音同步到听悟匣,并启动澜舟芯做摘要。听悟匣负责转写,澜舟芯负责归档。最后确认玄甲协议是否已开启。";
|
||||
const hotwords = ["听悟匣", "澜舟芯", "玄甲协议"] as const;
|
||||
const vocabularyJson = '{"听悟匣":4,"澜舟芯":4,"玄甲协议":4}';
|
||||
const contextText = "听悟匣 澜舟芯 玄甲协议";
|
||||
|
||||
const missingHotwords = (text: string): string[] => {
|
||||
const normalized = text.replace(/\s+/g, "");
|
||||
return hotwords.filter((word) => !normalized.includes(word));
|
||||
};
|
||||
|
||||
const outDir = makeE2eOutputDir(e2eLabelFromMetaUrl(import.meta.url));
|
||||
const outMp3 = join(outDir, "hotword-tts.mp3");
|
||||
const syn = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"synthesize",
|
||||
"--model",
|
||||
"cosyvoice-v3-flash",
|
||||
"--voice",
|
||||
"longxiaochun_v3",
|
||||
"--text",
|
||||
hotwordScript,
|
||||
"--out",
|
||||
outMp3,
|
||||
"--output",
|
||||
"json",
|
||||
]);
|
||||
expect(syn.exitCode, syn.stderr).toBe(0);
|
||||
const synBody = parseStdoutJson<{ audio_url?: string }>(syn.stdout);
|
||||
const audioUrl = synBody.audio_url;
|
||||
expect(audioUrl?.startsWith("http")).toBe(true);
|
||||
|
||||
const baselineOut = join(outDir, "asr-baseline.json");
|
||||
const baseline = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"recognize",
|
||||
"--model",
|
||||
"qwen-audio-3.0-asr-flash",
|
||||
"--url",
|
||||
audioUrl!,
|
||||
"--language",
|
||||
"zh",
|
||||
"--out",
|
||||
baselineOut,
|
||||
"--quiet",
|
||||
]);
|
||||
expect(baseline.exitCode, baseline.stderr).toBe(0);
|
||||
writeFileSync(join(outDir, "asr-baseline.txt"), baseline.stdout);
|
||||
const baselineMissing = missingHotwords(baseline.stdout);
|
||||
// soft:仅落盘对照,不 fail(无热词偶发也能认出专名)
|
||||
writeFileSync(
|
||||
join(outDir, "asr-baseline-missing.txt"),
|
||||
baselineMissing.length > 0 ? baselineMissing.join("\n") + "\n" : "(none)\n",
|
||||
);
|
||||
|
||||
const hotOut = join(outDir, "asr-hot.json");
|
||||
const hot = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"recognize",
|
||||
"--model",
|
||||
"qwen-audio-3.0-asr-flash",
|
||||
"--url",
|
||||
audioUrl!,
|
||||
"--language",
|
||||
"zh",
|
||||
"--vocabulary",
|
||||
vocabularyJson,
|
||||
"--context",
|
||||
contextText,
|
||||
"--out",
|
||||
hotOut,
|
||||
"--quiet",
|
||||
]);
|
||||
expect(hot.exitCode, hot.stderr).toBe(0);
|
||||
writeFileSync(join(outDir, "asr-hot.txt"), hot.stdout);
|
||||
expect(
|
||||
missingHotwords(hot.stdout),
|
||||
`expected hotwords in ASR text, got: ${hot.stdout.trim()}`,
|
||||
).toEqual([]);
|
||||
}, 420_000);
|
||||
},
|
||||
);
|
||||
|
||||
@@ -0,0 +1,551 @@
|
||||
import { mkdtempSync, rmSync, writeFileSync } from "node:fs";
|
||||
import { tmpdir } from "node:os";
|
||||
import { join } from "node:path";
|
||||
import { afterEach, describe, expect, test } from "vite-plus/test";
|
||||
import {
|
||||
e2eLabelFromMetaUrl,
|
||||
isBailianE2EMediaEnabled,
|
||||
isDashScopeE2EReady,
|
||||
makeE2eOutputDir,
|
||||
parseStdoutJson,
|
||||
runCommandHelp,
|
||||
runCommandE2e,
|
||||
} from "./helpers.ts";
|
||||
import { SPEECH_ROUTES } from "./topic-routes.ts";
|
||||
|
||||
/**
|
||||
* Speech vocabulary:help / dry-run / 确认闸门无密钥;真实 CRUD 需媒体 E2E + DashScope。
|
||||
*/
|
||||
|
||||
const tempDirs: string[] = [];
|
||||
|
||||
afterEach(() => {
|
||||
for (const tempDir of tempDirs.splice(0)) {
|
||||
rmSync(tempDir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
function makeTempJson(content: string): string {
|
||||
const tempDir = mkdtempSync(join(tmpdir(), "bl-vocab-e2e-"));
|
||||
tempDirs.push(tempDir);
|
||||
const filePath = join(tempDir, "hotwords.json");
|
||||
writeFileSync(filePath, content);
|
||||
return filePath;
|
||||
}
|
||||
|
||||
describe("e2e: speech vocabulary", () => {
|
||||
test("speech vocabulary --help 列出 5 个子命令", async () => {
|
||||
const { stderr, exitCode } = await runCommandHelp(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"--help",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
expect(stderr).toMatch(/create/i);
|
||||
expect(stderr).toMatch(/list/i);
|
||||
expect(stderr).toMatch(/get/i);
|
||||
expect(stderr).toMatch(/update/i);
|
||||
expect(stderr).toMatch(/delete/i);
|
||||
});
|
||||
|
||||
test("create --help 展示关键 flags 与静默失效 / weight 50 文案", async () => {
|
||||
const { stderr, exitCode } = await runCommandHelp(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--help",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
expect(stderr).toMatch(/--model/i);
|
||||
expect(stderr).toMatch(/--prefix/i);
|
||||
expect(stderr).toMatch(/--words/i);
|
||||
expect(stderr).toMatch(/--words-file/i);
|
||||
expect(stderr).toMatch(/silently ignored|静默失效/i);
|
||||
expect(stderr).toMatch(/50/);
|
||||
});
|
||||
|
||||
test("delete --help 展示 Risk 与 --yes", async () => {
|
||||
const { stderr, exitCode } = await runCommandHelp(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"delete",
|
||||
"--help",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
expect(stderr).toMatch(/--yes/i);
|
||||
expect(stderr).toMatch(/Risk|风险/i);
|
||||
});
|
||||
|
||||
test("update --help 展示 Risk、replace 语义与 --yes", async () => {
|
||||
const { stderr, exitCode } = await runCommandHelp(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"update",
|
||||
"--help",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
expect(stderr).toMatch(/--yes/i);
|
||||
expect(stderr).toMatch(/Risk|风险/i);
|
||||
expect(stderr).toMatch(/replaces|替换/i);
|
||||
});
|
||||
|
||||
test("create 缺少 --model 时退出为用法错误 (2)", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--prefix",
|
||||
"demo",
|
||||
"--words",
|
||||
'{"x":4}',
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/--model|Missing required/i);
|
||||
});
|
||||
|
||||
test("create 缺少 --prefix 时退出为用法错误 (2)", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--model",
|
||||
"fun-asr",
|
||||
"--words",
|
||||
'{"x":4}',
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/--prefix|Missing required/i);
|
||||
});
|
||||
|
||||
test("create 两个词表 flag 都不传时退出为用法错误 (2)", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--model",
|
||||
"fun-asr",
|
||||
"--prefix",
|
||||
"demo",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/--words|--words-file/i);
|
||||
});
|
||||
|
||||
test("create 同时传 --words 与 --words-file 时退出为用法错误 (2)", async () => {
|
||||
const filePath = makeTempJson('{"x":4}');
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--model",
|
||||
"fun-asr",
|
||||
"--prefix",
|
||||
"demo",
|
||||
"--words",
|
||||
'{"x":4}',
|
||||
"--words-file",
|
||||
filePath,
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/either|--words|--words-file/i);
|
||||
});
|
||||
|
||||
test("create 非法 JSON 时退出为用法错误 (2)", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--model",
|
||||
"fun-asr",
|
||||
"--prefix",
|
||||
"demo",
|
||||
"--words",
|
||||
"{bad json",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/not valid JSON|JSON/i);
|
||||
});
|
||||
|
||||
test("create 字符串权重时退出为用法错误 (2)", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--model",
|
||||
"fun-asr",
|
||||
"--prefix",
|
||||
"demo",
|
||||
"--words",
|
||||
'{"x":"4"}',
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/must be a number|number/i);
|
||||
});
|
||||
|
||||
test("create 空词表 {} 时退出为用法错误 (2)", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--model",
|
||||
"fun-asr",
|
||||
"--prefix",
|
||||
"demo",
|
||||
"--words",
|
||||
"{}",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/at least one|hot word/i);
|
||||
});
|
||||
|
||||
test("get 缺少 --id 时退出为用法错误 (2)", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"get",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/--id|Missing required/i);
|
||||
});
|
||||
|
||||
test("delete 缺少 --id 时退出为用法错误 (2)", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"delete",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode).toBe(2);
|
||||
expect(stderr).toMatch(/--id|Missing required/i);
|
||||
});
|
||||
|
||||
test("create object 形态 --dry-run 输出 speech-biasing 信封", async () => {
|
||||
const { stdout, stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--model",
|
||||
"fun-asr",
|
||||
"--prefix",
|
||||
"demo",
|
||||
"--words",
|
||||
'{"奋斗者":4,"鲸落":4}',
|
||||
"--dry-run",
|
||||
"--output",
|
||||
"json",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
const body = parseStdoutJson<{
|
||||
request?: {
|
||||
model?: string;
|
||||
input?: {
|
||||
action?: string;
|
||||
target_model?: string;
|
||||
prefix?: string;
|
||||
vocabulary?: Array<{ text?: string; weight?: number; lang?: string }>;
|
||||
};
|
||||
};
|
||||
}>(stdout);
|
||||
expect(body.request?.model).toBe("speech-biasing");
|
||||
expect(body.request?.input?.action).toBe("create_vocabulary");
|
||||
expect(body.request?.input?.target_model).toBe("fun-asr");
|
||||
expect(body.request?.input?.prefix).toBe("demo");
|
||||
expect(body.request?.input?.vocabulary).toEqual([
|
||||
{ text: "奋斗者", weight: 4 },
|
||||
{ text: "鲸落", weight: 4 },
|
||||
]);
|
||||
});
|
||||
|
||||
test("create array 形态 --dry-run 原样透传 lang", async () => {
|
||||
const { stdout, stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--model",
|
||||
"fun-asr",
|
||||
"--prefix",
|
||||
"demo",
|
||||
"--words",
|
||||
'[{"text":"x","weight":4,"lang":"zh"}]',
|
||||
"--dry-run",
|
||||
"--output",
|
||||
"json",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
const body = parseStdoutJson<{
|
||||
request?: {
|
||||
input?: { vocabulary?: Array<{ text?: string; weight?: number; lang?: string }> };
|
||||
};
|
||||
}>(stdout);
|
||||
expect(body.request?.input?.vocabulary).toEqual([{ text: "x", weight: 4, lang: "zh" }]);
|
||||
});
|
||||
|
||||
test("create object + --lang --dry-run 下发到每一条", async () => {
|
||||
const { stdout, stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--model",
|
||||
"fun-asr",
|
||||
"--prefix",
|
||||
"demo",
|
||||
"--words",
|
||||
'{"x":4}',
|
||||
"--lang",
|
||||
"zh",
|
||||
"--dry-run",
|
||||
"--output",
|
||||
"json",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
const body = parseStdoutJson<{
|
||||
request?: {
|
||||
input?: { vocabulary?: Array<{ text?: string; weight?: number; lang?: string }> };
|
||||
};
|
||||
}>(stdout);
|
||||
expect(body.request?.input?.vocabulary).toEqual([{ text: "x", weight: 4, lang: "zh" }]);
|
||||
});
|
||||
|
||||
test("create --words-file --dry-run 读取文件", async () => {
|
||||
const filePath = makeTempJson('{"from-file":4}');
|
||||
const { stdout, stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--model",
|
||||
"fun-asr",
|
||||
"--prefix",
|
||||
"demo",
|
||||
"--words-file",
|
||||
filePath,
|
||||
"--dry-run",
|
||||
"--output",
|
||||
"json",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
const body = parseStdoutJson<{
|
||||
request?: {
|
||||
input?: { vocabulary?: Array<{ text?: string; weight?: number }> };
|
||||
};
|
||||
}>(stdout);
|
||||
expect(body.request?.input?.vocabulary).toEqual([{ text: "from-file", weight: 4 }]);
|
||||
});
|
||||
|
||||
test("list --page 2 --dry-run 将 page_index 转为 1", async () => {
|
||||
const { stdout, stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"list",
|
||||
"--page",
|
||||
"2",
|
||||
"--dry-run",
|
||||
"--output",
|
||||
"json",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
const body = parseStdoutJson<{
|
||||
request?: { input?: { action?: string; page_index?: number } };
|
||||
}>(stdout);
|
||||
expect(body.request?.input?.action).toBe("list_vocabulary");
|
||||
expect(body.request?.input?.page_index).toBe(1);
|
||||
});
|
||||
|
||||
test("get --dry-run 使用 query_vocabulary", async () => {
|
||||
const { stdout, stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"get",
|
||||
"--id",
|
||||
"vocab-x",
|
||||
"--dry-run",
|
||||
"--output",
|
||||
"json",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
const body = parseStdoutJson<{
|
||||
request?: { input?: { action?: string; vocabulary_id?: string } };
|
||||
}>(stdout);
|
||||
expect(body.request?.input?.action).toBe("query_vocabulary");
|
||||
expect(body.request?.input?.vocabulary_id).toBe("vocab-x");
|
||||
});
|
||||
|
||||
test("update --dry-run 无 --yes 也能预览", async () => {
|
||||
const { stdout, stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"update",
|
||||
"--id",
|
||||
"vocab-x",
|
||||
"--words",
|
||||
'{"x":4}',
|
||||
"--dry-run",
|
||||
"--output",
|
||||
"json",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
const body = parseStdoutJson<{
|
||||
request?: {
|
||||
input?: {
|
||||
action?: string;
|
||||
vocabulary_id?: string;
|
||||
vocabulary?: Array<{ text?: string; weight?: number }>;
|
||||
};
|
||||
};
|
||||
}>(stdout);
|
||||
expect(body.request?.input?.action).toBe("update_vocabulary");
|
||||
expect(body.request?.input?.vocabulary_id).toBe("vocab-x");
|
||||
expect(body.request?.input?.vocabulary).toEqual([{ text: "x", weight: 4 }]);
|
||||
});
|
||||
|
||||
test("delete --dry-run 无 --yes 也能预览", async () => {
|
||||
const { stdout, stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"delete",
|
||||
"--id",
|
||||
"vocab-x",
|
||||
"--dry-run",
|
||||
"--output",
|
||||
"json",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
const body = parseStdoutJson<{
|
||||
request?: { input?: { action?: string; vocabulary_id?: string } };
|
||||
}>(stdout);
|
||||
expect(body.request?.input?.action).toBe("delete_vocabulary");
|
||||
expect(body.request?.input?.vocabulary_id).toBe("vocab-x");
|
||||
});
|
||||
});
|
||||
|
||||
describe("e2e: speech vocabulary high-risk confirmation", () => {
|
||||
test("delete 无 --yes 返回确认请求 (7)", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"delete",
|
||||
"--id",
|
||||
"vocab-x",
|
||||
"--api-key",
|
||||
"e2e-dummy-key",
|
||||
"--output",
|
||||
"json",
|
||||
]);
|
||||
expect(exitCode).toBe(7);
|
||||
expect(JSON.parse(stderr)).toMatchObject({
|
||||
error: { code: 7, type: "requires_confirmation" },
|
||||
});
|
||||
});
|
||||
|
||||
test("update 无 --yes 返回确认请求 (7)", async () => {
|
||||
const { stderr, exitCode } = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"update",
|
||||
"--id",
|
||||
"vocab-x",
|
||||
"--words",
|
||||
'{"x":4}',
|
||||
"--api-key",
|
||||
"e2e-dummy-key",
|
||||
"--output",
|
||||
"json",
|
||||
]);
|
||||
expect(exitCode).toBe(7);
|
||||
expect(JSON.parse(stderr)).toMatchObject({
|
||||
error: { code: 7, type: "requires_confirmation" },
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
describe.skipIf(!isBailianE2EMediaEnabled() || !isDashScopeE2EReady())(
|
||||
"e2e: speech vocabulary(DashScope 媒体)",
|
||||
() => {
|
||||
test("create → list/get → delete 完整链路", async () => {
|
||||
const outDir = makeE2eOutputDir(e2eLabelFromMetaUrl(import.meta.url));
|
||||
let vocabularyId = "";
|
||||
|
||||
try {
|
||||
const created = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"create",
|
||||
"--model",
|
||||
"fun-asr",
|
||||
"--prefix",
|
||||
"blcli",
|
||||
"--words",
|
||||
'{"奋斗者":4}',
|
||||
"--quiet",
|
||||
]);
|
||||
expect(created.exitCode, created.stderr).toBe(0);
|
||||
vocabularyId = created.stdout.trim();
|
||||
expect(vocabularyId.length).toBeGreaterThan(0);
|
||||
writeFileSync(join(outDir, "vocabulary-id.txt"), vocabularyId + "\n");
|
||||
|
||||
const listed = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"list",
|
||||
"--prefix",
|
||||
"blcli",
|
||||
"--output",
|
||||
"json",
|
||||
]);
|
||||
expect(listed.exitCode, listed.stderr).toBe(0);
|
||||
const listBody = parseStdoutJson<{
|
||||
output?: { vocabulary_list?: Array<{ vocabulary_id?: string; status?: string }> };
|
||||
}>(listed.stdout);
|
||||
const listedItem = listBody.output?.vocabulary_list?.find(
|
||||
(item) => item.vocabulary_id === vocabularyId,
|
||||
);
|
||||
expect(listedItem).toBeTruthy();
|
||||
expect(listedItem?.status).toBe("OK");
|
||||
|
||||
const got = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"get",
|
||||
"--id",
|
||||
vocabularyId,
|
||||
"--output",
|
||||
"json",
|
||||
]);
|
||||
expect(got.exitCode, got.stderr).toBe(0);
|
||||
const getBody = parseStdoutJson<{
|
||||
output?: { status?: string; target_model?: string };
|
||||
}>(got.stdout);
|
||||
expect(getBody.output?.status).toBe("OK");
|
||||
expect(getBody.output?.target_model).toBe("fun-asr");
|
||||
} finally {
|
||||
if (vocabularyId) {
|
||||
const deleted = await runCommandE2e(SPEECH_ROUTES, [
|
||||
"speech",
|
||||
"vocabulary",
|
||||
"delete",
|
||||
"--id",
|
||||
vocabularyId,
|
||||
"--yes",
|
||||
"--quiet",
|
||||
]);
|
||||
expect(deleted.exitCode, deleted.stderr).toBe(0);
|
||||
}
|
||||
}
|
||||
}, 120_000);
|
||||
},
|
||||
);
|
||||
@@ -72,6 +72,11 @@ export const VISION_ROUTES: E2eRouteExports = {
|
||||
export const SPEECH_ROUTES: E2eRouteExports = {
|
||||
"speech synthesize": "speechSynthesize",
|
||||
"speech recognize": "speechRecognize",
|
||||
"speech vocabulary create": "speechVocabularyCreate",
|
||||
"speech vocabulary list": "speechVocabularyList",
|
||||
"speech vocabulary get": "speechVocabularyGet",
|
||||
"speech vocabulary update": "speechVocabularyUpdate",
|
||||
"speech vocabulary delete": "speechVocabularyDelete",
|
||||
};
|
||||
|
||||
export const MCP_ROUTES: E2eRouteExports = {
|
||||
|
||||
@@ -0,0 +1,80 @@
|
||||
import { mkdtempSync, rmSync, writeFileSync } from "node:fs";
|
||||
import { tmpdir } from "node:os";
|
||||
import { join } from "node:path";
|
||||
import { describe, expect, test } from "vite-plus/test";
|
||||
import { parseStdoutJson, runCommandE2e } from "./helpers.ts";
|
||||
import { IMAGE_ROUTES, VIDEO_ROUTES, type E2eRouteExports } from "./topic-routes.ts";
|
||||
|
||||
interface WatermarkScenario {
|
||||
name: string;
|
||||
routes: E2eRouteExports;
|
||||
args: string[];
|
||||
}
|
||||
|
||||
const scenarios: WatermarkScenario[] = [
|
||||
{
|
||||
name: "image generate",
|
||||
routes: IMAGE_ROUTES,
|
||||
args: ["image", "generate", "--prompt", "A cat"],
|
||||
},
|
||||
{
|
||||
name: "image edit",
|
||||
routes: IMAGE_ROUTES,
|
||||
args: [
|
||||
"image",
|
||||
"edit",
|
||||
"--image",
|
||||
"https://example.com/input.png",
|
||||
"--prompt",
|
||||
"Blue background",
|
||||
],
|
||||
},
|
||||
{
|
||||
name: "video generate",
|
||||
routes: VIDEO_ROUTES,
|
||||
args: ["video", "generate", "--prompt", "A cat waves"],
|
||||
},
|
||||
{
|
||||
name: "video edit",
|
||||
routes: VIDEO_ROUTES,
|
||||
args: ["video", "edit", "--video", "https://example.com/input.mp4", "--prompt", "Warm colors"],
|
||||
},
|
||||
{
|
||||
name: "video ref",
|
||||
routes: VIDEO_ROUTES,
|
||||
args: [
|
||||
"video",
|
||||
"ref",
|
||||
"--image",
|
||||
"https://example.com/person.png",
|
||||
"--prompt",
|
||||
"Image 1 waves",
|
||||
],
|
||||
},
|
||||
];
|
||||
|
||||
describe("e2e: global watermark config", () => {
|
||||
for (const scenario of scenarios) {
|
||||
test(`${scenario.name} uses watermark=false from the selected Profile`, async () => {
|
||||
const configDir = mkdtempSync(join(tmpdir(), "bl-watermark-profile-"));
|
||||
try {
|
||||
writeFileSync(
|
||||
join(configDir, "config.json"),
|
||||
JSON.stringify({ media: { watermark: false } }, null, 2) + "\n",
|
||||
);
|
||||
const { stdout, stderr, exitCode } = await runCommandE2e(
|
||||
scenario.routes,
|
||||
[...scenario.args, "--config", "media", "--dry-run", "--output", "json"],
|
||||
{ BAILIAN_CONFIG_DIR: configDir },
|
||||
);
|
||||
expect(exitCode, stderr).toBe(0);
|
||||
const data = parseStdoutJson<{
|
||||
request?: { parameters?: { watermark?: boolean } };
|
||||
}>(stdout);
|
||||
expect(data.request?.parameters?.watermark).toBe(false);
|
||||
} finally {
|
||||
rmSync(configDir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
}
|
||||
});
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "bailian-cli-core",
|
||||
"version": "1.22.0",
|
||||
"version": "1.23.0",
|
||||
"description": "Core SDK for bailian-cli. See https://www.npmjs.com/package/bailian-cli for usage.",
|
||||
"homepage": "https://bailian.console.aliyun.com/cli",
|
||||
"bugs": {
|
||||
|
||||
@@ -1,16 +1,20 @@
|
||||
import { imageSyncPath, speechRecognizePath } from "./endpoints.ts";
|
||||
|
||||
import type { AsrContextMessage } from "../types/api.ts";
|
||||
|
||||
/**
|
||||
* DashScope ASR APIs differ by model family:
|
||||
*
|
||||
* - async file transcription (`.../audio/asr/transcription`):
|
||||
* fun-asr*, paraformer* (non-realtime), *-filetrans, sensevoice*
|
||||
* language via `parameters.language_hints`
|
||||
* language via `parameters.language_hints`; optional `input.context` /
|
||||
* `parameters.vocabulary` (model-dependent effective range)
|
||||
* - sync multimodal (`.../aigc/multimodal-generation/generation`):
|
||||
* - qwen3: `{ content: [{ audio }] }` + optional `asr_options.language`
|
||||
* (qwen3-asr-flash*)
|
||||
* (qwen3-asr-flash*) — no vocabulary / context fields in this body shape
|
||||
* - input-audio: `{ type: input_audio, input_audio.data }` +
|
||||
* `format`/`sample_rate` + optional `language_hints`
|
||||
* `format`/`sample_rate` + optional `language_hints` /
|
||||
* `vocabulary_id` / `vocabulary`; optional leading `input_text` for context
|
||||
* (fun-asr-flash*, qwen-audio-*-asr-flash*)
|
||||
* - realtime / streaming: WebSocket — not supported by `speech recognize`
|
||||
*/
|
||||
@@ -158,9 +162,18 @@ export interface BuildAsrFlashRequestOpts {
|
||||
language?: string;
|
||||
/** Precompiled hotword vocabulary ID; supported for input-audio Flash (fun-asr-flash* / qwen-audio-*-asr-flash). */
|
||||
vocabularyId?: string;
|
||||
/** Instant hot words (word → weight); input-audio Flash only (command layer rejects qwen3). */
|
||||
vocabulary?: Record<string, number>;
|
||||
/** Context enhancement text; prepended as input_text before input_audio. */
|
||||
context?: string;
|
||||
flashFamily: AsrFlashFamily;
|
||||
}
|
||||
|
||||
/** Wrap plain text as a single user context message for ASR. */
|
||||
export function buildAsrContextMessages(text: string): AsrContextMessage[] {
|
||||
return [{ role: "user", content: [{ type: "input_text", text }] }];
|
||||
}
|
||||
|
||||
/**
|
||||
* Build language fields for async ASR routes.
|
||||
* qwen3-asr-flash-filetrans* → `language`; other async models → `language_hints`.
|
||||
@@ -178,10 +191,10 @@ export function buildAsyncAsrLanguageFields(
|
||||
|
||||
/** Build a sync multimodal ASR request body for Flash models. */
|
||||
export function buildAsrFlashRequest(opts: BuildAsrFlashRequestOpts): Record<string, unknown> {
|
||||
const { model, audioUrl, language, vocabularyId, flashFamily } = opts;
|
||||
const { model, audioUrl, language, vocabularyId, vocabulary, context, flashFamily } = opts;
|
||||
|
||||
if (flashFamily === "input-audio") {
|
||||
// Match official Qwen-Audio / Fun-ASR-Flash docs: language_hints + vocabulary_id
|
||||
// Match official Qwen-Audio / Fun-ASR-Flash docs: language_hints + vocabulary(_id)
|
||||
const parameters: Record<string, unknown> = {
|
||||
format: inferAudioFormatHint(audioUrl),
|
||||
sample_rate: "16000",
|
||||
@@ -192,21 +205,28 @@ export function buildAsrFlashRequest(opts: BuildAsrFlashRequestOpts): Record<str
|
||||
if (vocabularyId) {
|
||||
parameters.vocabulary_id = vocabularyId;
|
||||
}
|
||||
if (vocabulary) {
|
||||
parameters.vocabulary = vocabulary;
|
||||
}
|
||||
// input_audio must be the last message; prepend context as input_text when present
|
||||
const messages: Array<Record<string, unknown>> = [];
|
||||
if (context) {
|
||||
for (const message of buildAsrContextMessages(context)) {
|
||||
messages.push(message as unknown as Record<string, unknown>);
|
||||
}
|
||||
}
|
||||
messages.push({
|
||||
role: "user",
|
||||
content: [
|
||||
{
|
||||
type: "input_audio",
|
||||
input_audio: { data: audioUrl },
|
||||
},
|
||||
],
|
||||
});
|
||||
return {
|
||||
model,
|
||||
input: {
|
||||
messages: [
|
||||
{
|
||||
role: "user",
|
||||
content: [
|
||||
{
|
||||
type: "input_audio",
|
||||
input_audio: { data: audioUrl },
|
||||
},
|
||||
],
|
||||
},
|
||||
],
|
||||
},
|
||||
input: { messages },
|
||||
parameters,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -89,6 +89,11 @@ export function speechRecognizePath(): string {
|
||||
return "/api/v1/services/audio/asr/transcription";
|
||||
}
|
||||
|
||||
// ---- Hot-word Vocabulary (ASR customization) ----
|
||||
export function speechVocabularyPath(): string {
|
||||
return "/api/v1/services/audio/asr/customization";
|
||||
}
|
||||
|
||||
// ---- Memory Profile (DashScope v2) ----
|
||||
export function profileSchemaPath(): string {
|
||||
return "/api/v2/apps/memory/profile_schemas";
|
||||
|
||||
@@ -21,6 +21,7 @@ export {
|
||||
responsesPath,
|
||||
speechRecognizePath,
|
||||
speechSynthesizePath,
|
||||
speechVocabularyPath,
|
||||
taskPath,
|
||||
userProfilePath,
|
||||
videoGeneratePath,
|
||||
@@ -41,6 +42,7 @@ export {
|
||||
type ImageSizeProfile,
|
||||
} from "./image-routes.ts";
|
||||
export {
|
||||
buildAsrContextMessages,
|
||||
buildAsrFlashRequest,
|
||||
buildAsyncAsrLanguageFields,
|
||||
collectAsrTranscriptionItems,
|
||||
|
||||
@@ -272,6 +272,7 @@ export function buildSettings(s: ResolutionSources): Settings {
|
||||
outputExplicit: Boolean(flags.output || env.DASHSCOPE_OUTPUT || file.output),
|
||||
outputDir: file.output_dir || undefined,
|
||||
timeout,
|
||||
watermark: file.watermark ?? true,
|
||||
defaultTextModel: file.default_text_model,
|
||||
defaultVideoModel: file.default_video_model,
|
||||
defaultImageToVideoModel: file.default_image_to_video_model,
|
||||
|
||||
@@ -35,6 +35,7 @@ export interface ConfigFile {
|
||||
output?: "text" | "json";
|
||||
output_dir?: string;
|
||||
timeout?: number;
|
||||
watermark?: boolean;
|
||||
default_text_model?: string;
|
||||
default_video_model?: string;
|
||||
default_image_to_video_model?: string;
|
||||
@@ -63,6 +64,7 @@ export const CONFIG_FILE_KEYS = [
|
||||
"output",
|
||||
"output_dir",
|
||||
"timeout",
|
||||
"watermark",
|
||||
"default_text_model",
|
||||
"default_video_model",
|
||||
"default_image_to_video_model",
|
||||
@@ -156,6 +158,7 @@ export function parseConfigFile(raw: unknown): ConfigFile {
|
||||
if (typeof obj.output_dir === "string" && obj.output_dir.length > 0)
|
||||
out.output_dir = obj.output_dir;
|
||||
if (typeof obj.timeout === "number" && obj.timeout > 0) out.timeout = obj.timeout;
|
||||
if (typeof obj.watermark === "boolean") out.watermark = obj.watermark;
|
||||
if (typeof obj.default_text_model === "string" && obj.default_text_model.length > 0)
|
||||
out.default_text_model = obj.default_text_model;
|
||||
if (typeof obj.default_video_model === "string" && obj.default_video_model.length > 0)
|
||||
@@ -224,6 +227,7 @@ export interface Settings {
|
||||
outputExplicit: boolean;
|
||||
outputDir?: string;
|
||||
timeout: number;
|
||||
watermark: boolean;
|
||||
defaultTextModel?: string;
|
||||
defaultVideoModel?: string;
|
||||
defaultImageToVideoModel?: string;
|
||||
|
||||
@@ -13,6 +13,7 @@ export * from "./files/index.ts";
|
||||
export * from "./dataset/index.ts";
|
||||
export * from "./finetune/index.ts";
|
||||
export * from "./deploy/index.ts";
|
||||
export * from "./speech/index.ts";
|
||||
export * from "./types/index.ts";
|
||||
export * from "./utils/index.ts";
|
||||
export * from "./telemetry/index.ts";
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
export {
|
||||
SPEECH_BIASING_MODEL,
|
||||
buildVocabularyRequest,
|
||||
createVocabulary,
|
||||
listVocabularies,
|
||||
queryVocabulary,
|
||||
updateVocabulary,
|
||||
deleteVocabulary,
|
||||
type VocabularyEntry,
|
||||
type VocabularyListItem,
|
||||
type VocabularyEnvelope,
|
||||
type VocabularyRequest,
|
||||
} from "./vocabulary.ts";
|
||||
export { parseInstantVocabulary, parseVocabularyEntries } from "./vocabulary-input.ts";
|
||||
@@ -0,0 +1,76 @@
|
||||
import { UsageError } from "../errors/base.ts";
|
||||
import type { VocabularyEntry } from "./vocabulary.ts";
|
||||
|
||||
/** Shared JSON decode + top-level shape guard for both hot-word flags. */
|
||||
function decodeVocabularyJson(raw: string, flagName: string): unknown {
|
||||
try {
|
||||
return JSON.parse(raw);
|
||||
} catch (error) {
|
||||
throw new UsageError(`${flagName} is not valid JSON — ${(error as Error).message}`);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Instant hot words for `recognize --vocabulary`.
|
||||
* The API field is a flat word→weight object, so an array is a usage error here.
|
||||
*/
|
||||
export function parseInstantVocabulary(raw: string): Record<string, number> {
|
||||
const parsed = decodeVocabularyJson(raw, "--vocabulary");
|
||||
if (!parsed || typeof parsed !== "object" || Array.isArray(parsed)) {
|
||||
throw new UsageError("--vocabulary must decode to a JSON object of word→weight.");
|
||||
}
|
||||
for (const [word, weight] of Object.entries(parsed)) {
|
||||
if (typeof weight !== "number" || !Number.isFinite(weight)) {
|
||||
throw new UsageError(`--vocabulary weight for "${word}" must be a number.`);
|
||||
}
|
||||
}
|
||||
return parsed as Record<string, number>;
|
||||
}
|
||||
|
||||
/**
|
||||
* Vocabulary create/update --words.
|
||||
* Accepts the same word→weight object as recognize, or the API entry array
|
||||
* when per-entry lang is needed. Array form ignores the optional lang param.
|
||||
*/
|
||||
export function parseVocabularyEntries(raw: string, lang?: string): VocabularyEntry[] {
|
||||
const flagName = "--words";
|
||||
const parsed = decodeVocabularyJson(raw, flagName);
|
||||
let entries: VocabularyEntry[];
|
||||
if (Array.isArray(parsed)) {
|
||||
entries = parsed.map((item, index) => {
|
||||
if (!item || typeof item !== "object" || Array.isArray(item)) {
|
||||
throw new UsageError(
|
||||
`${flagName} entry #${index} must be an object with string "text" and number "weight".`,
|
||||
);
|
||||
}
|
||||
const entry = item as Partial<VocabularyEntry>;
|
||||
if (typeof entry.text !== "string" || typeof entry.weight !== "number") {
|
||||
throw new UsageError(
|
||||
`${flagName} entry #${index} must have a string "text" and a number "weight".`,
|
||||
);
|
||||
}
|
||||
if (!Number.isFinite(entry.weight)) {
|
||||
throw new UsageError(`${flagName} entry #${index} weight must be a finite number.`);
|
||||
}
|
||||
return {
|
||||
text: entry.text,
|
||||
weight: entry.weight,
|
||||
...(typeof entry.lang === "string" ? { lang: entry.lang } : {}),
|
||||
};
|
||||
});
|
||||
} else if (!parsed || typeof parsed !== "object") {
|
||||
throw new UsageError(`${flagName} must decode to a JSON object or array.`);
|
||||
} else {
|
||||
entries = [];
|
||||
for (const [text, weight] of Object.entries(parsed)) {
|
||||
if (typeof weight !== "number" || !Number.isFinite(weight)) {
|
||||
throw new UsageError(`${flagName} weight for "${text}" must be a number.`);
|
||||
}
|
||||
entries.push({ text, weight, ...(lang ? { lang } : {}) });
|
||||
}
|
||||
}
|
||||
if (entries.length === 0) {
|
||||
throw new UsageError(`${flagName} must contain at least one hot word.`);
|
||||
}
|
||||
return entries;
|
||||
}
|
||||
@@ -0,0 +1,128 @@
|
||||
/**
|
||||
* Hot-word vocabulary HTTP API wrappers.
|
||||
*
|
||||
* Thin functions over `requestJson`. They return the parsed body verbatim
|
||||
* (snake_case) so callers can decide how to surface fields.
|
||||
*/
|
||||
import { speechVocabularyPath } from "../client/endpoints.ts";
|
||||
import type { Client } from "../client/client.ts";
|
||||
|
||||
/** Fixed model id for the hot-word customization endpoint. */
|
||||
export const SPEECH_BIASING_MODEL = "speech-biasing";
|
||||
|
||||
export interface VocabularyEntry {
|
||||
text: string;
|
||||
weight: number;
|
||||
lang?: string;
|
||||
}
|
||||
|
||||
export interface VocabularyListItem {
|
||||
vocabulary_id?: string;
|
||||
gmt_create?: string;
|
||||
gmt_modified?: string;
|
||||
/** OK | UNDEPLOYED — UNDEPLOYED vocabularies are silently ignored by ASR. */
|
||||
status?: string;
|
||||
}
|
||||
|
||||
export interface VocabularyEnvelope<T> {
|
||||
request_id?: string;
|
||||
output?: T;
|
||||
usage?: { count?: number };
|
||||
}
|
||||
|
||||
export interface VocabularyRequest {
|
||||
model: string;
|
||||
input: Record<string, unknown>;
|
||||
}
|
||||
|
||||
/** Shared request body builder for dry-run and live calls. */
|
||||
export function buildVocabularyRequest(
|
||||
action: string,
|
||||
input: Record<string, unknown>,
|
||||
): VocabularyRequest {
|
||||
return {
|
||||
model: SPEECH_BIASING_MODEL,
|
||||
input: { action, ...input },
|
||||
};
|
||||
}
|
||||
|
||||
async function callVocabularyApi<T>(
|
||||
client: Client,
|
||||
action: string,
|
||||
input: Record<string, unknown>,
|
||||
signal?: AbortSignal,
|
||||
): Promise<VocabularyEnvelope<T>> {
|
||||
return client.requestJson<VocabularyEnvelope<T>>({
|
||||
path: speechVocabularyPath(),
|
||||
method: "POST",
|
||||
body: buildVocabularyRequest(action, input),
|
||||
signal,
|
||||
});
|
||||
}
|
||||
|
||||
export function createVocabulary(
|
||||
client: Client,
|
||||
params: { targetModel: string; prefix: string; vocabulary: VocabularyEntry[] },
|
||||
signal?: AbortSignal,
|
||||
): Promise<VocabularyEnvelope<{ vocabulary_id?: string }>> {
|
||||
return callVocabularyApi(
|
||||
client,
|
||||
"create_vocabulary",
|
||||
{
|
||||
target_model: params.targetModel,
|
||||
prefix: params.prefix,
|
||||
vocabulary: params.vocabulary,
|
||||
},
|
||||
signal,
|
||||
);
|
||||
}
|
||||
|
||||
export function listVocabularies(
|
||||
client: Client,
|
||||
params: { prefix?: string; pageIndex?: number; pageSize?: number } = {},
|
||||
signal?: AbortSignal,
|
||||
): Promise<VocabularyEnvelope<{ vocabulary_list?: VocabularyListItem[] }>> {
|
||||
const input: Record<string, unknown> = {};
|
||||
if (params.prefix !== undefined) input.prefix = params.prefix;
|
||||
if (params.pageIndex !== undefined) input.page_index = params.pageIndex;
|
||||
if (params.pageSize !== undefined) input.page_size = params.pageSize;
|
||||
return callVocabularyApi(client, "list_vocabulary", input, signal);
|
||||
}
|
||||
|
||||
export function queryVocabulary(
|
||||
client: Client,
|
||||
vocabularyId: string,
|
||||
signal?: AbortSignal,
|
||||
): Promise<
|
||||
VocabularyEnvelope<{
|
||||
gmt_create?: string;
|
||||
gmt_modified?: string;
|
||||
status?: string;
|
||||
target_model?: string;
|
||||
vocabulary?: VocabularyEntry[];
|
||||
}>
|
||||
> {
|
||||
return callVocabularyApi(client, "query_vocabulary", { vocabulary_id: vocabularyId }, signal);
|
||||
}
|
||||
|
||||
export function updateVocabulary(
|
||||
client: Client,
|
||||
vocabularyId: string,
|
||||
vocabulary: VocabularyEntry[],
|
||||
signal?: AbortSignal,
|
||||
): Promise<VocabularyEnvelope<Record<string, never>>> {
|
||||
return callVocabularyApi(
|
||||
client,
|
||||
"update_vocabulary",
|
||||
{ vocabulary_id: vocabularyId, vocabulary },
|
||||
signal,
|
||||
);
|
||||
}
|
||||
|
||||
export function deleteVocabulary(
|
||||
client: Client,
|
||||
vocabularyId: string,
|
||||
signal?: AbortSignal,
|
||||
): Promise<VocabularyEnvelope<Record<string, never>>> {
|
||||
return callVocabularyApi(client, "delete_vocabulary", { vocabulary_id: vocabularyId }, signal);
|
||||
}
|
||||
@@ -580,11 +580,19 @@ export interface DashScopeTTSStreamChunk {
|
||||
|
||||
// ---- Speech Recognition / ASR (DashScope) ----
|
||||
|
||||
/** Context-enhancement message for async ASR `input.context` / sync Flash `input.messages`. */
|
||||
export interface AsrContextMessage {
|
||||
role: "user" | "assistant";
|
||||
content: Array<{ type: "input_text" | "text"; text: string }>;
|
||||
}
|
||||
|
||||
export interface DashScopeASRRequest {
|
||||
model: string;
|
||||
input: {
|
||||
file_urls?: string[];
|
||||
file_url?: string;
|
||||
/** Context enhancement for async filetrans (array of chat-style messages). */
|
||||
context?: AsrContextMessage[];
|
||||
};
|
||||
parameters?: {
|
||||
channel_id?: number[];
|
||||
@@ -595,6 +603,8 @@ export interface DashScopeASRRequest {
|
||||
diarization_enabled?: boolean;
|
||||
speaker_count?: number;
|
||||
vocabulary_id?: string;
|
||||
/** Instant hot words (word → weight); takes effect on Qwen-Audio-3.0-ASR-Flash series. */
|
||||
vocabulary?: Record<string, number>;
|
||||
};
|
||||
}
|
||||
|
||||
|
||||
@@ -49,6 +49,7 @@ export type {
|
||||
ChatRequest,
|
||||
ChatResponse,
|
||||
ChatTool,
|
||||
AsrContextMessage,
|
||||
DashScopeASRRequest,
|
||||
DashScopeASRTaskResult,
|
||||
DashScopeASRTranscriptionItem,
|
||||
|
||||
@@ -34,7 +34,7 @@ export function resolveBooleanFlag(
|
||||
return defaultWhenUnset;
|
||||
}
|
||||
|
||||
/** Resolve `--watermark` flag; default true when unset. */
|
||||
export function resolveWatermark(flagValue: unknown): boolean {
|
||||
return parseOptionalBooleanValue(flagValue, "watermark") ?? true;
|
||||
/** Resolve `--watermark`; command flag overrides config, then defaults to true. */
|
||||
export function resolveWatermark(flagValue: unknown, configuredValue?: boolean): boolean {
|
||||
return parseOptionalBooleanValue(flagValue, "watermark") ?? configuredValue ?? true;
|
||||
}
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
import { expect, test } from "vite-plus/test";
|
||||
import {
|
||||
buildAsrContextMessages,
|
||||
buildAsrFlashRequest,
|
||||
buildAsyncAsrLanguageFields,
|
||||
collectAsrTranscriptionItems,
|
||||
@@ -141,6 +142,47 @@ test("buildAsrFlashRequest shapes qwen3 and input-audio bodies", () => {
|
||||
});
|
||||
});
|
||||
|
||||
test("buildAsrContextMessages wraps plain text as a single user input_text message", () => {
|
||||
expect(buildAsrContextMessages("奋斗者号 鲸落")).toEqual([
|
||||
{ role: "user", content: [{ type: "input_text", text: "奋斗者号 鲸落" }] },
|
||||
]);
|
||||
});
|
||||
|
||||
test("buildAsrFlashRequest injects instant vocabulary and prepends context before input_audio", () => {
|
||||
const body = buildAsrFlashRequest({
|
||||
model: "qwen-audio-3.0-asr-flash",
|
||||
audioUrl: "https://example.com/a.wav",
|
||||
vocabulary: { 奋斗者: 4, 鲸落: 4 },
|
||||
context: "奋斗者号 鲸落",
|
||||
flashFamily: "input-audio",
|
||||
});
|
||||
|
||||
expect(body).toEqual({
|
||||
model: "qwen-audio-3.0-asr-flash",
|
||||
input: {
|
||||
messages: [
|
||||
{
|
||||
role: "user",
|
||||
content: [{ type: "input_text", text: "奋斗者号 鲸落" }],
|
||||
},
|
||||
{
|
||||
role: "user",
|
||||
content: [{ type: "input_audio", input_audio: { data: "https://example.com/a.wav" } }],
|
||||
},
|
||||
],
|
||||
},
|
||||
parameters: {
|
||||
format: "wav",
|
||||
sample_rate: "16000",
|
||||
vocabulary: { 奋斗者: 4, 鲸落: 4 },
|
||||
},
|
||||
});
|
||||
|
||||
const messages = (body.input as { messages: Array<{ content: Array<{ type?: string }> }> })
|
||||
.messages;
|
||||
expect(messages[messages.length - 1]?.content?.[0]?.type).toBe("input_audio");
|
||||
});
|
||||
|
||||
test("buildAsyncAsrLanguageFields maps language by async style", () => {
|
||||
expect(buildAsyncAsrLanguageFields("language_hints", "zh")).toEqual({
|
||||
language_hints: ["zh"],
|
||||
|
||||
@@ -79,6 +79,14 @@ test("default_speech_recognition_model 从配置文件进入运行时 Settings",
|
||||
expect(resolve({ file }).defaultSpeechRecognitionModel).toBe("qwen-audio-3.0-asr-flash");
|
||||
});
|
||||
|
||||
test("watermark 从配置文件进入 Settings,缺省时保持合规默认值 true", () => {
|
||||
expect(parseConfigFile({ watermark: false }).watermark).toBe(false);
|
||||
expect(parseConfigFile({ watermark: true }).watermark).toBe(true);
|
||||
expect(parseConfigFile({ watermark: "false" }).watermark).toBeUndefined();
|
||||
expect(resolve({ file: { watermark: false } }).watermark).toBe(false);
|
||||
expect(resolve({}).watermark).toBe(true);
|
||||
});
|
||||
|
||||
test("baseUrl:flag > env > file > 默认,所有来源统一归一化", () => {
|
||||
const flags = { baseUrl: "https://flag.example.com/compatible-mode/v1?source=flag" };
|
||||
const env = { DASHSCOPE_BASE_URL: "https://env.example.com/apps/anthropic#env" };
|
||||
|
||||
@@ -28,6 +28,7 @@ function makeSettings(configName?: string): Settings {
|
||||
output: "json",
|
||||
outputExplicit: false,
|
||||
timeout: 30,
|
||||
watermark: true,
|
||||
verbose: false,
|
||||
quiet: true,
|
||||
dryRun: false,
|
||||
|
||||
@@ -31,6 +31,7 @@ function testDeps(identity: Partial<Identity> = {}): {
|
||||
output: "json",
|
||||
outputExplicit: true,
|
||||
timeout: 30,
|
||||
watermark: true,
|
||||
verbose: false,
|
||||
quiet: true,
|
||||
dryRun: false,
|
||||
@@ -290,6 +291,9 @@ test("resolveWatermark uses flag or defaults to true", () => {
|
||||
expect(resolveWatermark("false")).toBe(false);
|
||||
expect(resolveWatermark("true")).toBe(true);
|
||||
expect(resolveWatermark(undefined)).toBe(true);
|
||||
expect(resolveWatermark(undefined, false)).toBe(false);
|
||||
expect(resolveWatermark("true", false)).toBe(true);
|
||||
expect(resolveWatermark("false", true)).toBe(false);
|
||||
});
|
||||
|
||||
test("resolveBooleanFlag uses flag or defaultWhenUnset", () => {
|
||||
|
||||
@@ -23,6 +23,7 @@ function testDeps(overrides?: Partial<Settings>): { identity: Identity; settings
|
||||
output: "json",
|
||||
outputExplicit: true,
|
||||
timeout: 5,
|
||||
watermark: true,
|
||||
verbose: false,
|
||||
quiet: true,
|
||||
dryRun: false,
|
||||
|
||||
@@ -0,0 +1,106 @@
|
||||
import { describe, expect, test } from "vite-plus/test";
|
||||
import {
|
||||
parseInstantVocabulary,
|
||||
parseVocabularyEntries,
|
||||
buildVocabularyRequest,
|
||||
SPEECH_BIASING_MODEL,
|
||||
UsageError,
|
||||
} from "../src/index.ts";
|
||||
|
||||
describe("parseInstantVocabulary", () => {
|
||||
test("接受 word→weight 对象", () => {
|
||||
expect(parseInstantVocabulary('{"奋斗者":4,"鲸落":5}')).toEqual({
|
||||
奋斗者: 4,
|
||||
鲸落: 5,
|
||||
});
|
||||
});
|
||||
|
||||
test("拒绝数组", () => {
|
||||
expect(() => parseInstantVocabulary('[{"text":"x","weight":4}]')).toThrow(UsageError);
|
||||
expect(() => parseInstantVocabulary('[{"text":"x","weight":4}]')).toThrow(/object/i);
|
||||
});
|
||||
|
||||
test("拒绝字符串权重", () => {
|
||||
expect(() => parseInstantVocabulary('{"x":"4"}')).toThrow(UsageError);
|
||||
expect(() => parseInstantVocabulary('{"x":"4"}')).toThrow(/must be a number/);
|
||||
});
|
||||
|
||||
test("拒绝非有限数字权重(null)", () => {
|
||||
expect(() => parseInstantVocabulary('{"x":null}')).toThrow(UsageError);
|
||||
expect(() => parseInstantVocabulary('{"x":null}')).toThrow(/must be a number/);
|
||||
});
|
||||
|
||||
test("拒绝 NaN / Infinity 字面量(非法 JSON)", () => {
|
||||
expect(() => parseInstantVocabulary('{"x":NaN}')).toThrow(UsageError);
|
||||
expect(() => parseInstantVocabulary('{"x":Infinity}')).toThrow(UsageError);
|
||||
});
|
||||
|
||||
test("拒绝顶层 null", () => {
|
||||
expect(() => parseInstantVocabulary("null")).toThrow(UsageError);
|
||||
});
|
||||
|
||||
test("非法 JSON 抛 UsageError", () => {
|
||||
expect(() => parseInstantVocabulary("{bad json")).toThrow(UsageError);
|
||||
expect(() => parseInstantVocabulary("{bad json")).toThrow(/not valid JSON/);
|
||||
});
|
||||
});
|
||||
|
||||
describe("parseVocabularyEntries", () => {
|
||||
test("对象形态转为条目数组", () => {
|
||||
expect(parseVocabularyEntries('{"奋斗者":4,"鲸落":5}')).toEqual([
|
||||
{ text: "奋斗者", weight: 4 },
|
||||
{ text: "鲸落", weight: 5 },
|
||||
]);
|
||||
});
|
||||
|
||||
test("对象形态叠加 --lang", () => {
|
||||
expect(parseVocabularyEntries('{"奋斗者":4}', "zh")).toEqual([
|
||||
{ text: "奋斗者", weight: 4, lang: "zh" },
|
||||
]);
|
||||
});
|
||||
|
||||
test("数组形态透传 lang", () => {
|
||||
expect(parseVocabularyEntries('[{"text":"奋斗者","weight":4,"lang":"zh"}]')).toEqual([
|
||||
{ text: "奋斗者", weight: 4, lang: "zh" },
|
||||
]);
|
||||
});
|
||||
|
||||
test("数组形态忽略第二参 lang", () => {
|
||||
expect(parseVocabularyEntries('[{"text":"奋斗者","weight":4}]', "zh")).toEqual([
|
||||
{ text: "奋斗者", weight: 4 },
|
||||
]);
|
||||
expect(parseVocabularyEntries('[{"text":"奋斗者","weight":4,"lang":"en"}]', "zh")).toEqual([
|
||||
{ text: "奋斗者", weight: 4, lang: "en" },
|
||||
]);
|
||||
});
|
||||
|
||||
test("数组缺 text 或 weight 拒绝", () => {
|
||||
expect(() => parseVocabularyEntries('[{"weight":4}]')).toThrow(UsageError);
|
||||
expect(() => parseVocabularyEntries('[{"text":"x"}]')).toThrow(UsageError);
|
||||
});
|
||||
|
||||
test("非 object 数组元素拒绝", () => {
|
||||
expect(() => parseVocabularyEntries('["x"]')).toThrow(UsageError);
|
||||
expect(() => parseVocabularyEntries("[null]")).toThrow(UsageError);
|
||||
});
|
||||
|
||||
test("空对象 / 空数组拒绝", () => {
|
||||
expect(() => parseVocabularyEntries("{}")).toThrow(UsageError);
|
||||
expect(() => parseVocabularyEntries("{}")).toThrow(/at least one/);
|
||||
expect(() => parseVocabularyEntries("[]")).toThrow(UsageError);
|
||||
expect(() => parseVocabularyEntries("[]")).toThrow(/at least one/);
|
||||
});
|
||||
});
|
||||
|
||||
describe("buildVocabularyRequest", () => {
|
||||
test("固定 model 为 speech-biasing", () => {
|
||||
const body = buildVocabularyRequest("create_vocabulary", {
|
||||
target_model: "fun-asr",
|
||||
prefix: "demo",
|
||||
vocabulary: [{ text: "奋斗者", weight: 4 }],
|
||||
});
|
||||
expect(body.model).toBe(SPEECH_BIASING_MODEL);
|
||||
expect(body.model).toBe("speech-biasing");
|
||||
expect(body.input.action).toBe("create_vocabulary");
|
||||
});
|
||||
});
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "knowledge-studio-cli",
|
||||
"version": "1.22.0",
|
||||
"version": "1.23.0",
|
||||
"description": "Lightweight RAG CLI for Aliyun Model Studio — focused on knowledge-base retrieval.",
|
||||
"keywords": [
|
||||
"alibaba-cloud",
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "bailian-cli-runtime",
|
||||
"version": "1.22.0",
|
||||
"version": "1.23.0",
|
||||
"description": "Runtime framework for bailian-cli (createCli, registry, args, output, pipeline). See https://www.npmjs.com/package/bailian-cli for usage.",
|
||||
"homepage": "https://bailian.console.aliyun.com/cli",
|
||||
"bugs": {
|
||||
|
||||
@@ -16,6 +16,22 @@ export interface PipelineEnv {
|
||||
settings: Settings;
|
||||
}
|
||||
|
||||
/** Media steps that inherit Profile `watermark` when the YAML omits the field. */
|
||||
export const PIPELINE_WATERMARK_STEPS = new Set(["image/generate", "image/edit", "video/generate"]);
|
||||
|
||||
/**
|
||||
* Fill Profile watermark into a planned/executed step input.
|
||||
* Explicit step `watermark` wins; otherwise use Settings (file → default true).
|
||||
*/
|
||||
export function applyProfileWatermarkToStepInput(
|
||||
stepType: string,
|
||||
input: Record<string, unknown>,
|
||||
settings: Settings,
|
||||
): Record<string, unknown> {
|
||||
if (!PIPELINE_WATERMARK_STEPS.has(stepType) || input.watermark !== undefined) return input;
|
||||
return { ...input, watermark: settings.watermark };
|
||||
}
|
||||
|
||||
/**
|
||||
* Build the in-process env for pipeline steps. Uses the same source resolution
|
||||
* as the CLI itself (env vars, config file; no CLI flags), but forces JSON
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
import { PipelineError, toPipelineError } from "./errors.ts";
|
||||
import { buildPipelineEnv } from "./bl-config.ts";
|
||||
import { applyProfileWatermarkToStepInput, buildPipelineEnv } from "./bl-config.ts";
|
||||
import { getDefaultStepDispatcher, type StepDispatcher } from "./dispatcher.ts";
|
||||
import {
|
||||
evaluateCondition,
|
||||
@@ -156,12 +156,17 @@ async function executePipelineInternal(
|
||||
if (options.dryRun) {
|
||||
for (const planStep of topologicalOrder(plan)) {
|
||||
const resolved = resolvePlannedStepInput(planStep.step, pipeline, normalizedRuntimeInput);
|
||||
const plannedInput = applyProfileWatermarkToStepInput(
|
||||
planStep.step.type,
|
||||
resolved.redacted,
|
||||
blEnv.settings,
|
||||
);
|
||||
const report: PipelineStepReport = {
|
||||
id: planStep.step.id,
|
||||
type: planStep.step.type,
|
||||
status: "planned",
|
||||
dependencies: planStep.dependencies,
|
||||
input: resolved.redacted,
|
||||
input: plannedInput,
|
||||
...(planStep.step.when !== undefined ? { condition: "pending" } : {}),
|
||||
};
|
||||
reports.push(report);
|
||||
@@ -170,14 +175,14 @@ async function executePipelineInternal(
|
||||
timestamp: now(),
|
||||
status: "planned",
|
||||
step: stepEvent(planStep),
|
||||
input: inputSummary(resolved.redacted, resolved.sensitiveKeys),
|
||||
input: inputSummary(plannedInput, resolved.sensitiveKeys),
|
||||
});
|
||||
await emit({
|
||||
type: "step.planned",
|
||||
timestamp: now(),
|
||||
status: "planned",
|
||||
step: stepEvent(planStep),
|
||||
input: inputSummary(resolved.redacted, resolved.sensitiveKeys),
|
||||
input: inputSummary(plannedInput, resolved.sensitiveKeys),
|
||||
...(planStep.step.when !== undefined ? { condition: "pending" as const } : {}),
|
||||
});
|
||||
}
|
||||
|
||||
@@ -12,6 +12,7 @@ import {
|
||||
speechRecognizePath,
|
||||
resolveAsrApi,
|
||||
buildAsrFlashRequest,
|
||||
buildAsrContextMessages,
|
||||
buildAsyncAsrLanguageFields,
|
||||
collectAsrTranscriptionItems,
|
||||
extractAsrFlashText,
|
||||
@@ -192,7 +193,8 @@ export async function imageGenerate(
|
||||
n,
|
||||
seed: input.seed,
|
||||
prompt_extend: promptExtend,
|
||||
watermark: resolveWatermark(input.watermark),
|
||||
// Step input overrides Profile; omit → use Profile / CLI default (true).
|
||||
watermark: resolveWatermark(input.watermark, env.settings.watermark),
|
||||
};
|
||||
|
||||
const body: DashScopeImageRequest =
|
||||
@@ -299,7 +301,8 @@ export async function imageEdit(
|
||||
n,
|
||||
seed: input.seed,
|
||||
prompt_extend: promptExtend,
|
||||
watermark: resolveWatermark(input.watermark),
|
||||
// Step input overrides Profile; omit → use Profile / CLI default (true).
|
||||
watermark: resolveWatermark(input.watermark, env.settings.watermark),
|
||||
};
|
||||
|
||||
let body: DashScopeImageRequest;
|
||||
@@ -460,7 +463,8 @@ export async function videoGenerate(
|
||||
ratio: input.ratio || undefined,
|
||||
duration: input.duration,
|
||||
prompt_extend: resolveBooleanFlag(input["prompt-extend"], undefined, "prompt-extend"),
|
||||
watermark: resolveWatermark(input.watermark),
|
||||
// Step input overrides Profile; omit → use Profile / CLI default (true).
|
||||
watermark: resolveWatermark(input.watermark, env.settings.watermark),
|
||||
seed: input.seed,
|
||||
},
|
||||
};
|
||||
@@ -563,6 +567,10 @@ export interface SpeechRecognizeInput {
|
||||
diarization?: boolean;
|
||||
"speaker-count"?: number;
|
||||
"vocabulary-id"?: string;
|
||||
/** Instant hot words (already structured; no JSON string parse needed). */
|
||||
vocabulary?: Record<string, number>;
|
||||
/** Context enhancement plain text. */
|
||||
context?: string;
|
||||
"channel-id"?: number;
|
||||
"poll-interval"?: number;
|
||||
}
|
||||
@@ -602,9 +610,11 @@ export async function speechRecognize(
|
||||
const unsupportedFlags: string[] = [];
|
||||
if (input.diarization) unsupportedFlags.push("diarization");
|
||||
if (input["speaker-count"] !== undefined) unsupportedFlags.push("speaker-count");
|
||||
// input-audio Flash supports vocabulary_id; qwen3 sync Flash does not
|
||||
if (route.flashFamily === "qwen3" && input["vocabulary-id"] !== undefined) {
|
||||
unsupportedFlags.push("vocabulary-id");
|
||||
// qwen3 sync Flash has no place for vocabulary_id / vocabulary / context in its body shape
|
||||
if (route.flashFamily === "qwen3") {
|
||||
if (input["vocabulary-id"] !== undefined) unsupportedFlags.push("vocabulary-id");
|
||||
if (input.vocabulary !== undefined) unsupportedFlags.push("vocabulary");
|
||||
if (input.context !== undefined) unsupportedFlags.push("context");
|
||||
}
|
||||
if (input["channel-id"] !== undefined) unsupportedFlags.push("channel-id");
|
||||
if (unsupportedFlags.length > 0) {
|
||||
@@ -648,6 +658,8 @@ export async function speechRecognize(
|
||||
audioUrl: fileUrls[0]!,
|
||||
language: input.language,
|
||||
vocabularyId: input["vocabulary-id"],
|
||||
vocabulary: input.vocabulary,
|
||||
context: input.context,
|
||||
flashFamily,
|
||||
});
|
||||
const response = await env.client.requestJson<Record<string, unknown>>({
|
||||
@@ -671,14 +683,19 @@ export async function speechRecognize(
|
||||
);
|
||||
const body: DashScopeASRRequest = {
|
||||
model,
|
||||
input:
|
||||
route.asyncInputStyle === "file_url" ? { file_url: fileUrls[0]! } : { file_urls: fileUrls },
|
||||
input: {
|
||||
...(route.asyncInputStyle === "file_url"
|
||||
? { file_url: fileUrls[0]! }
|
||||
: { file_urls: fileUrls }),
|
||||
...(input.context ? { context: buildAsrContextMessages(input.context) } : {}),
|
||||
},
|
||||
parameters: {
|
||||
channel_id: input["channel-id"] !== undefined ? [input["channel-id"]] : undefined,
|
||||
...languageFields,
|
||||
diarization_enabled: input.diarization,
|
||||
speaker_count: input["speaker-count"],
|
||||
vocabulary_id: input["vocabulary-id"],
|
||||
vocabulary: input.vocabulary,
|
||||
},
|
||||
};
|
||||
stripUndefined(body.parameters as Record<string, unknown>);
|
||||
|
||||
@@ -1,5 +1,9 @@
|
||||
import { registerStep } from "../dispatcher.ts";
|
||||
import { buildPipelineEnv, type PipelineEnv } from "../bl-config.ts";
|
||||
import {
|
||||
applyProfileWatermarkToStepInput,
|
||||
buildPipelineEnv,
|
||||
type PipelineEnv,
|
||||
} from "../bl-config.ts";
|
||||
import { isRecord } from "../utils.ts";
|
||||
import {
|
||||
textChat,
|
||||
@@ -153,9 +157,13 @@ async function executeDirectBlStep(
|
||||
input: Record<string, unknown>,
|
||||
ctx: StepContext,
|
||||
): Promise<StepResult> {
|
||||
const env = (ctx.blEnv as PipelineEnv | undefined) ?? buildPipelineEnv();
|
||||
|
||||
if (ctx.dryRun) {
|
||||
// Surface the Profile-effective watermark in dry-run so plans match runtime.
|
||||
const plannedInput = applyProfileWatermarkToStepInput(id, input, env.settings);
|
||||
return {
|
||||
metadata: { dryRun: true, step: id, plannedInput: input },
|
||||
metadata: { dryRun: true, step: id, plannedInput },
|
||||
warnings: [
|
||||
{ code: "dry_run_skipped", message: `Step ${id} was not executed in dry-run mode` },
|
||||
],
|
||||
@@ -167,7 +175,6 @@ async function executeDirectBlStep(
|
||||
throw new Error(`No direct API handler registered for step: ${id}`);
|
||||
}
|
||||
|
||||
const env = (ctx.blEnv as PipelineEnv | undefined) ?? buildPipelineEnv();
|
||||
const data = await handler(env, input, ctx);
|
||||
const builder = RESULT_BUILDERS[id];
|
||||
if (builder) {
|
||||
|
||||
@@ -1,8 +1,9 @@
|
||||
/** Shared --foo <bool> help text; keep wording consistent with actual CLI/request behavior. */
|
||||
|
||||
export const BOOL_FLAG_WATERMARK = {
|
||||
"en-US": "Enable watermark (true/false). Omit flag to use CLI default (true).",
|
||||
"zh-CN": "是否启用水印(true/false)。不传时使用 CLI 默认值 true。",
|
||||
"en-US":
|
||||
"Enable watermark (true/false). Omit flag to use the Profile setting or CLI default (true).",
|
||||
"zh-CN": "是否启用水印(true/false)。不传时使用 Profile 配置或 CLI 默认值 true。",
|
||||
};
|
||||
|
||||
/** CLI sends prompt_extend=true when flag omitted (qwen-image edit, etc.). */
|
||||
|
||||
@@ -68,6 +68,81 @@ test("pipeline speechRecognize routes input-audio flash to sync multimodal endpo
|
||||
});
|
||||
});
|
||||
|
||||
test("pipeline speechRecognize injects vocabulary and context on sync input-audio", async () => {
|
||||
const { env, captured } = makeEnv();
|
||||
await speechRecognize(
|
||||
env,
|
||||
{
|
||||
url: "https://example.com/a.wav",
|
||||
model: "qwen-audio-3.0-asr-flash",
|
||||
vocabulary: { 奋斗者: 4 },
|
||||
context: "奋斗者号",
|
||||
},
|
||||
makeCtx(),
|
||||
);
|
||||
|
||||
expect(captured[0]?.body).toMatchObject({
|
||||
parameters: { vocabulary: { 奋斗者: 4 } },
|
||||
input: {
|
||||
messages: [
|
||||
{ role: "user", content: [{ type: "input_text", text: "奋斗者号" }] },
|
||||
{
|
||||
role: "user",
|
||||
content: [{ type: "input_audio", input_audio: { data: "https://example.com/a.wav" } }],
|
||||
},
|
||||
],
|
||||
},
|
||||
});
|
||||
});
|
||||
|
||||
test("pipeline speechRecognize injects vocabulary and context on async filetrans", async () => {
|
||||
const { env, captured } = makeEnv(async (opts) => {
|
||||
if (opts.async || opts.method === "POST") {
|
||||
return { output: { task_id: "task-1", task_status: "PENDING" } };
|
||||
}
|
||||
return {
|
||||
output: { task_id: "task-1", task_status: "SUCCEEDED", results: [] },
|
||||
request_id: "r1",
|
||||
};
|
||||
});
|
||||
|
||||
await speechRecognize(
|
||||
env,
|
||||
{
|
||||
url: "https://example.com/a.wav",
|
||||
model: "qwen-audio-3.0-asr-flash-filetrans",
|
||||
vocabulary: { 鲸落: 4 },
|
||||
context: "鲸落 深海勇士",
|
||||
"poll-interval": 0,
|
||||
},
|
||||
makeCtx(),
|
||||
);
|
||||
|
||||
expect(captured[0]?.body).toMatchObject({
|
||||
input: {
|
||||
file_urls: ["https://example.com/a.wav"],
|
||||
context: [{ role: "user", content: [{ type: "input_text", text: "鲸落 深海勇士" }] }],
|
||||
},
|
||||
parameters: { vocabulary: { 鲸落: 4 } },
|
||||
});
|
||||
});
|
||||
|
||||
test("pipeline speechRecognize rejects vocabulary on qwen3 sync flash", async () => {
|
||||
const { env, captured } = makeEnv();
|
||||
await expect(
|
||||
speechRecognize(
|
||||
env,
|
||||
{
|
||||
url: "https://example.com/a.wav",
|
||||
model: "qwen3-asr-flash",
|
||||
vocabulary: { 奋斗者: 4 },
|
||||
},
|
||||
makeCtx(),
|
||||
),
|
||||
).rejects.toBeInstanceOf(PipelineError);
|
||||
expect(captured).toHaveLength(0);
|
||||
});
|
||||
|
||||
test("pipeline speechRecognize maps qwen3-filetrans language to parameters.language", async () => {
|
||||
const { env, captured } = makeEnv(async (opts) => {
|
||||
if (opts.async || opts.method === "POST") {
|
||||
|
||||
@@ -0,0 +1,129 @@
|
||||
import { expect, test } from "vite-plus/test";
|
||||
import type { Client } from "bailian-cli-core";
|
||||
import { applyProfileWatermarkToStepInput, type PipelineEnv } from "../src/pipeline/bl-config.ts";
|
||||
import { imageEdit, imageGenerate, videoGenerate } from "../src/pipeline/steps/bl-api.ts";
|
||||
import type { StepContext } from "../src/pipeline/types.ts";
|
||||
|
||||
type CapturedRequest = {
|
||||
path?: string;
|
||||
method?: string;
|
||||
body?: {
|
||||
parameters?: { watermark?: boolean };
|
||||
};
|
||||
async?: boolean;
|
||||
};
|
||||
|
||||
function makeEnv(watermark: boolean): {
|
||||
env: PipelineEnv;
|
||||
captured: CapturedRequest[];
|
||||
} {
|
||||
const captured: CapturedRequest[] = [];
|
||||
const client = {
|
||||
uploadFile: async (source: string) => source,
|
||||
requestJson: async (opts: CapturedRequest) => {
|
||||
captured.push(opts);
|
||||
if (opts.async) {
|
||||
return { output: { task_id: "task-wm", task_status: "PENDING" } };
|
||||
}
|
||||
// Sync image response shape (qwen-image / wan2.7-image).
|
||||
return {
|
||||
request_id: "req-wm",
|
||||
output: {
|
||||
choices: [{ message: { content: [{ image: "https://example.com/out.png" }] } }],
|
||||
},
|
||||
};
|
||||
},
|
||||
} as unknown as Client;
|
||||
|
||||
return {
|
||||
env: {
|
||||
client,
|
||||
settings: {
|
||||
quiet: true,
|
||||
output: "json",
|
||||
outputExplicit: true,
|
||||
timeout: 300,
|
||||
watermark,
|
||||
verbose: false,
|
||||
dryRun: false,
|
||||
telemetry: false,
|
||||
} as PipelineEnv["settings"],
|
||||
},
|
||||
captured,
|
||||
};
|
||||
}
|
||||
|
||||
function makeCtx(): StepContext {
|
||||
return { dryRun: false, signal: new AbortController().signal, timeoutSeconds: 1 };
|
||||
}
|
||||
|
||||
test("pipeline imageGenerate inherits Profile watermark=false when step omits it", async () => {
|
||||
const { env, captured } = makeEnv(false);
|
||||
await imageGenerate(env, { prompt: "A cat" }, makeCtx());
|
||||
expect(captured[0]?.body?.parameters?.watermark).toBe(false);
|
||||
});
|
||||
|
||||
test("pipeline imageGenerate keeps step watermark over Profile", async () => {
|
||||
const { env, captured } = makeEnv(false);
|
||||
await imageGenerate(env, { prompt: "A cat", watermark: true }, makeCtx());
|
||||
expect(captured[0]?.body?.parameters?.watermark).toBe(true);
|
||||
});
|
||||
|
||||
test("pipeline imageEdit inherits Profile watermark=false when step omits it", async () => {
|
||||
const { env, captured } = makeEnv(false);
|
||||
await imageEdit(env, { prompt: "Blue sky", image: "https://example.com/in.png" }, makeCtx());
|
||||
expect(captured[0]?.body?.parameters?.watermark).toBe(false);
|
||||
});
|
||||
|
||||
test("pipeline videoGenerate inherits Profile watermark=false when step omits it", async () => {
|
||||
const { env, captured } = makeEnv(false);
|
||||
// First call submits async task; pollTaskWithOptions will keep polling — mock SUCCEEDED quickly.
|
||||
const client = env.client as unknown as {
|
||||
requestJson: (opts: CapturedRequest) => Promise<unknown>;
|
||||
};
|
||||
let calls = 0;
|
||||
client.requestJson = async (opts: CapturedRequest) => {
|
||||
captured.push(opts);
|
||||
calls += 1;
|
||||
if (calls === 1) {
|
||||
return { output: { task_id: "task-wm", task_status: "PENDING" } };
|
||||
}
|
||||
return {
|
||||
output: {
|
||||
task_id: "task-wm",
|
||||
task_status: "SUCCEEDED",
|
||||
video_url: "https://example.com/out.mp4",
|
||||
},
|
||||
};
|
||||
};
|
||||
|
||||
await videoGenerate(
|
||||
env,
|
||||
{ prompt: "A cat walks", "poll-interval": 0 },
|
||||
{ ...makeCtx(), timeoutSeconds: 5 },
|
||||
);
|
||||
expect(captured[0]?.body?.parameters?.watermark).toBe(false);
|
||||
});
|
||||
|
||||
test("pipeline defaults watermark to true when Profile leaves the compliance default", async () => {
|
||||
const { env, captured } = makeEnv(true);
|
||||
await imageGenerate(env, { prompt: "A cat" }, makeCtx());
|
||||
expect(captured[0]?.body?.parameters?.watermark).toBe(true);
|
||||
});
|
||||
|
||||
test("applyProfileWatermarkToStepInput fills omitted watermark from Settings", () => {
|
||||
const settings = { watermark: false } as PipelineEnv["settings"];
|
||||
expect(
|
||||
applyProfileWatermarkToStepInput("image/generate", { prompt: "A cat" }, settings).watermark,
|
||||
).toBe(false);
|
||||
expect(
|
||||
applyProfileWatermarkToStepInput(
|
||||
"image/generate",
|
||||
{ prompt: "A cat", watermark: true },
|
||||
settings,
|
||||
).watermark,
|
||||
).toBe(true);
|
||||
expect(
|
||||
applyProfileWatermarkToStepInput("text/chat", { message: "hi" }, settings).watermark,
|
||||
).toBeUndefined();
|
||||
});
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
name: bailian-cli
|
||||
metadata:
|
||||
version: "1.22.0"
|
||||
version: "1.23.0"
|
||||
requires:
|
||||
bins: ["bl"]
|
||||
description: >-
|
||||
|
||||
@@ -88,10 +88,10 @@ bl config list --output json
|
||||
|
||||
#### Flags
|
||||
|
||||
| Flag | Type | Required | Description |
|
||||
| ----------------- | ------ | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `--key <key>` | string | yes | Config key (language, base*url, output, output_dir, timeout, api_key, api_key_capabilities, access_token, access_key_id, access_key_secret, security_token, default*\*\_model, workspace_id) |
|
||||
| `--value <value>` | string | yes | Value to set |
|
||||
| Flag | Type | Required | Description |
|
||||
| ----------------- | ------ | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `--key <key>` | string | yes | Config key (language, base*url, output, output_dir, timeout, watermark, api_key, api_key_capabilities, access_token, access_key_id, access_key_secret, security_token, default*\*\_model, workspace_id) |
|
||||
| `--value <value>` | string | yes | Value to set |
|
||||
|
||||
#### Examples
|
||||
|
||||
@@ -107,6 +107,10 @@ bl config set --key output --value json
|
||||
bl config set --key timeout --value 600
|
||||
```
|
||||
|
||||
```bash
|
||||
bl config set --key watermark --value false
|
||||
```
|
||||
|
||||
```bash
|
||||
bl config set --key base_url --value https://dashscope.aliyuncs.com
|
||||
```
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
name: bailian-finetune
|
||||
metadata:
|
||||
version: "1.22.0"
|
||||
version: "1.23.0"
|
||||
requires:
|
||||
bins: ["bl"]
|
||||
description: >-
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
name: bailian-gen
|
||||
metadata:
|
||||
version: "1.22.0"
|
||||
version: "1.23.0"
|
||||
requires:
|
||||
bins: ["bl"]
|
||||
description: >-
|
||||
@@ -45,8 +45,27 @@ Unless the user explicitly specifies a model, omit `--model` and let the CLI use
|
||||
|
||||
For ASR model selection, keep `fun-asr` (or other `*-filetrans`) for long recordings, repeated files, speaker diarization, or asynchronous task IDs. For one local or remote audio file up to about five minutes when the user asks for low-latency Flash models, use `--model fun-asr-flash-2026-06-15`, `--model qwen-audio-3.0-asr-flash`, or `--model qwen3-asr-flash`. Flash recognition is synchronous and accepts exactly one file per call.
|
||||
|
||||
To improve ASR accuracy with domain terms:
|
||||
|
||||
- Prefer instant `--vocabulary` / `--context` on `bl speech recognize` when the model is Qwen-Audio-3.0-ASR-Flash series (and Fun-ASR-Flash for `--context` only) — no pre-built vocabulary needed. Start weights at 4 (do not default everything to 5). `--context` must list the target words themselves; a topic description alone has little effect.
|
||||
- Use `bl speech vocabulary create` + `--vocabulary-id` for Fun-ASR / Paraformer, or whenever the same hot words must be reused across requests. The vocabulary `--model` must exactly match recognize `--model` (otherwise the vocabulary is silently ignored). Each account may have at most 10 vocabularies; delete unused ones.
|
||||
|
||||
Flags, usage, and examples: see [`reference/`](reference/index.md) or `bl <command> --help` — do not guess flags.
|
||||
|
||||
## Watermark configuration
|
||||
|
||||
Image generation and editing, video generation and editing, and reference-to-video enable watermarks by default. Change the default for the active Profile with:
|
||||
|
||||
```bash
|
||||
bl config set --key watermark --value false
|
||||
```
|
||||
|
||||
Set it to `true` to enable watermarks again. To update a named Profile without switching Profiles, add `--config <name>`:
|
||||
|
||||
```bash
|
||||
bl config set --config media --key watermark --value false
|
||||
```
|
||||
|
||||
## Local files (mandatory)
|
||||
|
||||
Any command that accepts a **file URL** also accepts a **local path**; the CLI uploads to DashScope temporary storage (`oss://`, 48h) automatically. If the user gives a local file, pass the path directly — never ask them to upload or host a URL first.
|
||||
@@ -68,6 +87,11 @@ bl vision describe --image ./photo.jpg --prompt "图里有什么?"
|
||||
bl vision describe --video ./clip.mp4 --prompt "总结视频内容"
|
||||
bl omni --message "Describe the video content" --video ./demo.mp4 --text-only
|
||||
bl speech synthesize --text "Hello, welcome to Bailian" --out hello.mp3
|
||||
bl speech recognize --url ./meeting.wav --model qwen-audio-3.0-asr-flash-filetrans \
|
||||
--vocabulary '{"奋斗者":4}' --context "奋斗者号"
|
||||
VOCAB=$(bl speech vocabulary create --model fun-asr --prefix demo --words '{"奋斗者":4}' --quiet)
|
||||
bl speech recognize --url ./meeting.wav --model fun-asr --vocabulary-id "$VOCAB"
|
||||
bl speech vocabulary delete --id "$VOCAB" --yes
|
||||
```
|
||||
|
||||
## Output language
|
||||
|
||||
@@ -36,7 +36,7 @@ Index: [index.md](index.md)
|
||||
| `--negative-prompt <text>` | string | no | Negative prompt to exclude unwanted content |
|
||||
| `--function <name>` | string | no | wanx\*-imageedit function (default: description_edit). Examples: stylization_all, description_edit |
|
||||
| `--prompt-extend <bool>` | boolean | no | Enable prompt extend (true/false). Omit flag to use CLI default (true). |
|
||||
| `--watermark <bool>` | boolean | no | Enable watermark (true/false). Omit flag to use CLI default (true). |
|
||||
| `--watermark <bool>` | boolean | no | Enable watermark (true/false). Omit flag to use the Profile setting or CLI default (true). |
|
||||
| `--out-dir <dir>` | string | no | Download images to directory |
|
||||
| `--out-prefix <prefix>` | string | no | Filename prefix (default: edited) |
|
||||
| `--async` | switch | no | Return async task id without waiting |
|
||||
@@ -99,7 +99,7 @@ bl image edit --image ./photo.png --prompt "Replace the background with a beach"
|
||||
| `--seed <n>` | number | no | Random seed for reproducible generation |
|
||||
| `--negative-prompt <text>` | string | no | Negative prompt to exclude unwanted content |
|
||||
| `--prompt-extend <bool>` | boolean | no | Enable prompt extend (true/false). Omit flag: true for qwen-image sync; parameter omitted on async models (API default). |
|
||||
| `--watermark <bool>` | boolean | no | Enable watermark (true/false). Omit flag to use CLI default (true). |
|
||||
| `--watermark <bool>` | boolean | no | Enable watermark (true/false). Omit flag to use the Profile setting or CLI default (true). |
|
||||
| `--async` | switch | no | Return async task id without waiting |
|
||||
| `--concurrent <n>` | number | no | Run N parallel requests (default: 1) |
|
||||
| `--out-dir <dir>` | string | no | Download images to directory |
|
||||
|
||||
@@ -9,29 +9,34 @@ Use this index for the skill-scoped quick index and global flags.
|
||||
|
||||
## Quick index
|
||||
|
||||
| Command | Authentication | Description | Detail |
|
||||
| ---------------------- | -------------- | -------------------------------------------------------------------------------------------------------------------- | ---------------------- |
|
||||
| `bl image edit` | API Key | Edit an existing image with text instructions (Qwen-Image / Wan 2.7) | [image.md](image.md) |
|
||||
| `bl image generate` | API Key | Generate images (Qwen-Image / wan2.x) | [image.md](image.md) |
|
||||
| `bl omni` | API Key | Multimodal chat with text + audio output (Qwen-Omni) | [omni.md](omni.md) |
|
||||
| `bl speech recognize` | API Key | Recognize speech from audio files (FunAudio-ASR / Qwen-ASR Flash) | [speech.md](speech.md) |
|
||||
| `bl speech synthesize` | API Key | Synthesize speech from text | [speech.md](speech.md) |
|
||||
| `bl video download` | API Key | Download a completed video by task ID | [video.md](video.md) |
|
||||
| `bl video edit` | API Key | Edit a video with happyhorse-1.0-video-edit (style transfer, object replacement, etc.) | [video.md](video.md) |
|
||||
| `bl video generate` | API Key | Generate a video from text or image (wan3.0-video / wan2.6-t2v / happyhorse-1.1-i2v) | [video.md](video.md) |
|
||||
| `bl video ref` | API Key | Reference-to-video generation (wan3.0-video / happyhorse-1.1-r2v / wan2.6-r2v): multi-subject, multi-shot with voice | [video.md](video.md) |
|
||||
| `bl video task get` | API Key | Query async task status | [video.md](video.md) |
|
||||
| `bl vision describe` | API Key | Describe an image or video using Qwen-VL | [vision.md](vision.md) |
|
||||
| Command | Authentication | Description | Detail |
|
||||
| ----------------------------- | -------------- | -------------------------------------------------------------------------------------------------------------------- | ---------------------- |
|
||||
| `bl image edit` | API Key | Edit an existing image with text instructions (Qwen-Image / Wan 2.7) | [image.md](image.md) |
|
||||
| `bl image generate` | API Key | Generate images (Qwen-Image / wan2.x) | [image.md](image.md) |
|
||||
| `bl omni` | API Key | Multimodal chat with text + audio output (Qwen-Omni) | [omni.md](omni.md) |
|
||||
| `bl speech recognize` | API Key | Recognize speech from audio files (FunAudio-ASR / Qwen-ASR Flash) | [speech.md](speech.md) |
|
||||
| `bl speech synthesize` | API Key | Synthesize speech from text | [speech.md](speech.md) |
|
||||
| `bl speech vocabulary create` | API Key | Create a precompiled hot-word vocabulary for ASR | [speech.md](speech.md) |
|
||||
| `bl speech vocabulary delete` | API Key | Delete a precompiled hot-word vocabulary | [speech.md](speech.md) |
|
||||
| `bl speech vocabulary get` | API Key | Get details of a precompiled hot-word vocabulary | [speech.md](speech.md) |
|
||||
| `bl speech vocabulary list` | API Key | List precompiled hot-word vocabularies | [speech.md](speech.md) |
|
||||
| `bl speech vocabulary update` | API Key | Replace the contents of a precompiled hot-word vocabulary | [speech.md](speech.md) |
|
||||
| `bl video download` | API Key | Download a completed video by task ID | [video.md](video.md) |
|
||||
| `bl video edit` | API Key | Edit a video with happyhorse-1.0-video-edit (style transfer, object replacement, etc.) | [video.md](video.md) |
|
||||
| `bl video generate` | API Key | Generate a video from text or image (wan3.0-video / wan2.6-t2v / happyhorse-1.1-i2v) | [video.md](video.md) |
|
||||
| `bl video ref` | API Key | Reference-to-video generation (wan3.0-video / happyhorse-1.1-r2v / wan2.6-r2v): multi-subject, multi-shot with voice | [video.md](video.md) |
|
||||
| `bl video task get` | API Key | Query async task status | [video.md](video.md) |
|
||||
| `bl vision describe` | API Key | Describe an image or video using Qwen-VL | [vision.md](vision.md) |
|
||||
|
||||
## By group
|
||||
|
||||
| Group | Commands | Reference |
|
||||
| -------- | ------------------------------------------------- | ---------------------- |
|
||||
| `image` | `edit`, `generate` | [image.md](image.md) |
|
||||
| `omni` | `(root)` | [omni.md](omni.md) |
|
||||
| `speech` | `recognize`, `synthesize` | [speech.md](speech.md) |
|
||||
| `video` | `download`, `edit`, `generate`, `ref`, `task get` | [video.md](video.md) |
|
||||
| `vision` | `describe` | [vision.md](vision.md) |
|
||||
| Group | Commands | Reference |
|
||||
| -------- | ----------------------------------------------------------------------------------------------------------------------------- | ---------------------- |
|
||||
| `image` | `edit`, `generate` | [image.md](image.md) |
|
||||
| `omni` | `(root)` | [omni.md](omni.md) |
|
||||
| `speech` | `recognize`, `synthesize`, `vocabulary create`, `vocabulary delete`, `vocabulary get`, `vocabulary list`, `vocabulary update` | [speech.md](speech.md) |
|
||||
| `video` | `download`, `edit`, `generate`, `ref`, `task get` | [video.md](video.md) |
|
||||
| `vision` | `describe` | [vision.md](vision.md) |
|
||||
|
||||
## Global flags
|
||||
|
||||
|
||||
@@ -7,10 +7,15 @@ Index: [index.md](index.md)
|
||||
|
||||
## Commands in this group
|
||||
|
||||
| Command | Authentication | Description |
|
||||
| ---------------------- | -------------- | ----------------------------------------------------------------- |
|
||||
| `bl speech recognize` | API Key | Recognize speech from audio files (FunAudio-ASR / Qwen-ASR Flash) |
|
||||
| `bl speech synthesize` | API Key | Synthesize speech from text |
|
||||
| Command | Authentication | Description |
|
||||
| ----------------------------- | -------------- | ----------------------------------------------------------------- |
|
||||
| `bl speech recognize` | API Key | Recognize speech from audio files (FunAudio-ASR / Qwen-ASR Flash) |
|
||||
| `bl speech synthesize` | API Key | Synthesize speech from text |
|
||||
| `bl speech vocabulary create` | API Key | Create a precompiled hot-word vocabulary for ASR |
|
||||
| `bl speech vocabulary delete` | API Key | Delete a precompiled hot-word vocabulary |
|
||||
| `bl speech vocabulary get` | API Key | Get details of a precompiled hot-word vocabulary |
|
||||
| `bl speech vocabulary list` | API Key | List precompiled hot-word vocabularies |
|
||||
| `bl speech vocabulary update` | API Key | Replace the contents of a precompiled hot-word vocabulary |
|
||||
|
||||
## Command details
|
||||
|
||||
@@ -25,20 +30,22 @@ Index: [index.md](index.md)
|
||||
|
||||
#### Flags
|
||||
|
||||
| Flag | Type | Required | Description |
|
||||
| --------------------------- | ------ | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `--url <url>` | array | yes | Audio file URL or local file path (repeatable, max 100) |
|
||||
| `--model <model>` | string | no | Model ID (default: configured Profile ASR model, otherwise fun-asr). Async: fun-asr / _-filetrans / paraformer-_; sync: qwen3-asr-flash* / fun-asr-flash* / qwen-audio-\*-asr-flash |
|
||||
| `--language <lang>` | string | no | Language hint (e.g. zh, en, ja). Classic async/input-audio: language_hints; qwen3-filetrans: language; qwen3 sync: asr_options.language |
|
||||
| `--diarization` | switch | no | Enable automatic speaker diarization |
|
||||
| `--speaker-count <n>` | number | no | Expected number of speakers (requires --diarization) |
|
||||
| `--vocabulary-id <id>` | string | no | Hot-word vocabulary ID for improved accuracy |
|
||||
| `--channel-id <n>` | number | no | Audio channel ID (default: 0) |
|
||||
| `--out <path>` | string | no | Save full transcription result to JSON file |
|
||||
| `--async` | switch | no | Return async task id without waiting |
|
||||
| `--poll-interval <seconds>` | number | no | Polling interval in seconds (default: 2) |
|
||||
| `--api-key <key>` | string | no | API key |
|
||||
| `--base-url <url>` | string | no | API base URL |
|
||||
| Flag | Type | Required | Description |
|
||||
| --------------------------- | ------ | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `--url <url>` | array | yes | Audio file URL or local file path (repeatable, max 100) |
|
||||
| `--model <model>` | string | no | Model ID (default: configured Profile ASR model, otherwise fun-asr). Async: fun-asr / _-filetrans / paraformer-_; sync: qwen3-asr-flash* / fun-asr-flash* / qwen-audio-\*-asr-flash |
|
||||
| `--language <lang>` | string | no | Language hint (e.g. zh, en, ja). Classic async/input-audio: language_hints; qwen3-filetrans: language; qwen3 sync: asr_options.language |
|
||||
| `--diarization` | switch | no | Enable automatic speaker diarization |
|
||||
| `--speaker-count <n>` | number | no | Expected number of speakers (requires --diarization) |
|
||||
| `--vocabulary-id <id>` | string | no | Pre-built hot-word vocabulary ID (create it via `speech vocabulary create`). Its target_model must exactly match --model, otherwise it is silently ignored. Wider model support than --vocabulary, including Fun-ASR and Paraformer |
|
||||
| `--vocabulary <json>` | string | no | Instant hot words as JSON object of word→weight, e.g. '{"Fendouzhe":4}'. Weight 1-5 (4 recommended; higher values can hurt other words), 50 for super hot word. No pre-built vocabulary needed. Takes effect only on Qwen-Audio-3.0-ASR-Flash models |
|
||||
| `--context <text>` | string | no | Context enhancement word list to improve accuracy on proper nouns; must contain the target words themselves (a topic description alone has little effect); max 400 chars. Takes effect only on Qwen-Audio-3.0-ASR-Flash and Fun-ASR-Flash models |
|
||||
| `--channel-id <n>` | number | no | Audio channel ID (default: 0) |
|
||||
| `--out <path>` | string | no | Save full transcription result to JSON file |
|
||||
| `--async` | switch | no | Return async task id without waiting |
|
||||
| `--poll-interval <seconds>` | number | no | Polling interval in seconds (default: 2) |
|
||||
| `--api-key <key>` | string | no | API key |
|
||||
| `--base-url <url>` | string | no | API base URL |
|
||||
|
||||
#### Examples
|
||||
|
||||
@@ -62,6 +69,14 @@ bl speech recognize --url https://example.com/audio.mp3 --language zh
|
||||
bl speech recognize --url https://example.com/audio.mp3 --vocabulary-id vocab-abc123
|
||||
```
|
||||
|
||||
```bash
|
||||
bl speech recognize --url https://example.com/audio.mp3 --model qwen-audio-3.0-asr-flash-filetrans --vocabulary '{"奋斗者":4,"鲸落":4}'
|
||||
```
|
||||
|
||||
```bash
|
||||
bl speech recognize --url https://example.com/audio.mp3 --model qwen-audio-3.0-asr-flash-filetrans --context "奋斗者号 鲸落 深海勇士"
|
||||
```
|
||||
|
||||
```bash
|
||||
bl speech recognize --url https://example.com/audio.mp3 --out result.json
|
||||
```
|
||||
@@ -148,3 +163,198 @@ bl speech synthesize --text "Hello" --voice <voice_id> --stream | afplay -
|
||||
```bash
|
||||
bl speech synthesize --text "Hello" --voice <voice_id> --stream | ffplay -nodisp -autoexit -f s16le -ar 24000 -ac 1 -
|
||||
```
|
||||
|
||||
### `bl speech vocabulary create`
|
||||
|
||||
| Field | Value |
|
||||
| ------------------ | --------------------------------------------------------------------------------------------------------------- |
|
||||
| **Name** | `speech vocabulary create` |
|
||||
| **Description** | Create a precompiled hot-word vocabulary for ASR |
|
||||
| **Authentication** | API Key |
|
||||
| **Usage** | `bl speech vocabulary create --model <model> --prefix <prefix> (--words <json> \| --words-file <path>) [flags]` |
|
||||
|
||||
#### Flags
|
||||
|
||||
| Flag | Type | Required | Description |
|
||||
| --------------------- | ------ | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `--model <model>` | string | yes | ASR model this vocabulary is built for (required). Must exactly match the --model passed to `speech recognize` later, otherwise the vocabulary is silently ignored |
|
||||
| `--prefix <prefix>` | string | yes | Custom vocabulary prefix (required). Digits and lowercase letters only, max 10 chars |
|
||||
| `--words <json>` | string | no | Hot words as JSON object of word→weight, e.g. '{"Fendouzhe":4}'; or the API entry array for per-entry lang. Weight 1-5 (4 recommended); when the vocabulary target_model is a Qwen-Audio-3.0-ASR-Flash series model, 50 is also allowed as super hot word. Or use --words-file |
|
||||
| `--words-file <path>` | string | no | JSON file with the hot words (use - for stdin) |
|
||||
| `--lang <code>` | string | no | Language code applied to every hot word when using object form (optional; ignored for array form). Paraformer: zh/en/ja/yue/ko/de/fr/ru; Fun-ASR: zh/en/ja |
|
||||
| `--api-key <key>` | string | no | API key |
|
||||
| `--base-url <url>` | string | no | API base URL |
|
||||
|
||||
#### Notes
|
||||
|
||||
- The --model must exactly match the --model used later with `speech recognize --vocabulary-id`; a mismatch causes silent failure with no error.
|
||||
- Each account may have at most 10 vocabularies; updates must be at least 5 minutes apart. See improve-asr-accuracy for full limits.
|
||||
- Hot-word vocabularies are not supported in Singapore sub-workspaces; the server error is passed through as-is.
|
||||
- Weight 1-5 (4 recommended); when the vocabulary target_model is a Qwen-Audio-3.0-ASR-Flash series model, 50 is also allowed as super hot word.
|
||||
|
||||
#### Examples
|
||||
|
||||
```bash
|
||||
bl speech vocabulary create --model fun-asr --prefix demo --words '{"Fendouzhe":4,"Jingluo":4}'
|
||||
```
|
||||
|
||||
```bash
|
||||
bl speech vocabulary create --model paraformer-v2 --prefix demo --words '[{"text":"Fendouzhe","weight":4,"lang":"zh"}]'
|
||||
```
|
||||
|
||||
```bash
|
||||
bl speech vocabulary create --model fun-asr --prefix demo --words '{"Fendouzhe":4}' --lang zh
|
||||
```
|
||||
|
||||
```bash
|
||||
bl speech vocabulary create --model fun-asr --prefix demo --words-file ./hotwords.json
|
||||
```
|
||||
|
||||
```bash
|
||||
bl speech vocabulary create --model fun-asr --prefix demo --words '{"Fendouzhe":4}' --quiet
|
||||
```
|
||||
|
||||
### `bl speech vocabulary delete`
|
||||
|
||||
| Field | Value |
|
||||
| ------------------ | -------------------------------------------------------------------------------- |
|
||||
| **Name** | `speech vocabulary delete` |
|
||||
| **Description** | Delete a precompiled hot-word vocabulary |
|
||||
| **Authentication** | API Key |
|
||||
| **Usage** | `bl speech vocabulary delete --id <id>` |
|
||||
| **Risk** | `high` |
|
||||
| **Risk message** | This permanently deletes the specified hot-word vocabulary and cannot be undone. |
|
||||
|
||||
> **Agent safety:** Never add `--yes` automatically. On `type="requires_confirmation"`, stop and ask for explicit user confirmation of the same action and scope.
|
||||
|
||||
#### Flags
|
||||
|
||||
| Flag | Type | Required | Description |
|
||||
| ------------------ | ------ | -------- | --------------------------------- |
|
||||
| `--id <id>` | string | yes | Hot-word vocabulary ID (required) |
|
||||
| `--yes` | switch | no | Confirm this high-risk operation |
|
||||
| `--api-key <key>` | string | no | API key |
|
||||
| `--base-url <url>` | string | no | API base URL |
|
||||
|
||||
#### Examples
|
||||
|
||||
```bash
|
||||
bl speech vocabulary delete --id vocab-demo-xxx --dry-run
|
||||
```
|
||||
|
||||
```bash
|
||||
# Only after explicit user confirmation:
|
||||
bl speech vocabulary delete --id vocab-demo-xxx --yes
|
||||
```
|
||||
|
||||
### `bl speech vocabulary get`
|
||||
|
||||
| Field | Value |
|
||||
| ------------------ | ------------------------------------------------ |
|
||||
| **Name** | `speech vocabulary get` |
|
||||
| **Description** | Get details of a precompiled hot-word vocabulary |
|
||||
| **Authentication** | API Key |
|
||||
| **Usage** | `bl speech vocabulary get --id <id>` |
|
||||
|
||||
#### Flags
|
||||
|
||||
| Flag | Type | Required | Description |
|
||||
| ------------------ | ------ | -------- | --------------------------------- |
|
||||
| `--id <id>` | string | yes | Hot-word vocabulary ID (required) |
|
||||
| `--api-key <key>` | string | no | API key |
|
||||
| `--base-url <url>` | string | no | API base URL |
|
||||
|
||||
#### Notes
|
||||
|
||||
- Use this command to confirm target_model before calling `speech recognize --vocabulary-id`; a model mismatch causes silent failure.
|
||||
|
||||
#### Examples
|
||||
|
||||
```bash
|
||||
bl speech vocabulary get --id vocab-demo-xxx
|
||||
```
|
||||
|
||||
```bash
|
||||
bl speech vocabulary get --id vocab-demo-xxx --quiet
|
||||
```
|
||||
|
||||
### `bl speech vocabulary list`
|
||||
|
||||
| Field | Value |
|
||||
| ------------------ | ------------------------------------------------------------------------------ |
|
||||
| **Name** | `speech vocabulary list` |
|
||||
| **Description** | List precompiled hot-word vocabularies |
|
||||
| **Authentication** | API Key |
|
||||
| **Usage** | `bl speech vocabulary list [--prefix <prefix>] [--page <n>] [--page-size <n>]` |
|
||||
|
||||
#### Flags
|
||||
|
||||
| Flag | Type | Required | Description |
|
||||
| ------------------- | ------ | -------- | --------------------------------------------------------------------------------- |
|
||||
| `--prefix <prefix>` | string | no | Filter by vocabulary prefix |
|
||||
| `--page <n>` | number | no | Page number, 1-based (default: 1). Mapped to API page_index (0-based) as page - 1 |
|
||||
| `--page-size <n>` | number | no | Results per page (default: 10) |
|
||||
| `--api-key <key>` | string | no | API key |
|
||||
| `--base-url <url>` | string | no | API base URL |
|
||||
|
||||
#### Notes
|
||||
|
||||
- List responses do not include target_model; use `speech vocabulary get` to inspect the model a vocabulary was built for.
|
||||
- Vocabularies with status UNDEPLOYED are silently ignored by ASR.
|
||||
|
||||
#### Examples
|
||||
|
||||
```bash
|
||||
bl speech vocabulary list
|
||||
```
|
||||
|
||||
```bash
|
||||
bl speech vocabulary list --prefix demo
|
||||
```
|
||||
|
||||
```bash
|
||||
bl speech vocabulary list --page 2 --page-size 20
|
||||
```
|
||||
|
||||
### `bl speech vocabulary update`
|
||||
|
||||
| Field | Value |
|
||||
| ------------------ | --------------------------------------------------------------------------------------------------------------- |
|
||||
| **Name** | `speech vocabulary update` |
|
||||
| **Description** | Replace the contents of a precompiled hot-word vocabulary |
|
||||
| **Authentication** | API Key |
|
||||
| **Usage** | `bl speech vocabulary update --id <id> (--words <json> \| --words-file <path>) [flags]` |
|
||||
| **Risk** | `high` |
|
||||
| **Risk message** | This fully replaces all hot words in the vocabulary. Entries not listed will be discarded and cannot be undone. |
|
||||
|
||||
> **Agent safety:** Never add `--yes` automatically. On `type="requires_confirmation"`, stop and ask for explicit user confirmation of the same action and scope.
|
||||
|
||||
#### Flags
|
||||
|
||||
| Flag | Type | Required | Description |
|
||||
| --------------------- | ------ | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `--id <id>` | string | yes | Hot-word vocabulary ID (required) |
|
||||
| `--words <json>` | string | no | Hot words as JSON object of word→weight, e.g. '{"Fendouzhe":4}'; or the API entry array for per-entry lang. Weight 1-5 (4 recommended); when the vocabulary target_model is a Qwen-Audio-3.0-ASR-Flash series model, 50 is also allowed as super hot word. Or use --words-file |
|
||||
| `--words-file <path>` | string | no | JSON file with the hot words (use - for stdin) |
|
||||
| `--lang <code>` | string | no | Language code applied to every hot word when using object form (optional; ignored for array form). Paraformer: zh/en/ja/yue/ko/de/fr/ru; Fun-ASR: zh/en/ja |
|
||||
| `--yes` | switch | no | Confirm this high-risk operation |
|
||||
| `--api-key <key>` | string | no | API key |
|
||||
| `--base-url <url>` | string | no | API base URL |
|
||||
|
||||
#### Notes
|
||||
|
||||
- update is a full replace, not an append. Prefer --dry-run first to preview the complete vocabulary that will be written.
|
||||
- Each account may have at most 10 vocabularies; updates must be at least 5 minutes apart. See improve-asr-accuracy for full limits.
|
||||
- Hot-word vocabularies are not supported in Singapore sub-workspaces; the server error is passed through as-is.
|
||||
- Weight 1-5 (4 recommended); when the vocabulary target_model is a Qwen-Audio-3.0-ASR-Flash series model, 50 is also allowed as super hot word.
|
||||
|
||||
#### Examples
|
||||
|
||||
```bash
|
||||
bl speech vocabulary update --id vocab-demo-xxx --words '{"Fendouzhe":4}' --dry-run
|
||||
```
|
||||
|
||||
```bash
|
||||
# Only after explicit user confirmation:
|
||||
bl speech vocabulary update --id vocab-demo-xxx --words '{"Fendouzhe":4,"Jingluo":4}' --yes
|
||||
```
|
||||
|
||||
@@ -56,26 +56,26 @@ bl video download --task-id 3b256896-xxxx --out video.mp4 --quiet
|
||||
|
||||
#### Flags
|
||||
|
||||
| Flag | Type | Required | Description |
|
||||
| -------------------------------- | ------- | -------- | --------------------------------------------------------------------------------------- |
|
||||
| `--model <model>` | string | no | Model ID (default: happyhorse-1.0-video-edit) |
|
||||
| `--video <url>` | string | yes | Input video URL or local file (mp4/mov, 2-10s) |
|
||||
| `--prompt <text>` | string | no | Edit instruction (e.g. "Convert the scene to a claymation style") |
|
||||
| `--ref-image <url>` | string | no | Reference image URL (up to 4, comma-separated) |
|
||||
| `--negative-prompt <text>` | string | no | Negative prompt to exclude unwanted content |
|
||||
| `--resolution <res>` | string | no | Resolution: 720P or 1080P (default: 1080P) |
|
||||
| `--ratio <ratio>` | string | no | Aspect ratio (16:9, 9:16, 1:1, 4:3, 3:4) |
|
||||
| `--duration <seconds>` | number | no | Output video duration in seconds (2-10) |
|
||||
| `--audio-setting <auto\|origin>` | string | no | Audio: auto (default) or origin (keep original) |
|
||||
| `--prompt-extend <bool>` | boolean | no | Enable prompt extend (true/false). Omit flag to omit the parameter (DashScope default). |
|
||||
| `--watermark <bool>` | boolean | no | Enable watermark (true/false). Omit flag to use CLI default (true). |
|
||||
| `--seed <n>` | number | no | Random seed for reproducible generation |
|
||||
| `--download <path>` | string | no | Save video to file on completion |
|
||||
| `--async` | switch | no | Return async task id without waiting |
|
||||
| `--concurrent <n>` | number | no | Run N parallel requests (default: 1) |
|
||||
| `--poll-interval <seconds>` | number | no | Polling interval when waiting (default: 15) |
|
||||
| `--api-key <key>` | string | no | API key |
|
||||
| `--base-url <url>` | string | no | API base URL |
|
||||
| Flag | Type | Required | Description |
|
||||
| -------------------------------- | ------- | -------- | ------------------------------------------------------------------------------------------ |
|
||||
| `--model <model>` | string | no | Model ID (default: happyhorse-1.0-video-edit) |
|
||||
| `--video <url>` | string | yes | Input video URL or local file (mp4/mov, 2-10s) |
|
||||
| `--prompt <text>` | string | no | Edit instruction (e.g. "Convert the scene to a claymation style") |
|
||||
| `--ref-image <url>` | string | no | Reference image URL (up to 4, comma-separated) |
|
||||
| `--negative-prompt <text>` | string | no | Negative prompt to exclude unwanted content |
|
||||
| `--resolution <res>` | string | no | Resolution: 720P or 1080P (default: 1080P) |
|
||||
| `--ratio <ratio>` | string | no | Aspect ratio (16:9, 9:16, 1:1, 4:3, 3:4) |
|
||||
| `--duration <seconds>` | number | no | Output video duration in seconds (2-10) |
|
||||
| `--audio-setting <auto\|origin>` | string | no | Audio: auto (default) or origin (keep original) |
|
||||
| `--prompt-extend <bool>` | boolean | no | Enable prompt extend (true/false). Omit flag to omit the parameter (DashScope default). |
|
||||
| `--watermark <bool>` | boolean | no | Enable watermark (true/false). Omit flag to use the Profile setting or CLI default (true). |
|
||||
| `--seed <n>` | number | no | Random seed for reproducible generation |
|
||||
| `--download <path>` | string | no | Save video to file on completion |
|
||||
| `--async` | switch | no | Return async task id without waiting |
|
||||
| `--concurrent <n>` | number | no | Run N parallel requests (default: 1) |
|
||||
| `--poll-interval <seconds>` | number | no | Polling interval when waiting (default: 15) |
|
||||
| `--api-key <key>` | string | no | API key |
|
||||
| `--base-url <url>` | string | no | API base URL |
|
||||
|
||||
#### Examples
|
||||
|
||||
@@ -117,7 +117,7 @@ bl video edit --video https://example.com/input.mp4 --prompt "Put clothes on the
|
||||
| `--ratio <ratio>` | string | no | Aspect ratio (e.g. 16:9, 9:16, 1:1) |
|
||||
| `--duration <seconds>` | number | no | Video duration in seconds (default: 5) |
|
||||
| `--prompt-extend <bool>` | boolean | no | Enable prompt extend (true/false). Omit flag to omit the parameter (DashScope default). |
|
||||
| `--watermark <bool>` | boolean | no | Enable watermark (true/false). Omit flag to use CLI default (true). |
|
||||
| `--watermark <bool>` | boolean | no | Enable watermark (true/false). Omit flag to use the Profile setting or CLI default (true). |
|
||||
| `--seed <n>` | number | no | Random seed for reproducible generation |
|
||||
| `--download <path>` | string | no | Save video to file on completion |
|
||||
| `--file <url-or-path>` | string | no | Reference file URL or local path for file-to-video (wan3.0-video only; mutually exclusive with --image/--last-frame) |
|
||||
@@ -172,7 +172,7 @@ bl video generate --prompt "A cat playing with a ball" --watermark false
|
||||
| `--ratio <ratio>` | string | no | Aspect ratio (16:9, 9:16, 1:1) |
|
||||
| `--duration <seconds>` | number | no | Video duration in seconds (default: 5) |
|
||||
| `--prompt-extend <bool>` | boolean | no | Enable prompt extend (true/false). Omit flag to omit the parameter (DashScope default). |
|
||||
| `--watermark <bool>` | boolean | no | Enable watermark (true/false). Omit flag to use CLI default (true). |
|
||||
| `--watermark <bool>` | boolean | no | Enable watermark (true/false). Omit flag to use the Profile setting or CLI default (true). |
|
||||
| `--seed <n>` | number | no | Random seed for reproducible generation |
|
||||
| `--download <path>` | string | no | Save video to file on completion |
|
||||
| `--async` | switch | no | Return async task id without waiting |
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
name: bailian-managed-agent
|
||||
metadata:
|
||||
version: "1.22.0"
|
||||
version: "1.23.0"
|
||||
requires:
|
||||
bins: ["bl"]
|
||||
description: >-
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
name: bailian-protocol
|
||||
metadata:
|
||||
version: "1.22.0"
|
||||
version: "1.23.0"
|
||||
requires:
|
||||
bins: ["bl"]
|
||||
description: >-
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
name: bailian-web-search
|
||||
metadata:
|
||||
version: "1.22.0"
|
||||
version: "1.23.0"
|
||||
requires:
|
||||
bins: ["bl"]
|
||||
description: >-
|
||||
|
||||
Reference in New Issue
Block a user