feat(token-plan): support local images with base64 data URIs

- convert local images to Base64 for Token Plan image and video commands
- preserve the existing OSS upload flow for standard API Key profiles
- use wan2.7-image as the default image model with the sync endpoint
- hide full Base64 image content in dry-run output
- add Token Plan compatibility tests and update related docs
This commit is contained in:
若麒
2026-07-22 09:54:54 +08:00
parent 4ca3e2de80
commit 4c566fd60e
22 changed files with 521 additions and 74 deletions
+23 -7
View File
@@ -80,7 +80,7 @@ CLI 应解析并保存以下配置:
"default_video_model": "happyhorse-1.1-t2v",
"default_image_to_video_model": "happyhorse-1.1-i2v",
"default_reference_to_video_model": "happyhorse-1.1-r2v",
"default_image_model": "qwen-image-2.0"
"default_image_model": "wan2.7-image"
}
}
```
@@ -245,7 +245,7 @@ default_text_model: qwen3.8-max-preview
default_video_model: happyhorse-1.1-t2v
default_image_to_video_model: happyhorse-1.1-i2v
default_reference_to_video_model: happyhorse-1.1-r2v
default_image_model: qwen-image-2.0
default_image_model: wan2.7-image
```
Token Plan Base URL 预设只在登录写入阶段提供最低优先级的缺省值:
@@ -259,7 +259,7 @@ Token Plan Base URL 预设只在登录写入阶段提供最低优先级的缺省
登录成功时应把显式 Base URL 或缺失的预设 Base URL,以及默认模型写入 Profile,使 `config show --config token-plan` 能看到完整配置。环境变量不复制进 Profile。运行时不再合并预设;如果手工删除字段,则按统一的环境变量、配置文件和系统默认值链继续解析。
默认模型采用更简单的固定策略:每次执行 `auth login --config token-plan`,都将 `default_text_model` 重置为 `qwen3.8-max-preview`,将 `default_video_model` 重置为 `happyhorse-1.1-t2v`,将 `default_image_to_video_model` 重置为 `happyhorse-1.1-i2v`,将 `default_reference_to_video_model` 重置为 `happyhorse-1.1-r2v`,将 `default_image_model` 重置为 `qwen-image-2.0`。登录不保留用户之前写入的其他 Profile 默认模型;用户需要临时调用其他 Token Plan 模型时,通过具体模型命令的 `--model` 覆盖,不修改这些内置默认值。
默认模型采用更简单的固定策略:每次执行 `auth login --config token-plan`,都将 `default_text_model` 重置为 `qwen3.8-max-preview`,将 `default_video_model` 重置为 `happyhorse-1.1-t2v`,将 `default_image_to_video_model` 重置为 `happyhorse-1.1-i2v`,将 `default_reference_to_video_model` 重置为 `happyhorse-1.1-r2v`,将 `default_image_model` 重置为 `wan2.7-image`。登录不保留用户之前写入的其他 Profile 默认模型;用户需要临时调用其他 Token Plan 模型时,通过具体模型命令的 `--model` 覆盖,不修改这些内置默认值。
预设建议通过集中 registry 表达,不在 resolver、命令和 Client 中散落名称判断:
@@ -271,7 +271,7 @@ const MODEL_PROFILE_PRESETS = {
defaultVideoModel: "happyhorse-1.1-t2v",
defaultImageToVideoModel: "happyhorse-1.1-i2v",
defaultReferenceToVideoModel: "happyhorse-1.1-r2v",
defaultImageModel: "qwen-image-2.0",
defaultImageModel: "wan2.7-image",
},
};
```
@@ -381,13 +381,29 @@ image: <base_url>/api/v1/services/aigc/.../generation
| 能力 | 默认模型 | 调用方式 |
| -------------- | --------------------- | ---------------------------------- |
| 文本生成和推理 | `qwen3.8-max-preview` | OpenAI Compatible Chat Completions |
| 图片生成和编辑 | `qwen-image-2.0` | DashScope 原生图片接口 |
| 图片生成和编辑 | `wan2.7-image` | DashScope 多模态图片接口 |
| 文生视频 | `happyhorse-1.1-t2v` | DashScope 原生视频接口 |
| 图生视频 | `happyhorse-1.1-i2v` | `bl video generate --image` |
| 参考生视频 | `happyhorse-1.1-r2v` | `bl video ref` |
Token Plan 当前模型快照中还包含其他文本、视觉理解、图片和视频模型,但该列表可能由后端调整。基础接入不维护阻断请求的本地白名单;用户可通过具体模型命令的 `--model` 临时覆盖本次请求,但再次登录时 Profile 默认模型仍重置为内置版本。
### 当前模型与本地图片兼容范围
| 模型 | 图片输入能力 | CLI 本地图片处理 |
| ----------------------------------- | -------------- | ---------------------------------------------------------- |
| `qwen3.8-max-preview` | 视觉理解 | Token Plan 下转换为 Base64 Data URI |
| `qwen3.7-plus` | 视觉理解 | Token Plan 下转换为 Base64 Data URI |
| `qwen3.7-max` | 纯文本 | 不涉及图片上传 |
| `qwen3.6-flash` | 视觉理解 | Token Plan 下转换为 Base64 Data URI |
| `wan2.7-image` / `wan2.7-image-pro` | 图片生成与编辑 | 文生图不需要输入图片;编辑本地图片时转换为 Base64 Data URI |
| `happyhorse-1.1-i2v` | 图生视频 | 首帧本地图片转换为 Base64 Data URI |
| `happyhorse-1.1-t2v` | 文生视频 | 不涉及图片上传 |
| `happyhorse-1.1-r2v` | 参考生视频 | 参考本地图片转换为 Base64 Data URI |
| `deepseek-v4-pro` / `glm-5.2` | 纯文本 | 不涉及图片上传 |
Token Plan 图片兼容只处理官方明确支持 Base64 的图片字段;参考视频和参考音频仍要求可访问 URL。普通 API Key 保持各命令既有行为:图片编辑和视频入口继续使用临时 OSS,视觉理解的小图继续使用原有 Base64 路径。
语音和音频不作为本阶段支持承诺。现有命令仍保持通用实现,但 Token Plan Profile 的验收不包含这些模态。
## 错误处理
@@ -443,7 +459,7 @@ feat(auth): support token-plan API key login
- 使用 Token Plan 预设文本模型验证 API Key。
- 登录验证前不写配置。
- 验证成功后一次写入 API Key、canonical Base URL 和默认模型。
- 每次登录都将默认模型重置为 `qwen3.8-max-preview`、`qwen-image-2.0`、`happyhorse-1.1-t2v`、`happyhorse-1.1-i2v` 和 `happyhorse-1.1-r2v`。
- 每次登录都将默认模型重置为 `qwen3.8-max-preview`、`wan2.7-image`、`happyhorse-1.1-t2v`、`happyhorse-1.1-i2v` 和 `happyhorse-1.1-r2v`。
- 验证失败不留下半配置。
- 补充一个最小 Token Plan 登录 E2E,覆盖命名 Profile 落盘、环境变量不复制、预设 Base URL 物化和默认模型重置;通用 API Key 登录 E2E 继续覆盖成功原子保存和失败不写半配置。
- 该 commit 暂不承诺自动归一化用户显式输入的 SDK Base URL。
@@ -546,7 +562,7 @@ fix(core): normalize model base URLs across all sources
- 显式 Base URL 覆盖预设并经过通用归一化。
- 登录验证失败不写入任何 Token Plan 半配置。
- 文本默认使用 `qwen3.8-max-preview`。
- 图片默认使用 `qwen-image-2.0`。
- 图片默认使用 `wan2.7-image`。
- 视频默认使用 `happyhorse-1.1-t2v`;图生和参考生入口分别使用 `happyhorse-1.1-i2v` 和 `happyhorse-1.1-r2v`。
- 文本、图片和视频均复用现有 `apiKey` Client。
- 管控命令继续使用 OpenAPI AK/SK,不受模型 Profile 影响。
+24 -6
View File
@@ -22,6 +22,7 @@ import {
resolveWatermark,
ASYNC_FLAG,
CONCURRENT_FLAG,
redactDataUri,
} from "bailian-cli-core";
import { poll } from "bailian-cli-runtime";
import { downloadFile } from "bailian-cli-runtime";
@@ -31,10 +32,15 @@ import { resolveImageSize } from "bailian-cli-runtime";
import { join } from "path";
import { BOOL_FLAG_PROMPT_EXTEND_CLI_TRUE, BOOL_FLAG_WATERMARK } from "bailian-cli-runtime";
const SYNC_MODEL_PREFIXES = ["qwen-image-2.0", "qwen-image-max"];
const SYNC_MODEL_PREFIXES = ["qwen-image-2.0", "qwen-image-max", "wan2.7-image"];
const PROMPT_EXTEND_DEFAULT_PREFIXES = ["qwen-image-2.0", "qwen-image-max"];
function isSyncModel(model: string): boolean {
return SYNC_MODEL_PREFIXES.some((p) => model.startsWith(p));
return SYNC_MODEL_PREFIXES.some((prefix) => model.startsWith(prefix));
}
function enablesPromptExtendByDefault(model: string): boolean {
return PROMPT_EXTEND_DEFAULT_PREFIXES.some((prefix) => model.startsWith(prefix));
}
const EDIT_FLAGS = {
@@ -98,7 +104,7 @@ const EDIT_FLAGS = {
type EditFlags = ParsedFlags<typeof EDIT_FLAGS>;
export default defineCommand({
description: "Edit an existing image with text instructions (Qwen-Image)",
description: "Edit an existing image with text instructions (Qwen-Image / Wan 2.7)",
auth: "apiKey",
usageArgs: "--image <url> --prompt <text> [flags]",
flags: EDIT_FLAGS,
@@ -107,6 +113,7 @@ export default defineCommand({
'--image https://example.com/logo.png --prompt "Change color to blue" --n 3',
'--image ./a.png --image ./b.png --prompt "Merge two images into one collage"',
'--image https://example.com/photo.png --prompt "Remove the person" --model qwen-image-2.0-pro',
'--image ./photo.png --prompt "Change the style" --model wan2.7-image',
'--image ./photo.png --prompt "Replace the background with a beach" --watermark false',
],
async run(ctx) {
@@ -125,13 +132,13 @@ export default defineCommand({
// Auto-upload local files (resolve all images in parallel)
const resolvedImages = await Promise.all(
rawImages.map((img) => ctx.client.uploadFile(img, model)),
rawImages.map((image) => ctx.client.resolveImageInput(image, model)),
);
const n = flags.n ?? 1;
const promptExtend = resolveBooleanFlag(
flags.promptExtend,
useSync ? true : undefined,
enablesPromptExtendByDefault(model) ? true : undefined,
"prompt-extend",
);
@@ -169,7 +176,18 @@ export default defineCommand({
const format = detectOutputFormat(settings.output);
if (settings.dryRun) {
emitResult({ request: body, mode: useSync ? "sync" : "async" }, format);
const previewBody = {
...body,
input: {
messages: body.input.messages.map((message) => ({
...message,
content: message.content.map((item) =>
item.image ? { ...item, image: redactDataUri(item.image) } : item,
),
})),
},
};
emitResult({ request: previewBody, mode: useSync ? "sync" : "async" }, format);
return;
}
@@ -31,11 +31,16 @@ import { BOOL_FLAG_PROMPT_EXTEND_IMAGE_GENERATE, BOOL_FLAG_WATERMARK } from "bai
import { join } from "path";
// qwen-image-2.0 series uses the sync multimodal-generation endpoint
const SYNC_MODEL_PREFIXES = ["qwen-image-2.0", "qwen-image-max"];
// Qwen-Image 2.0 and Wan 2.7 use the sync multimodal-generation endpoint.
const SYNC_MODEL_PREFIXES = ["qwen-image-2.0", "qwen-image-max", "wan2.7-image"];
const PROMPT_EXTEND_DEFAULT_PREFIXES = ["qwen-image-2.0", "qwen-image-max"];
function isSyncModel(model: string): boolean {
return SYNC_MODEL_PREFIXES.some((p) => model.startsWith(p));
return SYNC_MODEL_PREFIXES.some((prefix) => model.startsWith(prefix));
}
function enablesPromptExtendByDefault(model: string): boolean {
return PROMPT_EXTEND_DEFAULT_PREFIXES.some((prefix) => model.startsWith(prefix));
}
const GENERATE_FLAGS = {
@@ -121,7 +126,7 @@ export default defineCommand({
const promptExtend = resolveBooleanFlag(
flags.promptExtend,
useSync ? true : undefined,
enablesPromptExtendByDefault(model) ? true : undefined,
"prompt-extend",
);
@@ -13,6 +13,7 @@ import {
resolveWatermark,
ASYNC_FLAG,
CONCURRENT_FLAG,
redactDataUri,
} from "bailian-cli-core";
import { poll } from "bailian-cli-runtime";
import { downloadFile, formatBytes } from "bailian-cli-runtime";
@@ -113,7 +114,7 @@ export default defineCommand({
// Auto-upload local image file for i2v
let resolvedImageUrl: string | undefined;
if (imageUrl) {
resolvedImageUrl = await ctx.client.uploadFile(imageUrl, model);
resolvedImageUrl = await ctx.client.resolveImageInput(imageUrl, model);
}
const watermark = resolveWatermark(flags.watermark);
@@ -140,7 +141,16 @@ export default defineCommand({
};
if (settings.dryRun) {
emitResult({ request: body }, format);
const previewBody = resolvedImageUrl
? {
...body,
input: {
...body.input,
media: [{ type: "first_frame" as const, url: redactDataUri(resolvedImageUrl) }],
},
}
: body;
emitResult({ request: previewBody }, format);
return;
}
+24 -12
View File
@@ -13,6 +13,7 @@ import {
resolveWatermark,
ASYNC_FLAG,
CONCURRENT_FLAG,
redactDataUri,
} from "bailian-cli-core";
import { poll } from "bailian-cli-runtime";
import { downloadFile, formatBytes } from "bailian-cli-runtime";
@@ -124,16 +125,16 @@ export default defineCommand({
const media: DashScopeVideoRefRequest["input"]["media"] = [];
// Add reference images
for (let i = 0; i < images.length; i++) {
const resolved = await ctx.client.uploadFile(images[i]!, model);
for (let imageIndex = 0; imageIndex < images.length; imageIndex++) {
const resolved = await ctx.client.resolveImageInput(images[imageIndex]!, model);
const entry: DashScopeVideoRefRequest["input"]["media"][number] = {
type: "reference_image",
url: resolved,
};
// Pair voice by position
if (imageVoices[i]) {
const resolvedVoice = await ctx.client.uploadFile(imageVoices[i]!, model);
if (imageVoices[imageIndex]) {
const resolvedVoice = await ctx.client.uploadFile(imageVoices[imageIndex]!, model);
entry.reference_voice = resolvedVoice;
}
@@ -141,16 +142,16 @@ export default defineCommand({
}
// Add reference videos
for (let i = 0; i < refVideos.length; i++) {
const resolved = await ctx.client.uploadFile(refVideos[i]!, model);
for (let videoIndex = 0; videoIndex < refVideos.length; videoIndex++) {
const resolved = await ctx.client.uploadFile(refVideos[videoIndex]!, model);
const entry: DashScopeVideoRefRequest["input"]["media"][number] = {
type: "reference_video",
url: resolved,
};
// Pair voice by position
if (videoVoices[i]) {
const resolvedVoice = await ctx.client.uploadFile(videoVoices[i]!, model);
if (videoVoices[videoIndex]) {
const resolvedVoice = await ctx.client.uploadFile(videoVoices[videoIndex]!, model);
entry.reference_voice = resolvedVoice;
}
@@ -178,7 +179,18 @@ export default defineCommand({
};
if (settings.dryRun) {
emitResult({ request: body }, format);
const previewBody = {
...body,
input: {
...body.input,
media: body.input.media.map((item) => ({
...item,
url: redactDataUri(item.url),
reference_voice: item.reference_voice ? redactDataUri(item.reference_voice) : undefined,
})),
},
};
emitResult({ request: previewBody }, format);
return;
}
@@ -233,11 +245,11 @@ export default defineCommand({
);
const videos: Array<{ taskId: string; videoUrl: string }> = [];
for (let i = 0; i < results.length; i++) {
const result = results[i]!;
for (let resultIndex = 0; resultIndex < results.length; resultIndex++) {
const result = results[resultIndex]!;
const videoUrl =
result.output.video_url || (result.output.results && result.output.results[0]?.url);
if (videoUrl) videos.push({ taskId: taskIds[i]!, videoUrl });
if (videoUrl) videos.push({ taskId: taskIds[resultIndex]!, videoUrl });
}
if (videos.length === 0) {
@@ -8,18 +8,13 @@ import {
BailianError,
ExitCode,
isLocalFile,
imageFileToDataUri,
redactDataUri,
} from "bailian-cli-core";
import { emitResult, emitBare } from "bailian-cli-runtime";
import { readFileSync, existsSync } from "fs";
import { existsSync, statSync } from "fs";
import { extname } from "path";
const IMAGE_MIME_TYPES: Record<string, string> = {
".jpg": "image/jpeg",
".jpeg": "image/jpeg",
".png": "image/png",
".webp": "image/webp",
};
const VIDEO_EXTENSIONS = new Set([".mp4", ".mov", ".avi", ".mkv", ".webm", ".flv", ".wmv"]);
function isVideoInput(input: string): boolean {
@@ -35,18 +30,7 @@ async function toImageUrl(image: string): Promise<string> {
if (image.startsWith("data:")) return image;
if (image.startsWith("http://") || image.startsWith("https://")) return image;
if (image.startsWith("oss://")) return image;
// Local file → data URI (for small files < 10MB, fallback)
if (!existsSync(image)) throw new BailianError(`File not found: ${image}`, ExitCode.USAGE);
const ext = extname(image).toLowerCase();
const mime = IMAGE_MIME_TYPES[ext];
if (!mime)
throw new BailianError(
`Unsupported image format "${ext}". Supported: jpg, jpeg, png, webp`,
ExitCode.USAGE,
);
const buf = readFileSync(image);
return `data:${mime};base64,${buf.toString("base64")}`;
return imageFileToDataUri(image);
}
export default defineCommand({
@@ -86,7 +70,10 @@ export default defineCommand({
const { settings, flags } = ctx;
let image = flags.image;
const videoInputs = flags.video ?? [];
const model = flags.model || "qwen3-vl-plus";
const model =
flags.model ||
(ctx.client.usesTokenPlanEndpoint() ? settings.defaultTextModel : undefined) ||
"qwen3-vl-plus";
// Auto-detect: if --image was given a video file, treat it as --video
if (image && isVideoInput(image)) {
@@ -102,7 +89,14 @@ export default defineCommand({
if (settings.dryRun) {
emitResult(
{ request: { prompt, image, video: videoInputs.length ? videoInputs : undefined, model } },
{
request: {
prompt,
image: image ? redactDataUri(image) : undefined,
video: videoInputs.length ? videoInputs.map(redactDataUri) : undefined,
model,
},
},
format,
);
return;
@@ -132,10 +126,9 @@ export default defineCommand({
let finalImageUrl = imageUrl;
if (isLocalFile(image) && imageUrl.startsWith("data:")) {
const { statSync } = await import("fs");
const fileSize = statSync(image).size;
if (fileSize > 5 * 1024 * 1024) {
finalImageUrl = await ctx.client.uploadFile(image, model);
finalImageUrl = await ctx.client.resolveImageInput(image, model);
}
}
+2 -2
View File
@@ -271,7 +271,7 @@ describe("e2e: auth", () => {
default_video_model: "happyhorse-1.1-t2v",
default_image_to_video_model: "happyhorse-1.1-i2v",
default_reference_to_video_model: "happyhorse-1.1-r2v",
default_image_model: "qwen-image-2.0",
default_image_model: "wan2.7-image",
});
} finally {
await validationServer.close();
@@ -334,7 +334,7 @@ describe("e2e: auth", () => {
default_video_model: "happyhorse-1.1-t2v",
default_image_to_video_model: "happyhorse-1.1-i2v",
default_reference_to_video_model: "happyhorse-1.1-r2v",
default_image_model: "qwen-image-2.0",
default_image_model: "wan2.7-image",
});
expect((config["token-plan"] as Record<string, unknown>).base_url).not.toBe(
validationServer.baseUrl,
@@ -1,5 +1,6 @@
import { describe, expect, test } from "vite-plus/test";
import { join } from "path";
import { writeFileSync } from "node:fs";
import { join } from "node:path";
import {
e2eFixturesDir,
e2eLabelFromMetaUrl,
@@ -46,6 +47,55 @@ describe("e2e: image edit", () => {
expect(data.mode).toBe("async");
expect(data.request?.input?.messages?.length).toBeGreaterThan(0);
});
test("Token Plan 使用 Base64 传入 wan2.7-image 本地图片", async () => {
const configDir = makeE2eOutputDir("image-edit-token-plan-local-image");
writeFileSync(
join(configDir, "config.json"),
JSON.stringify({
"token-plan": {
api_key: "sk-sp-e2e-placeholder",
base_url: "https://token-plan.cn-beijing.maas.aliyuncs.com",
default_image_model: "wan2.7-image",
},
}),
);
const { stdout, stderr, exitCode } = await runCommandE2e(
IMAGE_ROUTES,
[
"image",
"edit",
"--config",
"token-plan",
"--image",
join(e2eFixturesDir, ".smoke-32.png"),
"--prompt",
"改成蓝色",
"--dry-run",
"--output",
"json",
],
{
BAILIAN_CONFIG_DIR: configDir,
DASHSCOPE_API_KEY: "",
DASHSCOPE_BASE_URL: "",
},
);
expect(exitCode, stderr).toBe(0);
const data = parseStdoutJson<{
mode?: string;
request?: {
model?: string;
input?: { messages?: Array<{ content?: Array<{ image?: string }> }> };
};
}>(stdout);
expect(data.mode).toBe("sync");
expect(data.request?.model).toBe("wan2.7-image");
expect(data.request?.input?.messages?.[0]?.content?.[0]?.image).toBe(
"data:image/png;base64,<omitted>",
);
});
});
describe.skipIf(!isBailianE2EMediaEnabled() || !isDashScopeE2EReady())("e2e: image edit", () => {
@@ -1,4 +1,6 @@
import { describe, expect, test } from "vite-plus/test";
import { writeFileSync } from "node:fs";
import { join } from "node:path";
import {
e2eLabelFromMetaUrl,
isBailianE2EMediaEnabled,
@@ -21,6 +23,44 @@ describe("e2e: image generate", () => {
expect(exitCode, stderr).toBe(0);
expect(stderr).toMatch(/generate|--prompt|--model/i);
});
test("Token Plan 默认使用 wan2.7-image 同步接口", async () => {
const configDir = makeE2eOutputDir("image-generate-token-plan-default");
writeFileSync(
join(configDir, "config.json"),
JSON.stringify({
"token-plan": {
api_key: "sk-sp-e2e-placeholder",
base_url: "https://token-plan.cn-beijing.maas.aliyuncs.com",
default_image_model: "wan2.7-image",
},
}),
);
const { stdout, stderr, exitCode } = await runCommandE2e(
IMAGE_ROUTES,
[
"image",
"generate",
"--config",
"token-plan",
"--prompt",
"一只猫",
"--dry-run",
"--output",
"json",
],
{
BAILIAN_CONFIG_DIR: configDir,
DASHSCOPE_API_KEY: "",
DASHSCOPE_BASE_URL: "",
},
);
expect(exitCode, stderr).toBe(0);
const data = parseStdoutJson<{ mode?: string; request?: { model?: string } }>(stdout);
expect(data.mode).toBe("sync");
expect(data.request?.model).toBe("wan2.7-image");
});
});
describe.skipIf(!isBailianE2EMediaEnabled() || !isDashScopeE2EReady())(
@@ -59,6 +59,8 @@ export const VIDEO_ROUTES: E2eRouteExports = {
"video download": "videoDownload",
};
export const VISION_ROUTES: E2eRouteExports = { "vision describe": "visionDescribe" };
export const SPEECH_ROUTES: E2eRouteExports = {
"speech synthesize": "speechSynthesize",
"speech recognize": "speechRecognize",
@@ -69,6 +69,49 @@ describe("e2e: video generate (i2v)", () => {
expect(data.request?.model).toBe("custom-image-to-video-model");
expect(data.request?.input?.media?.[0]?.type).toBe("first_frame");
});
test("Token Plan 图生视频将本地首帧转换为 Base64", async () => {
const configDir = makeE2eOutputDir("video-i2v-token-plan-local-image");
const imagePath = join(configDir, "first-frame.png");
writeFileSync(imagePath, Buffer.from([1, 2, 3]));
writeFileSync(
join(configDir, "config.json"),
JSON.stringify({
"token-plan": {
api_key: "sk-sp-e2e-placeholder",
base_url: "https://token-plan.cn-beijing.maas.aliyuncs.com",
default_image_to_video_model: "happyhorse-1.1-i2v",
},
}),
);
const { stdout, stderr, exitCode } = await runCommandE2e(
VIDEO_ROUTES,
[
"video",
"generate",
"--config",
"token-plan",
"--dry-run",
"--image",
imagePath,
"--prompt",
"让画面动起来",
"--output",
"json",
],
{
BAILIAN_CONFIG_DIR: configDir,
DASHSCOPE_API_KEY: "",
DASHSCOPE_BASE_URL: "",
},
);
expect(exitCode, stderr).toBe(0);
const data = parseStdoutJson<{
request?: { input?: { media?: Array<{ url?: string }> } };
}>(stdout);
expect(data.request?.input?.media?.[0]?.url).toBe("data:image/png;base64,<omitted>");
});
});
describe.skipIf(!isBailianE2EVideoEnabled() || !isDashScopeE2EReady())(
@@ -87,6 +87,49 @@ describe("e2e: video ref (r2v)", () => {
const data = parseStdoutJson<{ request?: { model?: string } }>(stdout);
expect(data.request?.model).toBe("custom-reference-to-video-model");
});
test("Token Plan 参考生视频将本地参考图转换为 Base64", async () => {
const configDir = makeE2eOutputDir("video-r2v-token-plan-local-image");
const imagePath = join(configDir, "reference.png");
writeFileSync(imagePath, Buffer.from([1, 2, 3]));
writeFileSync(
join(configDir, "config.json"),
JSON.stringify({
"token-plan": {
api_key: "sk-sp-e2e-placeholder",
base_url: "https://token-plan.cn-beijing.maas.aliyuncs.com",
default_reference_to_video_model: "happyhorse-1.1-r2v",
},
}),
);
const { stdout, stderr, exitCode } = await runCommandE2e(
VIDEO_ROUTES,
[
"video",
"ref",
"--config",
"token-plan",
"--dry-run",
"--image",
imagePath,
"--prompt",
"Image 1 waves",
"--output",
"json",
],
{
BAILIAN_CONFIG_DIR: configDir,
DASHSCOPE_API_KEY: "",
DASHSCOPE_BASE_URL: "",
},
);
expect(exitCode, stderr).toBe(0);
const data = parseStdoutJson<{
request?: { input?: { media?: Array<{ url?: string }> } };
}>(stdout);
expect(data.request?.input?.media?.[0]?.url).toBe("data:image/png;base64,<omitted>");
});
});
describe.skipIf(!isBailianE2EVideoEnabled() || !isDashScopeE2EReady())(
@@ -0,0 +1,45 @@
import { writeFileSync } from "node:fs";
import { join } from "node:path";
import { describe, expect, test } from "vite-plus/test";
import { makeE2eOutputDir, parseStdoutJson, runCommandE2e } from "./helpers.ts";
import { VISION_ROUTES } from "./topic-routes.ts";
describe("e2e: vision describe", () => {
test("Token Plan 默认使用支持视觉理解的文本模型", async () => {
const configDir = makeE2eOutputDir("vision-describe-token-plan-default");
writeFileSync(
join(configDir, "config.json"),
JSON.stringify({
"token-plan": {
api_key: "sk-sp-e2e-placeholder",
base_url: "https://token-plan.cn-beijing.maas.aliyuncs.com",
default_text_model: "qwen3.8-max-preview",
},
}),
);
const { stdout, stderr, exitCode } = await runCommandE2e(
VISION_ROUTES,
[
"vision",
"describe",
"--config",
"token-plan",
"--image",
"https://example.com/image.png",
"--dry-run",
"--output",
"json",
],
{
BAILIAN_CONFIG_DIR: configDir,
DASHSCOPE_API_KEY: "",
DASHSCOPE_BASE_URL: "",
},
);
expect(exitCode, stderr).toBe(0);
const data = parseStdoutJson<{ request?: { model?: string } }>(stdout);
expect(data.request?.model).toBe("qwen3.8-max-preview");
});
});
+27 -1
View File
@@ -4,7 +4,7 @@ import { BailianError } from "../errors/base.ts";
import { ExitCode } from "../errors/codes.ts";
import { request, requestJson, type HttpDeps, type RequestOpts } from "./http.ts";
import { buildAcsCanonicalQuery, signAcsRequest, type AcsQueryParams } from "./acs.ts";
import { isLocalFile, resolveFileUrl } from "../files/upload.ts";
import { imageFileToDataUri, isLocalFile, resolveFileUrl } from "../files/upload.ts";
import { McpClient } from "./mcp.ts";
import { callConsoleGateway } from "../console/gateway.ts";
import { refreshAccessToken } from "../auth/refresh-token.ts";
@@ -118,6 +118,32 @@ export class Client {
return resolveFileUrl(source, this.requireApi().token, model, opts);
}
/**
* Resolve an image input while keeping Token Plan's upload limitation isolated.
* Token Plan local images are sent as Data URIs; every other connection keeps
* the established temporary OSS upload flow. URLs and existing Data URIs pass through.
*/
resolveImageInput(
source: string,
model: string,
opts: { signal?: AbortSignal } = {},
): Promise<string> {
if (!isLocalFile(source)) return Promise.resolve(source);
if (this.usesTokenPlanEndpoint()) {
return Promise.resolve(imageFileToDataUri(source));
}
return this.uploadFile(source, model, { signal: opts.signal });
}
usesTokenPlanEndpoint(): boolean {
if (this.deps.settings.configName === "token-plan") return true;
try {
return /^token-plan\.[a-z0-9-]+\.maas\.aliyuncs\.com$/i.test(new URL(this.baseUrl).hostname);
} catch {
return false;
}
}
/** Open an MCP client. Accepts a path (prepended with the model baseUrl) or an absolute URL. */
mcp(pathOrUrl: string): McpClient {
const url = /^https?:\/\//.test(pathOrUrl) ? pathOrUrl : this.requireApi().baseUrl + pathOrUrl;
+1 -1
View File
@@ -14,7 +14,7 @@ const MODEL_PROFILE_PRESETS: Readonly<Record<string, ModelProfilePreset>> = {
defaultVideoModel: "happyhorse-1.1-t2v",
defaultImageToVideoModel: "happyhorse-1.1-i2v",
defaultReferenceToVideoModel: "happyhorse-1.1-r2v",
defaultImageModel: "qwen-image-2.0",
defaultImageModel: "wan2.7-image",
},
};
+7 -1
View File
@@ -1 +1,7 @@
export { uploadFile, isLocalFile, resolveFileUrl } from "./upload.ts";
export {
uploadFile,
isLocalFile,
resolveFileUrl,
imageFileToDataUri,
redactDataUri,
} from "./upload.ts";
+44 -1
View File
@@ -6,7 +6,7 @@
* X-DashScope-OssResourceResolve: enable
*/
import { existsSync, readFileSync, statSync } from "fs";
import { basename } from "path";
import { basename, extname } from "path";
import { BailianError } from "../errors/base.ts";
import { ExitCode } from "../errors/codes.ts";
import { trackingHeaders } from "../client/headers.ts";
@@ -112,6 +112,49 @@ export interface UploadOptions {
signal?: AbortSignal;
}
const IMAGE_MIME_TYPES: Readonly<Record<string, string>> = {
".bmp": "image/bmp",
".heic": "image/heic",
".jpe": "image/jpeg",
".jpeg": "image/jpeg",
".jpg": "image/jpeg",
".png": "image/png",
".tif": "image/tiff",
".tiff": "image/tiff",
".webp": "image/webp",
};
/** Encode a local image as a Data URI. */
export function imageFileToDataUri(filePath: string): string {
if (!existsSync(filePath)) {
throw new BailianError(`File not found: ${filePath}`, ExitCode.USAGE);
}
const stat = statSync(filePath);
if (!stat.isFile()) {
throw new BailianError(`Not a file: ${filePath}`, ExitCode.USAGE);
}
const extension = extname(filePath).toLowerCase();
const mimeType = IMAGE_MIME_TYPES[extension];
if (!mimeType) {
throw new BailianError(
`Unsupported image format "${extension || "unknown"}".`,
ExitCode.USAGE,
"Use an image file with a recognized extension.",
);
}
const encoded = readFileSync(filePath).toString("base64");
return `data:${mimeType};base64,${encoded}`;
}
/** Keep dry-run output readable and avoid echoing the complete inline image. */
export function redactDataUri(input: string): string {
const match = /^data:([^;,]+);base64,/i.exec(input);
return match ? `data:${match[1]};base64,<omitted>` : input;
}
/**
* Upload a local file to DashScope temporary storage and return the oss:// URL.
* The URL is valid for 48 hours.
+1 -1
View File
@@ -36,7 +36,7 @@ test("token-plan Profile 预设保持固定", () => {
defaultVideoModel: "happyhorse-1.1-t2v",
defaultImageToVideoModel: "happyhorse-1.1-i2v",
defaultReferenceToVideoModel: "happyhorse-1.1-r2v",
defaultImageModel: "qwen-image-2.0",
defaultImageModel: "wan2.7-image",
});
});
+91
View File
@@ -0,0 +1,91 @@
import { mkdtempSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { afterEach, describe, expect, test } from "vite-plus/test";
import { Client } from "../src/client/client.ts";
import { imageFileToDataUri, redactDataUri } from "../src/files/upload.ts";
import type { Settings } from "../src/config/schema.ts";
const tempDirs: string[] = [];
afterEach(() => {
for (const tempDir of tempDirs.splice(0)) {
rmSync(tempDir, { recursive: true, force: true });
}
});
function makeImage(extension = ".png", content = Buffer.from([1, 2, 3, 4])): string {
const tempDir = mkdtempSync(join(tmpdir(), "bailian-image-input-"));
tempDirs.push(tempDir);
const filePath = join(tempDir, `input${extension}`);
writeFileSync(filePath, content);
return filePath;
}
function makeSettings(configName?: string): Settings {
return {
configName,
output: "json",
outputExplicit: false,
timeout: 30,
verbose: false,
quiet: true,
dryRun: false,
telemetry: false,
};
}
function makeClient(baseUrl: string, configName?: string): Client {
return new Client({
identity: {
binName: "bl",
version: "test",
npmPackage: "bailian-cli",
clientName: "bailian-cli-test",
},
settings: makeSettings(configName),
baseUrl,
});
}
describe("Token Plan image input compatibility", () => {
test("encodes supported local images and redacts previews", () => {
const imagePath = makeImage(".png");
const dataUri = imageFileToDataUri(imagePath);
expect(dataUri).toBe("data:image/png;base64,AQIDBA==");
expect(redactDataUri(dataUri)).toBe("data:image/png;base64,<omitted>");
});
test("rejects files whose image MIME type cannot be inferred", () => {
const imagePath = makeImage(".unknown");
expect(() => imageFileToDataUri(imagePath)).toThrow(/Unsupported image format/);
});
test("uses Data URI for the token-plan profile even through a custom proxy", async () => {
const imagePath = makeImage(".webp");
const client = makeClient("https://proxy.example.com/bailian", "token-plan");
await expect(client.resolveImageInput(imagePath, "happyhorse-1.1-i2v")).resolves.toMatch(
/^data:image\/webp;base64,/,
);
});
test("uses Data URI for an official Token Plan endpoint under any profile name", async () => {
const imagePath = makeImage(".jpg");
const client = makeClient("https://token-plan.ap-southeast-1.maas.aliyuncs.com", "custom-plan");
await expect(client.resolveImageInput(imagePath, "wan2.7-image")).resolves.toMatch(
/^data:image\/jpeg;base64,/,
);
});
test("ordinary endpoints retain the existing upload path", () => {
const imagePath = makeImage(".png");
const client = makeClient("https://dashscope.aliyuncs.com", "default");
expect(() => client.resolveImageInput(imagePath, "wan2.7-image")).toThrow(
/model-domain API key/,
);
});
});
+1 -1
View File
@@ -68,7 +68,7 @@ The built-in `token-plan` profile defaults to:
- Base URL: `https://token-plan.cn-beijing.maas.aliyuncs.com`
- Text model: `qwen3.8-max-preview`
- Image model: `qwen-image-2.0`
- Image model: `wan2.7-image`
- Text-to-video model (`default_video_model`): `happyhorse-1.1-t2v`
- Image-to-video model (`default_image_to_video_model`): `happyhorse-1.1-i2v`
- Reference-to-video model (`default_reference_to_video_model`): `happyhorse-1.1-r2v`
+13 -9
View File
@@ -7,20 +7,20 @@ Index: [index.md](index.md)
## Commands in this group
| Command | Description |
| ------------------- | ---------------------------------------------------------- |
| `bl image edit` | Edit an existing image with text instructions (Qwen-Image) |
| `bl image generate` | Generate images (Qwen-Image / wan2.x) |
| Command | Description |
| ------------------- | -------------------------------------------------------------------- |
| `bl image edit` | Edit an existing image with text instructions (Qwen-Image / Wan 2.7) |
| `bl image generate` | Generate images (Qwen-Image / wan2.x) |
## Command details
### `bl image edit`
| Field | Value |
| --------------- | ---------------------------------------------------------- |
| **Name** | `image edit` |
| **Description** | Edit an existing image with text instructions (Qwen-Image) |
| **Usage** | `bl image edit --image <url> --prompt <text> [flags]` |
| Field | Value |
| --------------- | -------------------------------------------------------------------- |
| **Name** | `image edit` |
| **Description** | Edit an existing image with text instructions (Qwen-Image / Wan 2.7) |
| **Usage** | `bl image edit --image <url> --prompt <text> [flags]` |
#### Flags
@@ -61,6 +61,10 @@ bl image edit --image ./a.png --image ./b.png --prompt "Merge two images into on
bl image edit --image https://example.com/photo.png --prompt "Remove the person" --model qwen-image-2.0-pro
```
```bash
bl image edit --image ./photo.png --prompt "Change the style" --model wan2.7-image
```
```bash
bl image edit --image ./photo.png --prompt "Replace the background with a beach" --watermark false
```
+1 -1
View File
@@ -51,7 +51,7 @@ Use this index for the full quick index and global flags.
| `bl finetune logs` | Fetch training logs for a fine-tune job | [finetune.md](finetune.md) |
| `bl finetune text create` | Create a text model fine-tune job (sft \| sft-lora \| dpo \| dpo-lora \| cpt) | [finetune.md](finetune.md) |
| `bl finetune watch` | Probe a fine-tune job's status (default: single non-blocking fetch). Pass --follow to poll until terminal. | [finetune.md](finetune.md) |
| `bl image edit` | Edit an existing image with text instructions (Qwen-Image) | [image.md](image.md) |
| `bl image edit` | Edit an existing image with text instructions (Qwen-Image / Wan 2.7) | [image.md](image.md) |
| `bl image generate` | Generate images (Qwen-Image / wan2.x) | [image.md](image.md) |
| `bl knowledge chat` | Chat with a Bailian knowledge base (RAG Q&A with streaming) | [knowledge.md](knowledge.md) |
| `bl knowledge retrieve` | Retrieve from a Bailian knowledge base (deprecated, use `search` instead) | [knowledge.md](knowledge.md) |