mirror of
https://github.com/Pika-Labs/Pika-Plugins.git
synced 2026-09-20 17:47:27 +08:00
Sync updated skills from pika-claude-plugin
Update app-sizzle, app-store-screens, build-a-brand, explainer, founder-product-video, kiss-cam, ugc-ads to match latest pika-claude-plugin. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
+102
-25
@@ -4,13 +4,24 @@ description: >
|
||||
Generate cinematic 1080p iOS app teaser videos from real App Store screenshots,
|
||||
with a GPT-image-2 enhancement pass on each selected screen before generation.
|
||||
Output is a beat-driven cinematic teaser built from GPT-enhanced screenshots,
|
||||
ending with the brand logo/icon + "COMING SOON" title card. Screens sourced
|
||||
from Pika MCP App Store fetch, a live website (auto-captured), user-supplied files, or URLs.
|
||||
ending with the brand logo/icon plus a deterministic `COMING SOON` overlay.
|
||||
Screens sourced from Pika MCP App Store fetch, a live website (auto-captured), user-supplied files, or URLs.
|
||||
Starts by sourcing real screens and brand assets before any generation.
|
||||
Triggers on: app sizzle, app teaser, app promo, video promo, app video, product
|
||||
video, coming soon, seedance, motion graphics, make a promo, make a video for
|
||||
[app], gpt enhance promo.
|
||||
Triggers on: app sizzle, app teaser, app promo, iOS app promo video, app video,
|
||||
app product video, coming soon, seedance, motion graphics, make a promo for my
|
||||
app, make a video for [app], gpt enhance promo.
|
||||
NOT for: short-form consumer content like GRWM, vlogs, UGC, or non-app product
|
||||
ads (use content-video); app-sizzle is specifically for iOS app teaser videos
|
||||
sourced from App Store screens or real app UI.
|
||||
argument-hint: <app-name-or-url> [screens=<app-store-url|website-url|paths>] [logo=<path-or-url>] [aspect=16:9|9:16|1:1]
|
||||
required-capabilities:
|
||||
- mcp__pika__capture_website
|
||||
- mcp__pika__fetch_appstore_screens
|
||||
- mcp__pika__generate_image
|
||||
- mcp__pika__generate_reference_video
|
||||
- mcp__pika__edit_text_overlay
|
||||
- mcp__pika__task_status
|
||||
- mcp__pika__upload_asset
|
||||
---
|
||||
|
||||
# App Sizzle — GPT-Image-2 Enhanced iOS App Teaser
|
||||
@@ -25,11 +36,14 @@ The visual aesthetic is **derived from the app's personality** — not defaulted
|
||||
|
||||
## Mode: Reference-to-Video
|
||||
|
||||
Primary: `generate_reference_video(provider="seedance", resolution="1080p")` with 3–5 screenshots + the app icon/logo as the final reference.
|
||||
Primary: `mcp__pika__generate_reference_video(provider="seedance", resolution="1080p")` with 3–5 screenshots + the app icon/logo as the final reference.
|
||||
|
||||
Fallback to `provider="kling", quality_mode="pro"` (= 1080p) when:
|
||||
- Seedance returns `partner_validation_failed` (celebrity faces, screen-recording UI)
|
||||
- Seedance returns non-audio `partner_validation_failed` (celebrity faces, screen-recording UI)
|
||||
- Seedance returns `insufficient_balance`
|
||||
- Seedance stays queued/running until it returns a timeout such as `seedance timed out after ...`
|
||||
|
||||
Do not treat generated-audio moderation as an immediate Kling fallback. See the Seedance generated-audio moderation recovery runbook in Generate Video first.
|
||||
|
||||
Kling prompt uses `<<<image_1>>>` … `<<<image_5>>>` tokens instead of `@Image1` … `@Image5`. Drop the `resolution` param (Kling uses `quality_mode` instead). See Gotchas.
|
||||
|
||||
@@ -47,7 +61,7 @@ To make your app promo, I need:
|
||||
|
||||
2. Where should I pull the app screens from?
|
||||
— iOS App Store: give me the App Store URL or app name → I'll use
|
||||
`fetch_appstore_screens` to fetch screenshots, metadata, and icon
|
||||
`mcp__pika__fetch_appstore_screens` to fetch screenshots, metadata, and icon
|
||||
— Web app / website: give me the URL → I'll capture it with Pika MCP
|
||||
— Local files / URLs: drop the paths and I'll upload them
|
||||
|
||||
@@ -78,7 +92,7 @@ Before calling any generation tool, verify both assets are in hand:
|
||||
| Asset | Required | If missing |
|
||||
|-------|----------|------------|
|
||||
| Real app screenshots (≥1 actual sourced image) | Yes | Stop and ask for screenshots |
|
||||
| Brand logo OR app icon | Yes | Use the `fetch_appstore_screens` icon when App Store sourcing is used; otherwise stop and ask for a logo/icon |
|
||||
| Brand logo OR app icon | Yes | Use the `mcp__pika__fetch_appstore_screens` icon when App Store sourcing is used; otherwise stop and ask for a logo/icon |
|
||||
|
||||
If either is missing, tell the user exactly what's needed and wait. Real assets are what keep the teaser grounded; text-to-video placeholders make Seedance invent UI.
|
||||
|
||||
@@ -98,7 +112,7 @@ what happened and ask the user to provide the screens manually. Never invent the
|
||||
|
||||
### iOS App Store
|
||||
|
||||
Use Pika MCP `fetch_appstore_screens`; do not use a local scraper. It accepts a full App Store URL, numeric app ID, or app-name search term:
|
||||
Use Pika MCP `mcp__pika__fetch_appstore_screens`; do not use a local scraper. It accepts a full App Store URL, numeric app ID, or app-name search term:
|
||||
|
||||
```
|
||||
fetch_appstore_screens(
|
||||
@@ -123,7 +137,7 @@ Expected result shape:
|
||||
}
|
||||
```
|
||||
|
||||
If `fetch_appstore_screens` returns no screenshots, report the error and ask the user for 3-5 real screenshots plus a logo/icon. Do not fall back to Playwright/headless App Store capture and do not invent UI.
|
||||
If `mcp__pika__fetch_appstore_screens` returns no screenshots, report the error and ask the user for 3-5 real screenshots plus a logo/icon. Do not fall back to Playwright/headless App Store capture and do not invent UI.
|
||||
|
||||
After App Store assets are fetched, pick the 3–5 screens that show the core UI. Skip:
|
||||
- Pure text/splash screens (no UI)
|
||||
@@ -158,7 +172,7 @@ For each screenshot, record:
|
||||
- **What feature it represents** — e.g. "creation entry", "agent at work", "output/share"
|
||||
- **Emotional register** — is this the power moment, the ease moment, the aha moment?
|
||||
|
||||
Also pull the app metadata from the `fetch_appstore_screens` result, or from the user-provided description:
|
||||
Also pull the app metadata from the `mcp__pika__fetch_appstore_screens` result, or from the user-provided description:
|
||||
- App name, subtitle, one-line value prop
|
||||
- Category and target user
|
||||
|
||||
@@ -186,7 +200,7 @@ Every 15s promo needs a spine. Design the story arc before touching the prompt t
|
||||
| **Hook** | 0–3s | Grab attention — show the most dramatic UI moment or the problem being solved | The most visually striking screen |
|
||||
| **Build** | 3–10s | Feature walkthrough in logical user-journey order | 2–3 screens in sequence |
|
||||
| **Reveal** | 10–13s | Pull-back or product overview — the "so that's what it does" moment | Wide shot or most complete screen |
|
||||
| **Logo** | 13–15s | Brand lock — wordmark materializes, accent color pulse | Logo (@Image6 or last ref) |
|
||||
| **Logo** | 13–15s | Brand lock — wordmark materializes, accent color pulse | Logo (@Image6 or last ref). `COMING SOON` is added later as a post-generation text overlay. |
|
||||
|
||||
### Story arc types — pick one based on the app
|
||||
|
||||
@@ -205,7 +219,7 @@ Arc type: [Problem→Solution / Feature Parade / Journey / Transformation]
|
||||
Hook (0-3s): Screen [N] — [what happens] — camera: [extreme close-up on X]
|
||||
Build (3-10s): Screen [N] → [N] → [N] — [what each reveals] — camera: [whip pan / orbital / etc.]
|
||||
Reveal (10-13s): Screen [N] — [what it shows] — camera: [pull-back to show full product]
|
||||
Logo (13-15s): @Image[N] — wordmark materializes whole in a burst of [accent color] light. Below it, smaller text "COMING SOON" fades in. (NOT "assembles" — triggers per-glyph hallucination)
|
||||
Logo (13-15s): @Image[N] — wordmark materializes whole in a burst of [accent color] light and holds. Do not ask the video model to render the `COMING SOON` copy; it is added later as a post-generation text overlay.
|
||||
```
|
||||
|
||||
Do NOT write the Seedance prompt until this arc is defined.
|
||||
@@ -254,7 +268,7 @@ This is the proven template. It uses the BEAT structure directly — Seedance re
|
||||
BEAT 1 (Hook, 0–3s): [Camera action] — @Image1 is [exact UI description from Stage 1 feature map, as specific as possible, quoting actual UI text if visible]. [What happens — camera move + how the UI is framed or revealed].
|
||||
BEAT 2 (Build A, 3–8s): [Camera cuts to] — @Image2 is [exact UI description]. [What the beat reveals about the feature — show the output or the moment of delight].
|
||||
BEAT 3 (Build B / Reveal, 8–12s): [Camera sweeps to or pulls back] — @Image3 is [exact UI description]. [What the product overview or transformation moment shows].
|
||||
BEAT 4 (Logo, 12–15s): Hard cut to black — @Image[last] is the [brand] wordmark. It materializes whole in a burst of [accent color] light. Below it, smaller text "COMING SOON" fades in and holds.
|
||||
BEAT 4 (Logo, 12–15s): Hard cut to black — @Image[last] is the [brand] wordmark. It materializes whole in a burst of [accent color] light and holds for the final overlay.
|
||||
|
||||
Style: [aesthetic-specific — e.g. "dark cinematic thriller, self-luminous UI on absolute black, electric blue accent"]. No text, no words rendered in motion.
|
||||
```
|
||||
@@ -277,7 +291,7 @@ Style: [aesthetic-specific — e.g. "dark cinematic thriller, self-luminous UI o
|
||||
BEAT 1 (Hook): [Camera action] on @Image1 — [transformation: choose from vocabulary below].
|
||||
BEAT 2 (Build): [Camera action] cuts to @Image2 — [transformation]. Hard cut to @Image3 — [transformation].
|
||||
BEAT 3 (Reveal): Pull-back reveals [what the full product view shows].
|
||||
BEAT 4 (Logo): Hard cut to black — @Image[last] materializes whole in a burst of [accent color] light. Below, "COMING SOON" fades in.
|
||||
BEAT 4 (Logo): Hard cut to black — @Image[last] materializes whole in a burst of [accent color] light and holds for the final overlay.
|
||||
|
||||
Style: liquid glass morphism, Apple Vision Pro aesthetic, premium 3D depth, self-luminous forms on absolute black, [accent color] accent lighting. No text rendered in motion.
|
||||
```
|
||||
@@ -311,11 +325,31 @@ generate_reference_video(
|
||||
duration=15, # always
|
||||
sound=True, # always
|
||||
aspect_ratio="16:9", # or 9:16 / 1:1 per user request
|
||||
seed=<int>, # optional — use to retry on content policy failures
|
||||
seed=<int>, # set one; reuse it for content-policy recovery
|
||||
)
|
||||
```
|
||||
|
||||
**Fallback — Kling (partner_validation_failed or insufficient_balance):**
|
||||
### Seedance generated-audio moderation recovery
|
||||
|
||||
If Seedance finishes generation and then returns a 422 whose body includes `type: "content_policy_violation"`, `reason: "partner_validation_failed"`, `loc: ["body", "generated_video"]`, and `msg: "Output audio has sensitive content."`, treat it as a recoverable generated-audio moderation false positive.
|
||||
|
||||
1. Retry the exact same prompt and `reference_images` with `sound=False` and the same `seed`.
|
||||
2. If the silent probe succeeds, retry the exact same prompt/reference set with `sound=True` and the same seed.
|
||||
3. If the `sound=True` replay succeeds, route the recovered sound-on URL into Stage 4 as `generated_teaser_url`. Keep the silent probe URL only as debugging context.
|
||||
4. If the silent probe fails, treat the failure as video/reference moderation and use the Kling fallback.
|
||||
5. If the silent probe succeeds but the `sound=True` replay fails again, run the Kling fallback once. If Kling is unavailable, route the silent URL into Stage 4 as `generated_teaser_url` and explicitly note that generated-audio moderation remained flaky.
|
||||
|
||||
Do not change the prompt, references, aspect ratio, duration, or seed during this recovery path. Changing any of them turns the silent probe into a new generation instead of testing whether only generated audio triggered moderation.
|
||||
|
||||
### Seedance timeout recovery
|
||||
|
||||
If task status remains queued or running until Seedance returns a timeout such as `seedance timed out after 900s` or `seedance timed out after 1200s`, treat it as provider queue saturation, not a prompt/content failure.
|
||||
|
||||
When this happens, run the Kling fallback with the same selected references, same beat structure, `duration=15`, `sound=True`, and `quality_mode="pro"`. Convert `@ImageN` prompt tokens to `<<<image_N>>>` before calling Kling.
|
||||
|
||||
Do not keep retrying Seedance after a timeout unless the user explicitly asks to wait for Seedance. The timeout path has already spent the launch-demo wall-clock budget; switching provider is the documented recovery.
|
||||
|
||||
**Fallback — Kling (non-audio partner_validation_failed or insufficient_balance):**
|
||||
```python
|
||||
generate_reference_video(
|
||||
provider="kling",
|
||||
@@ -334,6 +368,18 @@ Seedance constraints: skip `fast=True` because it caps at 720p; skip `negative_p
|
||||
|
||||
Kling constraint: use `quality_mode="pro"` for 1080p; Kling rejects `resolution=`.
|
||||
|
||||
### Kling queued/handoff recovery
|
||||
|
||||
Kling fallback is async. If `generate_reference_video(provider="kling")` returns a `task_id`, follow the task until terminal.
|
||||
|
||||
If `task_status` returns status: `queued` with `statusMessage` containing `Worker handoff: task was requeued for retry on another worker.`, treat it as a worker restart handoff, not a failed render. Keep polling `mcp__pika__task_status(task_id)`; the next worker should reclaim the same task.
|
||||
|
||||
If `statusMessage` starts with `Kling is at capacity`, treat it as provider capacity wait. Keep polling the same task while `lastUpdatedAt` continues moving.
|
||||
|
||||
Do not submit a duplicate Kling request while the original task is still `queued` or `running`. Duplicates can burn provider quota and make artifact provenance unclear.
|
||||
|
||||
If `status` stays `queued` for more than 10 minutes with no `lastUpdatedAt` movement, capture the `task_id`, `status`, `statusMessage`, and `lastUpdatedAt`, then cancel the stalled original with `mcp__pika__task_cancel(task_id)` before retrying. Only after cancel returns `cancelled`, retry the exact same Kling request once with the same prompt, references, shots, aspect ratio, duration, and quality mode. If cancel fails because the task already completed or failed, inspect that terminal result instead of retrying. If the retry also stalls, stop and report both task IDs instead of changing the creative prompt.
|
||||
|
||||
---
|
||||
|
||||
## Asset Upload (local files → public URL)
|
||||
@@ -341,7 +387,7 @@ Kling constraint: use `quality_mode="pro"` for 1080p; Kling rejects `resolution=
|
||||
If the user provides local file paths, convert them to public URLs before calling generate:
|
||||
|
||||
1. Read the file size and MIME type.
|
||||
2. Call `upload_asset(filename, mime_type, size_bytes)`.
|
||||
2. Call `mcp__pika__upload_asset(filename, mime_type, size_bytes)`.
|
||||
3. Upload the bytes to the returned `presigned_url` using the host client's file-upload capability.
|
||||
4. Use the returned `public_url` as the reference URL in generation calls.
|
||||
|
||||
@@ -349,6 +395,34 @@ Supported mime types: `image/png`, `image/jpeg`, `image/webp`, `video/mp4`, `aud
|
||||
|
||||
---
|
||||
|
||||
## Stage 4 — Deterministic COMING SOON Overlay
|
||||
|
||||
Do not ask Seedance or Kling to render `COMING SOON`. Video models garble new
|
||||
typography, especially all-caps CTA text, so the final two seconds use a
|
||||
deterministic `COMING SOON` overlay as a post-generation text overlay.
|
||||
|
||||
After Seedance or Kling returns the 15s teaser URL, call:
|
||||
|
||||
```python
|
||||
edit_text_overlay(
|
||||
video_url=<generated_teaser_url>,
|
||||
text="COMING SOON",
|
||||
position="bottom_center",
|
||||
font_size=56,
|
||||
font_color="white",
|
||||
start_s=13,
|
||||
end_s=15,
|
||||
)
|
||||
```
|
||||
|
||||
If `edit_text_overlay` returns `{ task_id }`, poll `mcp__pika__task_status`
|
||||
until it reaches `completed`, `failed`, or `cancelled`, then unwrap the returned
|
||||
URL. Save the returned URL as `final_url`. If the overlay call fails, surface
|
||||
that failure and the unoverlaid teaser URL as a diagnostic preview; do not
|
||||
deliver a teaser whose only `COMING SOON` text was generated by the video model.
|
||||
|
||||
---
|
||||
|
||||
## Result Delivery
|
||||
|
||||
Return the final Pika CDN URL as the primary deliverable. If the host client requires local media markers, create that local preview outside this skill flow after confirming the CDN URL is reachable.
|
||||
@@ -440,15 +514,16 @@ Typical run time is 4-8 minutes:
|
||||
|
||||
| Step | Wall clock | Notes |
|
||||
|---|---:|---|
|
||||
| Asset sourcing | 10-60s | App Store via `fetch_appstore_screens`; website capture depends on page load |
|
||||
| Asset sourcing | 10-60s | App Store via `mcp__pika__fetch_appstore_screens`; website capture depends on page load |
|
||||
| Screen analysis + arc | 2-5 min | User confirmation can add time |
|
||||
| GPT-image-2 enhancement | 30-90s | Run selected screens in parallel |
|
||||
| Seedance generation | 3-5 min | Kling fallback can be slower |
|
||||
| Seedance generation | 3-5 min | Generated-audio moderation recovery adds one silent probe plus one same-seed sound replay |
|
||||
| Kling fallback | 5-15 min | Capacity wait or worker handoff may temporarily show `queued`; follow the Kling queued/handoff recovery runbook |
|
||||
| Download verification | <30s | Local sanity check before delivery |
|
||||
|
||||
## Engine Choice: Seedance Primary, Kling Fallback
|
||||
|
||||
Seedance is the default because it handles polished motion-graphics references and 1080p app teasers well. Kling is the fallback for moderation or balance failures because it is more permissive on some screen content and uses `quality_mode="pro"` for 1080p.
|
||||
Seedance is the default because it handles polished motion-graphics references and 1080p app teasers well. Kling is the fallback for moderation, balance, or Seedance timeout failures because it is more permissive on some screen content and uses `quality_mode="pro"` for 1080p.
|
||||
|
||||
## Failure Modes
|
||||
|
||||
@@ -456,13 +531,15 @@ Seedance is the default because it handles polished motion-graphics references a
|
||||
|---|---|---|
|
||||
| `fast=True` with `resolution="1080p"` | Seedance caps fast mode at 720p | Remove `fast`; keep `resolution="1080p"` |
|
||||
| `negative_prompt` rejected | Seedance does not accept this field | Use positive framing such as "smooth motion, stable camera" |
|
||||
| Seedance `partner_validation_failed` on audio | Often a false positive | Retry with `sound=False`; if video passes, retry with `sound=True` and the same seed |
|
||||
| Seedance generated-audio moderation: `content_policy_violation` / `partner_validation_failed`, `generated_video`, "Output audio has sensitive content." | Often a false positive on non-sensitive app-sizzle references | Follow the generated-audio recovery runbook: same-seed `sound=False` probe, then same-seed `sound=True` replay |
|
||||
| Seedance timeout such as `seedance timed out after ...` | Provider queue saturation or tail latency exceeded the tool budget | Run the Kling fallback; do not keep retrying Seedance unless the user explicitly asks to wait |
|
||||
| Seedance `partner_validation_failed` on video | Screen content includes recording UI, celebrity faces, or similar moderation triggers | Switch to `provider="kling"` and convert tokens to `<<<image_N>>>` |
|
||||
| Faces in screenshots trigger content policy | Screenshot includes real people | Crop faces out before upload, or use Kling |
|
||||
| 6+ reference images reduce quality | The model blends too many refs | Keep to 3-5 references, roughly one per 3 seconds |
|
||||
| Prompt tail ignored | Prompt exceeds about 200 words | Trim to the beat structure and the concrete UI details |
|
||||
| Text in output is garbled | Video model is asked to render new text | Keep text as existing reference-image content; overlay any new branding in post |
|
||||
| Logo reveal hallucinates letterforms | "assemble/build/construct" language triggers per-glyph rendering | Use "materializes whole", "crystallizes as a single form", or "fades in as a complete element" |
|
||||
| Task returns `{ task_id }` instead of inline | Long-running generation exceeded inline budget | Poll `task_status(task_id)` until `completed`/`done`, `failed`, or `cancelled`; unwrap `result.structuredContent` when present |
|
||||
| Task returns `{ task_id }` instead of inline | Long-running generation exceeded inline budget | Poll `mcp__pika__task_status(task_id)` until `completed`, `failed`, or `cancelled`; unwrap `result.structuredContent` when present |
|
||||
| Kling task returns status: `queued` after previously running | Worker handoff or provider capacity wait | Follow the Kling queued/handoff recovery runbook; do not duplicate-submit unless queued for more than 10 minutes with no `lastUpdatedAt` movement |
|
||||
| Kling rejects `resolution=` | Kling uses a different quality knob | Use `quality_mode="pro"` |
|
||||
| App Store icon URL points to promo art | App Store metadata fallback found feature artwork | Prefer the `icon.url` returned by `fetch_appstore_screens`; if missing, ask for a logo/icon file |
|
||||
| App Store icon URL points to promo art | App Store metadata fallback found feature artwork | Prefer the `icon.url` returned by `mcp__pika__fetch_appstore_screens`; if missing, ask for a logo/icon file |
|
||||
|
||||
@@ -9,7 +9,14 @@ description: >
|
||||
for [brand]", "I have a brand.md and screenshots, generate store assets", "screenshot set for
|
||||
app launch", "iOS store screens", "app store creative", "store listing visuals", "splashy app
|
||||
store screens", "app-store-screens".
|
||||
argument-hint: <brand-md-or-brand-spec-or-app-store-url> [product-screenshots-or-figma-url] [reference=<url-or-path>] [count=5|6]
|
||||
argument-hint: <brand-md-or-brand-spec-or-app-store-url> [product-screenshots-or-figma-url] [reference=<url-or-path>] [count=5|6] [--quick] [--config <path>]
|
||||
required-capabilities:
|
||||
- mcp__pika__analyze_media
|
||||
- mcp__pika__capture_website
|
||||
- mcp__pika__fetch_appstore_screens
|
||||
- mcp__pika__generate_image
|
||||
- mcp__pika__html_to_png
|
||||
- mcp__pika__upload_asset
|
||||
---
|
||||
|
||||
# App Store Screens
|
||||
@@ -37,21 +44,68 @@ Plus a contact sheet (`_preview.png`) showing all 6 at a glance.
|
||||
|
||||
### Step 0 — Intake and style choice
|
||||
|
||||
If invoked with empty args and no usable brand/screenshot context, print this menu verbatim and stop. Do not generate imagery or render HTML until the required inputs are present.
|
||||
If invoked with empty args and no usable brand/screenshot/demo context, print this menu verbatim and stop. Do not generate imagery or render HTML until the required inputs are present.
|
||||
|
||||
> **What App Store screenshot set should I make?** Required:
|
||||
>
|
||||
> - **Brand spec** — `brand.md` or equivalent brand notes with name, palette, fonts, voice, and imagery direction; or an App Store URL/app name if you want me to fetch the listing and draft the brand read first
|
||||
> - **Raw product screenshots** — exported PNGs, a Figma/source file that can export them, a folder of screenshots, or an App Store URL/app name I can fetch with Pika MCP
|
||||
> - **Or a fictional app demo brief** — only for launch demos/concepts where no real product exists; say `demo_mode: true` and provide what the fake app does
|
||||
>
|
||||
> Optional: reference App Store screenshot, moodboard, preferred screen count (5 or 6), output folder.
|
||||
|
||||
If the user has supplied partial input, ask only for the missing required items and stop. Required inputs:
|
||||
In interactive mode, if the user has supplied partial input, ask only for the missing required items and stop. If the non-interactive fast lane applies, use Step 0.5 instead. Required inputs:
|
||||
- `brand.md` or an equivalent brand spec with name, palette, fonts, voice, and imagery direction; if the user gives an App Store URL and asks you to "help create it", fetch the listing first and draft an inferred brand spec from the listing/icon/screenshots for approval
|
||||
- raw product screenshots, a Figma/source file that can export them, or an App Store URL/app name that can be fetched through Pika MCP
|
||||
- one product source: raw product screenshots, a Figma/source file that can export them, or an App Store URL/app name that can be fetched through Pika MCP
|
||||
- for fictional launch demos only, `demo_mode: true` with `demo_brief` is an alternative to real product screenshots; use the Fictional app demo mode path below
|
||||
- optional reference screenshot or moodboard if they want a specific App Store style
|
||||
|
||||
When the required inputs are present, open with a brief agenda and ask one upfront style question. This is the cheapest moment to learn whether the user wants the default or a specific reference.
|
||||
### Step 0.5 — Non-interactive fast lane
|
||||
|
||||
Use this path when the caller passes `--quick` or `--config <path>`, or when the
|
||||
caller states they are running from CI, a subagent, a batch job, or any other
|
||||
non-interactive harness.
|
||||
|
||||
This section has precedence over the interactive ask/wait instructions below.
|
||||
When it applies, use this fast lane and do not fall through to the multi-turn
|
||||
intake unless the required brand and either product screenshots or explicit demo
|
||||
mode inputs are truly unavailable.
|
||||
|
||||
- `--config <path>` points to a JSON file with pre-baked choices: `brand_spec`,
|
||||
`product_screenshots`, `app_store_url`, `website_url`, `reference`, `style`,
|
||||
`screen_count`, `narrative_arc`, `demo_mode`, `demo_brief`, and
|
||||
`output_folder`.
|
||||
- `--quick` means choose the default style unless a reference is supplied, infer
|
||||
the brand read from the provided brand spec or App Store listing, draft the
|
||||
5-6 screen arc yourself, and proceed.
|
||||
- For `--quick` or `--config`, do not stop for confirmation at the style choice,
|
||||
brand-read playback, reference-rule playback, or 5-6 screen strategy pitch.
|
||||
Record assumptions inline and continue to design/render.
|
||||
- If neither real product screenshots nor explicit `demo_mode: true` with
|
||||
`demo_brief` is available, stop once with a single compact missing-fields list
|
||||
instead of starting a multi-turn Q&A loop.
|
||||
|
||||
#### Fictional app demo mode
|
||||
|
||||
Use this only when the caller explicitly says this is a fictional app, fake app,
|
||||
launch demo, concept demo, or passes `demo_mode: true` in config. Do not use demo mode for real products.
|
||||
|
||||
Demo mode is allowed to create representative UI mocks in the brand voice when no
|
||||
real product screenshots exist, but the output must be labeled as demo-only
|
||||
concept work, not production assets. Treat the invented screens as a storyboarded
|
||||
QRT/demo artifact, not App Store Connect-ready evidence of a real product.
|
||||
|
||||
- Require a `demo_brief` or enough user-provided product concept detail to define
|
||||
the app's core job, audience, 3-5 features, and proof/CTA angle.
|
||||
- Generate UI states that are internally consistent with that brief; do not imply
|
||||
real customers, real reviews, real metrics, or real integrations unless the
|
||||
prompt explicitly provides them.
|
||||
- Add a demo-only disclosure in the delivery notes and contact sheet label:
|
||||
"Demo UI concept — not production assets and not real product screenshots."
|
||||
- In non-interactive mode, proceed only if config sets `demo_mode: true` and
|
||||
provides `demo_brief`; otherwise stop with a compact missing-fields list.
|
||||
|
||||
In interactive mode, when the required inputs are present, open with a brief agenda and ask one upfront style question. This is the cheapest moment to learn whether the user wants the default or a specific reference. In the non-interactive fast lane, choose the default style unless `reference` or `style` is supplied, record that assumption inline, and continue.
|
||||
|
||||
```
|
||||
here's how this works:
|
||||
@@ -80,7 +134,7 @@ If the user can't or won't supply a reference but still wants more than default,
|
||||
|
||||
#### App Store listing path
|
||||
|
||||
If the user supplies an `apps.apple.com` URL, numeric App Store app ID, or app-name search term as the product source, try Pika MCP `fetch_appstore_screens` first because it is faster, more stable, and returns hosted screenshot/icon assets ready for later render steps:
|
||||
If the user supplies an `apps.apple.com` URL, numeric App Store app ID, or app-name search term as the product source, try Pika MCP `mcp__pika__fetch_appstore_screens` first because it is faster, more stable, and returns hosted screenshot/icon assets ready for later render steps:
|
||||
|
||||
```
|
||||
fetch_appstore_screens(
|
||||
@@ -91,7 +145,7 @@ fetch_appstore_screens(
|
||||
)
|
||||
```
|
||||
|
||||
Use the returned `metadata`, `icon`, and `screenshots` as the product source. The returned screenshot `url` values are already Pika-hosted HTTPS assets and can be used directly in later `html_to_png` stages.
|
||||
Use the returned `metadata`, `icon`, and `screenshots` as the product source. The returned screenshot `url` values are already Pika-hosted HTTPS assets and can be used directly in later `mcp__pika__html_to_png` stages.
|
||||
|
||||
If the user provided a country-specific App Store URL, preserve that storefront country when calling the tool when possible. If the country is unclear, default to `"us"` unless the user asked for another storefront.
|
||||
|
||||
@@ -111,7 +165,29 @@ Expected result shape:
|
||||
}
|
||||
```
|
||||
|
||||
After fetching, infer only a draft brand read from the listing and visuals: app name, category, visible palette, likely type direction, voice from subtitle/description, and notable UI moments. Read it back as a provisional brand spec and ask the user to correct it before pitching the 5-6 screen arc.
|
||||
#### Thin App Store listing guardrail
|
||||
|
||||
After `mcp__pika__fetch_appstore_screens`, count usable screenshots that show real
|
||||
product UI. If the listing returns fewer than 3 real product screenshots, treat it
|
||||
as a thin App Store listing.
|
||||
|
||||
- Do not create a 5-6 screen campaign by hallucinating UI. Real product UI is the
|
||||
default requirement for device frames.
|
||||
- Interactive mode: stop before strategy and offer two choices:
|
||||
1. **Website-capture path** — use `mcp__pika__capture_website` or supplied
|
||||
website/onboarding URLs to capture real product surfaces, then continue with
|
||||
those captures as product screenshots.
|
||||
2. **Real screenshot path** — ask for simulator exports, Figma frames, or other
|
||||
product UI captures before continuing.
|
||||
Do not offer synthesized device UI for real products. If the user explicitly
|
||||
pivots to a fictional launch/concept demo, route to Fictional app demo mode and
|
||||
require `demo_mode: true` with `demo_brief`.
|
||||
- Non-interactive fast lane: prefer the website-capture path when `website_url`
|
||||
or an obvious product website is available. If there is no captureable product
|
||||
surface, stop once with a compact missing-fields list unless config explicitly
|
||||
sets `demo_mode: true` and provides `demo_brief`.
|
||||
|
||||
After fetching, infer only a draft brand read from the listing and visuals: app name, category, visible palette, likely type direction, voice from subtitle/description, and notable UI moments. In interactive mode, read it back as a provisional brand spec and ask the user to correct it before pitching the 5-6 screen arc. In the non-interactive fast lane, record it as the provisional brand spec and continue.
|
||||
|
||||
#### Reference-driven path: how to apply the user's brand to the reference's style
|
||||
|
||||
@@ -146,13 +222,13 @@ grain. Vertical 9:16 portrait."
|
||||
|
||||
This is how you get the reference's STYLE without copying its CONTENT. The hand+phone composition transfers; the lighting/cast/setting comes from the brand.
|
||||
|
||||
**Pitch the extracted rules back before designing.** Write them out as a short bullet list ("device 75% canvas tilted -8°, headline-on-yellow-pill above device, over-the-shoulder hero, ink callouts with curved arrows") and confirm with the user that you read the reference correctly. Then build all 6 screens applying those rules consistently. Don't deviate mid-campaign.
|
||||
**Pitch the extracted rules back before designing in interactive mode.** Write them out as a short bullet list ("device 75% canvas tilted -8°, headline-on-yellow-pill above device, over-the-shoulder hero, ink callouts with curved arrows") and confirm with the user that you read the reference correctly. In the non-interactive fast lane, record the extracted rules inline and continue. Then build all 6 screens applying those rules consistently. Don't deviate mid-campaign.
|
||||
|
||||
Then actually read:
|
||||
|
||||
- **`brand.md`** — extract: brand name, tagline, palette (with hex), display font + body font, voice adjectives, voice examples, forbidden words, photography/illustration direction, mood words. If the file is a different format (PDF, plain notes), parse what's there and ask about gaps — don't invent a brand.
|
||||
- **`brand.md`** — extract: brand name, tagline, palette (with hex), display font + body font, voice adjectives, voice examples, forbidden words, photography/illustration direction, mood words. If the file is a different format (PDF, plain notes), parse what's there and ask about gaps in interactive mode. In the non-interactive fast lane, record reasonable assumptions for non-critical gaps — don't invent a brand.
|
||||
- **Product screenshots** — open each one and form an honest mental model: *what does this app actually do?* Note the core UI patterns (feed, chat, canvas, list, map, etc.), the primary action surface, and any "wow moment" screens (a generative result, a beautiful state, a unique interaction).
|
||||
- If you can use `analyze_media` to inspect screenshots without loading them as images, do — it's faster for a quick scan.
|
||||
- If you can use `mcp__pika__analyze_media` to inspect screenshots without loading them as images, do — it's faster for a quick scan.
|
||||
|
||||
**Then read back what you found** (3-5 lines, conversational):
|
||||
|
||||
@@ -162,7 +238,7 @@ it does [X] — i see [specific UI cue 1] and [specific UI cue 2]. the standout
|
||||
N] because [reason]. correct me where i'm wrong, otherwise i'll pitch the 6-screen arc.
|
||||
```
|
||||
|
||||
Wait for confirmation before pitching strategy. Catches misreads cheaply.
|
||||
Interactive mode: wait for confirmation before pitching strategy. Catches misreads cheaply. In the non-interactive fast lane, treat the brand read as provisional, record the assumption inline, and continue to the strategy.
|
||||
|
||||
### Step 2 — Pitch the 6-screen strategy
|
||||
|
||||
@@ -188,17 +264,17 @@ Headline: "One tap. Then quiet."
|
||||
...
|
||||
```
|
||||
|
||||
Ask the user to sign off on the arc before you generate or composite anything. Iterate on copy here — it's the cheapest moment to fix it.
|
||||
Interactive mode: ask the user to sign off on the arc before you generate or composite anything. Iterate on copy here — it's the cheapest moment to fix it. In the non-interactive fast lane, generate from the configured or inferred arc without stopping for signoff.
|
||||
|
||||
### Step 3 — Generate brand-world imagery
|
||||
|
||||
For any screen that calls for splashy background/hero imagery, generate it with `generate_image`.
|
||||
For any screen that calls for splashy background/hero imagery, generate it with `mcp__pika__generate_image`.
|
||||
|
||||
**Defaults that work for App Store splash:**
|
||||
- `provider: gpt-image-2`
|
||||
- `quality: medium` for finished screenshots; `low` only when the user explicitly wants fast iteration
|
||||
- `aspect_ratio: 9:16` for full-bleed backgrounds (matches the iPhone canvas)
|
||||
- gpt-image-2 is 1K-class only. If a screen genuinely needs 2K (e.g. a hyper-detailed background that gets blown up), surface the tradeoff and ask before swapping to `seedream`
|
||||
- 1K is the default. If a screen genuinely needs higher res (e.g. a hyper-detailed background that gets blown up), bump `gpt-image-2` to its 2K/4K tier (9:16 backgrounds support both) or escalate to `seedream` — surface the tradeoff and ask first
|
||||
|
||||
**Prompt structure for brand-world imagery:**
|
||||
1. Subject + composition ("a wide cinematic shot of a single ceramic mug on weathered oak…")
|
||||
@@ -209,26 +285,37 @@ For any screen that calls for splashy background/hero imagery, generate it with
|
||||
|
||||
**Reserve upper third for the headline.** Tell the image model where the type will go so it leaves room.
|
||||
|
||||
Keep generated image URLs as HTTPS/CDN URLs for server-side rendering. If the user provides local product screenshots or background images, upload them with `upload_asset` first and use the returned `public_url`; `html_to_png` cannot read local `file://` paths.
|
||||
Keep generated image URLs as HTTPS/CDN URLs for server-side rendering. If the user provides local product screenshots or background images, upload them with `mcp__pika__upload_asset` first and use the returned `public_url`; `mcp__pika__html_to_png` cannot read local `file://` paths.
|
||||
|
||||
### Step 4 — Composite each screen at 1290×2796
|
||||
|
||||
Write one HTML stage per screen, then render to PNG via Pika MCP `html_to_png`. See `references/render-pipeline.md` for:
|
||||
- The exact `html_to_png` request shape
|
||||
Write one HTML stage per screen, then render to PNG via Pika MCP `mcp__pika__html_to_png`. See `references/render-pipeline.md` for:
|
||||
- The exact `mcp__pika__html_to_png` request shape
|
||||
- Server-side asset/font rules
|
||||
- Safe-zone guides
|
||||
- Common gotchas (font loading, retina text sharpness, mid-curve crop)
|
||||
|
||||
Each HTML stage should be exactly 1290×2796px. Use `@font-face` with HTTPS raw font URLs or inline `data:font/...` sources. Local brand-kit font files must be exposed through a public HTTPS URL or inlined; `upload_asset` does not accept font mime types.
|
||||
Each HTML stage should be exactly 1290×2796px. Use `@font-face` with HTTPS raw font URLs or inline `data:font/...` sources. Local brand-kit font files must be exposed through a public HTTPS URL or inlined; `mcp__pika__upload_asset` does not accept font mime types.
|
||||
|
||||
After each render, run a pre-delivery QA pass before accepting the PNG. Use
|
||||
`mcp__pika__analyze_media` on the rendered image and inspect the authored HTML
|
||||
positions when available. Reject and rerender any screen whose text block
|
||||
bounding boxes put load-bearing headline, eyebrow, subhead, CTA, or product UI
|
||||
outside the safe content area documented in `references/render-pipeline.md`.
|
||||
|
||||
### Step 5 — Build the contact sheet + deliver
|
||||
|
||||
Once all 6 PNGs render cleanly:
|
||||
|
||||
1. Build a `_preview.png` contact sheet with `html_to_png` — 6 thumbnails in a 3×2 grid at ~25% size, on a neutral background, labeled by role.
|
||||
1. Build a `_preview.png` contact sheet with `mcp__pika__html_to_png` — 6 thumbnails in a 3×2 grid at ~25% size, on a neutral background, labeled by role.
|
||||
2. Present the contact sheet URL plus each individual screenshot `file_url`. If you also saved local copies, include those local paths separately.
|
||||
3. Ask: anything to revise? Common revisions are copy tweaks (cheap) or layout swaps (medium) or new imagery (most expensive).
|
||||
|
||||
Do not report the screenshot set complete until `_preview.png` exists and is
|
||||
included in the delivery. If the contact sheet render fails, fix the sheet HTML
|
||||
or rerender the missing PNGs first; do not deliver only the individual PNGs and
|
||||
call the campaign finished.
|
||||
|
||||
## The 5 layout archetypes
|
||||
|
||||
Use one per screen. **Don't stack devices.** One bold idea per screen.
|
||||
@@ -266,7 +353,7 @@ If the headline could appear on 500 other apps without anyone noticing — rewri
|
||||
- **Typography hierarchy.** One thing dominates per screen — usually the device. Headline 100–140px Funnel-Display-class; subhead 36–44px; visible difference in weight and size. (Past 150px the headline starts fighting the device for the eye.)
|
||||
- **Skip SVG-stroke iPhone frames — they don't align with screenshot corners.** Stroke-based SVG bezels sit half-inside/half-outside their path, and the radius/Dynamic Island proportions never quite match a real iPhone. Prefer transparent screenshot exports from Figma/simulator or a real iPhone mockup PNG with a transparent cutout. If only rectangular source screenshots are available, use a consistent CSS `border-radius` + `overflow:hidden` wrapper and visually QA the corners. The old local `clean_phone_uniform()` PIL recipe is now a documented MCP gap, not a default dependency.
|
||||
- **Pick one: bleed OR full-frame. Never mid-crop the bottom curve.** Default is full-frame (device top at y=580 with width=1000, so the rounded bottom corners sit fully inside the canvas with ~48px breathing room below). Bleed variant (device top ≥ 819) pushes the rounded bottom corners entirely past y=2796 — useful for hook screens that want drama. Anything in between produces a visible half-cut curve where the canvas slices through the rounded corner mid-arc; this is the failure mode the rule prevents.
|
||||
- **Apple safe zones.** Don't put load-bearing type in the top 100px (status bar area) or bottom 100px (where App Store overlays its own UI on a tap-to-expand). Decoration is fine; critical claim is not.
|
||||
- **Apple safe zones.** Strict safe-zone rule: load-bearing text and product UI must stay inside y=180..2616. The top and bottom 100px bands are never for critical content, and the extra 80px buffer keeps App Store chrome, search-result cropping, and text ascenders from crowding the edge. Decoration is fine; critical claim and readable app UI are not.
|
||||
- **One bold device per screen.** Massive type *or* generated imagery *or* a quote — not all three. Splashy ≠ chaotic.
|
||||
- **Color from the brand palette only.** No new colors invented for the screenshots. If the palette feels too restrictive, that's the brand's problem to solve, not yours.
|
||||
|
||||
@@ -286,6 +373,7 @@ Before delivering, do the squint test:
|
||||
2. **First-2-screens thumbnail test.** Crop screens 1 and 2 to 25% size. Is the headline still legible? Is the hook still clear? If you have to lean in, the type is too small.
|
||||
3. **Brand-test.** Cover the device on each screen. Does the surrounding design still feel like the brand? (Color, type, mood, voice in headlines.) If not, the screenshot is generic.
|
||||
4. **Anti-generic test** (same as build-a-brand): could these screens belong to literally any other app in this category? If yes — rewrite the copy and rework the focal hierarchy.
|
||||
5. **Safe-zone audit.** Check every rendered PNG and reject and rerender any screen with load-bearing text or product UI outside the safe content area.
|
||||
|
||||
## Load-bearing phrases
|
||||
|
||||
@@ -301,7 +389,7 @@ These anchors keep the generated campaign legible and on-brand:
|
||||
|
||||
## Engine choice: HTML-first render, gpt-image-2 only for brand-world imagery
|
||||
|
||||
The final screenshots should be deterministic HTML/CSS composites rendered through `html_to_png`, because App Store copy, device masks, safe zones, and brand typography need exact control. Use `gpt-image-2` only for splash/background/brand-world imagery where generation adds visual richness; keep real product UI as screenshots. Use `seedream` only when a specific generated background needs higher resolution than gpt-image-2's 1K-class output.
|
||||
The final screenshots should be deterministic HTML/CSS composites rendered through `mcp__pika__html_to_png`, because App Store copy, device masks, safe zones, and brand typography need exact control. Use `gpt-image-2` only for splash/background/brand-world imagery where generation adds visual richness; keep real product UI as screenshots. Bump `gpt-image-2` to its 2K/4K tier (or escalate to `seedream`) only when a specific generated background genuinely needs higher resolution than 1K.
|
||||
|
||||
## Runtime Expectations
|
||||
|
||||
@@ -323,19 +411,23 @@ Typical run time is 10-25 minutes, depending on how much user confirmation is ne
|
||||
| Screens look like six variants of the same layout | Default skeleton was repeated without enough content contrast | Reassign layout archetypes and rotate focal weight, color, and screenshot choice |
|
||||
| Headlines are illegible in contact sheet | Type is too small for App Store thumbnail use | Increase headline size, shorten copy to 5-8 words, and rerender |
|
||||
| Generated background contains text or fake UI | Prompt over-described brand/product specifics | Regenerate with a no-text/no-logo guardrail and reserve the real UI for screenshots |
|
||||
| App Store listing has fewer than 3 usable screenshots | Thin App Store listing, often a companion app or early listing | Use the website-capture path, ask for real screenshots, or stop. For fictional launch-demo concepts only, use Fictional app demo mode with `demo_mode: true`, `demo_brief`, and the demo-only disclosure |
|
||||
| Fictional app has no real product screenshots | Launch-demo concept needs representative UI, but production rules forbid hallucinated UI | Use Fictional app demo mode only with `demo_mode: true`, a concrete `demo_brief`, representative UI mocks, and a demo-only disclosure |
|
||||
| Load-bearing text appears too close to the top or bottom edge | Safe-zone QA was skipped or checked only the rendered look, not text block bounding boxes | Reject and rerender with text block bounding boxes inside the strict safe-zone margin |
|
||||
| Device corners look uneven | Source screenshot already contains background in the rounded corner area, or CSS radius does not match the source | Ask for transparent/high-res source exports, use a real mockup cutout, or apply one consistent CSS mask and visually QA. A server-side uniform corner cleaner is still a tool gap. |
|
||||
| Brand feels generic after hiding the screenshots | Surrounding design ignores `brand.md` voice, palette, or imagery rules | Rebuild the screen shell from the brand spec before rerendering |
|
||||
| Render is blurry or scaled | Stage dimensions or raster options are wrong | Verify 1290x2796 stage size and `html_to_png` `viewport_px:1290x2796`, `device_scale:1` |
|
||||
| Render is blurry or scaled | Stage dimensions or raster options are wrong | Verify 1290x2796 stage size and `mcp__pika__html_to_png` `viewport_px:1290x2796`, `device_scale:1` |
|
||||
|
||||
## When to push back on the user
|
||||
|
||||
- If they want a "jazzed up" version without supplying a reference → push back. Ask for a screenshot of an App Store page they love, a figma file, or a moodboard. Explain that without a reference, you'll deliver default plus their brand color, and that's better than guessing at "exciting." (This rule comes from real campaign work — auto-jazzing produced amateur output every time.)
|
||||
- If they want a 10-screen campaign → push back. 5-6 is the sweet spot; more is fatigue.
|
||||
- If the brand voice in `brand.md` is generic ("modern, friendly, intuitive") → flag it. You can still produce screens, but tell them the headlines will only be as distinctive as the voice spec. Offer to sharpen the voice first (it lives in `brand.md`'s Voice & Tone section).
|
||||
- If they have no real product screenshots and did not explicitly request fictional app demo mode → don't invent UI. Ask for screenshots, website/onboarding capture targets, or explicit `demo_mode: true` with a concrete `demo_brief`.
|
||||
- If they don't have a brand at all → don't proceed on vibes. Route them to `build-a-brand` first, or have them write at minimum: name, palette hexes, display font, body font, one-line voice description.
|
||||
|
||||
## References
|
||||
|
||||
- `references/default-layout.md` — **the visual system the skill produces when the user supplies no reference.** Read this first for any run. Codifies the headline-top + device-dominant + full-frame composition, color rotation, and squint-test checklist. This is what the skill delivers well.
|
||||
- `references/layout-archetypes.md` — vocabulary list of layout patterns found in real App Store campaigns. Use only to analyze a user-supplied reference. Not a free-choice menu.
|
||||
- `references/render-pipeline.md` — `html_to_png` render request, server-side asset/font rules, safe-zone overlay, font-loading checklist, pre-delivery checklist.
|
||||
- `references/render-pipeline.md` — `mcp__pika__html_to_png` render request, server-side asset/font rules, safe-zone overlay, font-loading checklist, pre-delivery checklist.
|
||||
|
||||
@@ -6,7 +6,7 @@
|
||||
|
||||
Six layouts + techniques that show up in real App Store campaigns. If a user-supplied reference uses one of these, the description here helps you describe what you're seeing back to the user before designing.
|
||||
|
||||
All examples assume a 1290×2796 stage. Coordinates and font sizes are approximate; tune per brand.
|
||||
All examples assume a 1290×2796 stage. Coordinates and font sizes are approximate; tune per brand. The strict safe-zone rule from `render-pipeline.md` applies to every example: load-bearing text and readable product UI stay inside y=180..2616. Decorative background color, blur, and non-readable texture may extend outside that range; critical claims and usable app UI may not.
|
||||
|
||||
**Default device dimensions:** 1000px wide (≈ 78% of canvas), height ≈ 2168px at native iPhone 15/16 aspect (393/852). Fits fully within the canvas at top=580 (rounded bottom corners visible). For bleed variants, push top to ≥ 819. If you find yourself sizing devices below 800px wide, the layout is wrong, not the device — rework it.
|
||||
|
||||
@@ -69,14 +69,14 @@ The workhorse. Phone screen sits centered or slightly off-center on a splash bac
|
||||
|
||||
## 2. Full-Bleed UI
|
||||
|
||||
Phone canvas fills the entire stage (no device frame visible, or a faint one at the edges). Typography overlays at top or bottom in a contrasting color. Use for the **hook** screen — biggest, boldest, most "this is the app" moment.
|
||||
Phone canvas feels oversized and immersive, but the readable product UI remains inside the safe area. Use for the **hook** screen — biggest, boldest, most "this is the app" moment. If you want edge-to-edge drama, use a blurred or cropped duplicate as decorative background and place the real product UI inside y=180..2616.
|
||||
|
||||
```
|
||||
┌─────────────────────────────┐
|
||||
│ ┌─────────────────────┐ │
|
||||
│ │ │ │
|
||||
│ │ APP UI FILLS │ │ ← UI bleeds to ~90%
|
||||
│ │ THE FRAME │ │ of the canvas
|
||||
│ │ APP UI FILLS │ │ ← readable UI stays
|
||||
│ │ THE SAFE AREA │ │ inside y=180..2616
|
||||
│ │ │ │
|
||||
│ │ │ │
|
||||
│ │ ▓▓▓▓▓▓▓▓▓▓▓ │ │
|
||||
@@ -91,12 +91,18 @@ Phone canvas fills the entire stage (no device frame visible, or a faint one at
|
||||
|
||||
```html
|
||||
<div class="stage" style="width: 1290px; height: 2796px; background: var(--brand-bg); overflow: hidden;">
|
||||
<img src="https://cdn.pika.art/.../screenshot.png" style="
|
||||
position: absolute; top: 0; left: 0;
|
||||
width: 100%; height: 100%; object-fit: cover;
|
||||
<img class="decorative-bg" src="https://cdn.pika.art/.../screenshot.png" style="
|
||||
position: absolute; inset: -80px;
|
||||
width: calc(100% + 160px); height: calc(100% + 160px);
|
||||
object-fit: cover; filter: blur(28px); opacity: 0.18;
|
||||
">
|
||||
<img class="product-ui" src="https://cdn.pika.art/.../screenshot.png" style="
|
||||
position: absolute; top: 240px; left: 80px;
|
||||
width: 1130px; height: 1500px; object-fit: cover;
|
||||
border-radius: 88px;
|
||||
">
|
||||
<div class="overlay" style="
|
||||
position: absolute; bottom: 0; left: 0; right: 0;
|
||||
position: absolute; bottom: 180px; left: 0; right: 0;
|
||||
height: 720px;
|
||||
background: linear-gradient(to top, var(--brand-ink) 30%, transparent);
|
||||
display: flex; align-items: flex-end; padding: 0 80px 200px;
|
||||
@@ -187,7 +193,7 @@ Phone in front, a large branded shape or generated image floating behind it. Til
|
||||
```html
|
||||
<div class="stage" style="width: 1290px; height: 2796px; background: var(--brand-bg); position: relative;">
|
||||
<h1 style="
|
||||
position: absolute; top: 160px; left: 80px; right: 80px;
|
||||
position: absolute; top: 180px; left: 80px; right: 80px;
|
||||
text-align: center; font-size: 140px; line-height: 0.95;
|
||||
color: var(--brand-ink); margin: 0;
|
||||
">
|
||||
|
||||
@@ -93,6 +93,13 @@ App Store Connect overlays its own chrome on screenshots in certain views. Keep
|
||||
|
||||
Visuals can live in the top/bottom 100px bands. Load-bearing headlines and product UI cannot.
|
||||
|
||||
Strict text margin: text block bounding boxes for headlines, eyebrows, subheads,
|
||||
CTA text, labels, and any product UI must be fully inside y=180..2616. The top
|
||||
and bottom 100px bands are hard no-go zones; the extra 80px gives breathing room
|
||||
for status-bar chrome, App Store search-result cropping, and text ascenders. If
|
||||
a safe-zone audit finds load-bearing text with y < 180, or a block bottom past
|
||||
y=2616, reject and rerender.
|
||||
|
||||
## Device Compositing
|
||||
|
||||
Preferred source:
|
||||
@@ -210,7 +217,7 @@ Before handing files to the user, verify:
|
||||
- [ ] Every individual PNG is exactly 1290x2796.
|
||||
- [ ] Brand fonts loaded, not system fallbacks.
|
||||
- [ ] No black bleed at edges.
|
||||
- [ ] No load-bearing text in top/bottom 100px.
|
||||
- [ ] Safe-zone audit passes: no load-bearing text block bounding boxes at y < 180 or past y=2616; reject and rerender failures.
|
||||
- [ ] Device corners look consistent across all six.
|
||||
- [ ] First two screens are legible when viewed at 25%.
|
||||
- [ ] Contact sheet shows visual variety across the six.
|
||||
|
||||
+180
-153
@@ -9,22 +9,21 @@ description: >
|
||||
guidelines for [X]", "i want a brand book", "create a brand from scratch", "brand for [idea]",
|
||||
"i want a brand that feels like [X] + [Y]", "rebrand my [thing]", "visual identity for [thing]",
|
||||
"build-a-brand".
|
||||
argument-hint: "[brand idea, URL, photos, or reference brands]"
|
||||
argument-hint: <idea-or-url-or-reference-brands> [photos=<paths-or-urls>] [refresh=<existing-brand>] [--quick] [--config <path>]
|
||||
required-capabilities:
|
||||
- mcp__pika__generate_image
|
||||
- mcp__pika__html_to_pdf
|
||||
- mcp__pika__html_to_png
|
||||
- mcp__pika__analyze_media
|
||||
- mcp__pika__task_status
|
||||
- mcp__pika__upload_asset
|
||||
---
|
||||
|
||||
# Build a Brand
|
||||
|
||||
Take any input — an idea, a website, a list of reference brands, product photos, or an existing brand to refresh — and produce a complete brand identity, ending in a 14–16-page brand guidelines PDF.
|
||||
Take any input — an idea, a website, a list of reference brands, product photos, or an existing brand to refresh — and produce a complete brand identity, ending in a 15-page brand guidelines PDF.
|
||||
|
||||
This is the brand-building engine from `business-maker` without the commerce arm. Same uncompromising standards on strategy, design, copy. Output is **one** brand guidelines PDF for the chosen identity.
|
||||
|
||||
## Execution Model — Local First, Cloud Only Where Needed
|
||||
|
||||
This skill is local-first. Keep deterministic production work on the user's machine:
|
||||
- **Local:** workspace setup, downloaded fonts, generated image files after download, image compression, transparent-background cleanup, 16×16 favicon tests, HTML/CSS page builds, PDF rendering, PNG QA screenshots, crop-and-read QA files, logo asset assembly, token/prompt files, and final zip packaging.
|
||||
- **Cloud:** image generation only (`gpt-image-2` for symbols, mood images, photography, illustration, and ambient textures) and URL/source research when the brief requires it.
|
||||
|
||||
Do not use a cloud PDF renderer or upload PDFs by default. Save board PDFs, guidelines PDFs, and brand-kit zips to `~/Desktop` on Mac, or the project working directory when Desktop is unavailable. Only upload/share via CDN if the user explicitly asks for a hosted file.
|
||||
This is a standalone brand-building workflow focused on strategy, identity, design, and copy. Output is **one** brand guidelines PDF for the chosen identity.
|
||||
|
||||
## Full Workflow
|
||||
|
||||
@@ -42,16 +41,45 @@ If invoked with no input (no idea, no URL, no photos, no reference brands, and n
|
||||
|
||||
If the user already dropped one of the above, skip the menu and proceed straight to Step 1.
|
||||
|
||||
### Stage 0.5 — Non-interactive fast lane
|
||||
|
||||
Use this path when the caller passes `--quick` or `--config <path>`, or when the
|
||||
caller states they are running from CI, a subagent, a batch job, or any other
|
||||
non-interactive harness.
|
||||
|
||||
This section has precedence over the interactive ask/wait instructions below.
|
||||
When it applies, use this fast lane and do not fall through to the multi-turn
|
||||
intake unless a required input is truly missing.
|
||||
|
||||
- `--config <path>` points to a JSON file that pre-bakes intake answers:
|
||||
`input`, `photos`, `reference_brands`, `audience`, `positioning`, `assets_to_keep`,
|
||||
`references`, `chosen_direction`, `chosen_identity`, and `export_kit`.
|
||||
- `--quick` means use model judgment for all confirmation gates. Ask only if
|
||||
the original input is missing entirely; otherwise infer reasonable defaults,
|
||||
choose the strongest strategy direction and identity option, and continue.
|
||||
- For `--quick` or `--config`, do not stop for confirmation at the deliverable
|
||||
preview, strategy-direction choice, identity-option choice, or brand-kit
|
||||
export gate. Record the assumption inline, then proceed.
|
||||
- If a required asset is unavailable and cannot be inferred from the input,
|
||||
stop once with a single compact missing-fields list instead of starting a
|
||||
multi-turn Q&A loop.
|
||||
- Do not deliver a condensed or partial brand output just because the caller is
|
||||
non-interactive or the run is short on wall-clock time. A condensed 6-page
|
||||
deck is not an acceptable substitute for the required 15-page guidelines.
|
||||
If the run is out of wall-clock budget, save a resumable checkpoint with the
|
||||
pages/assets already completed and stop; do not mark the workflow complete.
|
||||
|
||||
### Step 1 — Read the Input
|
||||
|
||||
Inputs vary. **Before asking any questions, open with a brief agenda** so the user knows what's coming:
|
||||
|
||||
```
|
||||
here's how this works — 4 steps:
|
||||
here's how this works — 5 steps:
|
||||
1. **Read the input** — i ask a few questions, you answer, i play back what i'm hearing
|
||||
2. **3 visual brand boards** — i build a 3-page PDF with three complete brand directions, each with its own colors, fonts, photography, voice samples, logo concept. you pick one (or ask to mix elements).
|
||||
3. **Build the guidelines** — full 14–16-page brand book PDF for the chosen board
|
||||
4. **Export the brand kit** — `brand.md` spec + logo assets (symbol PNG sizes, wordmark SVG/PNG, lockup SVG/PNG) + fonts + tokens + AI prompts, zipped to your Desktop
|
||||
2. **Strategy directions** — 2-3 distinct positioning angles to choose from
|
||||
3. **Identity options** — 3 full brand identities (name, colors, voice, brand board PDF preview) within your chosen direction
|
||||
4. **Build the guidelines** — full 15-page brand book PDF for the chosen identity
|
||||
5. **Export the brand kit** — once you're happy with the guidelines, i'll bundle a `brand.md` spec + logo assets (transparent PNG + PDF wrapper for every mark; SVG when vector-authored) in each brand color as a zip you can use anywhere
|
||||
|
||||
let's start. [questions follow]
|
||||
```
|
||||
@@ -60,7 +88,7 @@ Then ask 3-5 targeted questions in a single message. Adapt to the input type:
|
||||
|
||||
**If they dropped an idea / description:**
|
||||
- What does this brand sell or do? (product / service / app / community / something else)
|
||||
- Who is this for — describe one specific person, not a demographic
|
||||
- Who is this for — describe the 2-3 audience segments this brand should serve, plus one vivid anchor persona inside the primary segment
|
||||
- Why does this exist? what's broken about the alternatives, or what feeling are you trying to deliver?
|
||||
- Do you have a name in mind, or is naming part of what you want help with?
|
||||
|
||||
@@ -70,19 +98,18 @@ Then ask 3-5 targeted questions in a single message. Adapt to the input type:
|
||||
- Who's the current customer vs. who you wish were the customer?
|
||||
|
||||
**If they dropped product photos:**
|
||||
- Use the `business-maker` Step 1 questions (how it's made, who's bought, price point, direction in mind).
|
||||
- Same playbook but the output is just guidelines, not a commerce launch.
|
||||
- Ask how the product is made, who has bought or used it, price point, current sales/channel context, and any direction they already have in mind.
|
||||
- Keep the output scoped to guidelines and a brand kit, not a commerce launch.
|
||||
|
||||
**If they dropped reference brands only ("I want a brand that feels like Aesop + Patagonia"):**
|
||||
- What's the product, service, or thing this brand will be attached to?
|
||||
- What about each reference brand specifically do you love? (the photography? the tone? the restraint?)
|
||||
- Who buys this — describe one specific person.
|
||||
- Who buys this — describe the 2-3 audience segments this brand should serve, plus one vivid anchor persona inside the primary segment.
|
||||
- Any constraints? (industry, regulation, location, price tier?)
|
||||
|
||||
**Always also ask (regardless of input type):**
|
||||
- Do you have any existing brand assets you want to keep or incorporate? (Logo, wordmark, symbol, name, colors, fonts, photography, packaging — anything you don't want to lose.)
|
||||
- Any specific references, inspirations, or moodboards you'd want this to draw from?
|
||||
- **Is this a digital product** (app, website, SaaS, web tool)? This determines whether the Icons page belongs in the guidelines — for non-digital brands (products, services, restaurants, fashion, etc.) the Icons page is skipped.
|
||||
|
||||
These two are essential — they prevent you from generating things the user already has, and they anchor the work in references the user actually likes. Always include them.
|
||||
|
||||
@@ -90,7 +117,7 @@ Keep it to a single message. Aim for 5-7 questions total (input-specific + the 2
|
||||
|
||||
**After answers**, analyze the input + answers together and read back:
|
||||
- **Aesthetic territory**: what visual world does this live in?
|
||||
- **Customer**: specific and vivid, not demographic
|
||||
- **Audience segments**: primary segment, secondary segment(s), and one vivid anchor persona inside the primary segment. Do not collapse the audience into one over-specific individual.
|
||||
- **Positioning**: what's the wedge — what does this stand for that competitors don't?
|
||||
- **Price tier / category fit**: where on the market shelf does this sit?
|
||||
- **Story hook**: what's the emotional reason someone cares?
|
||||
@@ -100,88 +127,92 @@ Keep it to a single message. Aim for 5-7 questions total (input-specific + the 2
|
||||
**Then preview the deliverable and invite specific guidance** — before moving to brand directions, show the user what'll be in the final guidelines so they can flag anything to add, change, or call out:
|
||||
|
||||
```
|
||||
here's what i'll build into the brand guidelines (14–16 pages depending on your brand):
|
||||
here's what i'll build into the 15-page brand guidelines:
|
||||
|
||||
1. Cover (brand name, tagline, hero mood)
|
||||
2. Strategy & positioning
|
||||
2. Strategy & positioning (primary/secondary audience segments + anchor persona)
|
||||
3. Brand foundation (mission, values, story)
|
||||
4. Logo (wordmark + symbol + variants)
|
||||
5. Logo don'ts
|
||||
6. Color palette
|
||||
7. Typography
|
||||
8. Icons — UI icon system + library guidance · ONLY IF this is a digital product (app / web / SaaS). Skipped for non-digital brands.
|
||||
8. Icons (UI icon system + library guidance)
|
||||
9. Voice & tone
|
||||
10. Imagery rules (photography and/or illustration, adapted to brand medium · splits to 2 pages if hybrid)
|
||||
10. Imagery rules (photography and/or illustration, adapted to brand medium)
|
||||
11. Visual world / lifestyle imagery
|
||||
12. Touchpoints (real photos showing the brand in use)
|
||||
13. Brand applications (mockups: business card, app icon, favicon, etc.)
|
||||
14. Digital + social
|
||||
15. Do & don't
|
||||
|
||||
plus a brand kit zip at the end with: `brand.md` spec, logo assets (symbol PNG sizes, wordmark SVG/PNG, lockup SVG/PNG), brand fonts (TTF), design tokens (CSS / JSON / Tailwind), AI prompts (system prompt + task-specific starters), and the icon SVGs (if applicable).
|
||||
plus a brand kit zip at the end with: `brand.md` spec, logo assets (transparent PNG + PDF wrapper for every mark; SVG when vector-authored), design tokens (CSS / JSON / Tailwind), AI prompts (system prompt + task-specific starters), and the icon SVGs.
|
||||
|
||||
anything you want to add, change, call out specifically, or want me to handle differently? if not, i'll move on to the 3 brand boards.
|
||||
anything you want to add, change, call out specifically, or want me to handle differently? if not, i'll move on to brand directions.
|
||||
```
|
||||
|
||||
Wait for response. Incorporate any specific user guidance (add a page, swap something, special focus on a particular section, exclude something) before moving to Step 2. This catches scope mismatches early — much cheaper than discovering them after the PDF is built.
|
||||
|
||||
### Step 2 — Generate 3 Visual Brand Boards (PDF)
|
||||
### Step 2 — Present 2-3 Brand Directions
|
||||
|
||||
**This step is the user's first visual touchpoint with the brand.** No text-only "directions" precede it. The boards ARE the directions, made visible. Each board contains a complete visual identity at-a-glance so the user can SEE the difference, not just read it.
|
||||
Based on your read, present 2-3 distinct brand directions. See `references/brand-directions.md` for structure.
|
||||
|
||||
Build a single 3-page PDF (one page per board, 1200×850 each) and save it to `~/Desktop/[brand-slug]-brand-boards.pdf`.
|
||||
Each direction must be a genuinely different business answer — not aesthetic variations. Differentiate on WHO and WHY, not WHAT.
|
||||
|
||||
**Each board must contain ALL of:**
|
||||
- Brand name + tagline (rendered in that board's display font)
|
||||
- 4-color palette with hex codes + role labels
|
||||
- Display + body font specimens with the actual font names labeled
|
||||
- A logo concept (see "Logo Pipeline" below — generate the symbol via gpt-image-2, NEVER hand-code an SVG)
|
||||
- One mood image (real photograph or illustration, generated via gpt-image-2 to match the brand's photography/illustration direction)
|
||||
- A voice sample (one quoted sentence in brand voice)
|
||||
- A 1-2 sentence brand story
|
||||
- 3-4 reference brand names
|
||||
Ask the user to pick one direction before proceeding.
|
||||
|
||||
**The 3 boards must be genuinely different brand directions** — not template recolors. Differentiate on WHO the brand is for and WHY it exists (see `references/brand-directions.md` for the differentiation rule and examples). Each board's layout, fonts, color logic, photography style, and voice must feel like 3 distinct brands.
|
||||
### Step 3 — Generate 3 Brand Identity Options
|
||||
|
||||
**Each board's LAYOUT must embody its design philosophy.** A magazine-cover board looks like a magazine cover (full-bleed photo, masthead-style). A soft consumer board looks like a homepage hero (rounded shapes, soft circles for swatches). An editorial board looks like a literary spread (huge italic centered, inset photo). See `references/brand-guidelines.md` "Brand Board Layout — Differentiate per Option" for rules.
|
||||
Once they choose a direction, generate 3 complete brand identity options within that direction. See `references/brand-identity.md` for structure.
|
||||
|
||||
**Mandatory before delivering:**
|
||||
1. **Render each board to PNG via Chrome headless** and READ each PNG. Verify no overlapping type, no broken layouts, no missing content. If anything looks sloppy — fix and re-render.
|
||||
2. **Verify the fonts feel chosen for THIS brand, not pulled from a mental shortlist.** No font is banned — but no font is a favorite either. If you reached for a font you used on a recent brand, you need a real reason this brand uniquely wants it. Otherwise pick a fresh option that fits as well. See `references/brand-guidelines.md` "No favorite fonts — every brand starts the search fresh" + "The mandatory research step" for the process.
|
||||
3. **Verify each board's mood image was generated with `provider="gpt-image-2"`** and contains no baked-in text.
|
||||
4. **Verify the 3 symbols differ in CONCEPT, not just style.** Don't ship three "literal mascot face" marks in three styles — that's one idea repeated. Push the symbols across different concept lanes (mascot / product-feature reference / abstract / monogram / hybrid / container). See `references/brand-identity.md` "Symbol concepts must DIFFER across the 3 brand options."
|
||||
5. **Verify each symbol passes the "Symbol output rules" check.** Symbols can be any style (flat / 3D / painted / chrome / photographic — no flat-vector requirement). But every gen'd symbol MUST satisfy ALL of: (a) conceptually linked to the brand (means something about what the brand IS/DOES, not just decorative), (b) feels unique (not generic — would fit ONLY this brand), (c) recognizable at 16×16 favicon size — mandatory test: PIL resize to 16×16 with LANCZOS, then upscale that 16-px image with NEAREST to 128×128 to see what a viewer sees at favicon scale, Read the test image. If the mark dissolves to mush, regenerate with stronger compositional weight (for hairline/lux marks: give the dominant element a solid filled mass while keeping framing details fine-lined — see DMV exercise pattern). Style is preserved, weight is adjusted. Details in `references/brand-identity.md` "Recognizable at small size," (d) no more than 3 dominant colors, (e) high res (2048×2048+ when shipped), (f) no text inside the image, (g) true transparent background (verified alpha=0 in PIL, key out near-white pixels if gpt-image-2 painted them). The symbol is shipped as a high-res transparent PNG — **NOT traced to SVG**. Only the wordmark gets vectorized in the brand kit. See `references/brand-identity.md` "Symbol output rules" + "Logo Pipeline."
|
||||
6. **Adjective audit — verify the 3 boards diverge across multiple dimensions, not just palette.** Even when the user's brief is unified (one era / one mood / one set of references), the 3 boards must feel like 3 different brands. Write the top-3 primary adjectives for each board. If 2 boards share 2+ primary adjectives ("both bright, both maximalist, both cartoon-y"), they're collapsing — push apart on density / saturation / layout philosophy / type energy / voice register. See `references/brand-guidelines.md` "Unified aesthetic briefs STILL require structural divergence" for dimensions. The user gave you 3 references for a reason — interpret each as a SUB-territory, not as overlapping inputs into one mood.
|
||||
7. **Brief delivery check: every board delivers on the brief at full intensity.** Whatever character the brief named (Y2K, kawaii, brutalist, etc.), every board delivers it 100%. Variety comes from sub-territories within the brief, not from stripping the brief. NEVER make a board muted/dusty/generic to differentiate — that breaks the brief in service of variety. Restraint is allowed (sparse editorial layouts are valid), but stripping character is forbidden. See `references/brand-guidelines.md` "Divergence ≠ subtraction. Every board delivers on the brief."
|
||||
8. **Era palette check.** If the brief names an era (Y2K, 90s, 80s, mid-century, etc.), all 3 boards must use era-appropriate palette signatures. "Muted dusty cream" is rarely an era's actual signature — Y2K = iMac gel / Tamagotchi candy / chrome / holographic / cyber, NOT cottage-core. See `references/brand-guidelines.md` "Era palettes are specific."
|
||||
9. **Texture check — generated, ambient, NOT a pattern.** Retro/era briefs need era-appropriate texture, but it must be ATMOSPHERIC (subtle grain / VHS noise / grainy gradient / soft film grain), NEVER a recognizable pattern (literal halftone dots, scattered glitter, visible scanlines). If a viewer can name the texture as a noun ("halftone dots! glitter!") it's too literal — they should describe it as atmosphere ("kind of grainy", "feels VHS-y"). **Generate via `generate_image` with `provider="gpt-image-2"`** — CSS gradients read as stylesheet. Prompt for "ambient grain / atmospheric noise / grainy gradient," NEVER for "pattern / dots / flakes." Apply on the BODY (full board), `mix-blend-mode: overlay/multiply/screen`, `opacity: 0.12–0.30`. See `references/brand-guidelines.md` "Texture is era signal — but it must be AMBIENT, not a pattern."
|
||||
Each option includes: name + tagline, color palette, typography direction, voice & tone, logo concept, brand story, photography direction (product + lifestyle/mood), UI/website direction, example brands.
|
||||
|
||||
If you can't tell which brand is which without reading the labels — the boards have failed. Regenerate.
|
||||
After presenting all 3 in text, **build a 3-page brand board PDF** (one page per option) so the user can see each identity before committing.
|
||||
|
||||
For board PDF assembly via Chrome + `pdfunite`, see `references/brand-identity.md` "Visuals (Required)".
|
||||
**Each option must include:**
|
||||
- **Wordmark** in the brand's display font — use the user's existing wordmark if they have one they like; propose a new one if they need a logo or don't like their current one. A new wordmark must have custom letter treatment: adjusted spacing, ligature, cut, terminal, case, underline, or other ownable detail. It is not just a Google Font typed in a color.
|
||||
- **Symbol/mark** — a standalone graphic that lives without the wordmark. Use the user's existing symbol if they have one they like; propose a new one otherwise. Even if the user keeps their wordmark, propose a symbol if they don't have one — favicons and app icons need a non-typographic mark. For new symbols, default to a generated PNG via `mcp__pika__generate_image` with `provider="gpt-image-2"` when visual quality, texture, detail, or originality matters. Ask for a clean isolated mark on transparent background, no baked-in letters, no watermark, no mockup, centered in a square. Use inline SVG only if the mark is intentionally simple, can be drawn cleanly by hand, and passes small-size QA. Must work at 16×16 AND 512×512.
|
||||
- **Seal / badge** — if the option uses a seal, stamp, badge, or monogram, it must be readable and ownable at small and medium sizes. It cannot be a generic circular font lockup, clip-art crest, or low-contrast decorative filler.
|
||||
- Tagline (8 words max)
|
||||
- Voice sample with visible "VOICE" label (one quoted sentence, 14 words max)
|
||||
- Compact board story (~35 words max, min 2 sentences)
|
||||
- Lifestyle world description with visible "WORLD" label (~22 words max, min 1 full sentence)
|
||||
- Lifestyle mood image (generated via gpt-image-2)
|
||||
- 4-color palette with hex + role labels
|
||||
- Display + body type specimens with named fonts
|
||||
|
||||
**Do not build the full guidelines PDF until the user picks an option.** Present the 3-board PDF and ask: "Which one feels right? Pick one, mix elements from two, or tell me to swing further in a direction."
|
||||
Brand board pages should look different enough that the user can tell which identity they are seeing before reading the labels. Don't use the same template recolored 3 times; each board's layout should embody the option's design philosophy. A magazine-cover option should look like a magazine cover (full-bleed photo, masthead-style); a soft consumer option should look like a homepage hero (rounded shapes, soft circles for swatches); an editorial option should look like a literary spread (huge italic centered, inset photo). See `references/brand-guidelines.md` "Brand Board Layout — Differentiate per Option" for examples.
|
||||
|
||||
### Step 3 — Build the Full Brand Guidelines PDF
|
||||
If you can't physically tell which brand you're looking at *without* reading the labels — regenerate.
|
||||
|
||||
Once user confirms a board, build the brand guidelines PDF for that direction.
|
||||
**Page size:** 1200×850px. Renderer: Pika MCP `mcp__pika__html_to_pdf` for PDF output and `mcp__pika__html_to_png` for PNG QA previews. Both use server-side Chromium; do not run local WeasyPrint or Chrome headless by default. The delivered PDF and QA PNGs must come from the MCP render path. If a local/sandbox/browser fallback is used for debugging, discard that output, rerender through MCP, and QA the final MCP PNGs. If the MCP and fallback renders disagree, fix the HTML/CSS; do not ship the fallback render.
|
||||
|
||||
**Page count is conditional:**
|
||||
- **14 pages** — non-digital brands (product, restaurant, fashion, service, etc.) with single-medium imagery. Icons page is skipped.
|
||||
- **15 pages** — digital brands (app / web / SaaS) with single-medium imagery. Includes Icons page.
|
||||
- **15 pages** — non-digital brands with hybrid imagery (photo + illustration). Imagery splits to 2 pages; Icons skipped.
|
||||
- **16 pages** — digital brands with hybrid imagery. Both Icons and split-Imagery present.
|
||||
**Board quality gate:** the 3-page preview must pass visual QA, not just render QA.
|
||||
|
||||
**Page structure (full list — apply conditionally per above):**
|
||||
1. Inspect each PNG preview before sending the PDF. Fail ugly density, weak hierarchy, muddy one-note palette, unreadable small text, empty mockup/image slots, clipped text, and body copy or non-masthead text intersecting icons, swatches, seals, photos, phone mockups, or decorative rules. Also inspect the symbol at 16×16 and 512×512; if the small-size read is illegible, muddy, too generic, or collapses into noise, regenerate or simplify before delivery. Website/social/app mockups are optional on brand boards; do not add them unless they contain real content. If included, flat color rectangles count as empty placeholders unless the section is explicitly a palette specimen.
|
||||
2. Run `mcp__pika__analyze_media` on each board PNG. The prompt must start: "Answer with PASS or FAIL on the first line, then explain." Ask about all issues listed above. Masthead wordmark/tagline/issue metadata overlays on photos are allowed only when they use deliberate negative space or a contrast scrim and pass contrast QA as defined in `references/brand-guidelines.md` Rule 5. Body copy on images is still forbidden.
|
||||
3. Handle tool state before interpreting the result. If `mcp__pika__analyze_media` returns `{task_id, status: "running"}`, poll `mcp__pika__task_status({task_id})` until terminal. Treat the QA as unavailable and halt with a manual-review warning if the tool is missing, raises `tool_not_found`, `provider_unavailable`, `unsupported_media_type`, `rate_limited`, `quota_exceeded`, `auth_error`, any HTTP 4xx/5xx error envelope, or a transport error, says it cannot analyze the image, or returns final text that does not match ``/^\s*[`*]{0,2}(PASS|FAIL)\b/``.
|
||||
4. Interpret the result with that regex only. Fix every captured FAIL before delivery. If the captured result is PASS but the explanation lists a blocking collision, clipping, unreadable text, or missing required board content, treat it as FAIL. Do not proceed silently.
|
||||
|
||||
**Font rule:** fresh fonts per brand AND fonts must have character. Never default to Inter / Karla / Outfit / DM Sans / Lato — they have no point of view as a display face. Explore the full Google Fonts library. See `references/brand-guidelines.md` "Must Have Character — Don't Default to Safe Fonts" for approved high-character options by vibe (Fraunces, Instrument Serif, Bricolage Grotesque, Funnel Display, Bodoni Moda, Reddit Mono, etc.). For MCP rendering, use HTTPS font URLs or inline `data:font/...` sources; local `file://` font paths are not available server-side.
|
||||
|
||||
Do not build the full 15-page guidelines PDF until the user picks an option — that wastes time on rejected identities.
|
||||
|
||||
Present all 3 clearly. Ask the user to pick one, or mix elements from different options.
|
||||
|
||||
### Step 4 — Build the Full Brand Guidelines PDF
|
||||
|
||||
Once user confirms an identity, build the 15-page brand guidelines PDF.
|
||||
|
||||
**Page structure (15 pages; becomes 16 if hybrid imagery split):**
|
||||
|
||||
1. **Cover** — Brand name, tagline, hero mood image. Full-bleed.
|
||||
2. **Strategy & Positioning** — Direction name. Positioning statement (one punchy sentence). Target customer described vividly (one specific person, not a demographic). 3-4 reference brands with "borrow this" notes.
|
||||
2. **Strategy & Positioning** — Direction name. Positioning statement (one punchy sentence). Target audience system: primary segment, secondary segment(s), and one vivid anchor persona inside the primary segment. Do not describe only one over-specific customer. 3-4 reference brands with "borrow this" notes.
|
||||
3. **Brand Foundation** — Mission. Brand values (3-5). The "why this exists" story (2-3 paragraphs of real copy in brand voice — not a template).
|
||||
4. **Logo** — Primary mark + all variants (horizontal, icon-only, reversed), usage rules (on dark / on light / on color), logo mark explanation, clear space rule.
|
||||
5. **Logo Don'ts** — Explicit misuse rendered in CSS: never stretch, never rotate, never wrong background, never recolor, never use drop shadow. Show each violation visually with a ✗ label.
|
||||
6. **Color** — All swatches with hex + RGB + CMYK, primary pairings, accessibility/contrast note, never-do combinations. Full-bleed color columns, not swatches floating on white.
|
||||
7. **Typography** — Full hierarchy (H1 through caption with exact px sizes), display/accent/body fonts, usage rules per context, type on color backgrounds, minimum sizes.
|
||||
8. **Icons** *(digital brands only — skip for product, fashion, restaurant, service brands)* — UI icon system: 8-12 essential icons (arrow-right, check, close, plus, settings, search, user, bell, menu, info, etc.) rendered in the brand's geometric style + stroke/corner/grid rules + library recommendation for icons beyond the set. See `references/brand-guidelines.md` "Icons Page — Structure & Rules" section.
|
||||
8. **Icons** — UI icon system: 8-12 essential icons (arrow-right, check, close, plus, settings, search, user, bell, menu, info, etc.) rendered in the brand's geometric style + stroke/corner/grid rules + library recommendation for icons beyond the set. See `references/brand-guidelines.md` "Icons Page — Structure & Rules" section.
|
||||
9. **Voice & Tone** — Tone adjectives, copy examples by context (headline, body, button, error state, social caption), forbidden words/phrases. Show actual brand copy, not generic example copy.
|
||||
10. **Imagery Rules** — adapts to the brand's medium. **Photography-led** → photography rules (subject/light/color/cast/texture/forbidden + 1 example photo). **Illustration-led** → illustration rules (style/color/line/character/composition/forbidden + 1 example illustration). **Hybrid** (both equally) → split into two pages, guidelines becomes 16 pages. See `references/brand-guidelines.md` "Imagery Rules Page — Adapts per Brand."
|
||||
11. **Visual World** — Full-bleed 4-column grid of 4 images matching the brand's medium mix (all photos, all illustrations, or mixed). Cast must be racially diverse for any people-featuring images.
|
||||
@@ -194,48 +225,24 @@ Once user confirms a board, build the brand guidelines PDF for that direction.
|
||||
14. **Digital / Social** — Website hero aesthetic (colors, fonts, layout feel), Instagram grid style (3×3 mockup with color palette + caption tone), story template (brand colors + logo placement), link-in-bio layout.
|
||||
15. **Do & Don't** — 5 dos and 5 don'ts, brand-specific and actionable. Not generic ("do use the logo correctly") — brand-specific ("do leave a full em-dash of space around the wordmark in social posts; never crop our tagline mid-word").
|
||||
|
||||
**ORDER IS MANDATORY:** Page 2 (strategy) comes before page 3 (foundation) which comes before logo/color/type. Strategy frames everything else.
|
||||
Keep Page 2 (strategy) before Page 3 (foundation), and keep both before logo/color/type. Strategy frames every visual decision that follows.
|
||||
|
||||
**ALL PAGES MUST APPEAR.** Never skip a page just because the brand "doesn't have packaging" — adapt the touchpoints page to the brand type instead.
|
||||
Include every page in the structure. If the brand has no packaging, adapt the touchpoints page to the brand type instead of skipping it; the guidelines should still show how the identity survives in real contexts.
|
||||
|
||||
All build rules in `references/brand-guidelines.md` apply: local rendering rules (Chrome headless preferred; WeasyPrint allowed when following its table-only constraints), explicit pixel dimensions, no opacity on images, no text on images, font `@font-face` with absolute `file://` paths, image compression, no duplicate generated images across deck, text contrast thresholds, and mandatory pre-send QA.
|
||||
**Completion gate:** do not deliver a condensed or partial guidelines PDF. A
|
||||
condensed 6-page deck, missing Visual World page, missing Touchpoints page, or
|
||||
missing Brand Applications page is a failed checkpoint, not a final deliverable.
|
||||
If time runs out, stop with a resumable checkpoint that lists completed pages,
|
||||
missing pages, generated asset URLs, and the next render step. Do not present
|
||||
the deck as done until all mandatory pages have rendered and passed QA.
|
||||
|
||||
**Mandatory pre-delivery checks for the full guidelines PDF (Step 3):**
|
||||
All build rules in `references/brand-guidelines.md` apply: server-side Chromium render contract, explicit 1200×850 page dimensions, HTTPS/data-URI assets, no load-bearing text on generated images, no duplicate generated images across deck, text contrast thresholds on dark backgrounds, and mandatory pre-send QA previews for every page.
|
||||
|
||||
1. **Product clarity gate — Page 1 (Cover) + Page 2 (Strategy) must make the product unambiguous.** A reader landing on either page must be able to answer "what does this brand sell?" without ambiguity. Cover hero photo must show the product IN USE (for digital products = phone/device showing the actual result, or person using it). Cover subtitle + Page 2 opening paragraph must communicate what the product literally IS, in the brand's voice (clarity through content — not robotic templates). See `references/brand-guidelines.md` Page-Specific Notes for Page 1 + Page 2.
|
||||
**Deliver as PDF.** `mcp__pika__html_to_pdf` returns a CDN `file_url`; give the user that URL. If exporting a local kit later, download that `file_url` into `~/Desktop/[brand-name]-brand-guidelines.pdf` or the kit folder as `brand-guidelines.pdf`. Do not use `mcp__pika__upload_asset` for PDFs; that tool still only accepts images/audio/video.
|
||||
|
||||
2. **Photography-in-use gate — hero photos show the product, not a representational object.** The cover, photography direction page, touchpoints page, and application page hero photos must show the actual product being used. The brand's *symbol* can be a stylized abstract mark (a heart, a charm shape, etc.) — that's logo language. But brand *photography* must show the actual product. Apply "the keychain test": if a stranger saw the cover with no other context, would they correctly guess what the brand sells? If they'd guess the brand sells the object shown — re-shoot. See `references/brand-guidelines.md` "Photography Must Show the Product IN USE."
|
||||
### Step 5 — Export the Brand Kit (after user confirms guidelines)
|
||||
|
||||
3. **Contrast gate — every text/background pair must be verifiably legible.** Audit every text-on-color combination. Auto-fails: pink-on-pink, lime-on-cream, dark-on-dark photo, any small text within 25% luminance of its container. 4.5:1 minimum for body, 3:1 for large display. Fix any failure before delivery — pink display text on a pink background block is a hard ship-block.
|
||||
|
||||
4. **Complex-fill scale gate — chrome / holographic / multi-stop gradients ONLY at hero scale.** When the brand wordmark appears at multiple sizes (e.g. cover hero AND small page-header breadcrumb), the cover uses the complex chrome fill BUT the small-scale instances use a solid color with optional stroke. Multi-stop gradients muddy at small sizes. Document both versions on the Logo page.
|
||||
|
||||
5. **Render-and-READ every single page individually before merge, including ACTUAL zoom crops.** Render every page in the 14–16-page deck to PNG. Open and READ each PNG full-size (thumbnail pass). Then, for each page, use PIL to crop into 800×800 regions around every text-bearing area (title, subtitle, every section header, every body block, every card pull-quote, every label, every hex code, every footer band). **Read each crop file.** Glancing at the full-page PNG is NOT zoom-reading — it's the thumbnail pass. Without the actual crop-and-read, contrast failures, shaded-font issues, footer overlaps, and text-wrap flaws all hide in plain sight. If you find yourself about to ship without having opened named `zoom-{page}-{region}.png` files — you haven't done the zoom pass. Go back. See `references/brand-guidelines.md` "Three-pass QA" for the full list of regions to crop.
|
||||
|
||||
**Deliver as PDF.** Save to `~/Desktop/[brand-name]-brand-guidelines.pdf` and tell the user the local path. Do not upload or host the PDF unless the user explicitly asks for a hosted copy. If the user is not on a Mac, save to the project working directory and mention the path in your reply. Either way: emit the file path so the user can open it — never attach the PDF as a file in the chat.
|
||||
|
||||
### Step 4 — Export the Brand Kit (HARD GATE — only after user explicitly approves the guidelines)
|
||||
|
||||
🛑 **STOP. This is a hard gate, not a soft suggestion.**
|
||||
|
||||
After delivering the 14–16-page guidelines PDF, you MUST wait for **explicit user approval** before exporting the brand kit. The kit codifies the final brand, so only build it once the brand is locked.
|
||||
|
||||
**Explicit approval looks like:**
|
||||
- "yes" / "approved" / "ship it" / "go" / "looks good, export the kit"
|
||||
- "this is great, do the kit"
|
||||
- Specific changes followed by approval after those changes are made
|
||||
|
||||
**Explicit approval does NOT look like:**
|
||||
- "just make it cute" (this is creative direction, not approval)
|
||||
- "ok" alone (ambiguous — could be acknowledging a previous message)
|
||||
- "i like it" without referring to the guidelines specifically
|
||||
- The user simply not objecting
|
||||
|
||||
**When you deliver the guidelines PDF, your message must end with an explicit ask** like: "ready to lock this in and export the brand kit, or want to adjust anything first?" Then **stop and wait**. Do not call any tool that begins the kit export. Do not stage the kit directory. Do not even copy files into a `kit/` folder. The brand is not approved until the user says so.
|
||||
|
||||
**Why this gate matters:** the kit zip locks in every decision — color tokens, voice prompts, logo variants, brand.md — into a portable artifact people will share and reference. Exporting before approval means we ship a kit that might need to change, which defeats the whole point of a kit. Worse, the user feels skipped.
|
||||
|
||||
If the user gives feedback on the guidelines (any feedback), iterate on the guidelines PDF first, then ask for approval again. Do not export the kit until they explicitly say go.
|
||||
After delivering the 15-page guidelines PDF, **wait for explicit user confirmation** that they're happy with the brand. Don't auto-export — the kit codifies the final brand, so only build it once the brand is locked.
|
||||
|
||||
Then build a comprehensive brand kit zip that lets the user produce on-brand work anywhere — in Claude, GPT, Figma, with a designer, with a developer.
|
||||
|
||||
@@ -244,29 +251,21 @@ Then build a comprehensive brand kit zip that lets the user produce on-brand wor
|
||||
```
|
||||
[brand-name]-brand-kit.zip
|
||||
├── brand.md # comprehensive machine-readable spec
|
||||
├── brand-guidelines.pdf # full 14–16-page guidelines PDF (the visual deliverable)
|
||||
├── brand-guidelines.pdf # full 15-page guidelines PDF (the visual deliverable)
|
||||
├── README.md # 1-page how-to-use guide
|
||||
├── logo/
|
||||
│ ├── symbol/ # standalone mark — RASTER ONLY, no SVG (symbol is gen'd PNG, not traced)
|
||||
│ │ ├── symbol-[color]-16.png # favicon size
|
||||
│ │ ├── symbol-[color]-32.png
|
||||
│ │ ├── symbol-[color]-64.png
|
||||
│ │ ├── symbol-[color]-128.png
|
||||
│ │ ├── symbol-[color]-256.png
|
||||
│ │ ├── symbol-[color]-512.png
|
||||
│ │ ├── symbol-[color]-1024.png
|
||||
│ │ └── symbol-[color]-2048.png # high-res master, transparent background
|
||||
│ ├── wordmark/ # the brand name — vectorized via text-as-paths
|
||||
│ │ ├── wordmark-[color].svg # Google Font converted to outlined paths (renders without font file)
|
||||
│ │ └── wordmark-[color].png # 1024-wide raster fallback
|
||||
│ └── lockup/ # symbol + wordmark together at locked measurements
|
||||
│ ├── symbol/ # standalone mark, one set per color variant
|
||||
│ │ ├── symbol-[color].png # transparent background, 1024×1024+
|
||||
│ │ ├── symbol-[color].svg # only when vector-authored
|
||||
│ │ └── symbol-[color].pdf # vector PDF or raster PDF wrapper
|
||||
│ ├── wordmark/ # the brand name styled
|
||||
│ │ └── wordmark-[color].{svg,png,pdf}
|
||||
│ └── lockup/ # symbol + wordmark together
|
||||
│ ├── horizontal/
|
||||
│ │ ├── lockup-h-[color].svg # symbol PNG embedded inline + wordmark as paths
|
||||
│ │ └── lockup-h-[color].png # assembled lockup as raster, 1024-wide
|
||||
│ │ └── lockup-h-[color].{png,pdf} + svg when fully vector-authored
|
||||
│ └── stacked/
|
||||
│ ├── lockup-s-[color].svg
|
||||
│ └── lockup-s-[color].png
|
||||
├── icons/ # digital brands only — UI icons from the Icons page as SVGs
|
||||
│ └── lockup-s-[color].{png,pdf} + svg when fully vector-authored
|
||||
├── icons/ # the 12 UI icons from page 8 as SVGs
|
||||
│ ├── arrow-right.svg
|
||||
│ ├── check.svg
|
||||
│ ├── close.svg
|
||||
@@ -300,7 +299,7 @@ Then build a comprehensive brand kit zip that lets the user produce on-brand wor
|
||||
|
||||
**brand.md** — see `references/brand-md-template.md` for the full structure. It must include:
|
||||
- Quick reference block (name, tagline, primary color, fonts, voice in one scannable section)
|
||||
- Positioning + customer
|
||||
- Positioning + audience segments
|
||||
- Mission, values, story
|
||||
- Voice & tone (adjectives, copy examples by context, forbidden words)
|
||||
- Colors (table with hex / RGB / CMYK / Pantone / role)
|
||||
@@ -313,12 +312,12 @@ Then build a comprehensive brand kit zip that lets the user produce on-brand wor
|
||||
- Reference brands with "borrow this" notes
|
||||
- How-to-use section telling downstream tools/people how to apply the spec
|
||||
|
||||
**Logo asset pipeline:**
|
||||
1. **Symbol assets** — the symbol is a generated raster mark, so export PNG only. Start from the approved 2048×2048 transparent master, verify true alpha, then generate `16/32/64/128/256/512/1024/2048` PNGs per needed color/background variant. Do **not** trace it to SVG and do **not** claim it is vector.
|
||||
2. **Wordmark assets** — render the brand name as real font text, then convert the chosen font text to outlined paths for `wordmark-[color].svg`; also export a 1024-wide transparent PNG fallback. The wordmark is reproducible because it is typography, not an image-generation artifact.
|
||||
3. **Lockup assets** — assemble the approved symbol PNG + outlined wordmark at the locked measurements. Export `lockup-[orientation]-[color].svg` with the PNG embedded inline and the wordmark as paths, plus a 1024-wide PNG fallback. Keep geometry identical across color variants.
|
||||
4. **Optional print PDF** — only add PDFs if the user or printer specifically asks. A PDF may embed the raster symbol plus vector wordmark, but it is not a pure-vector logo file.
|
||||
5. **Icon SVGs** — for digital brands only, write the icon set as standalone SVGs with `stroke="currentColor"`, `viewBox="0 0 24 24"`, and the brand's chosen stroke weight + corner style applied consistently. See `references/brand-guidelines.md` "Icons Page — Structure & Rules" for which icons to include.
|
||||
**Asset generation pipeline:**
|
||||
1. **Symbol master asset** — for new marks, prefer the generated PNG route: call `mcp__pika__generate_image` with `provider="gpt-image-2"` for a clean isolated mark on transparent background, centered, no text, no watermark, no mockup, no shadows. Use SVG only if the mark is deliberately simple and still reads at 16×16. Keep the best generated PNG as the source of truth when it is stronger than SVG.
|
||||
2. **Logo SVGs** — write SVG variants for wordmarks, lockups, and any symbol that is actually vector-authored. Do not trace a rich generated PNG into a weak SVG just to satisfy a vector preference. For raster-generated symbols, note in `brand.md` that the symbol master is PNG.
|
||||
3. **Logo PNGs** — export every symbol / wordmark / lockup as PNG on transparent background at `1024×1024+` for symbols and enough width for wordmarks. For SVG-derived assets, render SVG → PNG via `mcp__pika__html_to_png`: HTML wrapper with `<body style="margin:0;background:transparent;">` containing just the SVG, `raster_options.viewport_px:1024x1024`, `transparent_background:true` when available. For generated PNG symbols, preserve the original high-res transparent PNG and create color variants only when they remain crisp.
|
||||
4. **Logo PDFs** — render SVG logo wrappers via `mcp__pika__html_to_pdf` when vector source exists. For raster-generated symbols, create a PDF wrapper that embeds the high-res PNG at full resolution and label it as raster-source in the README; do not pretend it is vector.
|
||||
5. **Icon SVGs** — write each of the 12 icons as a standalone SVG with `stroke="currentColor"`, `viewBox="0 0 24 24"`, and the brand's chosen stroke weight + corner style applied consistently. See `references/brand-guidelines.md` "Icons Page — Structure & Rules" for which icons to include.
|
||||
6. **Design tokens** — generate all three files from the brand spec:
|
||||
- `tokens.css` — `:root` block with `--color-*`, `--font-*`, `--font-size-*`, `--line-height-*`, `--space-*`, `--radius-*`, `--shadow-*` custom properties
|
||||
- `tokens.json` — same content as JSON object with sections: `color`, `font`, `fontSize`, `lineHeight`, `spacing`, `radius`, `shadow`
|
||||
@@ -331,58 +330,69 @@ Then build a comprehensive brand kit zip that lets the user produce on-brand wor
|
||||
- `error-message.md` — task starter: how the brand handles error/empty/loading states in voice (warm not robotic, specific not vague)
|
||||
- **`photography.md`** — task starter for generating brand-style photography (gpt-image-2 etc.). Must include: master prompt template tailored to the brand's photo direction (subject, light, color grade, cast diversity, texture); explicit "what to AVOID in the prompt" list (studio strobes, stock terms, glass coworking spaces, "engineers at laptops," "professional," "premium," etc.); banned cliché concepts list (hourglasses, lightbulbs, handshakes, network nodes, glowing brains, etc.); subject substitutes for "person doing X"; quality requirements (butter accent, diversity, film grain, documentary); explicit no-text guardrail string; note about never naming real publications.
|
||||
- **`illustration.md`** — only if the brand uses illustration as a medium. Task starter for generating brand-style illustrations. Master prompt template with strict palette + style rules (flat vector / line art / etc), banned elements (gradients, drop shadows, 3D, photographic textures), when to use illustration vs photography. Skip this file entirely if the brand has no illustration in its visual world.
|
||||
8. **Brand fonts** — download the actual font files from Google Fonts (or wherever the brand fonts live) and include in `fonts/`:
|
||||
8. **Brand fonts** — local export step. Download the actual font files from Google Fonts (or wherever the brand fonts live) and include in `fonts/`:
|
||||
- Variable font files when available: `[FontName]-Variable.ttf` (single file, supports all weights)
|
||||
- Or static weights at the levels the brand uses
|
||||
- GitHub mirror pattern: `https://github.com/google/fonts/raw/main/ofl/[fontname]/[FontName][wght].ttf`
|
||||
- Add a `fonts/README.md` noting the license (OFL is common, allows redistribution) + Google Fonts URL for online installation
|
||||
9. **Brand guidelines PDF** — copy the 14–16-page guidelines PDF (the visual deliverable produced in Step 3) into the kit as `brand-guidelines.pdf`. The kit is incomplete without it.
|
||||
9. **Brand guidelines PDF** — download the `html_to_pdf.file_url` from Step 4 into the kit as `brand-guidelines.pdf`. The kit is incomplete without it.
|
||||
10. **README.md** — 1-page guide telling the user: what's in the kit, how to use brand.md with AI tools, which logo file for which context, where to install fonts (local TTFs or Google Fonts URLs), how to use the photography/illustration prompts.
|
||||
11. **Zip everything**: `zip -r [brand]-brand-kit.zip brand.md brand-guidelines.pdf README.md logo/ icons/ fonts/ tokens/ prompts/`
|
||||
|
||||
**Brand-kit completion gate:** when the user confirms export, or when
|
||||
`export_kit` is set in `--config`, the brand kit zip is a required deliverable.
|
||||
Do not mark the brand kit complete until the zip exists and contains
|
||||
`brand.md`, `brand-guidelines.pdf`, README, logo assets, icons, fonts, tokens,
|
||||
and prompts. If any required file cannot be produced, stop with a resumable
|
||||
checkpoint and list the missing files instead of shipping a partial zip.
|
||||
|
||||
**README.md** — 1-page guide telling the user:
|
||||
- What's in the kit
|
||||
- How to use `brand.md` with AI tools (paste into Claude/GPT to generate on-brand work)
|
||||
- Which logo file to use for which context (web favicon → symbol PNG; print collateral → high-res lockup PNG or optional PDF; web header → wordmark SVG; etc.)
|
||||
- Which logo file to use for which context (web favicon → symbol PNG; print collateral → PDF; web header → wordmark SVG; etc.)
|
||||
- Font installation links (Google Fonts URLs)
|
||||
|
||||
**Delivery:**
|
||||
- Save zip to `~/Desktop/[brand-name]-brand-kit.zip` for local Mac users.
|
||||
- Save to the project working directory if Desktop is unavailable.
|
||||
- Tell the user what's in it and link to the `brand.md` so they can preview without unzipping.
|
||||
- If the environment cannot download fonts or write a local zip, do not ship a
|
||||
partial kit. Stop with a resumable blocked checkpoint that lists the
|
||||
completed artifacts (`file_url`, `brand.md`, tokens, logo assets), the missing
|
||||
files, and the exact filesystem/network blocker.
|
||||
- Upload the completed zip to a CDN if needed.
|
||||
- Tell the user what's in the completed zip and link to the `brand.md` so they can preview without unzipping.
|
||||
|
||||
## Key Principles
|
||||
|
||||
- **The input is the brief.** Don't ask for lengthy intake forms. Read what's in front of you and ask 3-5 precise questions.
|
||||
- **Be specific about customers.** Vague audiences = weak brands. Push for vivid specificity — one person, not a demographic.
|
||||
- **3 boards at the main choice point.** Step 2 always gives the user three visual brand boards, not text-only directions or template recolors.
|
||||
- **Be specific about customers without narrowing the brand to one person.** Vague audiences = weak brands, but one hyper-specific individual can make the output unusably narrow. Define audience segments first: a primary segment, 1-2 secondary segments, and one anchor persona that makes the primary segment feel concrete.
|
||||
- **3 options at each choice point.** Direction (step 2), then identity (step 3). Always 3.
|
||||
- **Opinionated but collaborative.** Present your read confidently. They can push back.
|
||||
- Generate actual copy — don't give templates with [BRACKETS]. Write real words in the brand voice.
|
||||
- **All images must look real and crafted.** Generated lifestyle/touchpoint images need film grain, natural light, slight imperfections, editorial composition. Banned: perfect symmetry, gradient backgrounds, studio strobes, stock-photo energy, AI-smooth surfaces, floating objects on white. If it looks fake — regenerate.
|
||||
- **Primary deliverable is the guidelines PDF.** The brand kit is exported only after explicit approval. Not a press kit. Not a launch package. Not a social calendar.
|
||||
- **Single deliverable.** One brand guidelines PDF. Not a press kit. Not a launch package. Not a social calendar. Just the brand.
|
||||
|
||||
---
|
||||
|
||||
## Brand Quality Standards — Non-Negotiable
|
||||
## Brand Quality Standards
|
||||
|
||||
Every brand produced by this skill must meet the following standards without exception. Generic is a failure state.
|
||||
Every brand produced by this skill should meet the following standards. Generic output is a failure state because the deliverable is meant to guide real design decisions, not decorate a template.
|
||||
|
||||
### The Anti-Generic Test
|
||||
|
||||
Before delivering anything, ask: *Could this be a brand for literally anything else?* If yes — it's not done.
|
||||
|
||||
Strong brand = specific product/service + specific person + specific point of view. Weak brand = vibes + aesthetic mood board + empty tagline. Never deliver the second.
|
||||
Strong brand = specific product/service + clear audience model + specific point of view. Weak brand = vibes + aesthetic mood board + empty tagline. Never deliver the second.
|
||||
|
||||
### Copy Standards
|
||||
|
||||
**What good brand copy sounds like:**
|
||||
- It makes a specific claim: "Heavy wool. Made to last a decade." / "Built for one quiet hour a day."
|
||||
- It has a point of view: "Not trend-led. Not mass-made."
|
||||
- It talks to one person, not a demographic: "The app you reach for before checking your phone."
|
||||
- It can speak concretely to a reader inside a segment: "The app you reach for before checking your phone." This is copy style, not audience strategy; do not collapse the brand's audience model to only that reader.
|
||||
- It creates tension or contrast: "Handmade. Overused. On purpose."
|
||||
- It trusts the reader: no over-explaining, no "perfect for any occasion", no "cozy vibes"
|
||||
|
||||
**What bad brand copy sounds like (NEVER write this):**
|
||||
**What bad brand copy sounds like:**
|
||||
- "Crafted with love" / "Made with care" / "Designed with passion"
|
||||
- "Perfect for any occasion" / "A timeless addition"
|
||||
- "Quality you can feel" / "Designed to inspire"
|
||||
@@ -402,7 +412,7 @@ Strong brand = specific product/service + specific person + specific point of vi
|
||||
- Pages feel designed, not assembled
|
||||
- Whitespace is intentional, not default padding
|
||||
|
||||
**What generic brand design looks like (NEVER produce this):**
|
||||
**What generic brand design looks like:**
|
||||
- Equal-sized boxes arranged in a grid
|
||||
- Body copy the same size as everything else
|
||||
- Centered everything
|
||||
@@ -415,7 +425,7 @@ Strong brand = specific product/service + specific person + specific point of vi
|
||||
|
||||
### Photography & Diversity Standards
|
||||
|
||||
**All generated images featuring people must show racial diversity.** No exceptions.
|
||||
Generated image sets featuring people should show racial diversity. This avoids defaulting every brand world to the same narrow cast.
|
||||
- Default to a mixed cast across all 4+ lifestyle images: include Black, Asian, Latina, South Asian, Middle Eastern, or mixed-race subjects
|
||||
- Vary body types, not just skin tone
|
||||
- If only one person is shown, make a deliberate choice about who that person is — don't default to white/light-skinned
|
||||
@@ -429,8 +439,8 @@ Strong brand = specific product/service + specific person + specific point of vi
|
||||
|
||||
### Deck / Guidelines Design Standards
|
||||
|
||||
- Typography must load. Always use absolute paths for fonts in WeasyPrint. Always verify loaded fonts before signing off on a render. If fonts fall back to system defaults — the deck is broken, not deliverable.
|
||||
- See `references/brand-guidelines.md` for the full set of WeasyPrint technical rules.
|
||||
- Typography must load. Use HTTPS font URLs or inline `data:font/...` sources in MCP-rendered HTML; local `file://` paths are not available to server-side Chromium. Always verify loaded fonts before signing off on a render. If fonts fall back to system defaults — the deck is broken, not deliverable.
|
||||
- See `references/brand-guidelines.md` for the full MCP render contract and QA rules.
|
||||
- Every page must have a clear visual hierarchy — one thing to look at first.
|
||||
- Full-bleed photography pages should feel like magazine spreads, not slideshow slides.
|
||||
- Color palette pages: full-bleed color columns, not swatches floating on white.
|
||||
@@ -438,7 +448,7 @@ Strong brand = specific product/service + specific person + specific point of vi
|
||||
- Voice page: show actual brand copy, not generic example copy.
|
||||
- Touchpoints page: must include generated photographs of actual touchpoints — never CSS boxes.
|
||||
|
||||
**Deliver as PDF always.** Never as individual page images. Save the PDF locally and share the path.
|
||||
Deliver the guidelines as one PDF, not as individual page images. Return the `mcp__pika__html_to_pdf` CDN URL and save a local copy when practical for the brand-kit zip.
|
||||
|
||||
### The Taste Check
|
||||
|
||||
@@ -453,6 +463,20 @@ If any answer is "not sure" — improve it before delivering. Strong and specifi
|
||||
|
||||
---
|
||||
|
||||
## Load-bearing phrases
|
||||
|
||||
These are the anchors that keep this skill from drifting into generic brand-book output:
|
||||
|
||||
| Phrase | Where | Why load-bearing |
|
||||
|---|---|---|
|
||||
| `different business answer — not aesthetic variations` | Step 2 directions | Forces positioning variety before visual variety. |
|
||||
| `fonts must have character` | Step 3 identity options | Prevents safe-font defaults from making every brand feel interchangeable. |
|
||||
| `no load-bearing text on generated images` | Guidelines build rules | Keeps brand claims editable and legible in deterministic HTML/PDF. |
|
||||
| `film grain, natural light, slight imperfections` | Image quality standards | Pushes lifestyle/touchpoint images away from stock-photo smoothness. |
|
||||
| `Name a specific ethnicity per prompt` | Diverse-cast recovery | Fixes the model tendency toward all-white casts more reliably than generic diversity language. |
|
||||
|
||||
---
|
||||
|
||||
## Engine choice: gpt-image-2 (with caveats)
|
||||
|
||||
Default to `gpt-image-2` at `quality: "medium"` for all brand imagery. Why:
|
||||
@@ -460,7 +484,7 @@ Default to `gpt-image-2` at `quality: "medium"` for all brand imagery. Why:
|
||||
- Strongest no-text guardrail adherence — critical for touchpoint shots (hang tag / woven label / sticker) where any baked-in text would ruin the mockup.
|
||||
- Native 3:4 / 4:3 / 9:16 ratios crop cleanly on sharp subjects without weird stretching.
|
||||
|
||||
Avoid `nano-banana-pro` for this skill — it bakes magazine-cover-style text into product shots when prompts mention "editorial." Use `seedream` only when the brand needs 2K/4K print-tier touchpoint photos; otherwise the 1K from gpt-image-2 is plenty for a 1200×850 PDF page.
|
||||
Avoid `nano-banana-pro` for this skill — it bakes magazine-cover-style text into product shots when prompts mention "editorial." 1K from gpt-image-2 is plenty for a 1200×850 PDF page; bump to gpt-image-2's 2K tier (or escalate to `seedream` for higher) only if a specific touchpoint genuinely needs print-tier resolution. (4K on gpt-image-2 is 16:9 / 9:16 only — this skill's 3:4 / 4:3 ratios route to `seedream` if 4K is required.)
|
||||
|
||||
## Runtime expectations
|
||||
|
||||
@@ -469,10 +493,11 @@ Tell the user the rough total up front — long stages without status updates fe
|
||||
| Stage | Time | Notes |
|
||||
|---|---|---|
|
||||
| Stage 0 → Step 1 (Q&A loop) | 5–15 min | User-paced; questions in one message |
|
||||
| Step 2 (3 visual brand boards PDF) | 6–10 min | Per-board: gpt-image-2 symbol + 1 mood image + Chrome render. Then pdfunite all 3. |
|
||||
| Step 3 image gen (8 photos via gpt-image-2 in 2 parallel batches of 4) | 8–12 min | The longest stage; each batch ≈ 4–6 min |
|
||||
| Step 3 page build (14–16 HTMLs + Chrome render + pdfunite) | 2–3 min | Sequential render per page |
|
||||
| Step 4 brand kit zip | 3–5 min | Symbol PNG sizes + wordmark/lockup SVG+PNG + conditional icons + tokens + fonts + prompts |
|
||||
| Step 2 (3 directions, text) | 1–2 min | Pure model output |
|
||||
| Step 3 (3 identities + brand board PDF) | 5–7 min | 3 brand boards rendered via `mcp__pika__html_to_pdf` / `mcp__pika__html_to_png` QA |
|
||||
| Step 4 image gen (8 photos via gpt-image-2 in 2 parallel batches of 4) | 8–12 min | The longest stage; each batch ≈ 4–6 min |
|
||||
| Step 4 page build (15-page HTML + MCP render) | 2–5 min | `mcp__pika__html_to_pdf` async; `mcp__pika__html_to_png` previews for QA |
|
||||
| Step 5 brand kit zip | 3–5 min | 4 colors × 3 logo types × 3 formats + 12 icons + tokens + fonts + prompts |
|
||||
|
||||
Total: ~25–45 min wall-clock excluding user response time.
|
||||
|
||||
@@ -482,12 +507,14 @@ Total: ~25–45 min wall-clock excluding user response time.
|
||||
|
||||
| Symptom | Cause | Fix |
|
||||
|---|---|---|
|
||||
| Fonts render as Times / Arial in the PDF | `@import` from Google Fonts races in Chrome headless; WeasyPrint can't resolve relative font paths | Download TTFs to a local `fonts/` dir, declare via `@font-face` with `file://` absolute paths |
|
||||
| Fonts render as Times / Arial in the PDF | Font URL not reachable by server-side Chromium, or `@font-face` points to a local path | Use HTTPS raw font URLs or inline `data:font/...` sources; render a one-page `mcp__pika__html_to_png` preview before the full PDF |
|
||||
| Generated image has baked-in magazine title or watermark | Prompt mentioned "magazine cover," "Vogue," "TIME," "Bloomberg," or any real publication | Strip publication names from prompt; append the verbatim no-text guardrail; regenerate. Describe visual qualities, not publications |
|
||||
| Touchpoint / lifestyle photo shows only forehead / hand-only crop | 9:16 portrait source got cropped to a landscape cell | Regen with `aspect_ratio: "4:3"` or `"16:9"` to match the cell aspect, OR change the layout to a portrait cell |
|
||||
| Page overflows the 850px ceiling | Headline > 60px combined with > 3 body paragraphs on the same page | Cut content, drop headline to 48px, or split across two pages. Re-render and verify with a screenshot |
|
||||
| Board technically fits but looks ugly | Too much decorative styling, tiny text, muddy one-note palette, empty mockups, or weak hierarchy | Rewrite board copy to fit the budgets in `brand-identity.md`, remove decorative microtype, increase body text to 18px+, add negative space/contrast, and rerun PNG + visual QA |
|
||||
| Text overlaps icons/swatches/seals/mockups | Decorative or absolute-positioned elements share the same reading area as copy | Give text a clean reading column/card, move graphics behind non-text areas only, and rerender. Passing `scrollHeight` is not enough if a sibling graphic occludes text |
|
||||
| Brand board pages feel like recolored templates | Same template reused with palette swaps | Rebuild from `references/brand-guidelines.md` "Brand Board Layout — Differentiate per Option" — each board's layout must physically embody its design philosophy |
|
||||
| `pdfunite` is not on user's machine | Linux user without poppler-utils, or container env | Fall back to `qpdf --empty --pages file1.pdf file2.pdf -- out.pdf` or Python `pypdf.PdfWriter().append()` |
|
||||
| User picks a hybrid identity ("02's palette + 01's voice") | Skill assumes single-option pick | Build a hybrid spec brief before Step 3, confirm with user before rendering the guidelines |
|
||||
| User asks for a hosted PDF and upload fails | PDF upload paths often reject `application/pdf` | Keep the local PDF as canonical; if hosting is required, use a user-approved file host or deployment path |
|
||||
| Multi-page merge fails | Trying to stitch local page PDFs instead of using MCP | Prefer one `mcp__pika__html_to_pdf` call with native `@page`, or use `body_pages` + `shared_head` so the server merges pages |
|
||||
| User picks a hybrid identity ("02's palette + 01's voice") | Skill assumes single-option pick | Build a hybrid spec brief before Step 4, confirm with user before rendering 15 pages |
|
||||
| PDF upload to pika MCP returns "Unsupported file type" | `mcp__pika__upload_asset` allowlist is images/audio/video only — no PDFs | Don't use `mcp__pika__upload_asset` for PDFs. Use the `file_url` returned by `mcp__pika__html_to_pdf`; optionally save a local copy |
|
||||
| Lifestyle grid all-white-cast despite diverse-cast rule | gpt-image-2 defaults to lighter skin tone when ethnicity isn't named explicitly per prompt | Name a specific ethnicity per prompt (Black, mixed-race East-Asian-and-white, East Asian, Latina, South Asian, Middle Eastern) — vary across the 4 grid prompts |
|
||||
|
||||
@@ -1,35 +1,19 @@
|
||||
# Brand Guidelines — Build Guide
|
||||
|
||||
The brand guidelines PDF is the primary deliverable of this skill. It is 14–16 pages depending on brand type, built locally with Chrome headless or WeasyPrint, and delivered as a local PDF path.
|
||||
The brand guidelines PDF is the primary (and only) deliverable of this skill. 15 pages (16 if the imagery page splits into separate photo + illustration pages), rendered through Pika MCP `html_to_pdf`, delivered as a single PDF via CDN link.
|
||||
|
||||
This guide is the technical playbook: page layouts, image generation, font rules, local rendering, QA checklist. All non-negotiable.
|
||||
This guide is the technical playbook: page layouts, image generation, font rules, MCP render contract, QA checklist. All non-negotiable.
|
||||
|
||||
## Execution Model — Local First
|
||||
|
||||
Keep deterministic production local:
|
||||
- **Local:** workspace setup, downloaded fonts, compressed image files, transparent-background cleanup, favicon tests, HTML/CSS builds, PDF rendering, PNG screenshots, crop-and-read QA, logo asset assembly, and kit packaging.
|
||||
- **Cloud:** `gpt-image-2` image generation for symbols, photography, illustration, and textures; URL/source research when the brief requires it.
|
||||
|
||||
Do not upload PDFs by default. Save PDFs and zips to `~/Desktop` on Mac, or the project working directory if Desktop is unavailable. Only create a hosted/CDN copy when the user explicitly asks.
|
||||
|
||||
## Page Structure (14–16 pages depending on brand)
|
||||
|
||||
**Page count is conditional:**
|
||||
- **14 pages** — non-digital brand (product, restaurant, fashion, service) with single-medium imagery. Icons page skipped.
|
||||
- **15 pages** — digital brand (app/web/SaaS) with single-medium imagery. Icons page included.
|
||||
- **15 pages** — non-digital brand with hybrid imagery. Imagery splits to 2 pages; Icons skipped.
|
||||
- **16 pages** — digital brand with hybrid imagery. Both Icons + split-Imagery present.
|
||||
|
||||
Renumber pages contiguously based on what's included — don't leave gaps.
|
||||
## Page Structure (15 pages — all mandatory; 16 if hybrid imagery split)
|
||||
|
||||
1. **Cover** — Full-bleed brand-specific layout. Brand name + tagline + hero mood image.
|
||||
2. **Strategy & Positioning** — Direction, positioning statement, customer, 3-4 reference brands with "borrow this" notes.
|
||||
2. **Strategy & Positioning** — Direction, positioning statement, audience segments (primary segment, secondary segment(s), and anchor persona), 3-4 reference brands with "borrow this" notes.
|
||||
3. **Brand Foundation** — Mission, values (3-5), why this exists (story in brand voice).
|
||||
4. **Logo** — Primary mark + variants (horizontal, icon-only, reversed), usage rules, clear space.
|
||||
5. **Logo Don'ts** — Misuse rendered in CSS with ✗ labels.
|
||||
6. **Color** — Swatches with hex+RGB+CMYK, full-bleed color columns (not floating swatches).
|
||||
7. **Typography** — Full hierarchy with px sizes, display/body specimens, usage rules.
|
||||
8. **Icons** *(digital brands only)* — 8-12 essential UI icons in brand's geometric style + stroke/corner/grid rules + library recommendation. SKIP this page entirely for non-digital brands; renumber subsequent pages accordingly.
|
||||
8. **Icons** — 8-12 essential UI icons in brand's geometric style + stroke/corner/grid rules + library recommendation. See "Icons Page — Structure & Rules" section.
|
||||
9. **Voice & Tone** — Adjectives + actual brand copy examples by context.
|
||||
10. **Imagery Rules** — adapts to the brand's primary medium (see "Imagery Rules Page — Adapts per Brand" section below):
|
||||
- Photography-led brand → Photography Rules (subject / light / cast / treatment / forbidden) + 1 example photo
|
||||
@@ -43,32 +27,23 @@ Renumber pages contiguously based on what's included — don't leave gaps.
|
||||
|
||||
---
|
||||
|
||||
## Step 0 — Workspace Setup
|
||||
## Step 0 — Render Inputs
|
||||
|
||||
Images and fonts live in a persistent workspace path. `/tmp` is wiped between sessions on many systems, so don't put assets there.
|
||||
Server-side Chromium cannot read local `file://` paths. Every asset referenced by the HTML must be one of:
|
||||
|
||||
Pick any writable directory. The default below works on Mac, Linux, and inside containers (it resolves `$HOME`, so no hardcoded paths). Override by exporting `BUILD_A_BRAND_WS=/some/other/path` before running, if you have a preferred location (e.g. `~/Downloads/...` or a project-relative `./tmp/build-a-brand`).
|
||||
- HTTPS URL returned by Pika tools (`generate_image`, `upload_asset`, `html_to_png`, etc.)
|
||||
- Public HTTPS raw asset URL
|
||||
- Inline `data:` URI (best for small SVGs and font subsets)
|
||||
|
||||
```bash
|
||||
WS="${BUILD_A_BRAND_WS:-$HOME/build-a-brand-workspace}"
|
||||
mkdir -p "$WS/fonts" "$WS/images"
|
||||
```
|
||||
For local source files, call `upload_asset` first and use the returned `public_url` in HTML. `upload_asset` does not accept font files or PDFs, so fonts should use public HTTPS raw URLs or inline `data:font/...` sources.
|
||||
|
||||
In Python, expand the same way:
|
||||
|
||||
```python
|
||||
import os
|
||||
WS = os.environ.get('BUILD_A_BRAND_WS') or os.path.expanduser('~/build-a-brand-workspace')
|
||||
```
|
||||
|
||||
Use `$WS/images/lifestyle1.jpg` etc., and `file://$WS/images/lifestyle1.jpg` in HTML (resolve `$WS` to an absolute path before baking into HTML — `file://` URLs don't do shell expansion).
|
||||
Build the HTML as a separate `.py` script file (e.g., `$WS/build_guidelines.py`) — avoids f-string parsing errors with inline Python heredocs.
|
||||
Build the final guidelines as one HTML string or as `body_pages` fragments plus `shared_head`. You may keep local working copies for debugging and kit export, but local paths must not appear in render HTML.
|
||||
|
||||
---
|
||||
|
||||
## Step 1 — Generate Imagery (in parallel batches of 4)
|
||||
|
||||
You need **a minimum of 5 generated images**:
|
||||
You need **a minimum of 10 generated images**:
|
||||
- 1 hero mood image (for the cover)
|
||||
- 1 example image (for the photography rules page)
|
||||
- 4 lifestyle images (for the visual world grid)
|
||||
@@ -76,14 +51,19 @@ You need **a minimum of 5 generated images**:
|
||||
|
||||
That's 10 images. Run in parallel batches of 4 with `&` + `wait`. Never more than 4 at once (timeouts).
|
||||
|
||||
**After every generation, compress before using in PDFs:**
|
||||
```python
|
||||
from PIL import Image
|
||||
img = Image.open(path)
|
||||
img.thumbnail((1200, 1200), Image.LANCZOS)
|
||||
img.save(path, 'JPEG', quality=68, optimize=True)
|
||||
```
|
||||
Target: under 150KB per image. WeasyPrint silently drops images that are too large — most common cause of missing images.
|
||||
Before generating or placing those images, write a short **crop plan** and **pre-generation slot plan** for the deck. This is required working context, not final user-facing copy:
|
||||
|
||||
- **Crop plan**: for every generated image, name the destination page, slot aspect ratio, final pixel box, intended subject anchor (face/product/hands/object), expected `object-position`, and rounded frame risk. If rounded frame risk is high, use a softer radius, move the subject anchor lower/center, or regenerate for the actual slot ratio.
|
||||
- **Pre-generation slot plan**: list every slot expected to show imagery, including destination page, slot aspect ratio, planned medium, and prompt intent. Do not invent asset URLs before generation.
|
||||
- **Post-generation filled manifest / image slot manifest**: after image generation or upload, copy the final asset URLs/IDs into the same slot list and verify every planned slot is filled. The manifest must include the cover hero, imagery-rules example, visual-world grid, touchpoints grid, and the Digital / Social website hero, Instagram grid, and story template. Each required slot needs a real generated image or hosted image asset. Flat color rectangles count as empty unless the section is explicitly a palette specimen.
|
||||
- The Visual World page must use real generated images that match the brand's
|
||||
medium. CSS color blocks, gradients, caption-only placeholders, and empty
|
||||
rectangles are not image assets.
|
||||
- The Touchpoints page must use real generated photographs in believable
|
||||
physical or digital context. CSS color blocks, flat vector mockups, and
|
||||
captioned boxes are not substitutes for touchpoint photography.
|
||||
|
||||
Generated image URLs can be used directly in MCP-rendered HTML. Keep page images reasonably sized: request the smallest image that survives the target crop, avoid duplicating the same source across pages, and use CSS `object-fit` / `object-position` explicitly. If a user provides a huge local image, upload it only after resizing or replacing it with a Pika-generated/hosted equivalent; extremely large assets slow the server renderer.
|
||||
|
||||
**Photography prompt template (lifestyle):**
|
||||
```
|
||||
@@ -110,7 +90,7 @@ These are the failure modes that have burned us before. Apply EVERY prompt.
|
||||
|
||||
### 0. Default provider: gpt-image-2
|
||||
|
||||
**Every `generate_image` call must pass `provider="gpt-image-2"` unless the user explicitly names a different model.** This is a global preference, not a per-skill rule. Don't default to `nano-banana-pro` (Gemini) — it has worse instruction-following for our brand work and bakes in text more aggressively. Use gpt-image-2 with `quality="medium"` for the default balance of speed and fidelity.
|
||||
**Every `generate_image` call must pass `provider="gpt-image-2"` unless monica explicitly names a different model.** This is a global preference, not a per-skill rule. Don't default to `nano-banana-pro` (Gemini) — it has worse instruction-following for our brand work and bakes in text more aggressively. Use gpt-image-2 with `quality="medium"` for the default balance of speed and fidelity.
|
||||
|
||||
### 1. Never let text bake into the image
|
||||
|
||||
@@ -134,25 +114,7 @@ Before writing the prompt, decide WHERE this photo will appear in the layout and
|
||||
|
||||
Specify the subject's position in the prompt explicitly: "subject centered in frame, face occupying middle 50% of the image vertically."
|
||||
|
||||
### 3. Logos: generate a high-res symbol via gpt-image-2, ship as transparent PNG (no tracing)
|
||||
|
||||
Hand-coded SVG symbols look amateur. But also: **don't trace the gen'd symbol to SVG.** Keep the symbol as a high-resolution transparent PNG. Only the wordmark gets vectorized in the brand kit. The pipeline:
|
||||
|
||||
1. **Generate the symbol via `generate_image` with `provider="gpt-image-2"`, `quality="high"` (or `"medium"` for first drafts), 1:1 aspect ratio, 1024×1024 minimum.** The symbol can be any style that fits the brand — flat illustration, 3D-rendered, painted, photographic, gradient-rich, chrome, holographic, hand-drawn. No flat-vector requirement. The only constraints are in "Symbol output rules" in `brand-identity.md`. **Always append the no-text guardrail**: "absolutely no text, no letters, no typography, no words, no characters anywhere in the image." gpt-image-2 produces garbled fake text inside logos if you don't explicitly forbid it.
|
||||
2. **The gen'd symbol must satisfy ALL of:** conceptually linked to the brand (means something about what the brand IS/DOES), feels unique (not generic), recognizable at 16×16, no more than 3 dominant colors, high res (2048×2048+ for the shipped version), no text inside the image, true transparent background. If it fails any of these — regenerate. See `brand-identity.md` "Symbol output rules" for the full list.
|
||||
3. **Generate 2-3 variations** if appropriate. Show them to the user. Wait for approval before committing. **When generating across the 3 brand-board options, symbols must DIFFER IN CONCEPT, not just style.** Don't ship three "literal mascot face" logos in three styles. See `brand-identity.md` "Symbol concepts must DIFFER across the 3 brand options" for concept lanes (mascot / product-feature reference / abstract / monogram / hybrid / container).
|
||||
4. **Save as transparent PNG, verified.** True alpha=0 in transparent regions. gpt-image-2 frequently paints near-white pixels in the "transparent" area — verify by sampling corner pixels with PIL (`alpha == 0`). If not, key them out with PIL, or regenerate with a solid background matching the placement surface. See "gpt-image-2 transparent-background caveat" in the QA section. **Do NOT trace to SVG.** The symbol stays as a high-res PNG.
|
||||
5. **The wordmark is ALWAYS rendered as real text in a Google Font, never baked into a generated image.** Pick the font in the typography step. Render the wordmark via live HTML/CSS for the guidelines pages, and convert to text-as-paths SVG only when packaging the brand kit.
|
||||
6. **Lockup composition is perfectly measured and permanently fixed.** Pick one horizontal lockup geometry AND one stacked lockup geometry. For each, specify: symbol size (px or em), wordmark font size (px), gap between symbol and wordmark (px), vertical baseline alignment (which point of the symbol aligns with which baseline of the wordmark). **The measurements never change across color variants or contexts** — cream / pink / lime / on-photo / on-dark variants all use the IDENTICAL geometry. Document the exact measurements on the Logo page so a designer or developer can rebuild the lockup without guessing.
|
||||
|
||||
**In the brand kit:**
|
||||
- `symbol-[color].png` — high-res transparent raster at multiple sizes (16, 32, 64, 128, 256, 512, 1024, 2048). NOT vectorized.
|
||||
- `wordmark-[color].svg` — Google Font text converted to outlined paths (text-as-paths), so the SVG renders identically without the font file installed. PNG version at 1024 wide also included.
|
||||
- `lockup-[orientation]-[color].svg` — contains the symbol PNG embedded inline + the wordmark as paths, positioned at the locked measurements. PNG version of the assembled lockup at 1024 wide also included.
|
||||
|
||||
This rule applies to brand boards (Step 2), the guidelines logo page (Step 3), and the kit's exported logo files (Step 4). Always.
|
||||
|
||||
### 4. Verify by screenshot BEFORE delivering
|
||||
### 3. Verify by screenshot BEFORE delivering
|
||||
|
||||
After rendering ANY PDF or board, screenshot every page and read every screenshot. The QA checklist at the bottom of this doc is mandatory. Never deliver based on assumption that the layout worked. Specifically check:
|
||||
|
||||
@@ -169,6 +131,11 @@ If anything looks wrong, fix it before delivering. Never ask the user to spot pr
|
||||
|
||||
**Fonts are NOT hardcoded.** Select fonts that match the identity built in Step 3 of the main skill. If the fonts could work for a competitor, pick different ones.
|
||||
|
||||
Before selecting final type, make a **font shortlist** for the identity:
|
||||
- Include at least 2 display families and at least 2 body/accent candidates that fit the specific brand vibe.
|
||||
- Do not reuse the same display/body pair from the last 3 brand briefs unless the user explicitly asks for that exact pairing.
|
||||
- Pick the pair for this brief from the shortlist and state why it fits the brand's category, audience segments, and visual world.
|
||||
|
||||
### Must Have Character — Don't Default to Safe Fonts
|
||||
|
||||
If the brand's display font could appear on any random SaaS site without anyone noticing, it's wrong. Push for fonts with recognizable personality.
|
||||
@@ -184,40 +151,7 @@ If the brand's display font could appear on any random SaaS site without anyone
|
||||
- **Playful / loud / personality-forward** → Honk (chubby 3D), Tilt Warp, Tilt Neon, Bagel Fat One, Caveat (handwritten)
|
||||
- **Quiet / minimal-with-soul** → Public Sans, Hahmlet, Newsreader (light weights), Spectral (light weights)
|
||||
|
||||
### No favorite fonts — every brand starts the search fresh
|
||||
|
||||
There is no "house font" for this skill. No font is a default. No font is banned either — every font on Google Fonts is still in the running for the right brand, including ones used on previous brands. What's forbidden is the *pattern*: reaching for the same fonts across unrelated brands because they worked last time.
|
||||
|
||||
Run the selection from scratch every brand. The fact that Sansita, Fraunces, Funnel Display, Bagel Fat One, etc. fit a past brand doesn't make them the right pick for this one — and it doesn't disqualify them either. The question is always "what UNIQUELY fits this brand's vibe?", run fresh.
|
||||
|
||||
### The mandatory research step — do this every brand, every time
|
||||
|
||||
Google Fonts hosts ~1500 families. The failure mode is selecting from a tiny mental shortlist of ~20 fonts that worked before. Force a real search:
|
||||
|
||||
1. **Name the brand's vibe in 3-5 specific adjectives** ("warm, archival, slightly weird, with restraint"). The adjectives are the brief; the font search runs against them.
|
||||
2. **Brainstorm 5+ candidates from at least 3 different categories** — serif / sans / mono / display / script / slab / stencil. Don't pre-filter to fonts you remember liking; widen the net first, narrow after.
|
||||
3. **Pull at least 2 less-obvious options into the shortlist** — Eczar, Workbench, Bungee Shade, Climate Crisis, IM Fell DW Pica, Krona One, Suez One, Yeseva One, Tilt Prism, Italiana, Stardos Stencil, Inria Serif, Ribeye Marrow, etc. These don't get picked because they're weird; they get picked when the brand is the one that wants them.
|
||||
4. **Cross-brand variety check.** Before locking the choice, ask: "Have I reached for this font on a recent brand?" If yes, you need a real reason this brand specifically wants the same font — not just "it worked before." If you can't articulate why this brand uniquely wants it, pick a different fresh option that fits as well.
|
||||
5. **Sanity check:** would 5 brands in unrelated industries reach for this font? If yes, it's too generic — keep looking.
|
||||
6. **Pick** the one that UNIQUELY fits this brand's vibe.
|
||||
|
||||
### Wider library by vibe — go beyond the top tier
|
||||
|
||||
Use these as starting points only, not automatic defaults. The point is to widen your candidate pool every brand.
|
||||
|
||||
- **Editorial / archival / literary** → Vollkorn, EB Garamond, Crimson Pro, Newsreader, Spectral, Cardo, Eczar, Inria Serif, Petrona, Faustina, Libre Caslon Text, Libre Baskerville, IM Fell DW Pica, Old Standard TT, PT Serif
|
||||
- **Magazine / cover energy / bold display** → Bricolage Grotesque, Big Shoulders Display, Familjen Grotesk, Karantina, Yeseva One, Suez One, Krona One, Italiana, Workbench
|
||||
- **Heavy display / chunky / 3D** → Honk, Bungee, Bungee Shade, Bungee Inline, Climate Crisis, Lilita One, Bowlby One, Modak, Alfa Slab One, Workbench, Ribeye, Ribeye Marrow
|
||||
- **Slab / vintage / sturdy** → Aleo, Bitter, Arvo, Rokkitt, Zilla Slab, Roboto Slab, Saira Stencil One, Stardos Stencil, Sancreek (western), Rye, Smokum
|
||||
- **Friendly / soft / consumer** → Lilita One, Albert Sans, Plus Jakarta Sans, Onest, Quicksand, Comfortaa, Nunito Sans, Mulish, Mona Sans
|
||||
- **Tech / mono / digital-native** → Reddit Mono, Geist Mono, IBM Plex Mono, Space Mono, DM Mono, Fragment Mono, Anonymous Pro, B612 Mono, Major Mono Display, Cutive Mono, Inconsolata, Fira Code, Cousine
|
||||
- **Playful / weird / personality-forward** → Honk, Tilt Warp, Tilt Neon, Tilt Prism, Workbench, Iceland, Limelight, Knewave, Climate Crisis, Permanent Marker, Caveat, Indie Flower, Cinzel Decorative
|
||||
- **Old style / classical / serious** → IM Fell DW Pica, Cinzel, Cardo, Old Standard TT, Vollkorn, Italiana, EB Garamond, Inria Serif, Stardos Stencil
|
||||
- **Script / handwriting** → Lobster, Lobster Two, Pacifico, Sacramento, Caveat, Indie Flower, Permanent Marker, Allura, Great Vibes, Berkshire Swash, Yellowtail
|
||||
- **Quiet / minimal-with-soul** → Public Sans, Hahmlet, Newsreader light, Spectral light, Inria Sans, Albert Sans, Onest
|
||||
|
||||
### Pairing rules
|
||||
|
||||
**Pairing rules:**
|
||||
- Display font must have character. Body font can be quieter but should still feel intentional.
|
||||
- Never pair two characterless fonts (Inter + DM Sans = no point of view).
|
||||
- Display + body should feel related but distinct.
|
||||
@@ -238,35 +172,29 @@ Use these as starting points only, not automatic defaults. The point is to widen
|
||||
- Bold condensed headline → lightweight sans (DM Sans, Lato Light, Inter)
|
||||
- Geometric sans headline → same family lighter weight, or Inter
|
||||
|
||||
### Download fonts to workspace
|
||||
### Font loading in MCP renders
|
||||
|
||||
```bash
|
||||
WS="${BUILD_A_BRAND_WS:-$HOME/build-a-brand-workspace}"
|
||||
mkdir -p "$WS/fonts"
|
||||
# example — download the fonts you've selected:
|
||||
curl -sL "https://github.com/google/fonts/raw/main/ofl/syne/Syne%5Bwght%5D.ttf" -o "$WS/fonts/Syne.ttf" &
|
||||
curl -sL "https://github.com/google/fonts/raw/main/ofl/playfairdisplay/PlayfairDisplay%5Bwght%5D.ttf" -o "$WS/fonts/PlayfairDisplay.ttf" &
|
||||
curl -sL "https://github.com/google/fonts/raw/main/ofl/karla/Karla%5Bwght%5D.ttf" -o "$WS/fonts/Karla.ttf" &
|
||||
wait
|
||||
Use `@font-face` with server-reachable sources:
|
||||
|
||||
```css
|
||||
@font-face {
|
||||
font-family: 'SyneBrand';
|
||||
src: url('https://github.com/google/fonts/raw/main/ofl/syne/Syne%5Bwght%5D.ttf') format('truetype');
|
||||
font-weight: 300 900;
|
||||
}
|
||||
@font-face {
|
||||
font-family: 'KarlaBrand';
|
||||
src: url('data:font/ttf;base64,...') format('truetype');
|
||||
}
|
||||
```
|
||||
|
||||
Declare in CSS via `@font-face` only (`@import` or `<link>` cause 30s+ render timeouts in WeasyPrint and unreliable loading in Chrome headless). Resolve `$WS` to an absolute path in your HTML-generator script, then bake it in via f-string:
|
||||
|
||||
```python
|
||||
# In your HTML-builder script:
|
||||
css = f"""
|
||||
@font-face {{ font-family: 'Syne'; src: url('file://{WS}/fonts/Syne.ttf'); }}
|
||||
@font-face {{ font-family: 'Karla'; src: url('file://{WS}/fonts/Karla.ttf'); }}
|
||||
"""
|
||||
```
|
||||
|
||||
Always use **absolute file:// paths**. Never relative. Never `@import url(...)` from Google Fonts.
|
||||
Avoid `@import` and `<link>` because the renderer prefetches explicit asset URLs more reliably than CSS import chains. Give each brand face a unique family name. Render one page with `html_to_png` before the full PDF; if the type falls back to Times/Arial, fix the font source.
|
||||
|
||||
---
|
||||
|
||||
## Step 3 — Build HTML
|
||||
|
||||
Write a fresh Python script (e.g., `$WS/build_guidelines.py`) that emits the HTML to `$WS/guidelines.html`. Never copy old deck HTML — always write fresh.
|
||||
Write fresh HTML for the chosen identity. Never copy old deck HTML — always write fresh.
|
||||
|
||||
**Critical CSS (required in every guidelines doc):**
|
||||
```css
|
||||
@@ -274,62 +202,67 @@ Write a fresh Python script (e.g., `$WS/build_guidelines.py`) that emits the HTM
|
||||
.page { width: 1200px; height: 850px; overflow: hidden; page-break-after: always; display: block; }
|
||||
```
|
||||
|
||||
### WeasyPrint Hard Rules — Read Before Writing a Single Div
|
||||
### MCP Chromium Render Rules — Read Before Writing a Single Div
|
||||
|
||||
WeasyPrint's flexbox engine is severely broken. `display:flex` fails silently — columns collapse to zero width and content disappears with no error message. This is the #1 cause of blank pages.
|
||||
The renderer is server-side Chromium. Flexbox, grid, absolute positioning, and CSS transforms are supported. Keep fixed-format pages explicit so QA is deterministic:
|
||||
|
||||
#### Rule 1: No flex anywhere
|
||||
- **NEVER** `display:flex` on `.page`, on header bars, on two-column rows. Nowhere.
|
||||
- **For all multi-column layouts: use `<table>` elements** with explicit `width` and `height` on every `<td>`.
|
||||
- **For `.page`: use `display:block`**.
|
||||
- **Header bar template**: `<table style="width:1200px;height:68px;border-collapse:collapse;">` with two `<td>` cells.
|
||||
- **Two-column content template**: `<table style="width:1200px;height:782px;border-collapse:collapse;table-layout:fixed;">`.
|
||||
- **Column widths must add up to 1200px exactly.** Always verify.
|
||||
- **CSS `display:grid` is acceptable** for internal item arrangements (swatches, mockup grids, type specimens). NOT for page structure.
|
||||
#### Rule 1: Every page is a fixed canvas
|
||||
- Use `@page { size: 1200px 850px; margin: 0; }`.
|
||||
- Every `.page` must be `width:1200px;height:850px;overflow:hidden;page-break-after:always;position:relative;`.
|
||||
- Use explicit pixel dimensions for key regions. Flex/grid are fine, but don't let page height be content-driven.
|
||||
- Avoid viewport units (`vh`, `vw`) inside pages; they couple layout to the browser window rather than the page box.
|
||||
|
||||
#### Rule 2: `vertical-align:middle` is unreliable
|
||||
Even with real `<table>` elements + explicit heights, `vertical-align:middle` often renders content at the top.
|
||||
#### Rule 2: Use server-reachable assets only
|
||||
- No `file://` URLs.
|
||||
- Local images must be uploaded via `upload_asset` first.
|
||||
- Small SVGs and font subsets can be inlined as `data:` URIs.
|
||||
- External HTTPS assets are server-fetched and inlined by the renderer. If an asset is private or blocks server requests, upload/replace it.
|
||||
|
||||
**For full-page content centering**: nested table pattern — outer table 782px, inner table auto-height, `vertical-align:middle` on outer td. Works when inner content has NO explicit height.
|
||||
#### Rule 3: Use layout systems intentionally
|
||||
- CSS grid is preferred for swatches, icon sets, mockup grids, type specimens, and contact sheets.
|
||||
- Flexbox is fine for compact rows and centered stacks.
|
||||
- Use absolute positioning for full-bleed editorial pages where overlap and crop are intentional.
|
||||
- Give repeated tiles fixed dimensions so badges, labels, and icons cannot resize the layout.
|
||||
|
||||
**For shapes (logo mockups, hang tags, circles, labels)**: never use `vertical-align:middle`. Always use explicit `padding-top`:
|
||||
```
|
||||
padding-top = (shape_height - estimated_content_height) / 2
|
||||
```
|
||||
#### Rule 4: Text containers must wrap naturally
|
||||
- Text cards, sidebars, and copy columns must have an explicit readable width. Reserve at least `320px` for body copy via `min-width:320px`, grid `minmax(320px, ...)`, or an equivalent fixed px floor. You may use `ch` only as a `max-width` line-length cap, never as the width floor. Never let a right-column card collapse until each line becomes one word.
|
||||
- In flex/grid layouts, set `min-width:0` on text children so copy wraps inside its assigned track, and separately give the track/card a real width floor (`width`, `flex-basis`, or grid track minmax).
|
||||
- Body-copy containers should include `box-sizing:border-box; overflow-wrap:break-word; word-break:normal; hyphens:none;`.
|
||||
- Never use `word-break:break-all`, `overflow-wrap:anywhere`, or a narrow absolute-positioned card squeezed by an illustration/phone mockup for readable prose.
|
||||
- If an illustration, phone, seal, swatch, or decorative element sits near a copy card, the card owns a clean rectangle above it in z-order and geometry. Do not depend on the visual QA pass to catch preventable overlap.
|
||||
|
||||
#### Rule 3: `position:absolute` is unreliable
|
||||
Inside fixed-height shapes, `position:absolute` does not render reliably — elements appear in document flow. Replace with table rows or `padding-top`.
|
||||
|
||||
#### Rule 4: Image rules
|
||||
#### Rule 5: Image rules
|
||||
- Always explicit px dimensions: `style="width:300px;height:400px;object-fit:cover;display:block;"`
|
||||
- Never `height:100%` or `width:100%` — WeasyPrint cannot resolve percentage heights
|
||||
- Avoid percentage heights unless the parent has an explicit pixel height
|
||||
- Never `opacity:` on any `<img>` — images always at full opacity
|
||||
- Never text/overlay on images — put captions in an adjacent column or block
|
||||
- Never body text/labels/rules on images — put captions in an adjacent column or block. Masthead brand boards may overlay wordmark/tagline/issue metadata on a full-bleed photo only when the type sits on intentional negative space or a contrast scrim and passes contrast QA.
|
||||
- Contrast QA for masthead overlays means the masthead text remains readable in the full-page PNG preview and the `mcp__pika__analyze_media` board QA result does not flag low contrast, muddy overlay, or unreadable type. If uncertain, run a targeted follow-up prompt: "CONTRAST: PASS or FAIL. Is the masthead wordmark/tagline/issue metadata readable against the photo at full-page size without hiding the photo subject?"
|
||||
- Never duplicate an image src across the deck — each file appears at most once
|
||||
- Compress to under 150KB before render
|
||||
- Use `object-position` deliberately and verify the crop in PNG previews
|
||||
|
||||
#### Rule 5: Text contrast thresholds on dark backgrounds
|
||||
#### Rule 6: Text contrast thresholds on dark backgrounds
|
||||
On graphite (#2E2E2E) or any dark background:
|
||||
- Body text minimum: `rgba(248,243,236,.7)`
|
||||
- Sub-descriptions minimum: `rgba(248,243,236,.55)`
|
||||
- Decorative / ghost text minimum: `rgba(248,243,236,.45)` — below this, remove the element entirely
|
||||
- `.3` opacity on dark = invisible. Never use for any visible text.
|
||||
|
||||
#### Rule 6: Page overflow prevention
|
||||
#### Rule 7: Page overflow prevention
|
||||
- Every page is 850px tall. All content MUST fit.
|
||||
- If a page has a headline >60px AND more than 3 body paragraphs, it will overflow. Cut or split.
|
||||
- Never more than ~220 words of body text on a single page.
|
||||
- Padding: 64px top/bottom max on content pages. Don't stack multiple padded sections.
|
||||
|
||||
#### Rule 6b: Content must not bleed into the footer (Chrome --print-to-pdf path)
|
||||
When the brand guidelines deck is rendered via Chrome `--print-to-pdf` (not WeasyPrint), the footer is typically `position: absolute; bottom: 18px` and the main content is in normal flow. If main content extends past the available content height, it visually overlaps the footer text — Monica caught this on the koalacore Voice page (2026-05-22).
|
||||
#### Rule 8: Good-looking layout prevention
|
||||
- Passing render QA is not enough. A page can have no clipped text and still be bad if it looks crowded, muddy, or amateur.
|
||||
- Body copy, labels, and load-bearing informational text must never collide with swatches, icons, photos, decorative rules, grain, seals, or background imagery. If an element sits on top of body copy, the page fails even when the text technically remains inside its box. Masthead boards may overlay wordmark/tagline/issue metadata on a full-bleed photo only when the type sits on intentional negative space or a contrast scrim and passes contrast QA; body copy still gets its own clean reading area.
|
||||
- Body copy on brand boards and guidelines must be readable at the full-page screenshot size. Use 18px minimum for body copy, 14px minimum for labels, and 10px minimum only for decorative metadata that is not load-bearing.
|
||||
- Decorative microtype is optional. If small labels, faux archival notations, issue numbers, or specimen marks make the page noisy, remove them before reducing the real content.
|
||||
- Keep one primary visual focal point per page while required board content stays secondary and grouped. If the viewer's eye has to choose between a huge wordmark, a dense paragraph block, six swatches, a seal, a photo, and a pull quote at once, simplify hierarchy and grouping; do not drop required content.
|
||||
- Do not use a one-note dark brown/green/slate page unless the brief specifically demands it. Add contrast through scale, image light, accent color, or negative space; do not let the whole page collapse into one muddy value range.
|
||||
- Empty placeholders are a hard fail. Website/social/app mockups are optional on brand boards; do not add them unless they contain real content. If a website hero, grid, story template, image slot, or app mockup is included, either render the real content or remove/redesign the slot. Flat color rectangles in a Digital/Social mockup count as empty placeholders unless the section is explicitly a palette specimen.
|
||||
|
||||
**Mandate for Chrome `--print-to-pdf` page builds:**
|
||||
- `.content` (the main page area between header and footer) MUST have an explicit `height: calc(850px - HEADER_HEIGHT - FOOTER_HEIGHT)` (or a similar max-height) AND `overflow: hidden` as a safety net.
|
||||
- This way, even if a content block grows unexpectedly, the layout clips at the content boundary instead of bleeding into the footer.
|
||||
- Treat it as a guardrail, not a target — the goal is still to fit content within the available height. But the overflow:hidden prevents accidental overlap when content density is hard to predict.
|
||||
|
||||
#### Rule 7: Vertical centering critical gotcha
|
||||
#### Rule 9: Vertical centering critical gotcha
|
||||
When using `<table><tr><td style="vertical-align:middle;">` to center, the inner content div **must NOT have explicit height**. If the inner div has `height:782px` (same as td), the td has nothing to center → appears top-aligned. Set `height` only on the outer `<td>`, never on the inner content div.
|
||||
|
||||
### Reusable Templates
|
||||
@@ -348,74 +281,77 @@ When using `<table><tr><td style="vertical-align:middle;">` to center, the inner
|
||||
<table style="width:1200px;height:782px;border-collapse:collapse;table-layout:fixed;">
|
||||
<tr>
|
||||
<td style="width:480px;height:782px;vertical-align:top;padding:0;overflow:hidden;">
|
||||
<img src="file:///..." style="width:480px;height:782px;object-fit:cover;display:block;">
|
||||
<img src="https://..." style="width:480px;height:782px;object-fit:cover;display:block;">
|
||||
</td>
|
||||
<td style="width:720px;height:782px;vertical-align:middle;padding:52px;background:#2E2E2E;">
|
||||
<td style="width:720px;height:782px;vertical-align:middle;padding:52px;background:#2E2E2E;box-sizing:border-box;min-width:0;overflow-wrap:break-word;word-break:normal;hyphens:none;">
|
||||
<!-- right column text content -->
|
||||
<div style="max-width:560px;min-width:320px;box-sizing:border-box;overflow-wrap:break-word;word-break:normal;hyphens:none;">
|
||||
<!-- body copy goes here -->
|
||||
</div>
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
</div>
|
||||
```
|
||||
|
||||
**Full-bleed lifestyle grid (page 10):**
|
||||
**Full-bleed lifestyle grid (page 11):**
|
||||
```html
|
||||
<table style="width:1200px;height:782px;border-collapse:collapse;table-layout:fixed;">
|
||||
<tr>
|
||||
<td style="width:300px;height:782px;padding:0;overflow:hidden;"><img src="file:///..." style="width:300px;height:782px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:300px;height:782px;padding:0;overflow:hidden;"><img src="file:///..." style="width:300px;height:782px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:300px;height:782px;padding:0;overflow:hidden;"><img src="file:///..." style="width:300px;height:782px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:300px;height:782px;padding:0;overflow:hidden;"><img src="file:///..." style="width:300px;height:782px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:300px;height:782px;padding:0;overflow:hidden;"><img src="https://..." style="width:300px;height:782px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:300px;height:782px;padding:0;overflow:hidden;"><img src="https://..." style="width:300px;height:782px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:300px;height:782px;padding:0;overflow:hidden;"><img src="https://..." style="width:300px;height:782px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:300px;height:782px;padding:0;overflow:hidden;"><img src="https://..." style="width:300px;height:782px;object-fit:cover;display:block;"></td>
|
||||
</tr>
|
||||
</table>
|
||||
```
|
||||
|
||||
**Touchpoints 2×2 grid (page 11):**
|
||||
This grid is image-only. If captions or labels are needed, put them in adjacent Rule-4 text containers; do not overlay body copy on the photos.
|
||||
|
||||
**Touchpoints 2×2 grid (page 12):**
|
||||
```html
|
||||
<table style="width:1200px;height:782px;border-collapse:collapse;table-layout:fixed;">
|
||||
<tr>
|
||||
<td style="width:599px;height:390px;padding:0;overflow:hidden;"><img src="file:///..." style="width:599px;height:390px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:599px;height:390px;padding:0;overflow:hidden;"><img src="https://..." style="width:599px;height:390px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:1px;background:#fff;"></td>
|
||||
<td style="width:600px;height:390px;padding:0;overflow:hidden;"><img src="file:///..." style="width:600px;height:390px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:600px;height:390px;padding:0;overflow:hidden;"><img src="https://..." style="width:600px;height:390px;object-fit:cover;display:block;"></td>
|
||||
</tr>
|
||||
<tr><td colspan="3" style="height:2px;background:#fff;padding:0;"></td></tr>
|
||||
<tr>
|
||||
<td style="width:599px;height:390px;padding:0;overflow:hidden;"><img src="file:///..." style="width:599px;height:390px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:599px;height:390px;padding:0;overflow:hidden;"><img src="https://..." style="width:599px;height:390px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:1px;background:#fff;"></td>
|
||||
<td style="width:600px;height:390px;padding:0;overflow:hidden;"><img src="file:///..." style="width:600px;height:390px;object-fit:cover;display:block;"></td>
|
||||
<td style="width:600px;height:390px;padding:0;overflow:hidden;"><img src="https://..." style="width:600px;height:390px;object-fit:cover;display:block;"></td>
|
||||
</tr>
|
||||
</table>
|
||||
```
|
||||
|
||||
This touchpoints grid is image-only. If a label is required, use a separate caption strip or adjacent Rule-4 text container; do not place readable prose inside the image cells.
|
||||
|
||||
### Page-Specific Notes
|
||||
|
||||
**Page 1 (Cover) — must make the product unambiguous.** The cover hero photo + subtitle together must answer "what is this?" with zero ambiguity. For digital products (apps, services, software): the hero photo must show the product IN USE — a phone or device displaying the actual result of using the product, OR a person in the moment of using it. NEVER use a representational object (a charm, a token, a packaging mockup, a logo-shaped artifact) as the cover hero — it makes the brand look like it sells that object instead of the actual product. The cover subtitle should communicate what the thing IS in the brand's voice — clarity through content, never robotic templates.
|
||||
|
||||
**Page 2 (Strategy) — product clarity sentence required, in brand voice.** Within the first 80 words of the Strategy page, the reader must understand what the product literally IS. Communicated in the brand's tone, not a robotic format. "Koalacore is an app that…" is robotic; "It's an app. You pick a photo. We drop a koala in. The end." is on-brand. Either form works — what's required is that a reader who lands on this page knows what the product is by the end of the opening paragraph. NEVER substitute "what we believe" or "what we stand for" for "what we are."
|
||||
|
||||
**Page 4 — Logo applications:** Left half (~420px) = large logo mark centered with generous whitespace. Right half (~780px) = 5 CSS mockups in a 2-row grid (3 top, 2 bottom), `gap:32px`, each cell min 160×180px. Don't flex-wrap — use proper grid.
|
||||
|
||||
**Page 5 — Logo Don'ts:** 5 violation tiles in a row, each with the wrong-usage logo + a small ✗ label + a one-line caption explaining the violation.
|
||||
|
||||
**Page 9 — Photography Rules:** Left 2/3 (~780px) = 3 example images in a grid with explicit pixel dimensions. Right 1/3 (~420px) = sidebar of 4-5 specific rules (surface, light, propping, editing, mood). Small label caps + body text.
|
||||
**Page 10 — Photography Rules:** Left 2/3 (~780px) = 3 example images in a grid with explicit pixel dimensions. Right 1/3 (~420px) = sidebar of 4-5 specific rules (surface, light, propping, editing, mood). Small label caps + body text.
|
||||
|
||||
**Page 11 — Touchpoints:** The page **must show real generated photos** of the brand in context. Never substitute CSS vector mockups. Generate the 4 images before building HTML. For physical product brands: hang tag, woven label macro, kraft mailer, flat lay. For digital brands: phone showing app, laptop showing site, sticker, tote bag. For service brands: business card in hand, signage, branded notebook, swag.
|
||||
**Page 12 — Touchpoints:** The page **must show real generated photos** of the brand in context. Never substitute CSS vector mockups. Generate the 4 images before building HTML. For physical product brands: hang tag, woven label macro, kraft mailer, flat lay. For digital brands: phone in hand showing app, laptop on desk showing site, sticker on water bottle, tote bag in a real scene. For service brands: business card in hand, signage, branded notebook, swag.
|
||||
|
||||
**Page 12 — Brand Applications (CSS mockups):** Each mockup is a small physical-object representation rendered in CSS. Use explicit `padding-top` centering — never `vertical-align:middle` inside fixed-height shapes.
|
||||
**Page 13 — Brand Applications (CSS mockups):** Each mockup is a small physical-object representation rendered in CSS. Use fixed dimensions and verify the visual center in PNG previews.
|
||||
|
||||
| Shape | Dimensions | Suggested padding-top |
|
||||
| Shape | Dimensions | Suggested centering |
|
||||
|---|---|---|
|
||||
| Hang tag | 120×168px | ~31px in body row |
|
||||
| Woven label | 200×80px | ~13px |
|
||||
| Avatar circle | 100×100px | ~22px |
|
||||
| Sticker rounded square | 100×100px | ~18px |
|
||||
| Business card | 200×120px | ~35px |
|
||||
| Hang tag | 120×168px | CSS grid/flex center, then visual QA |
|
||||
| Woven label | 200×80px | CSS grid/flex center, then visual QA |
|
||||
| Avatar circle | 100×100px | CSS grid/flex center, then visual QA |
|
||||
| Sticker rounded square | 100×100px | CSS grid/flex center, then visual QA |
|
||||
| Business card | 200×120px | CSS grid/flex center, then visual QA |
|
||||
|
||||
### Image hard constraints
|
||||
|
||||
- **No opacity on images.** Never `opacity:` on any `<img>` — full brightness always.
|
||||
- **No text on images.** No `position:absolute` to overlay text/elements on images. Captions go in an adjacent column.
|
||||
- **No gradient overlays.** No `background:linear-gradient(...)` divs positioned over images.
|
||||
- **No body text on images.** Body copy, captions, labels, and rules never overlay images. Captions go in an adjacent column. Masthead brand boards may overlay the wordmark/tagline/issue metadata on a full-bleed photo only when the type sits on intentional negative space or a contrast scrim and passes contrast QA.
|
||||
- **No gradient overlays on ordinary content images.** Do not use decorative gradient overlays on photos. Masthead covers may use one controlled linear scrim/gradient mask behind wordmark/tagline/issue metadata text to preserve contrast: strongest stop <= 50% opacity, one edge direction only, max scrim height <= 40% of image height, and no product/detail/focal subject hidden under the scrim.
|
||||
- **No duplicate images.** Each image file appears at most once across the deck. If you run out, replace with typographic or color design elements (large CG italic quote, big page number, color field) — never reuse.
|
||||
- **Remove all price stickers / shelf labels / tags** from products before use. If source has a Goodwill sticker or similar, regenerate clean.
|
||||
- **Header logo on every page uses the brand's actual logo font/style** — never a generic fallback.
|
||||
@@ -423,105 +359,86 @@ When using `<table><tr><td style="vertical-align:middle;">` to center, the inner
|
||||
|
||||
---
|
||||
|
||||
## Step 4 — Render PDF Locally
|
||||
## Step 4 — Render PDF
|
||||
|
||||
**Preferred: Chrome headless.** Chrome handles modern CSS, flexbox, grid, and `@font-face file://` font declarations reliably when run locally. Use it when available:
|
||||
Render with Pika MCP `html_to_pdf`.
|
||||
|
||||
```bash
|
||||
CHROME="/Applications/Google Chrome.app/Contents/MacOS/Google Chrome"
|
||||
WS="${BUILD_A_BRAND_WS:-$HOME/build-a-brand-workspace}"
|
||||
"$CHROME" --headless --disable-gpu --no-sandbox --hide-scrollbars \
|
||||
--virtual-time-budget=8000 --no-pdf-header-footer \
|
||||
--print-to-pdf="$WS/guidelines.pdf" \
|
||||
"file://$WS/guidelines.html"
|
||||
Preferred single-HTML mode:
|
||||
|
||||
```
|
||||
html_to_pdf(
|
||||
html: guidelines_html,
|
||||
format: "pdf",
|
||||
mode: "async",
|
||||
wait_for: "domcontentloaded",
|
||||
pdf_options: {
|
||||
paper_size: { width: 1200, height: 850, unit: "px" },
|
||||
margins: { top: 0, right: 0, bottom: 0, left: 0 },
|
||||
print_background: true
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
**Fallback: WeasyPrint.** Use this only when Chrome is unavailable or the HTML was written to the table-only constraints above:
|
||||
Use `body_pages` + `shared_head` only when keeping each page as a separate body fragment is cleaner. In that mode the server renders pages in parallel and merges them; do not expect CSS counters or script state to cross page boundaries.
|
||||
|
||||
```python
|
||||
import weasyprint, warnings, os
|
||||
warnings.filterwarnings('ignore')
|
||||
Example completed PDF response:
|
||||
|
||||
WS = os.environ.get('BUILD_A_BRAND_WS') or os.path.expanduser('~/build-a-brand-workspace')
|
||||
pdf = weasyprint.HTML(filename=f'{WS}/guidelines.html',
|
||||
base_url=f'file://{WS}/').write_pdf()
|
||||
open(f'{WS}/guidelines.pdf', 'wb').write(pdf)
|
||||
print(f'guidelines: {os.path.getsize(f"{WS}/guidelines.pdf")//1024}KB')
|
||||
```
|
||||
|
||||
Use the Python API, not the CLI — CLI has font resolution issues.
|
||||
Pick one engine and stick with it for the whole 14–16-page build.
|
||||
{"status":"completed","file_url":"https://cdn.pika.art/v2/files/agent/b4ab5d48-443b-4ee2-ab9d-c1690d19ff72/5ce95210-47f9-4f8a-98bb-12d986bfa71e.pdf","format":"pdf","page_count":1,"byte_size":6279}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Verify Every Page (Mandatory — sub-step within Step 3 guidelines build)
|
||||
## Step 5 — Verify Every Page (Mandatory)
|
||||
|
||||
Before delivery, screenshot every page and verify against the QA checklist.
|
||||
Before delivery, render PNG previews and verify against the QA checklist. For full-deck QA, either call `html_to_pdf(format:"png", ...)` when page raster output is desired, or call `html_to_png` per page/body fragment.
|
||||
|
||||
```python
|
||||
# Screenshot every page — run this after every render
|
||||
import fitz, os
|
||||
WS = os.environ.get('BUILD_A_BRAND_WS') or os.path.expanduser('~/build-a-brand-workspace')
|
||||
doc = fitz.open(f'{WS}/guidelines.pdf')
|
||||
for i in range(len(doc)):
|
||||
doc[i].get_pixmap(matrix=fitz.Matrix(1.8,1.8)).save(f'{WS}/qa_p{i+1}.png')
|
||||
print(f'{len(doc)} pages — now read each one')
|
||||
### Full-Deck Visual QA
|
||||
|
||||
Run **full-deck visual QA** on every final guidelines page, not only the 3-page brand-board preview. Use `mcp__pika__analyze_media` page-by-page / per-page on the final PNG previews before delivering the PDF.
|
||||
|
||||
The prompt must start: "Answer with PASS or FAIL on the first line, then explain." Ask it to check for clipped text, accidental overlay, text/image collisions, missing image slots, empty placeholders, bad crops, rounded-frame crop damage, broken image loads, unreadable font fallback, weak hierarchy, and muddy one-note pages. Fix every captured FAIL before delivery. If the tool is unavailable or returns an ambiguous first line, halt with a manual-review warning instead of shipping silently.
|
||||
|
||||
```
|
||||
html_to_png(
|
||||
html: page_html,
|
||||
format: "png",
|
||||
mode: "sync",
|
||||
wait_for: "domcontentloaded",
|
||||
raster_options: {
|
||||
viewport_px: { width: 1200, height: 850 },
|
||||
device_scale: 1
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
```python
|
||||
# Corner-sample verification — catches blank columns
|
||||
import fitz, os
|
||||
WS = os.environ.get('BUILD_A_BRAND_WS') or os.path.expanduser('~/build-a-brand-workspace')
|
||||
doc = fitz.open(f'{WS}/guidelines.pdf')
|
||||
for i, page in enumerate(doc):
|
||||
pix = page.get_pixmap(matrix=fitz.Matrix(1,1))
|
||||
w, h = pix.width, pix.height
|
||||
s = pix.samples
|
||||
def px(x,y): return (s[(y*w+x)*3], s[(y*w+x)*3+1], s[(y*w+x)*3+2])
|
||||
corners = [px(10,10), px(w-10,10), px(10,h-10), px(w-10,h-10)]
|
||||
print(f'P{i+1}: {corners}')
|
||||
# white (255,255,255) corner on a dark page = blank column bug
|
||||
Example completed PNG response:
|
||||
|
||||
```
|
||||
{"status":"completed","file_url":"https://cdn.pika.art/v2/files/agent/4d944981-9897-40b6-9e37-533c2a90b541/5a863672-0835-4901-87a8-df8933d69cd4.png","format":"png","page_count":1,"byte_size":3806}
|
||||
```
|
||||
|
||||
### Pre-Send QA Checklist
|
||||
|
||||
Read every screenshot. Verify every item. If any check fails, fix it. No exceptions. Never ask the user to spot problems the agent should have caught.
|
||||
|
||||
**Three-pass QA is mandatory for multi-page brand guidelines (Step 3). Glancing at the full-page PNG is NOT zoom-reading — it's pass 1.**
|
||||
1. **Render-every-page pass** — render EVERY page in the 14–16-page deck to PNG individually (not just one or two). Open and READ each PNG full-size. If ANY page is blank, has overlapping text, missing content, or wrong rendering, fix and re-render that page before continuing. A blank page in the deliverable is a hard ship-block. **Each page is verified individually before the merge.**
|
||||
2. **Thumbnail pass** — full-page screenshots to catch layout-level issues (blank columns, missing content, wrong colors). The 1200×850 PNG read at standard screen size is the thumbnail pass. Do NOT confuse this with zoom-reading.
|
||||
3. **Zoom pass — MANDATORY CROP-AND-READ, NOT GLANCE.** Use PIL to crop the rendered PNG into 800×800 regions around every text-bearing area: page title, subtitle, every section header, every body block, every card pull-quote, every label, every hex code, every footer. **Read each crop individually.** This is the only way to catch borderline contrast, shade/3D font issues, footer overlap with content, and text-wrap failures — those issues all compress out of view at thumbnail scale. If you find yourself ready to ship without having opened named `zoom-{page}-{region}.png` files, you HAVEN'T done the zoom pass. Go back. (The agent has shipped twice claiming to have zoom-read when really it only glanced at the full PNG. The rule is now: zoom-read = generate crops + Read each crop file. No shortcut.)
|
||||
**Two-pass QA is mandatory:**
|
||||
1. **Thumbnail pass** — full-page screenshots to catch layout-level issues (blank columns, missing content, wrong colors).
|
||||
2. **Zoom pass** — crop and read 800×800px regions around every small UI element: pill markers, badges, avatars, captions, swatches, logo lockups, type specimens, business cards, app icons. Thumbnail screenshots compress all detail-level mistakes out of view. Without a zoom pass you WILL ship visible alignment, centering, or sizing flaws.
|
||||
|
||||
**Specifically zoom-read (crop into 800×800 and Read each):**
|
||||
- Cover hero copy + subtitle
|
||||
- Every page's top breadcrumb header (the brand wordmark at small scale)
|
||||
- Every page's footer band (look for footer-text overlapping page content)
|
||||
- Every "pull" or display-pull-quote block on every page
|
||||
- Every card title + body on cards with saturated backgrounds (pink cards especially)
|
||||
- Every text-on-photo region
|
||||
- Every swatch label + hex code
|
||||
- Every IG grid quote tile
|
||||
- Every small mockup label (app icon caption, splash subtitle, etc.)
|
||||
|
||||
**Contrast gate — every text/background pair must be verifiably legible.** Before delivery, audit every text-on-color combination across all pages. Auto-fails: pink text on pink background, lime text on cream background, dark text on dark photo, any small text whose color is within 25% luminance of its container background. When in doubt, run the pair through a contrast checker (4.5:1 minimum for body text, 3:1 for large display type). If even one text block fails contrast — fix it. Pink-on-pink is the most common failure; resolve by either (a) darkening the text to a deep accent, (b) lightening the background to cream, or (c) putting a contrasting card behind the text.
|
||||
|
||||
**Complex fills (chrome / holographic / multi-stop gradients) only at hero scale.** Multi-stop chrome / holographic / iridescent gradients on display type read clean at hero scale (60px+) but **muddy at small scale** (under ~24px) — the gradient stops compress into smudge and the type becomes hard to read. When a type style appears at multiple scales in the system (e.g. the brand wordmark on the cover hero AND in the small page-header breadcrumb), use the complex fill ONLY at hero scale, and switch to a solid color (with optional stroke) at small scale. Document both versions on the Logo page so the system has the right tool at every size.
|
||||
|
||||
**Same rule for shaded / 3D / decorative display fonts** — `BungeeShade`, `Honk`, `Bungee Inline`, `Bungee Outline`, `Workbench`, `Tilt Prism`, any font with built-in shadow / 3D / chrome / outline detail baked into the glyph design. These fonts have visible "shade" or "3D depth" portions inside each letter; at small/medium scale (under ~40px) those portions compress into mud, AND when used on a saturated background, the shade portion can render as a color that blends INTO the background — creating an apparent contrast failure where parts of letters disappear (pink BungeeShade on pink reads as pink-on-pink even when the letter face is cream). **Use shaded display fonts ONLY at hero scale (40px+) AND prefer cream/black backgrounds at any scale.** At smaller scales OR on saturated brand backgrounds, switch to the brand's flat sans (Onest Black, Plus Jakarta 900, etc.) for the same visual hierarchy with clean contrast.
|
||||
|
||||
**Test before shipping:** if the brand has a chunky display font (BungeeShade, Honk, etc.) and you're using it under 40px or on a hot-color background, render the page and zoom in on the letterforms. If you can see the shade/3D detail picking up the background color in any portion of any letter — switch to flat sans.
|
||||
|
||||
**Zoom-read every text-containing region, not just logos.** Specifically: numbered/lettered pill markers, button states, favicon-size logo renderings, color-swatch hex labels, type specimens, business card layouts, single-character badges, **hero taglines, decorative strips, callout banners, brand-name headers, voice-sample blocks, and any badge/sticker with text.** Anywhere text sits in a fixed-width container is a wrap risk — verify it stayed on the line you intended. A trailing word, comma, or star alone on its own line is a layout failure.
|
||||
|
||||
**Logo lockup centering check.** When a generated logo sits inside a circular or rounded frame: confirm the symbol is visually centered, not pushed to a corner by `object-fit: cover` revealing an off-center source comp. If the source image has a strong asymmetric mass (e.g. ears on top, body below), `object-fit: cover` will cut it badly. Use `object-fit: contain` on a transparent PNG, or remove the frame and let the source image's own composition speak.
|
||||
|
||||
**gpt-image-2 transparent-background caveat.** When generating logos with "transparent background" in the prompt, gpt-image-2 frequently outputs a literal painted checker pattern or near-white pixels in the "transparent" areas — NOT true alpha=0 transparency. Before placing on a colored card, **verify true transparency**: sample corner pixels with PIL and confirm `alpha == 0`. If not, either (a) key out the near-white pixels with PIL to true transparent, or (b) regenerate with a solid color background that matches the surface you'll place the logo on.
|
||||
Specifically zoom on: numbered/lettered pill markers, button states, favicon-size logo renderings, color-swatch labels, type specimens, business card layouts, and any badge that contains a single character. These are the highest-risk regions for centering bugs.
|
||||
|
||||
| Check | What to look for |
|
||||
|---|---|
|
||||
| Fonts loaded | Headlines render in the chosen font, not a system fallback (Times, Arial) |
|
||||
| No blank columns | Every column has content — no white/solid blocks where text should be |
|
||||
| No text on images | No overlaid headlines, labels, or divs sitting on top of any image |
|
||||
| No empty placeholders | Optional website heroes, grids, story templates, app mockups, image slots, and cards contain real content or are redesigned away. Flat color blocks in Digital/Social mockups count as empty unless they are explicitly palette specimens |
|
||||
| No text/image collisions | Body copy, captions, labels, and rules do not sit on top of images. Masthead wordmark/tagline/issue metadata overlays are allowed only with deliberate negative space or a contrast scrim and must pass contrast QA |
|
||||
| No text collisions | Text does not overlap or sit underneath icons, swatches, seals, decorative lines, photos, phone mockups, or other graphic elements |
|
||||
| No clipped or occluded text | Text is not cut off by its own container, page edge, rounded shape, sibling graphic, or z-index layer |
|
||||
| Board looks good, not just valid | Thumbnail read has one focal point, clear hierarchy, enough negative space, and no muddy one-note palette |
|
||||
| analyze_media PASS | For brand board PNG previews and every final guidelines page, run `mcp__pika__analyze_media` per `SKILL.md` Board quality gate and the full-deck visual QA above. Every FAIL or ambiguous first line must be fixed before delivery. If the tool is unavailable, halt with a manual-review warning |
|
||||
| Load-bearing copy is readable | Body copy is readable in the full-page PNG; do not hide key content in 10px decorative microtype |
|
||||
| **No baked-in text in generated images** | **Open each generated image and look for ANY text — magazine titles, watermarks, brand names, captions, headers. If you see any, regenerate with stronger no-text guardrails.** |
|
||||
| **Subjects survive their crop** | **For every generated image used in a layout: is the intended subject visible after the CSS crop? No forehead-only portraits, no hand-only kitchen scenes. If the subject got cut off by `object-fit:cover`, `object-position`, or a rounded/arched frame, fix the layout or regenerate the image.** |
|
||||
| **Rounded shapes don't eat content** | **Any `border-radius` ≥ ½ the element width creates a dome that crops content underneath. If the photo's subject sits in the top portion of the source, a dome top will hide it. Soften the radius or reposition the subject.** |
|
||||
@@ -531,11 +448,11 @@ Read every screenshot. Verify every item. If any check fails, fix it. No excepti
|
||||
| Text contrast | All text on dark (#2E2E2E) backgrounds at sufficient opacity |
|
||||
| Decorative text legible | Ghost / watermark text at ≥ .45 opacity — if lower, remove entirely |
|
||||
| Images load | No broken images — every img has explicit px width+height |
|
||||
| Page count = 14–16, conditionally correct | Non-digital/single-medium = 14; digital/single-medium = 15; non-digital/hybrid = 15; digital/hybrid = 16; no blank extras |
|
||||
| Page count = 15 (or 16 hybrid) | Right number of pages, no blank extras |
|
||||
| Touchpoints are real photos | Page 12 shows generated photographs, not CSS vector boxes |
|
||||
| Diverse cast | Lifestyle grid (page 11) shows racial diversity across subjects |
|
||||
| Icons consistent | Page 8 icons all use the same stroke weight + corner style + line caps |
|
||||
| Imagery medium matches brand | Page 10 reflects the brand's medium (photo / illustration / hybrid) — don't ship Photography Rules for an illustration brand |
|
||||
| Imagery medium matches brand | Imagery Rules page (page 10) reflects the brand's medium (photo / illustration / hybrid) — don't ship Photography Rules for an illustration brand |
|
||||
| No concept clichés | No hourglass-for-time / lightbulb-for-ideas / handshake-for-trust etc. — apply the three cheesiness tests |
|
||||
|
||||
Only deliver after all checks pass.
|
||||
@@ -544,20 +461,17 @@ Only deliver after all checks pass.
|
||||
|
||||
## Step 6 — Deliver
|
||||
|
||||
Save the final PDF to `~/Desktop/[brand-slug]-brand-guidelines.pdf` and send the local path in a single message. Do not upload or host the PDF unless the user explicitly asks for a hosted copy.
|
||||
|
||||
```bash
|
||||
cp "$WS/guidelines.pdf" "$HOME/Desktop/[brand-slug]-brand-guidelines.pdf"
|
||||
```
|
||||
Return the `file_url` from `html_to_pdf`. Also save a local copy when practical so the brand-kit zip can include `brand-guidelines.pdf`; the CDN URL remains the canonical deliverable.
|
||||
|
||||
Reply shape:
|
||||
|
||||
```
|
||||
**[BRAND NAME] — brand guidelines**
|
||||
Saved to: ~/Desktop/[brand-slug]-brand-guidelines.pdf ([actual page count] pages · 1200×850 · ~5MB)
|
||||
PDF: https://cdn.pika.art/...
|
||||
Local copy: ~/Desktop/[brand-slug]-brand-guidelines.pdf (15 pages · 1200×850)
|
||||
```
|
||||
|
||||
If the user is not on a Mac, save to the project working directory and emit that path instead. Never attach the PDF as a file in the chat — link by path only.
|
||||
If a local copy is not practical, the CDN URL is still the canonical deliverable. Do not use `upload_asset` for PDFs; `html_to_pdf` already returns the PDF URL.
|
||||
|
||||
Done.
|
||||
|
||||
@@ -613,15 +527,7 @@ Default to photography when:
|
||||
|
||||
## Icons Page — Structure & Rules
|
||||
|
||||
**⚠️ Conditional page — only build this for DIGITAL brands.** Skip entirely if the brand is a physical product, restaurant, fashion line, service business, or any non-software brand. For those, a UI icon system is irrelevant noise — the page would feel forced.
|
||||
|
||||
**Include the Icons page if the brand is:** an app, a web tool, a SaaS product, a software platform, a digital community, a website-as-product, or anything where users interact via UI.
|
||||
|
||||
**Skip the Icons page if the brand is:** a physical product, fashion, food/beverage, restaurant, hospitality, service business, consultancy, agency, or any brand whose surfaces are mostly physical / printed / packaging-driven.
|
||||
|
||||
If you're not sure, ask the user in Step 1: "Is this a digital product?" (this question should already be in your intake).
|
||||
|
||||
When included, page 8 defines the brand's UI icon system.
|
||||
Page 8 defines the brand's UI icon system. Every digital brand needs one — even non-tech brands benefit from consistent icons for navigation, social, and product surfaces.
|
||||
|
||||
### Page layout
|
||||
|
||||
@@ -702,18 +608,18 @@ Each SVG with `stroke="currentColor"` so it inherits color from the application
|
||||
|
||||
## Imagery Rules Page — Adapts per Brand
|
||||
|
||||
Page 9 of the guidelines defines the brand's imagery rules. **Its title and content adapt to whatever medium the brand actually uses** — don't ship a "Photography Rules" page for an illustration-led brand and don't ship "Illustration Rules" for a brand that lives in photos.
|
||||
Page 10 of the guidelines defines the brand's imagery rules. **Its title and content adapt to whatever medium the brand actually uses** — don't ship a "Photography Rules" page for an illustration-led brand and don't ship "Illustration Rules" for a brand that lives in photos.
|
||||
|
||||
### How to decide which page(s) to build
|
||||
|
||||
Look at the chosen identity option's specs (set in Step 3) and the lifestyle world you've defined:
|
||||
|
||||
| Brand visual medium | Page 9 setup |
|
||||
| Brand visual medium | Page 10 setup |
|
||||
|---|---|
|
||||
| All photography (no illustration anywhere) | Page 9 = Photography Rules (single page) |
|
||||
| All illustration (no photography anywhere) | Page 9 = Illustration Rules (single page) |
|
||||
| Hybrid — both matter equally in the brand world | Page 9a = Photography Rules, Page 9b = Illustration Rules (two pages → guidelines totals 15) |
|
||||
| Hybrid — one medium dominates but the other appears occasionally | Page 9 = the dominant medium's rules, with a short "minor medium" callout block at the bottom |
|
||||
| All photography (no illustration anywhere) | Page 10 = Photography Rules (single page) |
|
||||
| All illustration (no photography anywhere) | Page 10 = Illustration Rules (single page) |
|
||||
| Hybrid — both matter equally in the brand world | Page 10a = Photography Rules, Page 10b = Illustration Rules (two pages → guidelines totals 16) |
|
||||
| Hybrid — one medium dominates but the other appears occasionally | Page 10 = the dominant medium's rules, with a short "minor medium" callout block at the bottom |
|
||||
|
||||
### Photography Rules page structure
|
||||
|
||||
@@ -739,9 +645,9 @@ If the brand uses illustration:
|
||||
- **One example illustration** filling 2/3 of the page (real generated illustration in the brand's style)
|
||||
- **Reference brands' illustration** — 2-3 brands whose illustration style is close (e.g. "like Notion's homepage illustrations but a notch more grown-up")
|
||||
|
||||
### Visual World page (page 10) also adapts
|
||||
### Visual World page (page 11) also adapts
|
||||
|
||||
The 4-image lifestyle grid on page 10 should match the brand's medium mix:
|
||||
The 4-image lifestyle grid on page 11 should match the brand's medium mix:
|
||||
- All-photo brand → 4 photos
|
||||
- All-illustration brand → 4 illustrations (varied scenes, not 4 of the same composition)
|
||||
- Hybrid → mix of photos and illustrations in the proportions that match the brand world (e.g. 3 photos + 1 illustration if photography dominates)
|
||||
@@ -808,18 +714,6 @@ After generating, ask: *"If I saw this image without knowing the brand, would I
|
||||
|
||||
---
|
||||
|
||||
## Photography Must Show the Product IN USE — never a representational object
|
||||
|
||||
**For the full guidelines PDF (Step 3), the hero photography on the Cover, Photography Direction, Touchpoints, and Application pages must show the product being USED, not a representational object that stands in for the product.** This is the single most common reason a brand reads as "selling a thing it doesn't actually sell." A keychain charm shown on the cover of an app brand makes the brand look like it sells keychains. A coffee bag mockup for a streaming service makes it look like the brand sells coffee. A bottle photo for a software brand makes it look like the brand sells bottles.
|
||||
|
||||
**For digital products (apps, web services, SaaS):** hero photography must show the actual experience of using the product — selfies/photos taken with the app, screen captures of the result, people in the moment of using it (phone in hand showing the actual UI/result, video calls showing the feature in action, etc.). The brand's *symbol* can be a stylized heart-with-ears or any abstract mark — that's logo language. But brand *photography* must show the product itself, not the symbol.
|
||||
|
||||
**For physical products:** the product itself in real use — being worn, eaten, used, lived with. Not just unboxing or packshot.
|
||||
|
||||
**For services:** the moment of being served, the artifact produced, or the relationship in action.
|
||||
|
||||
**The keychain test:** ask yourself, "if a stranger saw this cover photo with no other context, would they correctly guess what we sell?" If the photo is of a physical-looking representational object that ISN'T the product, the answer is no — they'd guess the brand sells that object. Re-shoot.
|
||||
|
||||
## Photography Must Occupy Distinct Visual Territories
|
||||
|
||||
Each of the 3 mood images on the brand board must show a **genuinely different visual territory** — not three variations on the same subject. If all 3 images are "engineer at desk with laptop" with slightly different color grading, the photography has failed. The user reads three identical concepts and concludes the directions aren't really different.
|
||||
@@ -851,120 +745,12 @@ Each prompt should be **structurally different**: different subject, different s
|
||||
|
||||
---
|
||||
|
||||
## Brand Board Layout — Differentiate per Option (Step 2 of main skill)
|
||||
## Brand Board Layout — Differentiate per Option (Step 3 Preview)
|
||||
|
||||
**Critical rule for the 3-page brand board PDF built in Step 2:** do NOT use the same template for all 3 boards recolored. Each board's layout must physically embody its option's design philosophy. If you can swap colors and fonts and the layouts feel identical, the board has failed — viewers read the differences as cosmetic, not structural.
|
||||
|
||||
### Mandatory layout discipline — what makes a board NOT sloppy
|
||||
|
||||
The boards have repeatedly come back with overlapping type, mis-aligned elements, and broken hierarchy. The fixes below are non-negotiable.
|
||||
|
||||
1. **Plan the grid before writing CSS.** Sketch a 12-column × 12-row grid mentally for each board. Place each piece of content (wordmark, palette, type specimen, mood image, voice quote, brand story, references) into specific grid cells. Two pieces of content can NEVER occupy the same cell.
|
||||
|
||||
2. **Build a content inventory per board, with measured space allocations.** Example: "Wordmark = top-left 6 cols × 3 rows. Palette = top-right 6 cols × 2 rows. Mood image = middle 12 cols × 5 rows. Type specimen = bottom-left 6 cols × 2 rows. Voice quote + brand story = bottom-right 6 cols × 2 rows." If the sum of allocations exceeds 144 cells, cut content — don't squeeze.
|
||||
|
||||
3. **No text on top of text. Ever.** A headline overlapping a swatch label, a wordmark drifting into a voice quote, a tagline crashing into the brand story — all forbidden. Reserve a buffer (min 32px) around every text block.
|
||||
|
||||
4. **No text on top of busy parts of images.** Text overlaid on a mood image is allowed ONLY if the image area under the text is a flat color (e.g., a sky, a fade gradient zone). If text sits on a textured/detailed photo area, it becomes illegible. Either move the text into a flat zone, add a solid color band behind it, or lift the text out of the photo entirely.
|
||||
|
||||
5. **Render and READ each board PNG before delivering.** This is mandatory, not optional. Pipeline:
|
||||
```bash
|
||||
for i in 01 02 03; do
|
||||
"$CHROME" --headless --disable-gpu --no-sandbox --hide-scrollbars \
|
||||
--window-size=1200,850 \
|
||||
--screenshot="$WS/boards/qa-$i.png" "file://$WS/boards/$i-"*.html
|
||||
done
|
||||
```
|
||||
Then open each PNG and check: (a) no overlapping text, (b) no text on busy image regions, (c) every required content block visible, (d) layout matches the brand's design philosophy, not a template. If any check fails — fix the HTML and re-render, don't ship.
|
||||
|
||||
6. **One element per board has visual priority.** Decide before building which element is the loudest: usually the wordmark, sometimes the mood image, occasionally a giant pull quote. Everything else sizes down from that. If two elements are competing visually, the board reads as cluttered.
|
||||
|
||||
7. **Whitespace is content.** A board with 30% empty space, intentionally placed, reads as confident. A board with content crammed into every corner reads as sloppy.
|
||||
|
||||
### The differentiation rule
|
||||
**Critical rule for the 3-page brand board preview built in Step 3 of the main skill:** do NOT use the same template for all 3 boards recolored. Each board's layout must physically embody its option's design philosophy. If you can swap colors and fonts and the layouts feel identical, the board has failed — viewers read the differences as cosmetic, not structural.
|
||||
|
||||
Each brand board must answer: *what would this brand's actual hero page look like?* Build the answer.
|
||||
|
||||
### Unified aesthetic briefs STILL require structural divergence
|
||||
|
||||
When the user gives a unified aesthetic brief — one era ("Y2K"), one mood ("nostalgic"), one set of references ("Tamagotchi + PowerPuff + Bratz") — the easy failure mode is to deliver 3 variations of the same energy: all bright, all maximalist, all cartoon-y. **This is wrong.** A unified brief doesn't mean a unified output. The references the user lists are usually pointing at *sub-territories within the era*, not at a single shared visual. Pull them apart.
|
||||
|
||||
**Test before building:** when you read the brief, ask "are these references actually the same thing?" Tamagotchi is sparse 90s LCD mono. PowerPuff is loud 2001 cartoon outline. Bratz is glossy 2003 fashion editorial. These are three different visual languages from the same decade — each board should live in ONE of them, not blend all three into a single bright stew.
|
||||
|
||||
**The 3 boards must diverge across at least 4 of these dimensions** (not just color palette):
|
||||
|
||||
- **Density** — sparse breathing room / medium / dense maximalist
|
||||
- **Color saturation** — muted/desaturated / mid / neon-bright
|
||||
- **Color temperature** — warm cream-and-pink / cool blue-and-mint / monochrome / high-contrast neutrals
|
||||
- **Layout philosophy** — centered symmetric / asymmetric editorial / grid-based / full-bleed image
|
||||
- **Type energy** — quiet restrained / loud heavy display / technical mono / handwritten/script
|
||||
- **Composition focal point** — wordmark-dominant / image-dominant / quote-dominant / palette-dominant
|
||||
- **Photography subject scale** — macro detail / full scene / portrait / still life / abstract texture
|
||||
- **Photography mood** — quiet documentary / saturated cartoon / glossy editorial / lo-fi grainy
|
||||
- **Voice register** — deadpan / hype/loud / warm-poetic / technical
|
||||
- **Era within the era** — pick a specific year and pin the board to it (1998 LCD vs 2001 cartoon vs 2003 magazine)
|
||||
|
||||
**Adjective audit before delivering:** write the top-3 primary adjectives for each of the 3 boards. If 2 boards share 2+ primary adjectives (e.g. board 1: bright/maximalist/cartoon, board 2: bright/maximalist/sticker-y), they're too similar — push at least one of them further. The point of 3 options is to give the user genuinely different takes on what their brand could be, not three saturations of the same take.
|
||||
|
||||
### Divergence ≠ subtraction. Every board delivers on the brief.
|
||||
|
||||
**The brief is non-negotiable.** Whatever character the user named — Y2K, kawaii, brutalist, art-deco, cottage-core, cyberpunk — every single one of the 3 boards must deliver on it at full intensity. Variety comes from sub-territories within the brief, not from departing from the brief. Never strip a board back to differentiate it from another; never make a board boring or generic to make it "different."
|
||||
|
||||
When pushing the 3 boards into different sub-territories, the failure mode is interpreting "different from loud" as "quiet / muted / stripped back." Don't. Each board should MAXIMALLY express its own sub-territory's character — AND each board must deliver the brief's core aesthetic. If one territory is "minimal editorial within Y2K," it's still minimal-editorial-WITH-Y2K-character (chrome, holographic foil, glitter accents); it is NOT minimal-editorial-stripped-of-Y2K. If another is "lo-fi digital nostalgia within Y2K," it's still saturated-Tamagotchi-candy-and-LCD-greens; it is NOT dusty-cottage-cream.
|
||||
|
||||
**The 3 boards should feel like 3 fun things, each delivering the brief**, not "one loud + two stripped." Differentiation is on STRUCTURE (density, layout, type energy, voice register, composition), and on SUB-TERRITORY (different reference points within the same era/mood) — not on amount-of-character. Three boards, each at 100% of its own thing AND 100% of the brief.
|
||||
|
||||
**The trap to avoid:** "Board 2 is loud and saturated, so I'll make Board 1 muted and Board 3 minimal to differentiate." This sacrifices the brief for the sake of variety. The right pattern: "Board 2 is loud-PowerPuff-saturated, Board 1 is loud-Tamagotchi-candy-LCD, Board 3 is loud-Bratz-glossy-glittered-editorial." All three are still loud Y2K — they differ in *which Y2K* they live in.
|
||||
|
||||
### Era palettes are specific — don't substitute "muted" for "era-appropriate"
|
||||
|
||||
When the user names an era (Y2K, 90s, 80s, 60s, mid-century, art deco, etc.), the palette signatures of that era are non-negotiable. "Muted" is rarely the right answer; era-specific saturation profiles are. Research the era's actual palette before picking colors.
|
||||
|
||||
**Y2K palette signatures** (early 2000s, 1998–2004):
|
||||
- iMac G3 gel colors: Bondi blue, tangerine, grape, lime, strawberry, blueberry — translucent saturated
|
||||
- Tamagotchi: candy pink, mint, lavender, butter yellow, pastel egg colors
|
||||
- Bratz / mall culture: hot magenta #FF2BB8, chrome silver, lime #C8FF32, butter yellow, baby blue
|
||||
- Holographic / iridescent: shifting rainbow on chrome base
|
||||
- Cyber: lime green on black, hot pink + cyan, chrome metallic
|
||||
- Frosted: white-on-cream with chrome accents
|
||||
- NOT Y2K: dusty mauve, muted putty, cottage cream, faded sage — those are 2010s rustic/cottage-core, NOT Y2K
|
||||
|
||||
**90s grunge / 90s alt**: washed-out, desaturated, photocopied texture, off-register print
|
||||
**80s Memphis**: primary colors + black + cyan, geometric shapes, pastel accents
|
||||
**70s**: harvest gold, avocado, burnt orange, brown, mustard, warm earth tones
|
||||
**60s mod**: bold flat color blocks, op-art black/white, saturated psych
|
||||
**50s**: pastel mint, pink, baby blue, chrome + cream
|
||||
|
||||
If you've shifted a palette to "muted" or "dusty" to differentiate it from a brighter board, you've likely left the era. The era is non-negotiable; the energy varies through density, composition, and texture instead.
|
||||
|
||||
### Texture is era signal — but it must be AMBIENT, not a pattern
|
||||
|
||||
A flat-vector board with no texture reads as generic 2020s vector, not as a specific era. But the opposite failure is just as bad: a board with a recognizable repeating motif (literal halftone dots, scattered glitter flakes, visible scanline rows) reads as a graphic element stamped onto the layout, not as authentic era-character. Both fail. The right answer is **ambient atmospheric texture**: subtle grain, soft noise, grainy gradient, faded VHS shimmer — texture that the user *feels* without recognizing it as a graphic motif.
|
||||
|
||||
**Reference frame: VHS noise / film grain / paper grain / grainy gradient.** Not patterns. The user should look at the board, feel the era, and not be able to point at "the texture." If they can name what the texture is (halftone! glitter! scanlines!) it's too literal.
|
||||
|
||||
**Always generate textures via `generate_image` with `provider="gpt-image-2"`, never via CSS gradients.** CSS halftone dot patterns, conic-gradient chrome, and repeating-linear-gradient scanlines are too clean — they read as a stylesheet, not as material. Generated textures have the small imperfections (real grain, soft gradient drift, atmospheric noise) that make a board read as authentic-retro. The only acceptable CSS-rendered "texture" is something that's also a UI element (e.g. a thin holographic strip used as a divider).
|
||||
|
||||
**Pipeline for generating a texture:**
|
||||
1. `generate_image` with `provider="gpt-image-2"`, `quality="medium"`, aspect ratio 4:3 or 16:9 to cover full-canvas (not 1:1 — tileable patterns will repeat visibly).
|
||||
2. Prompt for **ambient grain / atmospheric noise**, NOT a pattern. Required prompt language: "subtle ambient grain," "atmospheric noise field," "grainy gradient," "soft film grain," "barely visible," "smooth not patterned," "NO recognizable motifs," "NO repeating elements," "NO visible patterns." Add the era's specific texture vocabulary as ATMOSPHERE not as object (e.g. "subtle VHS noise atmosphere" not "scanline pattern"; "ambient holographic shimmer" not "glitter flakes"; "soft newsprint paper grain" not "halftone dots").
|
||||
3. Download and place in `images/textures/`.
|
||||
4. Apply via CSS `background-image` on the **body** (full-board coverage), with `mix-blend-mode: overlay` or `multiply` or `screen`, at **low opacity (0.12–0.30 max)**. The texture should be felt across the whole board ambient-style, not stuck onto one card as a focal element.
|
||||
|
||||
**Era-texture prompts as ATMOSPHERE, not pattern:**
|
||||
- **Y2K LCD/handheld** → "subtle VHS noise atmosphere with very faint horizontal screen-line drift, low-contrast warm vintage screen feel, NOT a scanline pattern"
|
||||
- **Y2K cartoon/comic** → "soft newsprint paper grain, ambient print noise, barely-visible warm-cream paper roughness, NOT halftone dots"
|
||||
- **Y2K glam/Bratz** → "subtle iridescent grainy gradient, ambient holographic shimmer atmosphere, soft pink/lime color drift, NOT glitter flakes or stars"
|
||||
- **Mid-century print** → "fine paper grain atmosphere, soft warm tonal noise"
|
||||
- **80s digital** → "subtle CRT phosphor glow atmosphere, faint color drift, NOT pixel grid"
|
||||
- **90s grunge** → "ambient Xerox roughness, soft photocopy grain drift, NOT visible dust spots"
|
||||
- **70s organic** → "warm paper grain, soft fiber atmosphere"
|
||||
- **Film/photography** → "fine silver-halide grain, atmospheric noise, NOT visible particles"
|
||||
|
||||
**Application rule of thumb:** if a viewer can describe the texture as a noun ("halftone dots", "glitter flakes", "scanlines"), it's too literal. If they describe it as an atmosphere ("kind of grainy", "feels VHS-y", "soft retro feel"), it's right. The goal is era-character through ambience, not graphic elements through stamping.
|
||||
|
||||
Don't ship a "retro" board with zero texture. And don't ship a "retro" board with literal pattern overlays — generate ambient grain, apply widely, keep subtle.
|
||||
|
||||
### Examples of differentiated layouts
|
||||
|
||||
**Magazine-cover brand** — Full-bleed photo as background covering the entire page. Massive wordmark overlaid in display type. Tagline overlaid in small text. Issue/edition tag in corner ("Vol. 01 / Cover Story"). Bottom-margin strip showing color swatches and typography credit, like a magazine masthead. Whole page should look like a Bloomberg Businessweek or NYT Magazine cover.
|
||||
@@ -973,22 +759,37 @@ Don't ship a "retro" board with zero texture. And don't ship a "retro" board wit
|
||||
|
||||
**Editorial / literary essay** — Bone or cream full-bleed background. Massive italic display type centered with extreme whitespace. Photo as a small inset rectangle, not full-bleed. Pull-quote on the margin. Color swatches as a tiny ribbon at the bottom. Should feel like the opening page of a Frank Ocean visual essay or an Anthropic announcement.
|
||||
|
||||
**Tech-doc / dev-tool aesthetic** — Monochrome grid. Tight type. Code-like layout with bracket marks or syntax highlighting. Mono font everywhere. Color swatches as inline `code` blocks with hex strings. Should feel like Linear's changelog or Stripe's docs.
|
||||
**Tech-doc / dev-tool aesthetic** — Monochrome grid. Tight type. Code-like layout with bracket marks or syntax highlighting. Mono font everywhere. Color swatches as inline `code` blocks with hex strings. Should feel like Stripe's docs or GitHub's changelog.
|
||||
|
||||
### Required content per board (regardless of layout)
|
||||
|
||||
Every brand board page must include ALL of the following. The layout differentiation rule above does NOT mean cutting content — visually distinct layouts must still fit ALL the text. If a layout doesn't have room for the content, redesign the layout, don't drop content.
|
||||
|
||||
- Brand wordmark (set in the brand's display font)
|
||||
- **A standalone logo symbol/mark — rendered inline as SVG (preferred) or generated PNG. NOT just typography.** The symbol must work as a favicon, app icon, social avatar.
|
||||
- **A standalone logo symbol/mark — preferably a generated PNG via `mcp__pika__generate_image` with `provider="gpt-image-2"` when SVG would look simplistic, generic, or illegible. SVG is allowed only if it is intentionally simple and passes small-size QA. NOT just typography.** The symbol must work as a favicon, app icon, social avatar, and exported asset on transparent background at 1024×1024+.
|
||||
- **A distinctive wordmark and any seal/badge treatment — custom letter spacing, ligature/cut/terminal detail, stamp geometry, or other ownable touch. Not just a Google Font in a circle, not a generic monogram seal, and not decorative filler.**
|
||||
- Tagline
|
||||
- Voice sample (one sentence in brand voice, quoted, with a "VOICE" label)
|
||||
- Brand story (2-3 sentences in brand voice)
|
||||
- **Lifestyle world description (1-2 sentences describing the brand's visual territory — where it visually lives, who's in the frame, time of day, color temperature)** — labeled "WORLD" or similar
|
||||
- Brand story (~35 words max, min 2 sentences, one compact paragraph in brand voice)
|
||||
- **Lifestyle world description (~22 words max, min 1 full sentence describing the brand's visual territory — where it visually lives, who's in the frame, time of day, color temperature)** — labeled "WORLD" or similar
|
||||
- Lifestyle mood image (generated via gpt-image-2 — see "Photography Must Occupy Distinct Visual Territories" rule below)
|
||||
- 4-color palette with hex codes + role labels
|
||||
- Display + body type specimens with named fonts
|
||||
|
||||
### Content budgets per board
|
||||
|
||||
The board is a visual decision aid, not the final brand book. Keep each board sharp enough to sell the direction at a glance:
|
||||
|
||||
- Tagline: 8 words max.
|
||||
- Voice sample: 14 words max.
|
||||
- Brand story: 35 words max, min 2 sentences. Use one compact paragraph, not 2-3 full paragraphs.
|
||||
- World description: 22 words max, min 1 full sentence.
|
||||
- Palette: 4 colors max on the board. Full extended palettes belong in the final guidelines.
|
||||
- Type specimen: one display sample and one body sample. Do not add full hierarchy tables to boards.
|
||||
- Essential body copy: 18px minimum. If it needs to be smaller to fit, rewrite the copy.
|
||||
|
||||
If all required content cannot fit within those budgets, the content is too verbose for a board. Rewrite it; do not shrink, stack, or layer it until it becomes technically present but visually bad.
|
||||
|
||||
### What NOT to do
|
||||
|
||||
- Same left-column-color-block-right-column-photo template recolored 3 times
|
||||
@@ -996,6 +797,10 @@ Every brand board page must include ALL of the following. The layout differentia
|
||||
- Same type specimen "Aa" treatment on every page
|
||||
- Three different colors and three different fonts laid onto identical layouts
|
||||
- Wordmark with no separate symbol — the logo isn't complete without a mark
|
||||
- Dense archival/specimen styling where decorative rules, labels, swatches, and paragraphs intersect. If it looks like a broken certificate rather than a brand board, simplify.
|
||||
- Body copy crossing through color swatches, icons, seals, or decorative overlays. Text must own a clean reading area.
|
||||
- Empty mockup boxes or blank social grids. A placeholder reads as missing output, not restraint.
|
||||
- Full boards that are almost entirely one muddy value range. Use image light, accent color, or negative space to create hierarchy.
|
||||
- **Asymmetric rounded corners on color blocks** (e.g. only `border-top-left-radius` on a big shape) — these read as a clipping bug, not a design choice. If you want softness, use **symmetric** rounded corners (whole left edge rounded, or all four corners rounded), a **clean rectangular split**, or a deliberately organic shape via SVG/clip-path. Half-rounding a single corner of a big block looks like a mistake every time.
|
||||
- **Inline pill backgrounds on display text (40px+)** — they overlap into adjacent lines because line-height is usually tighter than the rendered character box. A `background: var(--color); padding: 0 12px; border-radius: 12px;` on big headline text WILL bleed into the line above or below. Two safer options:
|
||||
1. **Highlighter-underline gradient** (recommended): `background: linear-gradient(to bottom, transparent 0%, transparent 58%, var(--accent) 58%, var(--accent) 92%, transparent 92%); padding: 0 6px; -webkit-box-decoration-break: clone; box-decoration-break: clone;` — creates a marker-highlight band that only occupies the bottom of the line, never extends beyond.
|
||||
|
||||
@@ -7,7 +7,7 @@ genuinely different in name personality, color mood, and voice — not just pale
|
||||
|
||||
```
|
||||
### Option [1/2/3]: [Brand Name]
|
||||
**Tagline:** [Short punchy line — under 8 words]
|
||||
**Tagline:** [Short punchy line — 8 words max]
|
||||
|
||||
**Colors:**
|
||||
- [Name]: #[hex] (role: primary/accent/background/text)
|
||||
@@ -26,71 +26,22 @@ A complete brand identity has BOTH a wordmark and a symbol — they're different
|
||||
- **Symbol** = a standalone graphic mark that lives WITHOUT the wordmark. Used for app icon, favicon, social avatar, browser tab — anywhere the wordmark is too long.
|
||||
- **Lockup** = how the two combine (horizontal, stacked, symbol-only).
|
||||
|
||||
### Logo Pipeline — generate a high-res symbol via gpt-image-2, ship it as transparent PNG (no tracing)
|
||||
|
||||
**Hand-coded SVG symbols look amateur.** Do not write `<path d="M..."/>` strings to build the brand symbol. But also: **don't trace the gen'd symbol to SVG.** Keep the symbol as a high-resolution transparent PNG. Only the wordmark gets vectorized (in the brand kit). Pipeline:
|
||||
|
||||
1. **Generate the symbol via `generate_image` with `provider="gpt-image-2"`, `quality="high"` (or `"medium"` for first drafts), 1:1 aspect ratio, 1024×1024 minimum.** The symbol's visual style is a brand-personality decision — pick the style that fits the brand's voice and design language, not a default.
|
||||
- **Flat is correct** when the brand reads as flat — sticker-style, screen-print, comic/cartoon, vintage print, Memphis-era, modernist, anything that lives in a 2D vocabulary. Most brands land here.
|
||||
- **Dimensional / 3D / glossy / painted / photographic** is correct when the brand actually reads that way — luxury beauty, Y2K product-render, 3D-rendered toy, ceramic craft, anything where the visual world has volumetric depth as part of its character.
|
||||
- **Neither is the default.** The trace requirement is gone, but that doesn't mean dimensional is "better." Style follows brand. If you removed the trace constraint and immediately reached for dimensional, ask: would the brand prefer flat? Often the answer is yes.
|
||||
- The only style-style constraints are the ones in "Symbol output rules" below (≤3 colors, recognizable at 16×16, etc.). Within those, you're free.
|
||||
2. **The prompt MUST include the no-text guardrail**: "absolutely no text, no letters, no typography, no words, no characters anywhere in the image." gpt-image-2 produces garbled fake text inside logos if you don't explicitly forbid it.
|
||||
3. **Show the user the gen'd symbol(s).** Generate 2-3 variations if appropriate. Wait for user approval before committing.
|
||||
4. **Once approved, save as a high-res transparent PNG.** Standard: 2048×2048 PNG with true alpha-channel transparency (verify with PIL — see `brand-guidelines.md` "gpt-image-2 transparent-background caveat"; key out the painted near-white pixels if needed, or regenerate with a solid bg matching the placement surface). Use this PNG everywhere the symbol appears — guidelines pages, brand kit, etc. **Do NOT trace to SVG.** The symbol stays raster.
|
||||
5. **Wordmark = Google Font, NEVER baked into the image.** Pick a Google Font that matches the brand vibe (see `brand-guidelines.md` "Must Have Character"). The wordmark is rendered as live text in HTML/CSS for the guidelines pages, and converted to text-as-paths SVG in the brand kit (see step 7).
|
||||
6. **Lockup composition is perfectly measured + permanently fixed.** Pick one horizontal lockup geometry AND one stacked lockup geometry. For each, specify: symbol size (px or em), wordmark font size (px), gap between symbol and wordmark (px), vertical baseline alignment (which point of the symbol aligns with which baseline of the wordmark). **The measurements never change across color variants or contexts.** Document the exact measurements on the Logo page so a designer or developer can rebuild the lockup without guessing.
|
||||
7. **In the brand kit:** the symbol ships as `symbol-[color].png` (high-res transparent raster) at multiple sizes (16, 32, 64, 128, 256, 512, 1024, 2048). The wordmark ships as `wordmark-[color].svg` with the Google Font text converted to outlined paths (so the SVG renders identically without the font file installed). The lockup ships as `lockup-[orientation]-[color].svg` containing the symbol PNG embedded inline + the wordmark as paths, positioned at the locked measurements.
|
||||
|
||||
### Symbol output rules — what the gen'd symbol must satisfy
|
||||
|
||||
The symbol doesn't need to be flat or traceable anymore — but it does still need to function as a logo. Every gen'd symbol must satisfy ALL of:
|
||||
|
||||
- **Conceptually linked to the product / brand.** The mark must MEAN something about what the brand IS or DOES. Not a decorative shape, not a random pretty thing — a mark that connects to the brand's reason for existing. The koalacore heart-with-koala-ears doesn't just look cute; it says "a koala you wear like an accessory." A camera app's symbol could be a shutter aperture (product-feature), an abstract eye (sense of seeing), or a literal frame (what the app does to a photo). Whatever the concept lane (see "Symbol concepts must differ across the 3 brand options"), there has to be a real conceptual link.
|
||||
- **Feels unique. Makes sense. Doesn't feel generic.** If you can describe the mark in a way that would fit 50 other unrelated brands ("a heart," "a smiling face," "an arrow"), push further — add specificity that ties it to this brand's actual personality. The uniqueness comes from the SPECIFIC interpretation, not the broad category.
|
||||
- **Recognizable at small size — mandatory favicon test after every gen.** The mark must read as the brand at 16×16 (favicon size). Test procedure (run AFTER every symbol gen, before any approval ask): use PIL to resize the symbol to 16×16 with LANCZOS, then upscale that 16-px version with NEAREST to ~128×128 (pixel-doubling, so you can actually see what a viewer sees at favicon scale). Read the test image. If the dominant elements have collapsed into mush or vanished entirely, the mark fails — **regenerate**, don't ship.
|
||||
- **The fix when a mark fails the test depends on what went wrong:**
|
||||
- **All-hairline mark dissolves at small scale** (common for luxury / heritage / engraved aesthetics). Fix: give the *dominant central element* a solid filled mass while keeping the framing details (outer ring, decorative crest, laurel, ornamental flourishes) in hairline weight. The DMV exercise demonstrated this: hairline car silhouette → vanished at 16 px; same composition with the car silhouette filled solid → reads clearly at every scale. The lux feel survives because the engraved fine-line frame is still there at large sizes — at small sizes, the solid central mass carries the recognition load.
|
||||
- **Too many tiny detail elements** (e.g. many separate ornaments, fine pinstripes, small leaves). Fix: simplify — keep one or two dominant elements, remove the rest.
|
||||
- **Weak silhouette / outline blends with background.** Fix: thicken outlines OR pick colors with stronger luminance contrast against likely placement backgrounds.
|
||||
- **Light-on-light or low-contrast composition.** Fix: increase the contrast between the mark's dominant color and its background.
|
||||
- **Style is preserved, weight is adjusted.** A richly-rendered 3D mark, a painted mark, a photographic mark, a hairline engraving — all can pass the favicon test by ensuring the dominant readable element has enough mass at 16 px. The fix is compositional, not a style downgrade.
|
||||
- **No more than 3 colors.** Counting the dominant color regions, not individual gradient stops. The mark can be a gradient FROM one color TO another (that counts as 2). Or a 3-tone illustration. But not a full rainbow / not a richly multicolored scene. 3 dominant colors max keeps the mark memorable and reproducible across material applications (embroidery, print, etc.).
|
||||
- **High resolution.** Generate at gpt-image-2's highest available res, then upscale or regenerate larger if needed. The final shipped PNG should be at least 2048×2048 so it scales cleanly to billboard size without pixelation.
|
||||
- **NO TEXT inside the symbol image.** Ever. Text goes in the wordmark only. The symbol is purely a visual mark.
|
||||
- **Transparent background, verified.** True alpha=0 in the transparent region. If gpt-image-2 paints near-white pixels in the "transparent" area (it often does), key them out with PIL or regenerate with a solid bg that matches the placement surface.
|
||||
|
||||
If the gen'd output has text, more than 3 dominant colors, or fails the 16×16 recognizability test, **regenerate** — don't ship it.
|
||||
|
||||
### Symbol concepts must DIFFER across the 3 brand options
|
||||
|
||||
The three identity options are different brands, not the same brand in three skins. Their symbols should differ in *concept*, not only in visual style. **Do not generate three "literal brand mascot face" logos in three styles** — that's one idea repeated.
|
||||
|
||||
Pick a different concept lane for each option. Possible concept lanes:
|
||||
- **Literal mascot / character** — the brand's animal/object rendered as the mark (most expected).
|
||||
- **Product-feature reference** — symbol references what the product *does* (e.g. a camera shutter for a camera app, a needle-and-thread for a tailor, a flame for a delivery app).
|
||||
- **Abstract / geometric mark** — a non-representational shape with meaning (e.g. an arrow, an arc, a spiral, a chevron).
|
||||
- **Monogram** — the brand initial(s) drawn distinctively.
|
||||
- **Hybrid** — two concepts fused into one shape (e.g. a heart whose top is animal ears, an arrow that's also a leaf).
|
||||
- **Container / frame** — a window, badge, stamp, or seal that holds the brand's signature element.
|
||||
|
||||
When proposing 3 options, force the symbols across at least 2–3 different concept lanes. The brand-board comparison is more useful when the symbols argue different ideas about what the brand IS.
|
||||
|
||||
### Whether to generate new ones depends on what the user has
|
||||
|
||||
**Whether to generate new ones depends on what the user has:**
|
||||
- If the user has an existing wordmark or symbol they like — USE it. Document the existing asset in the guidelines.
|
||||
- If the user is asking for a new logo, or has said they don't like their current one — propose a new wordmark and/or symbol as part of the identity option.
|
||||
- If the user has a wordmark but no symbol — propose just the symbol. Brands need a non-typographic mark for app icons, favicons, etc., so a symbol is worth proposing even when the wordmark is kept.
|
||||
- Whatever the source, document BOTH in the guidelines. They're both part of complete brand documentation, even when only one is new.
|
||||
|
||||
For each identity option, fill in:
|
||||
- **Wordmark:** [Which Google Font (or commercial font) is the wordmark set in? Why does it match the brand vibe? Reference existing if kept; describe new if proposed. Spec the font weight + letterspacing + baseline.]
|
||||
- **Symbol / mark:** [The standalone graphic mark — shape, reference, what it evokes. Generated via gpt-image-2 (NOT hand-coded SVG). Must read at 16×16 AND 512×512. Reference existing if kept; describe new if proposed.]
|
||||
- **Lockup:** [How wordmark + symbol combine — horizontal (symbol left, name right at exact x-offset), stacked (symbol above name with measured spacing), symbol-only at small sizes. Lock the placement and don't vary it across color variants.]
|
||||
- **Wordmark:** [How the brand name is typeset — typeface choice, custom letter treatment, spacing, ligature/cut/terminal detail, and lockup rhythm. Reference existing if kept; describe new if proposed. A wordmark is not just a Google Font typed in a brand color.]
|
||||
- **Symbol / mark:** [The standalone graphic mark — shape, reference, what it evokes. For new marks, prefer a generated PNG via `mcp__pika__generate_image` with `provider="gpt-image-2"` when image generation gives a richer, more ownable mark than hand-written SVG. Ask for a clean isolated mark on transparent background, centered, no text/watermark/mockup. Use SVG only if the idea is simple enough to draw cleanly. Must read at 16×16 AND 512×512. Reference existing if kept; describe new if proposed.]
|
||||
- **Lockup:** [How wordmark + symbol combine — horizontal (symbol left, name right), stacked (symbol above name), symbol-only at small sizes.]
|
||||
|
||||
**Brand story:**
|
||||
[2-3 sentences for the About page. Written in brand voice. Real copy, not a template.]
|
||||
|
||||
Board-preview copy is intentionally shorter: when this story appears on the 3-option brand board, rewrite it to the board budget (~35 words max, min 2 sentences, one compact paragraph). The 2-3 sentence version belongs in the text option/About-page snippet, not the preview board. For the full guidelines foundation page, expand the story to 2-3 paragraphs of real copy in brand voice.
|
||||
|
||||
**Product / hero subject photography direction:**
|
||||
[How the main subject of brand photography should look — for product brands: the product itself.
|
||||
For service / app / community brands: the hero subject of the brand (the person using it, the
|
||||
@@ -125,56 +76,46 @@ After presenting all 3 identity options in text, **render a 3-page brand board P
|
||||
|
||||
### How to generate:
|
||||
|
||||
Build one self-contained HTML file per option, render each to PDF via Chrome headless, then merge into a single 3-page PDF. Each page MUST have a layout that physically embodies its option's design philosophy (not three template recolours — see `brand-guidelines.md` "Brand Board Layout — Differentiate per Option" for rules and examples).
|
||||
Build three 1200×850 HTML body fragments and render them through Pika MCP:
|
||||
|
||||
```bash
|
||||
WS="${BUILD_A_BRAND_WS:-$HOME/build-a-brand-workspace}"
|
||||
mkdir -p "$WS/boards"
|
||||
1. Use `html_to_pdf` with `body_pages` + `shared_head` so the server renders and merges the three boards into one PDF.
|
||||
2. Use `html_to_png` once per board for QA previews at `viewport_px: { width: 1200, height: 850 }`.
|
||||
3. Read every PNG preview before sending the PDF URL to the user. The check is not only "does it fit?" It must also look like a good brand board: one primary visual focal point with supporting required content grouped clearly, readable text, no empty placeholders, no muddy one-note palette, and no body copy intersecting decorative elements.
|
||||
4. Run `mcp__pika__analyze_media` on each PNG:
|
||||
- Prompt: "Answer with PASS or FAIL on the first line, then explain. Does this brand board look polished enough to send? Check for ugly density, unreadable small text, text overlap, missing/empty mockups, muddy palette, weak hierarchy, clipped text, occluded text, and required board content. Treat flat color rectangles in website/social/app mockups as empty placeholders unless they are explicitly palette specimens. Masthead wordmark/tagline/issue metadata overlays on photos are allowed only with deliberate negative space or a contrast scrim and must pass contrast QA as defined in `brand-guidelines.md` Rule 5; body copy must not overlap images."
|
||||
- If the tool returns `{task_id, status: "running"}`, poll `mcp__pika__task_status({task_id})` until terminal before judging.
|
||||
- Treat the QA as unavailable and halt with a manual-review warning if the tool is missing, raises `tool_not_found`, `provider_unavailable`, `unsupported_media_type`, `rate_limited`, `quota_exceeded`, `auth_error`, any HTTP 4xx/5xx error envelope, or a transport error, says it cannot analyze the image, or returns final text that does not match ``/^\s*[`*]{0,2}(PASS|FAIL)\b/``.
|
||||
- Interpret only the regex capture. Fix every captured FAIL before delivery. If the captured result is PASS but the explanation lists a blocking collision, clipping, unreadable text, or missing required board content, treat it as FAIL. Do not proceed silently.
|
||||
|
||||
# 1. Write 3 HTML files: $WS/boards/01-<option-slug>.html ... 03-<option-slug>.html
|
||||
# (Each with @page { size: 1200px 850px; margin: 0; } in CSS, fonts via @font-face file:// — see Step 0 of brand-guidelines.md.)
|
||||
Each page MUST have a layout that physically embodies its option's design philosophy (not three template recolours — see `brand-guidelines.md` "Brand Board Layout — Differentiate per Option" for rules and examples).
|
||||
|
||||
# 2. Render each to PDF via Chrome headless
|
||||
CHROME="/Applications/Google Chrome.app/Contents/MacOS/Google Chrome" # macOS path
|
||||
# Linux: CHROME=google-chrome | Windows-WSL: CHROME=/mnt/c/Program\ Files/Google/Chrome/Application/chrome.exe
|
||||
Assets in the HTML must be server-reachable: use HTTPS URLs from Pika tools, public raw URLs, or inline `data:` URIs. Local files and fonts must not appear as `file://` references.
|
||||
|
||||
for i in 01 02 03; do
|
||||
"$CHROME" --headless --disable-gpu --no-sandbox --hide-scrollbars \
|
||||
--virtual-time-budget=8000 --no-pdf-header-footer \
|
||||
--print-to-pdf="$WS/boards/$i.pdf" \
|
||||
"file://$WS/boards/$i-"*.html
|
||||
done
|
||||
|
||||
# 3. Merge into one PDF
|
||||
pdfunite "$WS/boards/01.pdf" "$WS/boards/02.pdf" "$WS/boards/03.pdf" "$WS/boards/brand-boards.pdf"
|
||||
|
||||
# 4. QA: screenshot each page, read every one before sending
|
||||
for i in 01 02 03; do
|
||||
"$CHROME" --headless --disable-gpu --no-sandbox --hide-scrollbars \
|
||||
--virtual-time-budget=8000 --window-size=1200,850 \
|
||||
--screenshot="$WS/boards/qa-$i.png" "file://$WS/boards/$i-"*.html
|
||||
done
|
||||
```
|
||||
|
||||
If `pdfunite` (from poppler-utils) isn't installed: `brew install poppler` on macOS, or fall back to Python `pypdf.PdfWriter().append()`. Chrome ships on every Mac at `/Applications/Google Chrome.app/...` — no install needed.
|
||||
|
||||
**Delivery:** save `$WS/boards/brand-boards.pdf` to `~/Desktop/[brand-slug]-brand-boards.pdf` and emit the local path in the chat. Do not upload or host the PDF unless the user explicitly asks for a hosted copy.
|
||||
**Delivery:** return the `file_url` from `html_to_pdf`. Save a local copy only when a downstream brand-kit export needs one.
|
||||
|
||||
**Design rules for each board:**
|
||||
- Each option panel uses its own background color from that option's palette
|
||||
- Brand name displayed large in that option's display typeface (loaded via `@font-face file://` from `$WS/fonts/`)
|
||||
- Brand name displayed large in that option's display typeface (loaded via HTTPS font URL or inline `data:font/...`)
|
||||
- Color swatches shown as circles or rectangles with color names beneath
|
||||
- Typography is clean and editorial — no generic fonts (no Inter / Karla / DM Sans default)
|
||||
- Layout philosophy differs per board — a tabloid board looks like a newspaper, a fashion-house board looks like a lookbook spread, an archival board looks like a book frontispiece. See `brand-guidelines.md` "Brand Board Layout — Differentiate per Option" for the differentiation rule and `brand-guidelines.md` "Required content per board" for the per-page checklist.
|
||||
- No generic AI mood-board collages. Each board may use one purpose-built `gpt-image-2` mood image that matches that board's photography/illustration direction.
|
||||
- No AI-generated imagery on the boards beyond the single required lifestyle mood image per board and an optional generated PNG symbol when that is the strongest logo route. No generic AI mood-board collages, AI app mockups, AI product mockups, AI seals, or extra AI filler assets.
|
||||
- The whole thing should look like something a real brand studio would produce
|
||||
|
||||
**Board copy budgets:**
|
||||
- Tagline: 8 words max.
|
||||
- Voice sample: 14 words max.
|
||||
- Brand story: 35 words max, min 2 sentences.
|
||||
- World description: 22 words max, min 1 full sentence.
|
||||
- Body text: 18px minimum. Labels: 14px minimum unless purely decorative.
|
||||
- If copy needs to be smaller than that, rewrite. Do not shrink, layer, or cram until the board technically contains everything but looks bad.
|
||||
|
||||
## Quality Bar
|
||||
|
||||
- **Names**: Should be memorable, say-able, and googleable. Avoid made-up words unless they're genuinely good.
|
||||
- **Colors**: Give them real names (not "Dark Blue" — try "Ink", "Dusk", "Bone"). Specify roles.
|
||||
- **Voice examples**: Write an actual sentence in the brand's voice (a headline, a button label, an error message), not a description of the voice.
|
||||
- **Logo concepts**: Cover BOTH wordmark and symbol in every identity option — they're different things doing different jobs. Describe each visually — shape, reference, style. (e.g. wordmark: "a custom hand-drawn serif, slightly imperfect, like a signature"; symbol: "a thin-line greyhound silhouette, drawn mid-stride, in a single continuous line"). What you GENERATE depends on what the user has: use existing assets if the user wants to keep them, propose new ones if the user needs a logo or doesn't like their current one. If the user has a wordmark but no symbol, still propose a symbol — favicons and app icons need a non-typographic mark.
|
||||
- **Logo concepts**: Cover BOTH wordmark and symbol in every identity option — they're different things doing different jobs. Describe each visually — shape, reference, style. (e.g. wordmark: "a custom hand-drawn serif, slightly imperfect, like a signature"; symbol: "a thin-line greyhound silhouette, drawn mid-stride, in a single continuous line"). What you GENERATE depends on what the user has: use existing assets if the user wants to keep them, propose new ones if the user needs a logo or doesn't like their current one. For a new symbol, prefer the generated PNG route when a hand-authored SVG would look simplistic or illegible; use `mcp__pika__generate_image` with `provider="gpt-image-2"` and a transparent-background, no-text prompt. If the user has a wordmark but no symbol, still propose a symbol — favicons and app icons need a non-typographic mark.
|
||||
- **Fonts must have character.** Don't default to Inter / Karla / Outfit / DM Sans / Lato — they have no point of view. Explore the full Google Fonts library. See `brand-guidelines.md` "Font Selection — Must Have Character" for approved high-character options.
|
||||
- **Brand story**: Should make someone feel something. Name the founder's origin if appropriate.
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# brand.md — Template Reference
|
||||
|
||||
This template defines the structure of the `brand.md` file delivered in Step 4's brand kit zip. It's a comprehensive, machine-readable brand spec that lets users (or downstream AI tools) produce on-brand work without needing the 14–16-page guidelines PDF.
|
||||
This template defines the structure of the `brand.md` file delivered in Step 5's brand kit zip. It's a comprehensive, machine-readable brand spec that lets users (or downstream AI tools) produce on-brand work without needing the 15-page guidelines PDF.
|
||||
|
||||
## Why this format
|
||||
|
||||
@@ -157,44 +157,28 @@ This template defines the structure of the `brand.md` file delivered in Step 4's
|
||||
|
||||
### Wordmark
|
||||
|
||||
[Description of the wordmark — which Google Font (or commercial font) it's set in, weight, letterspacing, why it matches the brand vibe. The wordmark is ALWAYS real font rendering, never a generated image.]
|
||||
[Description of the wordmark — typeface, treatment, custom touches, what it feels like.]
|
||||
|
||||
Files (in `logo/wordmark/`):
|
||||
- `wordmark-[primary].svg` — Google Font converted to text-as-paths (renders without the font file installed)
|
||||
- `wordmark-[primary].png` — 1024-wide raster fallback
|
||||
- `wordmark-on-dark.svg/.png` — variants for dark backgrounds
|
||||
- `wordmark-[primary].svg/.png/.pdf` — for use on background color
|
||||
- `wordmark-on-dark.svg/.png/.pdf` — for use on dark backgrounds
|
||||
- [etc — list all variants]
|
||||
|
||||
### Symbol
|
||||
|
||||
[Description of the symbol/mark — shape, what it evokes, how it relates to the brand. The symbol was generated via gpt-image-2 and ships as a high-resolution transparent PNG (NOT traced to SVG — only the wordmark gets vectorized). Style can be flat, dimensional, painted, photographic, gradient-rich — whatever fits the brand. Concept lane (mascot / product-feature / abstract / monogram / hybrid / container): [name the lane and explain how it links to the brand].]
|
||||
|
||||
Symbol output checks (all must pass):
|
||||
- Conceptually linked to brand
|
||||
- Feels unique (would fit ONLY this brand)
|
||||
- Recognizable at 16×16 favicon size (passed mandatory PIL favicon test)
|
||||
- No more than 3 dominant colors: [list them with hex]
|
||||
- Shipped at 2048×2048+ master resolution
|
||||
- No text inside the image
|
||||
- True alpha=0 transparency verified
|
||||
[Description of the symbol/mark — shape, what it evokes, how it relates to the brand metaphor.]
|
||||
|
||||
Files (in `logo/symbol/`):
|
||||
- `symbol-[primary]-16.png` through `symbol-[primary]-2048.png` — primary color at 8 sizes, transparent background
|
||||
- `symbol-on-dark-[sizes].png` — variant for dark backgrounds (if needed)
|
||||
- [etc — variants ship as PNG only; symbol is NEVER vectorized to SVG]
|
||||
- `symbol-[primary].svg/.png/.pdf` — primary color on transparent
|
||||
- `symbol-on-dark.svg/.png/.pdf` — reversed for dark backgrounds
|
||||
- [etc]
|
||||
|
||||
### Lockup
|
||||
|
||||
[Description of how wordmark + symbol combine. Primary lockup (stacked or horizontal). Secondary lockups. When to use each. Measurements are PERFECTLY MEASURED and PERMANENTLY FIXED across all color variants — never let the wordmark drift between cream / pink / dark versions.]
|
||||
|
||||
**Locked measurements (specify both lockups):**
|
||||
- Horizontal: symbol [WxH px], gap [X px], wordmark font-size [X px], alignment [optical center / baseline]
|
||||
- Stacked: symbol [WxH px], vertical gap [X px], wordmark font-size [X px]
|
||||
[Description of how wordmark + symbol combine. Primary lockup (stacked or horizontal). Secondary lockups. When to use each.]
|
||||
|
||||
Files (in `logo/lockup/horizontal/` and `logo/lockup/stacked/`):
|
||||
- `lockup-[orientation]-[color].svg` — symbol PNG embedded inline + wordmark as text-as-paths at locked measurements
|
||||
- `lockup-[orientation]-[color].png` — assembled lockup as raster, 1024-wide
|
||||
- [List all color variants]
|
||||
- [List variants]
|
||||
|
||||
**Clear space rule:** [Minimum space around the lockup — e.g. "Minimum one symbol-height of space on all sides."]
|
||||
|
||||
@@ -204,8 +188,6 @@ Files (in `logo/lockup/horizontal/` and `logo/lockup/stacked/`):
|
||||
|
||||
## Icons
|
||||
|
||||
*(This section only applies to digital brands — app, web, SaaS, software platform. For non-digital brands (product, fashion, restaurant, service), delete this section entirely and skip the `/icons/` folder in the kit.)*
|
||||
|
||||
**Stroke weight:** [e.g. 1.5px / 2px / 3px]
|
||||
**Corner radius:** [sharp / 2px / round]
|
||||
**Line caps:** [butt / round / square]
|
||||
@@ -375,6 +357,6 @@ These prompts encode the brand voice rules from this spec into instructions the
|
||||
|
||||
## What NOT to include
|
||||
|
||||
- Don't include the full 14–16-page guidelines verbatim. The brand.md is the SPEC, not the manual. Keep it tight.
|
||||
- Don't include the full 15-page guidelines verbatim. The brand.md is the SPEC, not the manual. Keep it tight.
|
||||
- Don't include marketing fluff. Every section should be either directly usable copy or actionable rules.
|
||||
- Don't include build mechanics (WeasyPrint quirks, prompt templates, etc.). Those are skill-internal.
|
||||
|
||||
@@ -395,7 +395,7 @@ If `len(zoom_keyframes) > 0`, call `mcp__pika__edit_animate_zoom` with `video_ur
|
||||
| **`pika`** (default) | ~2–5 min | Slightly more dramatic, naturalistic | Default for most runs — fast iteration, watchable output, ~10× faster than kling |
|
||||
| `kling` (opt-in) | ~5–30 min | Minimal, face-centered, presenter-style | High-stakes renders where the avatar must read like a polished presenter; tolerate the long pole |
|
||||
|
||||
Most runs complete inline. If lipsync returns an async status handle, follow the MCP tool's status flow until it reaches a terminal state. On success, capture `lipsync_url`. On failure or cancellation, fall back to the **other** provider (kling ↔ pika) per the failover note below.
|
||||
Server-side-await covers the call inline; if the response shape is `{task_id, status: "queued"}`, poll `mcp__pika__task_status` in a tight loop (no sleep) until the status reaches a terminal state (`completed`, `failed`, or `cancelled`). On `completed`, capture `lipsync_url`. On `failed` / `cancelled`, fall back to the **other** provider (kling ↔ pika) per the failover note below.
|
||||
|
||||
**Failover:**
|
||||
- If `pika` fails (rare — parrot a2v is robust at typical explainer audio lengths) → retry once with `provider: "kling"`.
|
||||
|
||||
@@ -8,7 +8,24 @@ description: >-
|
||||
"60s pitch video", "make a video of [founder] for [URL]", "talking founder explainer". Requires
|
||||
Pika MCP. Uses a supplied brand kit folder (`brand.json` or an exported build-a-brand kit with
|
||||
`brand.md`, tokens, logo assets); if no kit exists, run build-a-brand first.
|
||||
argument-hint: <product-url> --founder "<name, role>" --photo <path|url|generate> [brand-kit=<path>] [aspect=16:9|9:16|1:1]
|
||||
argument-hint: <product-url> --founder "<name, role>" [--photo <path|url|generate>] [brand-kit=<path>] [aspect=16:9|9:16|1:1] [--quick] [--config <path>]
|
||||
required-capabilities:
|
||||
- mcp__pika__add_captions
|
||||
- mcp__pika__analyze_brief
|
||||
- mcp__pika__analyze_media
|
||||
- mcp__pika__edit_audio_mix
|
||||
- mcp__pika__edit_concat
|
||||
- mcp__pika__edit_pip
|
||||
- mcp__pika__edit_text_overlay
|
||||
- mcp__pika__extract_frame
|
||||
- mcp__pika__generate_image
|
||||
- mcp__pika__generate_music
|
||||
- mcp__pika__generate_reference_video
|
||||
- mcp__pika__generate_slide_animation
|
||||
- mcp__pika__html_to_png
|
||||
- mcp__pika__render_html_animation
|
||||
- mcp__pika__task_status
|
||||
- mcp__pika__upload_asset
|
||||
---
|
||||
|
||||
# founder-product-video
|
||||
@@ -30,24 +47,48 @@ No cutaways. No website CSS extraction. AI generation, deterministic render, cap
|
||||
>
|
||||
> Example: `/founder-product-video https://scrapegraphai.com --founder "Eli Kim, CEO" --photo ~/Pictures/eli.jpg`
|
||||
|
||||
**If args carry partial input**, skip the menu and gather the missing required fields by asking one at a time — ask, wait, ask the next. Don't bundle questions into one block. If the user supplies a field unprompted (e.g. they pasted a URL in the trigger message), skip that question and confirm the value back to them once at the end. Don't start the pipeline until all required fields are answered.
|
||||
**If args carry partial input in interactive mode**, skip the menu and gather the missing required fields by asking one at a time — ask, wait, ask the next. Don't bundle questions into one block. If the user supplies a field unprompted (e.g. they pasted a URL in the trigger message), skip that question and confirm the value back to them once at the end. Don't start the pipeline until all required fields are answered. If the non-interactive fast lane applies, use step [0.5] instead.
|
||||
|
||||
### [0.5] Non-interactive fast lane
|
||||
|
||||
Use this path when the caller passes `--quick` or `--config <path>`, or when the
|
||||
caller states they are running from CI, a subagent, a batch job, or any other
|
||||
non-interactive harness.
|
||||
|
||||
This section has precedence over the interactive ask/wait instructions below.
|
||||
When it applies, use this fast lane and do not fall through to the multi-turn
|
||||
intake unless `url` or founder name/role is truly missing.
|
||||
|
||||
- `--config <path>` points to a JSON file with pre-baked values for the canonical
|
||||
input contract: `url`, `brand_kit_path` or `build_brand`, `founder_name`,
|
||||
`founder_role`, `founder_photo`, `assets`, `music_url`, `aspect_ratio`,
|
||||
`location_image_url`, `voice_style`, `product_type`, and `lower_third`.
|
||||
- `--quick` means use defaults for optional extras, auto-build the brand kit with
|
||||
`build-a-brand --quick` if `brand-kit` is omitted, and use `founder_photo =
|
||||
"generate"` when no photo is supplied.
|
||||
- For `--quick` or `--config`, do not stop for confirmation at the brand-kit
|
||||
branch, founder-photo generation prompt, optional-extras prompt, script
|
||||
choices, or end-card/caption defaults. Record assumptions inline and continue.
|
||||
- If `url` or founder name/role cannot be found in args or config, stop once with
|
||||
a single compact missing-fields list instead of starting a multi-turn Q&A loop.
|
||||
|
||||
**1. Product URL** *(required)* — `https://...`. Used to (a) derive the brief in step [1] and (b) feed the brand-kit branch below.
|
||||
|
||||
**2. Brand kit** *(required)* — ask: *"Do you already have a brand kit folder, or should I build one first?"*
|
||||
- If a path → use it (`state.brand_kit_path = <path>`). Accept either `brand.json` or an exported `build-a-brand` kit containing `brand.md`, `tokens/tokens.json`, and logo assets.
|
||||
- If "build" → invoke the `build-a-brand` skill on the URL/brief and wait for the exported brand kit. This is a full identity workflow and may pause for user choices; surface those prompts rather than trying to bypass them.
|
||||
- After either branch, set `state.brand_kit_path` and confirm the path back to the user before continuing.
|
||||
**2. Brand kit** *(required)* — interactive mode: ask *"Do you already have a brand kit folder, or should I build one first?"*
|
||||
- If a path -> use it (`state.brand_kit_path = <path>`). Accept either `brand.json` or an exported `build-a-brand` kit containing `brand.md`, `tokens/tokens.json`, and logo assets.
|
||||
- If "build" -> invoke the `build-a-brand` skill on the URL/brief and wait for the exported brand kit. This is a full identity workflow and may pause for user choices; surface those prompts in interactive mode.
|
||||
- Fast lane: if config provides `brand_kit_path`, use it. If config sets `build_brand` or `--quick` omits `brand-kit`, invoke `build-a-brand --quick` on the URL/brief and wait for the exported brand kit; do not surface build-a-brand prompts or stop for identity choices. After either branch, set `state.brand_kit_path`. Only stop with a single compact missing-fields list if there is no path and the brand kit cannot be built.
|
||||
|
||||
**3. Founder identity** *(required)* — ask all three together:
|
||||
**3. Founder identity** *(required)* — interactive mode: ask all three together:
|
||||
- `founder_name` — e.g. "Semi"
|
||||
- `founder_role` — e.g. "CEO, Pika"
|
||||
- `founder_photo` — local path / https URL / OR the literal string `generate` to auto-create a portrait. If `generate`, prompt the user for a 1-line vibe ("warm, casual smart attire" / "Pixar-style 3D animation" / etc.) — this becomes the seed prompt for `generate_image` in step [4].
|
||||
- `founder_photo` — local path / https URL / OR the literal string `generate` to auto-create a portrait. If `generate`, prompt the user for a 1-line vibe ("warm, casual smart attire" / "Pixar-style 3D animation" / etc.) — this becomes the seed prompt for `mcp__pika__generate_image` in step [4].
|
||||
- Fast lane: use founder values from args/config. If `founder_photo` is omitted, set `founder_photo = "generate"` and use a neutral founder-portrait vibe derived from the product tone; do not stop for a separate photo-vibe prompt.
|
||||
|
||||
Default to no lower-third so the happy path stays MCP-only. If the user explicitly asks for a lower-third, record `state.lower_third = true` and confirm that one local ffmpeg overlay pass will be used.
|
||||
Default to no lower-third so the happy path stays MCP-only. If the user explicitly asks for a lower-third, record `state.lower_third = true`; in interactive mode confirm that one local ffmpeg overlay pass will be used, and in the fast lane record that assumption inline.
|
||||
|
||||
**4. Optional extras** — offer these once as a single message, then proceed without waiting if no answer comes back in the same turn:
|
||||
- *Custom imagery* — list of `assets` (product photos / app screenshots) shown on the founder's phone. **Default if omitted:** use screenshots from `brand.json.screenshots` when present, otherwise look for obvious screenshots or product images inside the brand kit. If none exist, ask once whether to proceed without imagery or wait for uploads.
|
||||
**4. Optional extras** — interactive mode: offer these once as a single message, then proceed without waiting if no answer comes back in the same turn. Fast lane: use the defaults below without asking.
|
||||
- *Custom imagery* — list of `assets` (product photos / app screenshots) shown on the founder's phone. **Default if omitted:** use screenshots from `brand.json.screenshots` when present, otherwise look for obvious screenshots or product images inside the brand kit. If none exist in interactive mode, ask once whether to proceed without imagery or wait for uploads; in the fast lane, proceed without custom imagery and record the assumption.
|
||||
- *Music* — local path / https URL / OR `generate` (instrumental, ~60s). Default: `generate` via MiniMax.
|
||||
- *Lower-third* — optional. Default: off. If enabled, the render is MCP but the overlay onto body video uses one local ffmpeg pass.
|
||||
- *Aspect ratio* — `16:9` (default), `9:16`, `1:1`.
|
||||
@@ -73,14 +114,14 @@ After Stage 0, these are the fields downstream steps consume:
|
||||
|
||||
## State
|
||||
|
||||
Keep a simple `state` object as you work and save every CDN URL there so a partial run can be resumed. Treat `task_status` values `completed` and `done` as successful terminal states, then unwrap `result.structuredContent` when present. The final video lives on Pika's CDN; no local workspace is required unless you trigger the one lower-third alpha-overlay fallback in step [8b].
|
||||
Keep a simple `state` object as you work and save every CDN URL there so a partial run can be resumed. Treat `mcp__pika__task_status` value `completed` as the successful terminal state (`failed` and `cancelled` are the failure terminals), then unwrap `result.structuredContent` when present. The final video lives on Pika's CDN; no local workspace is required unless you trigger the one lower-third alpha-overlay fallback in step [8b].
|
||||
|
||||
## Pipeline overview
|
||||
|
||||
```
|
||||
[Stage 0] intake you (Claude): ask user for url + brand-kit (path or build) + founder (name/role/photo) + optional extras
|
||||
→ [0.5] brand-kit auto-build (only if user said "build")
|
||||
invoke `build-a-brand` and wait for exported brand kit
|
||||
invoke `build-a-brand`; in non-interactive mode use `build-a-brand --quick`
|
||||
→ [1] analyze_brief pika MCP: product name + tagline + features + tone + CTA
|
||||
→ [2] analyze_media × N pika MCP: understand what each user asset shows
|
||||
→ [3] write script you (Claude): 4 acts × 15s; map assets to acts
|
||||
@@ -92,6 +133,7 @@ Keep a simple `state` object as you work and save every CDN URL there so a parti
|
||||
→ [8] captions / lower-third pika MCP: add_captions for subtitles; render lower-third as transparent .mov. If lower-third is enabled, use the alpha-overlay fallback only because MCP has no arbitrary alpha-overlay tool yet.
|
||||
→ [9] render_html_animation pika MCP: 5s end card — author inline HTML, brand-kit fonts inlined, aspect matches body, no corner clutter, CSS @keyframes (NOT GSAP)
|
||||
→ [10] edit_concat + audio_mix pika MCP: concat body + end card, then mix music over the full ~65s
|
||||
→ [10.5] final duration probe pika MCP: analyze final_url and enforce the 55s duration floor before delivery
|
||||
→ [11] final_url save the MCP returned final_url; upload only if a local fallback created the final MP4
|
||||
→ [12] deliver
|
||||
```
|
||||
@@ -103,6 +145,7 @@ Keep the main workflow focused on sequencing. Historical server validation detai
|
||||
- Use a unique `seed` per SeeDance act (101, 202, 303, 404). Identical generation params can replay cached failures.
|
||||
- MiniMax music length is controlled by five `lyrics` sections (`intro` / `verse` / `bridge` / `chorus` / `outro`) with instrumental parentheticals; prompt prose alone is not a reliable length control.
|
||||
- If SeeDance rejects a real-person founder photo, re-roll the founder ref with stronger stylization rather than retrying the same rejected reference.
|
||||
- For local brand-kit logos, upload only logo-appropriate raster assets (`image/png`, `image/jpeg`, or `image/webp`). Do not send SVGs to `mcp__pika__upload_asset`; choose the PNG export from `build-a-brand` or rasterize first.
|
||||
- Use CSS `background-image: url(...)` for CDN-hosted logo/photo assets in end-card HTML; `<img crossorigin>` is blocked by CDN CORS.
|
||||
- Use server-side deterministic tools for captions, concat, and mix. Local ffmpeg is only the fallback for arbitrary transparent lower-third overlay composition, and that fallback is used only when `state.lower_third = true`.
|
||||
- Decompose every 15s act into 3 time-coded sub-shots. Single-shot acts look static.
|
||||
@@ -120,7 +163,7 @@ Save the result as `brief`. You'll reference `brief.product_name`, `brief.taglin
|
||||
|
||||
## [2] Analyze each user-provided asset + derive `product_type`
|
||||
|
||||
For each entry in `assets`, run `analyze_media` to extract content + visual style + **asset type**. Run all in parallel in one tool batch:
|
||||
For each entry in `assets`, run `mcp__pika__analyze_media` to extract content + visual style + **asset type**. Run all in parallel in one tool batch:
|
||||
|
||||
```
|
||||
analyze_media(
|
||||
@@ -487,8 +530,8 @@ The fix in every case is **plain language about what the camera sees**, not film
|
||||
Prepare `character_url` before any SeeDance call:
|
||||
|
||||
- If `founder_photo` is an HTTPS URL, set `founder_photo_url = character_url = founder_photo`.
|
||||
- If `founder_photo` is a local path, upload it with `upload_asset`, then set `founder_photo_url = character_url = public_url`.
|
||||
- If `founder_photo` is `generate`, call `generate_image` and use the returned URL:
|
||||
- If `founder_photo` is a local path, upload it with `mcp__pika__upload_asset`, then set `founder_photo_url = character_url = public_url`.
|
||||
- If `founder_photo` is `generate`, call `mcp__pika__generate_image` and use the returned URL:
|
||||
|
||||
```
|
||||
generate_image(
|
||||
@@ -501,15 +544,15 @@ generate_image(
|
||||
Handle location only when the user supplied a custom location:
|
||||
|
||||
- If `location_image_url` is an HTTPS URL, set `location_url = location_image_url`.
|
||||
- If it is a local path, upload it with `upload_asset` and set `location_url = public_url`.
|
||||
- If it is a text description, generate a custom location reference with `generate_image`.
|
||||
- If it is a local path, upload it with `mcp__pika__upload_asset` and set `location_url = public_url`.
|
||||
- If it is a text description, generate a custom location reference with `mcp__pika__generate_image`.
|
||||
- If no custom location was supplied, do nothing here. Step [4.5] generates the default brand-accent backdrop after `state.brand` exists.
|
||||
|
||||
Save the resulting URLs into `state`. If SeeDance later rejects the founder ref on content policy, see "Known infra quirks" — re-roll with stronger stylization.
|
||||
|
||||
## [4.5] Brand-kit ingestion (always — Stage 0 guarantees `brand_kit_path`)
|
||||
|
||||
`brand_kit_path` is required by Stage 0 — either user-supplied or built first with `build-a-brand`. Parse it once and reuse across the end card and the lower-third. If the folder is missing, ask the user to provide or rebuild the brand kit before continuing.
|
||||
`brand_kit_path` is required by Stage 0 — either user-supplied or built first with `build-a-brand`. Parse it once and reuse across the end card and the lower-third. If the folder is missing in interactive mode, ask the user to provide or rebuild the brand kit before continuing. In the non-interactive fast lane, try the `build-a-brand --quick` branch first; if that cannot produce a kit, stop once with a single compact missing-fields list.
|
||||
|
||||
Preferred source is `brand.json` when present. Otherwise extract from a `build-a-brand` export:
|
||||
- `brand.md` for name, tagline, voice, typography names, and logo descriptions.
|
||||
@@ -521,8 +564,8 @@ Extract into `state.brand`:
|
||||
| `state.brand` field | Source | Notes |
|
||||
|---|---|---|
|
||||
| `name` | `brand.json.name` or `brand.md` quick reference | brand display name |
|
||||
| `wordmark_path` | `logo.wordmark.path` or best `logo/wordmark/*.{svg,png}` | upload local SVG/PNG via `upload_asset`, save `public_url` as `state.brand.wordmark_url` |
|
||||
| `icon_url` | `logo.icon_mark.path` or best `logo/symbol/*.{svg,png}` | upload + capture `state.brand.icon_url` |
|
||||
| `wordmark_path` | `logo.wordmark.path` or best raster `logo/wordmark/*.{png,jpg,jpeg,webp}` | upload local raster asset via `mcp__pika__upload_asset`, save `public_url` as `state.brand.wordmark_url`; do not upload SVG |
|
||||
| `icon_url` | `logo.icon_mark.path` or best raster `logo/symbol/*.{png,jpg,jpeg,webp}` | upload local raster asset via `mcp__pika__upload_asset`, save `public_url` as `state.brand.icon_url`; do not upload SVG |
|
||||
| `colors.primary` | palette role `ink_primary`, `surface_dark`, or `tokens.color.text` | text and border color |
|
||||
| `colors.surface` | palette role `surface_page_bg`, `surface_white`, or `tokens.color.background` | page/background color |
|
||||
| `colors.accent` | CTA/primary brand color from palette or `tokens.color.primary` | end-card CTA pill bg + lower-third accent |
|
||||
@@ -531,9 +574,9 @@ Extract into `state.brand`:
|
||||
| `fonts.text_family` | typography body/text token or `brand.md` | fall back to system sans |
|
||||
| `fonts.mono_family` | typography mono token if present | fall back to Space Mono |
|
||||
|
||||
**Upload step is required** when the brand-kit assets are local files. Without public wordmark/icon URLs, the HTML rendered by `render_html_animation` can't reach them. Use the MCP `upload_asset` flow and save the returned `public_url` values on `state.brand`.
|
||||
**Upload step is required** when the brand-kit assets are local files. Without public wordmark/icon URLs, the HTML rendered by `mcp__pika__render_html_animation` can't reach them. Use the MCP `mcp__pika__upload_asset` flow with raster logo files only and save the returned `public_url` values on `state.brand`. `mcp__pika__upload_asset` rejects `image/svg+xml`; if the best logo is an SVG, pick the sibling PNG export from the brand kit or rasterize the SVG to PNG via `mcp__pika__html_to_png` by inlining the SVG inside an HTML `<svg>` block and using the returned PNG `public_url`.
|
||||
|
||||
**Brand-accent backdrop default location** — when `location_url` was not set by Step [4], render a solid-color PNG via `html_to_png` using `state.brand.colors.accent`. Match the requested video aspect so the reference is not cropped later:
|
||||
**Brand-accent backdrop default location** — when `location_url` was not set by Step [4], render a solid-color PNG via `mcp__pika__html_to_png` using `state.brand.colors.accent`. Match the requested video aspect so the reference is not cropped later:
|
||||
|
||||
| `aspect_ratio` | Backdrop size |
|
||||
|---|---|
|
||||
@@ -680,6 +723,24 @@ Notes:
|
||||
|
||||
Save the 4 returned URLs in submission order as `act_urls = [act1, act2, act3, act4]`.
|
||||
|
||||
### Duration floor and partial-act recovery
|
||||
|
||||
All 4 act_urls are required before step [6]. Do not concat a partial act list.
|
||||
Three completed acts plus the end card produce a ~50s asset, which misses the
|
||||
55s duration floor and must not be reported as a successful founder video.
|
||||
|
||||
If one SeeDance act times out, stalls past the run's wall budget, or reaches a
|
||||
failure terminal while other acts completed:
|
||||
- Retry the missing act once with the same prompt, `reference_images`,
|
||||
`duration`, `sound`, `resolution`, and `aspect_ratio`, but a new seed
|
||||
(`original_seed + 1000`). Do not rerun successful acts.
|
||||
- If the retry completes, insert that URL into the original act slot and
|
||||
continue with `act_urls = [act1, act2, act3, act4]`.
|
||||
- If the retry cannot complete, stop and surface the upstream SeeDance timeout.
|
||||
You may return completed act URLs as a diagnostic preview, but do not deliver
|
||||
a partial concat as `final_url`, do not call it production-ready, and do not
|
||||
proceed to step [6].
|
||||
|
||||
## [6] Stitch acts into 60s base
|
||||
|
||||
```
|
||||
@@ -732,7 +793,7 @@ How it works (load-bearing):
|
||||
|
||||
Save as `music_url`. Read `result.duration_seconds`:
|
||||
- If `>= 50s` → mix it. Expected path with the 5-section structure.
|
||||
- If `< 50s` → re-roll with the same call. After 2 attempts, accept whatever returned — `edit_audio_mix` plays the music for its duration then leaves silence; dialogue carries the rest.
|
||||
- If `< 50s` → re-roll with the same call. After 2 attempts, accept whatever returned — `mcp__pika__edit_audio_mix` plays the music for its duration then leaves silence; dialogue carries the rest.
|
||||
|
||||
**Banned anti-patterns** (each empirically caused a failure):
|
||||
- ❌ Bare `lyrics: "[instrumental]"` — MiniMax sings the literal word "instrumental" for ~10s.
|
||||
@@ -745,24 +806,24 @@ Save as `music_url`. Read `result.duration_seconds`:
|
||||
|
||||
> **Pipeline ordering note** — music mix happens in step [10], AFTER end-card concat. Mixing music into the body before the end card is concatenated leaves the end card silent (the music track ends at the cut). Always: overlays on body → end card → concat → THEN mix music over the full assembled clip.
|
||||
|
||||
Use MCP tools first. `add_captions` handles subtitle timing and burn-in server-side; `render_html_animation` handles authored HTML motion. The only remaining local fallback is arbitrary transparent lower-third overlay, because the current MCP surface has no general alpha-overlay/compose tool and `edit_pip` is not sized for a full-width 800×220 lower-third.
|
||||
Use MCP tools first. `mcp__pika__add_captions` handles subtitle timing and burn-in server-side; `mcp__pika__render_html_animation` handles authored HTML motion. The only remaining local fallback is arbitrary transparent lower-third overlay, because the current MCP surface has no general alpha-overlay/compose tool and `mcp__pika__edit_pip` is not sized for a full-width 800×220 lower-third.
|
||||
|
||||
### Default paths
|
||||
|
||||
| Requested layer | Default action |
|
||||
|---|---|
|
||||
| No lower-third, no subtitles | `body_with_overlays_url = base_url` |
|
||||
| Subtitles only | call `add_captions(video_url: base_url, caption_mode:"auto", style:"classic", position:"bottom", font:"inter")`; save returned `url` as `body_with_overlays_url` |
|
||||
| Lower-third only | render lower-third `.mov` via `render_html_animation`, then use one local ffmpeg overlay pass; upload the result with `upload_asset` and save `body_with_overlays_url` |
|
||||
| Lower-third + subtitles | render and overlay the lower-third first, upload that body checkpoint, then call `add_captions` on the checkpoint URL |
|
||||
| Subtitles only | call `mcp__pika__add_captions(video_url: base_url, caption_mode:"auto", style:"classic", position:"bottom", font:"inter")`; save returned `url` as `body_with_overlays_url` |
|
||||
| Lower-third only | render lower-third `.mov` via `mcp__pika__render_html_animation`, then use one local ffmpeg overlay pass; upload the result with `mcp__pika__upload_asset` and save `body_with_overlays_url` |
|
||||
| Lower-third + subtitles | render and overlay the lower-third first, upload that body checkpoint, then call `mcp__pika__add_captions` on the checkpoint URL |
|
||||
|
||||
If `state.lower_third` is false or unset, skip [8a] and [8b]. This keeps the default path fully MCP-native.
|
||||
|
||||
Do not call local Whisper/caption scripts or chained `edit_text_overlay` for captions. If exact original-script spelling matters, pass manual `subtitles[]` only when you already have exact timed segments from a trusted source; otherwise prefer the `add_captions` auto waterfall.
|
||||
Do not call local Whisper/caption scripts or chained `mcp__pika__edit_text_overlay` for captions. If exact original-script spelling matters, pass manual `subtitles[]` only when you already have exact timed segments from a trusted source; otherwise prefer the `mcp__pika__add_captions` auto waterfall.
|
||||
|
||||
### [8a] Render the lower-third (only if `state.lower_third = true`)
|
||||
|
||||
Skip this sub-step unless `state.lower_third = true`. Render via `render_html_animation` with `format: "mov"` (ProRes 4444 with yuva420p — preserves alpha). **Do NOT use `format: "webm"`** — HyperFrames currently emits webm as VP9 `pix_fmt=yuv420p` with no alpha channel, so "transparent" areas come out as literal black pixels and the composited LT shows a black box outside the pill. Discovered 2026-05-18 on the ScrapeGraphAI v6 run; ffprobe on the returned webm confirmed `pix_fmt=yuv420p` (no alpha). The `.mov` ProRes path is the only reliable alpha path right now.
|
||||
Skip this sub-step unless `state.lower_third = true`. Render via `mcp__pika__render_html_animation` with `format: "mov"` (ProRes 4444 with yuva420p — preserves alpha). **Do NOT use `format: "webm"`** — HyperFrames currently emits webm as VP9 `pix_fmt=yuv420p` with no alpha channel, so "transparent" areas come out as literal black pixels and the composited LT shows a black box outside the pill. Discovered 2026-05-18 on the ScrapeGraphAI v6 run; ffprobe on the returned webm confirmed `pix_fmt=yuv420p` (no alpha). The `.mov` ProRes path is the only reliable alpha path right now.
|
||||
|
||||
- Native dimensions: 800×220 (matches the placement size on a 1280×720 frame, so no scaling artifacts)
|
||||
- Pill: `state.brand.colors.primary` bg (default `#0d0d0d`), `state.brand.colors.highlight` border (default `#fefbcf`), `state.brand.colors.accent` drop shadow (default `#cfc3ff`), 18px border-radius
|
||||
@@ -782,9 +843,9 @@ Fallback contract:
|
||||
- Overlay the 800×220 lower-third at `x=50`, `y=video_height - 220 - 100`, enabled for `t=0..5s`.
|
||||
- Preserve the original body audio without re-encoding so lip-sync stays exact.
|
||||
- Use visually lossless H.264 settings for the local checkpoint.
|
||||
- Upload the checkpoint with `upload_asset` and save the returned `public_url` as `body_with_lower_third_url`.
|
||||
- Upload the checkpoint with `mcp__pika__upload_asset` and save the returned `public_url` as `body_with_lower_third_url`.
|
||||
|
||||
If subtitles are requested too, call `add_captions(video_url: body_with_lower_third_url, ...)` and save its returned `url` as `body_with_overlays_url`. If not, `body_with_overlays_url = body_with_lower_third_url`.
|
||||
If subtitles are requested too, call `mcp__pika__add_captions(video_url: body_with_lower_third_url, ...)` and save its returned `url` as `body_with_overlays_url`. If not, `body_with_overlays_url = body_with_lower_third_url`.
|
||||
|
||||
### [8c] Captions via MCP
|
||||
|
||||
@@ -808,7 +869,7 @@ Save returned `url` as `body_with_overlays_url`. The returned `transcript` is us
|
||||
|
||||
## [9] Animated end card (5s) — author inline HTML, render via HyperFrames
|
||||
|
||||
We do NOT use `generate_slide_animation` here. That tool delegates HTML authoring to a slide-card LLM, which routinely adds corner clutter (top-left wordmarks, bottom-right URLs), picks wrong aspects, and produces animations that don't reliably play through HyperFrames' per-frame seek. Instead, the orchestrator authors the end-card HTML directly and renders it via `render_html_animation`. Same engine the lower-third uses.
|
||||
We do NOT use `mcp__pika__generate_slide_animation` here. That tool delegates HTML authoring to a slide-card LLM, which routinely adds corner clutter (top-left wordmarks, bottom-right URLs), picks wrong aspects, and produces animations that don't reliably play through HyperFrames' per-frame seek. Instead, the orchestrator authors the end-card HTML directly and renders it via `mcp__pika__render_html_animation`. Same engine the lower-third uses.
|
||||
|
||||
### Hard rules — empirically verified, do not deviate
|
||||
|
||||
@@ -832,9 +893,17 @@ These are NOT stylistic preferences. Each was discovered by rendering, extractin
|
||||
|
||||
9. **Don't write CSS `font-family` fallback chains for brand designs.** If the brand font fails to load, a fallback chain hides the failure — you ship Helvetica thinking it's Telka. Use `font-family: "telka-700"` alone (no fallback). Then a font load failure renders Chrome's default serif, which is visually obvious and triggers a fix.
|
||||
|
||||
10. **Composition contract** — the HyperFrames contract: `<div id="stage" data-composition-id="main" data-start="0" data-duration="5" data-width="W" data-height="H">` wraps a SINGLE direct child `<div id="card" data-start="0" data-duration="5" data-track-index="0">` which contains everything else. Multi-tracked direct children of `#stage` interact poorly with frame seeking. (Fixed by mirroring the working lower-third structure.)
|
||||
10. **Composition contract** — the HyperFrames contract: `<div id="stage" data-composition-id="main" data-start="0" data-duration="5" data-width="W" data-height="H">` wraps a SINGLE direct child `<div id="card" class="clip" data-start="0" data-duration="5" data-track-index="0">` which contains everything else. Visible timed elements must include `class="clip"` because HyperFrames uses it for visibility control, and the clip must be nested inside the composition root, not a sibling. Multi-tracked direct children of `#stage` interact poorly with frame seeking. (Fixed by mirroring the working lower-third structure.)
|
||||
|
||||
11. **Always extract frames at t=0, t=1s, t=2s after rendering and visually compare.** If frames 0 and 2 look identical, the entrance animation isn't running. If the tagline looks like a serif, the brand font didn't load. Don't trust the URL alone. Don't ship without this check. (User caught these failures three renders in a row before frame extraction was added.)
|
||||
11. **Runtime readiness hook** — include a small compatibility hook before `</body>`:
|
||||
`window.__hf = { duration: 5, seek: (t) => { document.documentElement.style.setProperty("--hf-time", String(t)); } };`.
|
||||
CSS `@keyframes` still drive the visual animation, but the hook makes the
|
||||
prod frame-capture path ready when it probes for `window.__hf`. If the
|
||||
worker reports `window.__hf not ready after 45000ms`, treat the HTML as
|
||||
invalid for `render_html_animation`; fix the composition contract or hook
|
||||
and rerender. Do not fall back to a static PNG.
|
||||
|
||||
12. **Always extract frames at t=0, t=1s, t=2s after rendering and visually compare.** If frames 0 and 2 look identical, the entrance animation isn't running. If the tagline looks like a serif, the brand font didn't load. Don't trust the URL alone. Don't ship without this check. (User caught these failures three renders in a row before frame extraction was added.)
|
||||
|
||||
### Build steps
|
||||
|
||||
@@ -866,7 +935,7 @@ else:
|
||||
# tagline, CTA, palette values, and dimensions W/H.
|
||||
|
||||
# 5. Render
|
||||
end_card_url = render_html_animation(html=filled, fps=30, quality="standard")
|
||||
end_card_url = render_html_animation(html=filled, fps=30, quality="standard", format="mp4")
|
||||
```
|
||||
|
||||
### Layout (centered stack — no corners)
|
||||
@@ -904,15 +973,15 @@ All implemented as CSS `animation: name duration easing delay forwards` on the c
|
||||
| 2.80–4.80 | accent-top | `opacity:1 → 0.55 → 1`, alternate (gentle shimmer) |
|
||||
| 4.80–5.00 | hold | (final readable state) |
|
||||
|
||||
Reference implementation pattern: the v6 centered-stack HTML recipe verified on 2026-05-02. Author the filled HTML inline in the orchestrator and pass it directly to `render_html_animation`; do not call a bundled helper script or rely on a separate `presets/` directory.
|
||||
Reference implementation pattern: the v6 centered-stack HTML recipe verified on 2026-05-02. Author the filled HTML inline in the orchestrator and pass it directly to `mcp__pika__render_html_animation`; do not call a bundled helper script or rely on a separate `presets/` directory.
|
||||
|
||||
If the brand-kit lacks fonts (`state.brand.fonts` is null) — fall back to system `-apple-system, sans-serif` for tagline/CTA but DROP the title down to a system-display weight. Don't render brand-typography end cards with fallback fonts; they always look wrong. Flag this in the deliver step so the user knows the brand-kit is incomplete.
|
||||
|
||||
Save the returned MP4 URL as `end_card_url`. Download it only if you need local visual QA frames or a local fallback assembly.
|
||||
Save the returned MP4 URL as `end_card_url`. `end_card_url` must be an MP4 video segment, not a static PNG, because step [10] concatenates it with the body video. Download it only if you need local visual QA frames or a local fallback assembly.
|
||||
|
||||
## [10] Assemble body + end card + music via MCP
|
||||
|
||||
Use the server-side deterministic edit tools for final assembly. The current MCP server `edit_concat` normalizes mismatched inputs before concat, and `edit_audio_mix` preserves the original video audio while mixing the music track.
|
||||
Use the server-side deterministic edit tools for final assembly. The current MCP server `mcp__pika__edit_concat` normalizes mismatched inputs before concat, and `mcp__pika__edit_audio_mix` preserves the original video audio while mixing the music track.
|
||||
|
||||
```
|
||||
assembled = edit_concat(video_urls=[body_with_overlays_url, end_card_url])
|
||||
@@ -925,11 +994,30 @@ else:
|
||||
final_url = assembled_url
|
||||
```
|
||||
|
||||
Mix music after concat, never before, so the score continues through the end card. If `edit_audio_mix` fails because the music file is too short or malformed, deliver `assembled_url` and surface the music issue; do not rerun expensive SeeDance acts.
|
||||
Mix music after concat, never before, so the score continues through the end card. If `mcp__pika__edit_audio_mix` fails because the music file is too short or malformed, set `final_url = assembled_url`, surface the music issue, and still run step [10.5] before delivery; do not rerun expensive SeeDance acts.
|
||||
|
||||
### [10.5] Final duration floor
|
||||
|
||||
Before reporting `final_url` to the user, probe the assembled result:
|
||||
|
||||
```
|
||||
mcp__pika__analyze_media(
|
||||
media: final_url,
|
||||
query: "Return JSON with duration_seconds for this video."
|
||||
)
|
||||
```
|
||||
|
||||
Save the result as `final_duration_seconds`. It must be `>= 55` and `<= 75`
|
||||
before reporting `final_url` as the completed deliverable.
|
||||
If `final_duration_seconds` is under 55s, treat the run as a failed partial
|
||||
assembly: do not deliver the URL as final, do not mark the skill complete, and
|
||||
return to the missing-act recovery above. If all 4 acts were present but the
|
||||
probe is still under 55s, stop and surface the concat/provider truncation for
|
||||
investigation instead of padding with unrelated footage.
|
||||
|
||||
## [11] Final URL
|
||||
|
||||
Save `final_url` into `state`. If a local fallback assembly produced the final MP4, upload that checkpoint through `upload_asset` and replace `final_url` with the returned `public_url`.
|
||||
Save `final_url` into `state`. If a local fallback assembly produced the final MP4, upload that checkpoint through `mcp__pika__upload_asset` and replace `final_url` with the returned `public_url`.
|
||||
|
||||
## [12] Asset bundle (optional)
|
||||
|
||||
@@ -953,7 +1041,7 @@ Why this matters:
|
||||
|
||||
Report `final_url` to the user. Include:
|
||||
- Asset bundle URL/path only if the user asked for one
|
||||
- Total duration (~65s = 60s body + 5s end card)
|
||||
- Total duration from `final_duration_seconds` (~65s = 60s body + 5s end card)
|
||||
- Intermediate URLs or local files useful for reruns: `base_url`, optional `body_with_overlays_url`, `music_url`, `end_card_url`, `final_url`, and each `act_urls[i]`
|
||||
- The brief (`brief.product_name` / `brief.tagline`) so the user can confirm the model picked up the right product
|
||||
- A 1-line summary of which asset went into which act, so the user can confirm placement
|
||||
@@ -966,32 +1054,35 @@ Report `final_url` to the user. Include:
|
||||
| Asset analyses | One JSON object per asset, each with non-empty `content_description` and `best_for_act` |
|
||||
| Script | 4 acts; total dialogue 100–180 words; each asset is referenced by at least one act, OR explicitly noted as unused |
|
||||
| Refs | `character_url` and `location_url` are https URLs |
|
||||
| Acts | All 4 act_urls returned (each with a unique seed); for acts containing shots with non-null `asset_index`, vision-check ONE such act with `analyze_media` to confirm the asset is actually visible AND the reveal pattern is correct for `product_type` (e.g. for `physical_apparel`, verify the founder is HOLDING/WEARING the actual t-shirt with the right print — NOT showing it on a phone screen) |
|
||||
| Acts | All 4 act_urls returned (each with a unique seed); for acts containing shots with non-null `asset_index`, vision-check ONE such act with `mcp__pika__analyze_media` to confirm the asset is actually visible AND the reveal pattern is correct for `product_type` (e.g. for `physical_apparel`, verify the founder is HOLDING/WEARING the actual t-shirt with the right print — NOT showing it on a phone screen) |
|
||||
| Stitch | `base_url` returned |
|
||||
| Music | URL returned; `duration_seconds >= 50` (5-section `lyrics` structure should produce 60–80s; retry once if under 50, accept after 2) |
|
||||
| Overlays | `body_with_overlays_url` returned (or `body_with_overlays_url = base_url` if step [8] skipped) |
|
||||
| End card | `end_card_url` returned; inspect early frames with `extract_frame`/`analyze_media`: frame 0 should show the empty background before entrance, and a later first-second frame should show partial entrance — confirms CSS @keyframes are firing, not static |
|
||||
| Final assembly | `assembled_url` returned by `edit_concat`; `final_url` returned by `edit_audio_mix` when music is present, otherwise `final_url = assembled_url` |
|
||||
| End card | `end_card_url` returned as an MP4; inspect early frames with `mcp__pika__extract_frame`/`mcp__pika__analyze_media`: frame 0 should show the empty background before entrance, and a later first-second frame should show partial entrance — confirms CSS @keyframes are firing, not static |
|
||||
| Final assembly | `assembled_url` returned by `mcp__pika__edit_concat`; `final_url` returned by `mcp__pika__edit_audio_mix` when music is present, otherwise `final_url = assembled_url` |
|
||||
| Final duration | `mcp__pika__analyze_media` reports `final_duration_seconds >= 55` and `<= 75` before delivery |
|
||||
| Local fallback upload | Only if local fallback assembly was used: `final_url` replaced with an uploaded `public_url` |
|
||||
|
||||
## Failure modes
|
||||
|
||||
Stop and surface on first verification failure. Don't auto-retry expensive calls (SeeDance acts run 3–8 min each — repeated failures burn credits).
|
||||
Except for the one missing-act retry documented in "Duration floor and partial-act recovery", stop and surface on first verification failure. Don't auto-retry expensive calls (SeeDance acts run 3–8 min each — repeated failures burn credits).
|
||||
|
||||
| Symptom | Cause | Fix |
|
||||
|---|---|---|
|
||||
| `generate_reference_video` returns 402 "insufficient balance" with a familiar UUID | Idempotency cache replaying an old failed result for identical params | Pass a unique `seed` per call (101 / 202 / 303 / 404 for 4 acts) so the hash differs. |
|
||||
| `mcp__pika__generate_reference_video` returns 402 "insufficient balance" with a familiar UUID | Idempotency cache replaying an old failed result for identical params | Pass a unique `seed` per call (101 / 202 / 303 / 404 for 4 acts) so the hash differs. |
|
||||
| SeeDance returns 422 "may contain likenesses of real people" on founder ref | Content-policy filter tripped (intermittent — same photo may pass next attempt) | Re-roll the founder portrait with stronger stylization ("Pixar / Disney 3D animation aesthetic"). Don't auto-retry the same ref — burns credits. |
|
||||
| Founder shirt changes between acts | @Image1 read fresh each act, no wardrobe lock | Add a `WARDROBE LOCK:` line to every act prompt, identical sentence verbatim. |
|
||||
| All 4 acts have the same physical backdrop | Opening line says "inside the location matching @Image2" — read literally | Open with "in a setting whose visual style, palette, lighting and materials match @Image2" + add per-shot `Background context:` line. |
|
||||
| Founder looks frozen / no body language | Beats are facial-only; act is single-shot | Add 3 time-coded sub-shots per act with explicit camera-motion `Transition:` lines; every beat needs a hand/torso/head action (see [3c.1]). |
|
||||
| Final video is under 55s | One SeeDance act timed out or final concat trimmed the body, producing a partial run | Do not deliver it as final. Retry the missing act once with a new seed while preserving successful acts; if that cannot complete, stop and surface the upstream SeeDance timeout. |
|
||||
| Music returns < 50s | MiniMax non-determinism, or `lyrics` field omitted | Confirm `lyrics` has 5 `[section]` tags with `(instrumental — …)` parentheticals. Retry once; accept after 2 attempts. |
|
||||
| Music sings the word "instrumental" | `lyrics: "[instrumental]"` bare tag | Use the 5-section structure with parenthetical cues (see step [7]). |
|
||||
| Captions misspell product names | Auto transcription normalized the spoken audio | Use manual `subtitles[]` only if you already have trusted timestamped segments; otherwise surface the transcript limitation instead of running local Whisper by default. |
|
||||
| Lower-third overlay shows a black box outside the pill | webm format encoded without alpha (yuv420p) | Re-render with `format: "mov"` (ProRes 4444 yuva). Verify with `ffprobe \| grep pix_fmt` showing `yuva*`. |
|
||||
| Final video audio shorter than video | Local fallback concat used `-c copy` with mismatched audio params | Prefer MCP `edit_concat`. If local concat is unavoidable, normalize all inputs to aac/44100/stereo/192k before concat. |
|
||||
| Final video audio shorter than video | Local fallback concat used `-c copy` with mismatched audio params | Prefer MCP `mcp__pika__edit_concat`. If local concat is unavoidable, normalize all inputs to aac/44100/stereo/192k before concat. |
|
||||
| End card renders identical at t=0 and t=2s (no entrance animation) | GSAP `tl.from()` used instead of CSS `@keyframes` | Convert entrance to CSS `@keyframes`; keep the GSAP shim only for duration seeking. |
|
||||
| End-card tagline renders as serif fallback | woff2 font failed to load in HyperFrames Chrome | Use unique `font-family` names per face (not weight-matching); subset + base64-inline the woff2; verify by extracting frame 30 before shipping. |
|
||||
| `render_html_animation` fails with `window.__hf not ready after 45000ms` | End-card HTML did not expose the runtime readiness hook or valid nested `class="clip"` composition | Add/fix the `window.__hf` hook and `class="clip"` child, then rerender the MP4. Do not fall back to a static PNG and do not proceed to concat until `end_card_url` is a video URL. |
|
||||
|
||||
## Load-bearing phrases
|
||||
|
||||
@@ -1013,9 +1104,9 @@ These strings go into the SeeDance prompt (or HTML render) verbatim. Each was em
|
||||
- **Don't describe the character in prompt prose** — @Image1 carries identity. Prose conflicts produce phantom figures or wrong outfits. Exception: the `WARDROBE LOCK:` line.
|
||||
- **Don't use film-industry shot terms** — "Two-shot" / "Three-shot" / "Over-shoulder" / "OTS" / "Master shot" trigger SeeDance phantom-subject artifacts. Describe what the camera sees in plain language.
|
||||
- **Don't render the lower-third as webm** — alpha not preserved (HyperFrames emits yuv420p). Use `format: "mov"` (ProRes 4444 yuva). Verify with `ffprobe \| grep pix_fmt`.
|
||||
- **Don't use `generate_slide_animation` for the end card** — that tool's slide-card LLM adds corner clutter and produces animations that don't seek deterministically. Author inline HTML and render via `render_html_animation`.
|
||||
- **Don't chain pika MCP `edit_text_overlay` / overlay calls for pixel composition** — that cascades quality loss and can introduce lip-sync drift. Use `add_captions` for captions and the single local ffmpeg lower-third fallback only when the lower-third is enabled.
|
||||
- **Don't use local `ffmpeg concat -c copy` for final assembly unless MCP is unavailable** — the old audio-drop bug was in local concat behavior. Default to MCP `edit_concat` + `edit_audio_mix`.
|
||||
- **Don't use `mcp__pika__generate_slide_animation` for the end card** — that tool's slide-card LLM adds corner clutter and produces animations that don't seek deterministically. Author inline HTML and render via `mcp__pika__render_html_animation`.
|
||||
- **Don't chain pika MCP `mcp__pika__edit_text_overlay` / overlay calls for pixel composition** — that cascades quality loss and can introduce lip-sync drift. Use `mcp__pika__add_captions` for captions and the single local ffmpeg lower-third fallback only when the lower-third is enabled.
|
||||
- **Don't use local `ffmpeg concat -c copy` for final assembly unless MCP is unavailable** — the old audio-drop bug was in local concat behavior. Default to MCP `mcp__pika__edit_concat` + `mcp__pika__edit_audio_mix`.
|
||||
- **Don't fire SeeDance with identical params across acts** — the MCP idempotency cache hashes to the same task ID and replays old results (sometimes failures). Pass unique `seed` per act.
|
||||
- **Don't omit the `lyrics` field on MiniMax music** — output drops to ~17–30s. Don't put bare `[instrumental]` either — model sings the word. Use the 5-section structure with parenthetical cues.
|
||||
- **Don't copy the example music sound for every brand** — the recipe is the pattern (5 sections + parentheticals), not the specific instrumentation. Pick a register that matches `brief.tone` (see step [7] table).
|
||||
@@ -1023,9 +1114,9 @@ These strings go into the SeeDance prompt (or HTML render) verbatim. Each was em
|
||||
|
||||
## Engine choice: seedance-only (with caveats)
|
||||
|
||||
SeeDance (`fal-seedance-2-i2v` via `generate_reference_video` `provider: "seedance"`) is the sole video engine. Picked over alternatives after testing:
|
||||
SeeDance (`fal-seedance-2-i2v` via `mcp__pika__generate_reference_video` `provider: "seedance"`) is the sole video engine. Picked over alternatives after testing:
|
||||
|
||||
- **vs Kling v3-omni**: Kling has a true `shots[]` hard-cut array (cleaner multi-shot) but rejects the `seed` parameter (cache-busting harder), and 4 × pro 1080p outputs sum >50MB and exceed the `edit_concat` upload cap (forces local concat). Kling does have more permissive content policy for real-person photos — it's a worth keeping in mind as a fallback if SeeDance's intermittent 422 becomes a hard block.
|
||||
- **vs Kling v3-omni**: Kling has a true `shots[]` hard-cut array (cleaner multi-shot) but rejects the `seed` parameter (cache-busting harder), and 4 × pro 1080p outputs sum >50MB and exceed the `mcp__pika__edit_concat` upload cap (forces local concat). Kling does have more permissive content policy for real-person photos — it's a worth keeping in mind as a fallback if SeeDance's intermittent 422 becomes a hard block.
|
||||
- **vs Happy Horse `happyhorse-1.0-r2v`** (Alibaba DashScope): produced clean 1080p with native lip-sync but the multi-shot prompt direction was weaker. Validated 2026-05-18 (v5) but visibly less cinematic than SeeDance v6/v7.
|
||||
- **SeeDance wins because**: native `<<<voice_1>>>` lip-sync, accepts `seed` (cache-busting), permissive enough on real-person photos that 95%+ runs pass content filter, single 15s prompt with time-coded sub-shots gives enough variation for a talking-head register.
|
||||
|
||||
@@ -1067,9 +1158,9 @@ Wall-clock budget per step. Total run is ~12–18 minutes, dominated by the para
|
||||
- `physical_object` → founder holds the product up; assets passed to all shots where product is visible
|
||||
- `consumable` → founder uses/eats/drinks; same pattern
|
||||
- `service` → no asset reveal; environment + dialogue only
|
||||
- Music: target ~60–80s instrumental — pass `lyrics` with 5 `[section]` tags (`intro` / `verse` / `bridge` / `chorus` / `outro`), each containing a `(instrumental — …)` parenthetical production cue. Section count drives length; the parenthetical guarantees no vocals. See step [7] for the canonical call. Retry once if under 50s, then accept what you got — `edit_audio_mix` plays the music for its duration and leaves silence beyond.
|
||||
- 5s end card via `render_html_animation` — author inline HTML per step [9], inline brand-kit fonts as base64. Sources brand from `state.brand` (set in step [4.5]) → real logo, real palette, real fonts.
|
||||
- **Captions via `add_captions`.** Use server-side word-level caption burn-in by default. Font choices are the tool-supported set (`inter`, `bebas-neue`, `noto-cjk`); use brand accent colors for highlight/outline instead of local custom font drawtext.
|
||||
- **Lower-third fallback.** Off by default. If `state.lower_third = true`, render a 5s branded pill bottom-left via `render_html_animation(format:"mov")`; the final overlay onto the body uses one local ffmpeg pass only until MCP exposes a general alpha-overlay/compose tool.
|
||||
- **Final assembly via MCP.** Use `edit_concat` for body + end card, then `edit_audio_mix` for music. Local concat/mix is a fallback, not the canonical path.
|
||||
- Music: target ~60–80s instrumental — pass `lyrics` with 5 `[section]` tags (`intro` / `verse` / `bridge` / `chorus` / `outro`), each containing a `(instrumental — …)` parenthetical production cue. Section count drives length; the parenthetical guarantees no vocals. See step [7] for the canonical call. Retry once if under 50s, then accept what you got — `mcp__pika__edit_audio_mix` plays the music for its duration and leaves silence beyond.
|
||||
- 5s end card via `mcp__pika__render_html_animation` — author inline HTML per step [9], inline brand-kit fonts as base64. Sources brand from `state.brand` (set in step [4.5]) → real logo, real palette, real fonts.
|
||||
- **Captions via `mcp__pika__add_captions`.** Use server-side word-level caption burn-in by default. Font choices are the tool-supported set (`inter`, `bebas-neue`, `noto-cjk`); use brand accent colors for highlight/outline instead of local custom font drawtext.
|
||||
- **Lower-third fallback.** Off by default. If `state.lower_third = true`, render a 5s branded pill bottom-left via `mcp__pika__render_html_animation(format:"mov")`; the final overlay onto the body uses one local ffmpeg pass only until MCP exposes a general alpha-overlay/compose tool.
|
||||
- **Final assembly via MCP.** Use `mcp__pika__edit_concat` for body + end card, then `mcp__pika__edit_audio_mix` for music. Local concat/mix is a fallback, not the canonical path.
|
||||
- Provider: `seedance` only. Reference tokens are `@Image1` / `@Image2` / `@Image3`. Native lip-sync via `<<<voice_1>>>...<<<voice_1>>>` tokens per sub-shot. Real-person founder photos pass the content filter the vast majority of the time; intermittent 422 → re-roll with stronger stylization.
|
||||
|
||||
@@ -9,6 +9,7 @@ Operational notes moved out of `SKILL.md` so the main skill stays workflow-shape
|
||||
| SeeDance idempotency | Calls with identical prompt, refs, duration, and sound can replay a cached task. Use unique per-act seeds. |
|
||||
| SeeDance moderation | Real-person references can intermittently trip the likeness filter. Re-roll the founder ref with stronger stylization before retrying. |
|
||||
| MiniMax music length | Five `lyrics` sections drive 60-80s instrumental length more reliably than prompt prose. Avoid bare `[instrumental]`, which may be sung literally. |
|
||||
| Brand logo uploads | `mcp__pika__upload_asset` rejects SVG (`image/svg+xml`). Use PNG/JPG/WebP logo exports from `build-a-brand`, or rasterize SVGs to PNG with `mcp__pika__html_to_png` before uploading. |
|
||||
| CDN assets in HTML | Pika CDN does not send permissive CORS headers. Use CSS `background-image` for remote logos/photos in rendered HTML. |
|
||||
| Captions | Prefer `add_captions` word-level timing over local transcription for default subtitle burn-in. Use manual subtitles only when exact timestamps already exist. |
|
||||
| Final assembly | Prefer MCP `edit_concat` + `edit_audio_mix`; local concat/mix is only a fallback when the tool surface is unavailable. |
|
||||
|
||||
@@ -42,7 +42,7 @@ Confirm back in one line ("Generating a Madison Square Garden Kiss Cam moment fo
|
||||
|
||||
The kiss cam graphic + scoreboard + retro frame get baked into the still at frame 0 — load-bearing, so Kling treats the entire decorative UI as pixel-locked burned-in UI in Step 2 instead of animating it mid-clip.
|
||||
|
||||
**Why gpt-image-2 (and no fallback):** sharper LED panel detail (scoreboard numerals, kiss cam typography, retro decorative edges) and stronger reference-likeness lock than alternative providers; the LED-sharpness + likeness combo is what sells the trend. On a `moderation_blocked` response, re-roll the same call instead of swapping providers — alternatives produced softer likeness and softer LED detail in earlier trials. Trade-off: gpt-image-2's 16:9 native is 1792×1024 (1K-class grid), which is sufficient since Kling pro outputs 1080p downstream.
|
||||
**Why gpt-image-2 (and no fallback):** sharper LED panel detail (scoreboard numerals, kiss cam typography, retro decorative edges) and stronger reference-likeness lock than alternative providers; the LED-sharpness + likeness combo is what sells the trend. On a `moderation_blocked` response, re-roll the same call instead of swapping providers — alternatives produced softer likeness and softer LED detail in earlier trials. We call gpt-image-2 at 1K 16:9 (1792×1024); the 2K/4K variants (AGNT-336) don't help here since Kling pro outputs 1080p downstream.
|
||||
|
||||
**Why a Jumbotron-POV phone shot (and not a TV broadcast overlay):** the first iteration produced a TV broadcast cutaway with a pink-heart kiss cam graphic on the feed — user feedback was "the kiss cam graphics is ugly, look how real kiss cam moments look in real videos." Real viral kiss-cam clips online are virtually all spectator phone shots OF the Jumbotron (Obama-era USA Basketball kiss cam, Sarah Hyland / Wells Adams kiss cam, etc.). The Jumbotron-shot framing hits the aesthetic users actually associate with "real kiss cam" — retro red border + sparkly hearts + cursive Kiss Cam script + adjacent LED scoreboard panels + arena darkness + fans filming with phones.
|
||||
|
||||
|
||||
@@ -176,7 +176,7 @@ Call `generate_reference_video`:
|
||||
|
||||
For `provider=kling`: convert the multi-beat prose into `shots: [{prompt, duration}, ...]` (5 shots × 3s = 15s sum), plus a top-level `prompt` summarizing the ad. References use `<<<image_1>>>` / `<<<image_2>>>` instead of `@Image1` / `@Image2`.
|
||||
|
||||
If generation completes asynchronously, follow the MCP tool's returned status handle until it reaches a terminal state. Capture the result URL → `video_url` and proceed to step 8.
|
||||
If the call returns `{ task_id, status: "queued" }`, poll `task_status(task_id)` in a tight loop (no Bash, no sleep) until terminal (`completed | failed | cancelled`). On `completed`, capture `result.url` → `video_url` and proceed to step 8.
|
||||
|
||||
**7b. On rejection — auto-cartoonize the avatar**
|
||||
|
||||
|
||||
Reference in New Issue
Block a user