Files
OpenMontage/docs/comfyui-adapter-plan.md

474 lines
18 KiB
Markdown
Raw Normal View History

# ComfyUI Provider Adapter for OpenMontage
**RFC: Native ComfyUI backend for image and video generation**
---
## Motivation
OpenMontage's local GPU tools (`wan_video`, `hunyuan_video`, `cogvideo_video`,
`local_diffusion`) use HuggingFace `diffusers` directly. This works on x86 +
consumer GPUs but breaks on newer hardware where the PyTorch ecosystem hasn't
caught up:
| Issue | Detail |
|-------|--------|
| **NVIDIA Blackwell (sm_121)** | No stable PyTorch wheels for aarch64 + CUDA 13.0. Requires NGC containers or nightly builds. |
| **Flash Attention** | Does not support sm_121. Must be replaced with SageAttention v3 or native SDPA. |
| **Unified Memory (GB10/DGX Spark)** | `nvidia-smi` cannot report VRAM. Diffusers' memory estimation breaks. |
| **Model format mismatch** | Diffusers expects HF repos. Production deployments use `.safetensors` checkpoints with quantized variants (NVFP4, FP8) that diffusers doesn't natively load. |
ComfyUI already solves all of these. NVIDIA ships official ComfyUI containers
for DGX Spark. The community has optimized workflows for Blackwell (SageAttention,
NVFP4 quantization, LightX2V 4-step LoRAs). Models like WAN 2.2, FLUX 2,
and ACE-Step run reliably through ComfyUI on hardware where diffusers cannot.
A ComfyUI adapter gives OpenMontage access to any model ComfyUI supports,
on any hardware ComfyUI runs on, without shipping or maintaining PyTorch builds.
---
## Design
### Architecture
```
OpenMontage Agent
|
v
video_selector / image_selector
|
v
comfyui_video comfyui_image (new tools)
| |
v v
ComfyUI REST API (POST /prompt, GET /history, GET /view)
|
v
GPU (any hardware ComfyUI supports)
```
### Integration model
Two new `BaseTool` subclasses plus one shared client library:
```
tools/
_comfyui/
__init__.py
client.py # Shared ComfyUI REST client
workflows/ # Bundled workflow templates
flux2-txt2img.json
wan22-t2v-4step.json
wan22-i2v-4step.json
graphics/
comfyui_image.py # capability="image_generation", provider="comfyui"
video/
comfyui_video.py # capability="video_generation", provider="comfyui"
```
### Registry and selector integration
The tools declare `capability` and `provider` as class attributes.
`tool_registry.discover()` picks them up automatically via `pkgutil.walk_packages`.
`video_selector` and `image_selector` find them via `registry.get_by_capability()`.
The only selector change is operation-specific filtering in `video_selector` so
ComfyUI is not selected for `image_to_video` when only the text-to-video bundled
models are installed, or vice versa.
---
## Shared Client: `tools/_comfyui/client.py`
Encapsulates the ComfyUI REST API pattern proven in production (used by the
Bard project's Airflow DAGs for thousands of generations):
The endpoint contract was checked against current ComfyUI server documentation
and the April 2026 third-party developer guide:
- Official routes: `POST /prompt`, `GET /history/{prompt_id}`, `GET /view`,
`POST /upload/image`, `GET /object_info/{node_class}`, `GET /models/{folder}`,
`GET /system_stats`, and `WS /ws` are documented server routes.
- `/prompt` accepts the workflow in API format under the `prompt` key and
returns `prompt_id`, `number`, and `node_errors` on validation.
- `/history/{prompt_id}` returns completed node outputs; artifact records include
`filename`, `subfolder`, and `type`. The client passes all three through to
`/view` instead of assuming `type=output`.
- Workflows must be exported in ComfyUI API format, not the regular visual
canvas workflow format.
References:
- https://docs.comfy.org/development/comfyui-server/comms_routes
- https://www.runflow.io/blog/comfyui-api-developer-guide
```python
class ComfyUIClient:
"""Thin client for the ComfyUI REST API."""
def __init__(self, server_url: str | None = None):
self.server_url = server_url or os.environ.get(
"COMFYUI_SERVER_URL", "http://localhost:8188"
)
def is_available(self) -> bool:
"""Health check -- can we reach the server?"""
def submit(self, workflow: dict) -> str:
"""POST /prompt. Returns prompt_id. Raises on node_errors."""
def poll(self, prompt_id: str, timeout: int = 600, interval: int = 5) -> dict:
"""GET /history/{prompt_id} until complete. Returns outputs dict."""
def download(self, filename: str, subfolder: str, dest: Path) -> Path:
"""GET /view?filename=...&type=output. Writes bytes to dest."""
def upload_image(self, local_path: Path, name: str) -> str:
"""POST /upload/image. Returns server-side filename for LoadImage nodes."""
def generate(self, workflow: dict, output_node: str, dest: Path,
timeout: int = 600) -> Path:
"""Full cycle: submit -> poll -> download. Returns artifact path."""
```
**Why a shared client?** The submit/poll/download cycle is identical across
image and video generation. The only differences are: which workflow template,
which nodes to customize, and which output node to read from.
---
## Tool Specifications
### `comfyui_image` -- Image Generation
| Field | Value |
|-------|-------|
| capability | `image_generation` |
| provider | `comfyui` |
| runtime | `LOCAL_GPU` |
| tier | `GENERATE` |
| stability | `EXPERIMENTAL` |
| capabilities | `text_to_image`, `image_to_image` |
| dependencies | (runtime: ComfyUI server reachable) |
| fallback_tools | `flux_image`, `local_diffusion`, `openai_image` |
| cost | `$0.00` (local compute) |
**Bundled workflow:** `flux2-txt2img.json`
Loads FLUX 2 Dev (NVFP4) with Mistral text encoder. Templated nodes:
| Node | Class | Templated field |
|------|-------|-----------------|
| 4 | CLIPTextEncode | `text` (prompt) |
| 6 | EmptyFlux2LatentImage | `width`, `height` |
| 7 | RandomNoise | `noise_seed` |
| 10 | Flux2Scheduler | `steps` |
| 13 | SaveImage | `filename_prefix` |
**Input schema:**
```yaml
prompt: string # required
width: integer # default 1024
height: integer # default 1024
steps: integer # default 20
seed: integer # optional (random if omitted)
guidance: number # default 3.5
output_path: string # where to save the image
workflow_json: string # optional custom workflow; requires output_node
workflow_path: string # optional path to workflow JSON; requires output_node
output_node: string # required for custom workflows
workflow_name: string # optional custom workflow provenance label
workflow_model: string # optional custom model/provenance label
workflow_model_stack: [] # optional custom dependency provenance
```
**get_status():** Pings ComfyUI server and checks bundled FLUX model names via
`/object_info`. Returns `AVAILABLE` when the server and bundled model set are
ready, `DEGRADED` when the server is reachable but bundled models are missing,
and `UNAVAILABLE` when the server cannot be reached.
**execute() flow:**
1. Deep-copy workflow template
2. Inject prompt, seed, dimensions, steps into templated nodes
3. `client.generate(workflow, output_node="13", dest=output_path)`
4. Return `ToolResult` with artifact path, seed, model info
For custom workflows, the caller must provide `workflow_json` or `workflow_path`
plus `output_node`. The tool does not assume bundled node IDs for custom
workflows, and provenance is reported as user-supplied unless the caller provides
`workflow_model`. Results also include the final workflow SHA-256 hash and, for
bundled workflows, the known model stack.
---
### `comfyui_video` -- Video Generation
| Field | Value |
|-------|-------|
| capability | `video_generation` |
| provider | `comfyui` |
| runtime | `LOCAL_GPU` |
| tier | `GENERATE` |
| stability | `EXPERIMENTAL` |
| capabilities | `text_to_video`, `image_to_video` |
| dependencies | (runtime: ComfyUI server reachable) |
| fallback_tools | `wan_video`, `hunyuan_video`, `ltx_video_local` |
| cost | `$0.00` (local compute) |
**Bundled workflows:**
1. **`wan22-i2v-4step.json`** -- Image-to-video (WAN 2.2 14B, fp8, 4-step LightX2V LoRA)
2. **`wan22-t2v-4step.json`** -- Text-to-video (WAN 2.2 14B, fp8, 4-step LightX2V LoRA)
These bundled WAN 2.2 14B FP8 workflows are the high-quality profile and
recommend roughly 16GB VRAM. That is not a ComfyUI-wide requirement. The
`comfyui_video` tool's top-level `resource_profile` is an 8GB provider floor so
preflight does not imply ComfyUI itself requires 16GB. Low-VRAM users should use
custom workflows such as Wan 2.1 1.3B, LTX-Video/LTXV FP8 or quantized graphs,
or Wan 2.2 GGUF/quantized community workflows, with shorter frame counts and
lower resolutions as needed.
**I2V workflow -- templated nodes:**
| Node | Class | Templated field |
|------|-------|-----------------|
| 93 | CLIPTextEncode | `text` (positive prompt) |
| 97 | LoadImage | `image` (server filename from upload) |
| 98 | WanImageToVideo | `width`, `height`, `length` |
| 86 | KSamplerAdvanced | `noise_seed` |
| 108 | SaveVideo | `filename_prefix` |
**Input schema:**
```yaml
prompt: string # required
operation: string # "text_to_video" | "image_to_video" (default: t2v)
reference_image_path: string # local path (for i2v)
reference_image_url: string # URL (for i2v, downloaded first)
width: integer # default 640
height: integer # default 640
num_frames: integer # default 81 (5s at 16fps)
seed: integer # optional
output_path: string # where to save the video
workflow_json: string # optional custom workflow; requires output_node
workflow_path: string # optional path to workflow JSON; requires output_node
output_node: string # required for custom workflows
workflow_name: string # optional custom workflow provenance label
workflow_model: string # optional custom model/provenance label
workflow_model_stack: [] # optional custom dependency provenance
```
**execute() flow (i2v):**
1. Upload reference image via `client.upload_image()`
2. Deep-copy i2v workflow template
3. Inject prompt, uploaded image name, seed, dimensions
4. `client.generate(workflow, output_node="108", dest=output_path, timeout=900)`
5. Return `ToolResult`
**execute() flow (t2v):**
1. Deep-copy t2v workflow template
2. Inject prompt, seed, dimensions
3. `client.generate(workflow, output_node="16", dest=output_path, timeout=900)`
4. Return `ToolResult`
`comfyui_video` publishes `operation_statuses` in `get_info()` and implements
`is_operation_available(operation)` for selector routing. This keeps partial
ComfyUI installs useful for the installed mode without advertising unavailable
operation modes as ready. `video_selector` also applies this readiness check
when `operation="rank"` by using `target_operation`, so preflight rankings do
not promote ComfyUI for an operation whose bundled models are missing.
---
### `comfyui_music` -- Music Generation (not shipped)
We explored adding a `comfyui_music` tool using the ACE-Step 3.5B model.
The model runs well in ComfyUI, but the ComfyUI node interface for
ACE-Step is not standardized -- there are multiple custom node packs with
different class names (`AceStepModelLoader` vs native `TextEncodeAceStepAudio`,
etc.). Shipping a workflow that only works with one specific custom node
pack would break for most users.
**Future path:** ACE-Step support should be revisited once OpenMontage decides
the music-generation routing shape and a portable ComfyUI audio workflow
contract. Current image/video workflow overrides are intentionally scoped to
image and video artifacts, not arbitrary audio workflows.
---
## Workflow Override Mechanism
The image and video tools accept either `workflow_json` or `workflow_path`.
When provided, the custom workflow replaces the bundled template entirely and
the caller must also provide `output_node`. This stricter contract is required
because community workflows use arbitrary node IDs.
- Using newer model checkpoints without code changes
- Custom sampling strategies (different schedulers, step counts, LoRAs)
- Community workflows dropped in as-is
- A/B testing different generation approaches
The agent can also read workflow files from `tools/_comfyui/workflows/` and
modify them programmatically before passing to `execute()`.
Custom workflow result metadata reports `workflow_provenance.source` as
`user_supplied` and uses `workflow_model`, `model`, or `workflow_name` as the
model label when provided. If no custom label is supplied, the model is reported
as `custom-comfyui-workflow` instead of one of the bundled model names. The
provenance payload also records `workflow_hash_sha256`. For user-supplied
workflows, callers should provide `workflow_model_stack` with base model, text
encoder, VAE, LoRAs and strengths, scheduler, steps, and guidance when known.
---
## Agent Skill and Setup Contract
Both ComfyUI tools advertise the Layer 3 `comfyui` skill. Agents must read
`.agents/skills/comfyui/SKILL.md` before calling either tool so they know how to
load community workflows, identify output nodes, handle LoRA loader chains, and
record custom workflow provenance.
Unavailable ComfyUI tools expose a structured `setup_offer` in `get_info()`,
`provider_menu()`, and `provider_menu_summary().setup_offers[]`:
```yaml
kind: local_server
env_var: COMFYUI_SERVER_URL
default_url: http://localhost:8188
health_check: GET /system_stats
```
When bundled models are missing, the tool returns a machine-readable
`data.missing_models[]` list with filename, role, destination hint, and download
URL when OpenMontage knows the canonical source. Agents should surface that
payload rather than parsing prose error text.
---
## Configuration
**Environment variables:**
```bash
# .env
COMFYUI_SERVER_URL=http://localhost:8188 # ComfyUI API endpoint
COMFYUI_POLL_INTERVAL=5 # seconds between status checks
COMFYUI_POLL_TIMEOUT=600 # max wait for image gen
COMFYUI_VIDEO_TIMEOUT=900 # max wait for video gen
```
**For Docker Compose setups** (ComfyUI in a container):
```bash
COMFYUI_SERVER_URL=http://host.docker.internal:8188
# or
COMFYUI_SERVER_URL=http://comfyui:8188 # if on same docker network
```
---
## Provider Selection Behavior
When the adapter is available, selectors will rank it alongside other providers
using OpenMontage's 7-dimension scoring:
| Dimension | ComfyUI score | Rationale |
|-----------|---------------|-----------|
| Task fit | High | Supports t2i, i2v, t2v |
| Quality | High | Latest models (FLUX 2, WAN 2.2 14B) |
| Control | Highest | Full workflow customization |
| Reliability | High | Proven in production |
| Cost | $0 | Local compute |
| Latency | Medium | GPU-bound, no network round-trip |
| Continuity | High | Deterministic with seeds |
When ComfyUI is unavailable (server down), selectors fall through to other
available providers. When only one video operation is configured, `video_selector`
uses the tool's operation-specific readiness to avoid selecting ComfyUI for the
missing mode.
---
## What This Unlocks
### Immediate (with existing models)
- **FLUX 2 Dev NVFP4** image generation -- Blackwell-optimized, ~60s per image
- **WAN 2.2 14B FP8 high-quality profile** i2v with 4-step acceleration -- ~3.5 min per 5s clip, about 16GB VRAM recommended
- **WAN 2.2 14B FP8 high-quality profile** t2v (models downloaded, workflow included), about 16GB VRAM recommended
### Low-VRAM profile
ComfyUI can still be useful on 8GB-12GB GPUs when the user supplies an
appropriate `workflow_json` or `workflow_path`. Good candidates include:
- Wan 2.1 1.3B workflows for lower-memory text-to-video.
- LTX-Video/LTXV FP8 or quantized workflows for fast short clips.
- Wan 2.2 GGUF/quantized community workflows at lower resolution and frame count.
OpenMontage should treat those as custom workflow profiles until a blessed
low-VRAM workflow is bundled. For custom workflows, resource requirements are
workflow-supplied rather than inferred from the bundled WAN 2.2 14B profile.
### Future (add models to ComfyUI, no code changes to OpenMontage)
- Newer checkpoints (WAN 3.x, FLUX 3, etc.) -- just update workflow JSON
- ControlNet, IP-Adapter, AnimateDiff -- supported via ComfyUI custom nodes
- Upscaling, inpainting, outpainting -- ComfyUI nodes exist
- Any model the ComfyUI ecosystem supports
### Hardware portability
The same adapter works on:
- NVIDIA DGX Spark (GB10, aarch64, CUDA 13.0)
- Consumer GPUs (RTX 3090/4090, x86)
- Cloud instances (A100, H100)
- Multi-GPU setups (ComfyUI handles device placement)
No PyTorch version pinning, no architecture-specific wheels, no CUDA
compatibility matrices. ComfyUI is the abstraction layer.
---
## Implementation Scope
| Component | Files | Estimated size |
|-----------|-------|----------------|
| Shared client | `tools/_comfyui/client.py` | ~180 lines |
| Shared metadata | `tools/_comfyui/metadata.py` | setup, model stack, provenance helpers |
| Image tool | `tools/graphics/comfyui_image.py` | ~140 lines |
| Video tool | `tools/video/comfyui_video.py` | ~190 lines |
| Layer 3 skill | `.agents/skills/comfyui/SKILL.md` | usage contract |
| Registry summary | `tools/tool_registry.py` | setup offer surfacing |
| Selector readiness filter | `tools/video/video_selector.py` | small operation-readiness check |
| Workflow templates | `tools/_comfyui/workflows/*.json` | 3 files |
| Tests | `tests/contracts/test_comfyui_tools.py` | ~200 lines |
| Docs | `docs/comfyui-adapter-plan.md` | This file |
**Total:** ~500 lines of Python + 3 workflow JSONs.
No changes to: `base_tool.py`, existing non-ComfyUI generation providers, any
pipeline definition, or any schema.
---
## Open Questions
1. **Workflow versioning:** Should workflow JSONs live in the repo or be
user-provided via a config directory? Bundling gives reproducibility;
external gives flexibility.
2. **Async generation:** ComfyUI supports websocket connections for real-time
progress. Worth implementing for long video generations, or is polling
sufficient?
3. **Multi-server:** Should the adapter support multiple ComfyUI instances
(e.g., one for images, one for video) via per-capability URLs?
4. **Music generation:** ACE-Step works in ComfyUI but OpenMontage needs a
dedicated music-generation routing contract before adding `comfyui_music`.
The follow-up should decide selector integration, audio artifact schemas, and
a portable workflow/output-node contract rather than treating music as a
hidden image/video workflow override.