Files

894 lines
36 KiB
Python
Raw Permalink Normal View History

feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
#!/usr/bin/env python3
"""Structural validation of generated harness artifacts.
Approximates harness round-trip without installing each CLI. For every adapter output,
parse and validate against documented schemas. Surface issues with file paths and
remediation hints (OpenAI harness-engineering: lint errors carry their fix).
Usage:
python tools/validate_generated.py # validate all four
python tools/validate_generated.py --harness codex
python tools/validate_generated.py --strict # exit nonzero if any warning
"""
from __future__ import annotations
import argparse
import json
import re
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
import sys
import tomllib
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
from dataclasses import dataclass, field
from pathlib import Path
# Allow `python tools/validate_generated.py` from repo root.
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from tools.adapters.base import WORKTREE, parse_frontmatter
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
from tools.adapters.capabilities import supported_harnesses
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
# ── Findings ─────────────────────────────────────────────────────────────────
@dataclass
class Finding:
severity: str # 'error' | 'warning' | 'info'
harness: str
path: Path
message: str
remediation: str = ""
def render(self) -> str:
rel = self.path.relative_to(WORKTREE) if self.path.is_relative_to(WORKTREE) else self.path
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
tail = f"\n fix: {self.remediation}" if self.remediation else ""
return f"[{self.severity}] {self.harness}: {rel}: {self.message}{tail}"
@dataclass
class Report:
findings: list[Finding] = field(default_factory=list)
def add(self, **kwargs) -> None:
self.findings.append(Finding(**kwargs))
def errors(self) -> list[Finding]:
return [f for f in self.findings if f.severity == "error"]
def warnings(self) -> list[Finding]:
return [f for f in self.findings if f.severity == "warning"]
def infos(self) -> list[Finding]:
return [f for f in self.findings if f.severity == "info"]
# ── Helpers ──────────────────────────────────────────────────────────────────
def _check_nonempty_str_field(
report: Report,
fm: dict,
field: str,
harness: str,
path: Path,
label: str = "",
) -> bool:
"""Validate that `field` exists in `fm`, is a str, and is non-empty.
Adds one Finding per violation. Returns True if the field passes all checks.
``label`` is only used in remediation hints (defaults to ``field``).
"""
label = label or field
if field not in fm:
report.add(
severity="error",
harness=harness,
path=path,
message=f"missing required `{field}` field in frontmatter",
remediation=f"Each {harness} {label} needs a `{field}`.",
)
return False
if not isinstance(fm[field], str):
report.add(
severity="error",
harness=harness,
path=path,
message=f"`{field}` field must be a string",
remediation=f"Set `{field}` to a plain string in the {label} frontmatter.",
)
return False
if not fm[field].strip():
report.add(
severity="error",
harness=harness,
path=path,
message=f"`{field}` field is empty",
remediation=f"Provide a non-empty `{field}` in the {label} frontmatter.",
)
return False
return True
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
# ── Codex validators ─────────────────────────────────────────────────────────
def validate_codex(report: Report) -> None:
root = WORKTREE / ".codex"
if not root.is_dir():
return # nothing generated yet
# 1. Every agent .toml parses and has required fields.
required_agent_fields = {"name", "description", "developer_instructions"}
for toml_path in (root / "agents").glob("*.toml") if (root / "agents").is_dir() else []:
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
try:
data = tomllib.loads(toml_path.read_text(encoding="utf-8"))
except tomllib.TOMLDecodeError as e:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="codex",
path=toml_path,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"TOML parse error: {e}",
remediation="Regenerate via `make generate HARNESS=codex PLUGIN=<name>`.",
)
continue
missing = required_agent_fields - set(data.keys())
if missing:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="codex",
path=toml_path,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"missing required TOML fields: {sorted(missing)}",
remediation="The source agent markdown is missing fields. Check `description:` in frontmatter.",
)
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
if "sandbox_mode" in data and data["sandbox_mode"] not in {
"read-only",
"workspace-write",
"danger-full-access",
}:
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="codex",
path=toml_path,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"invalid sandbox_mode: {data['sandbox_mode']!r}",
remediation="Must be one of: read-only, workspace-write, danger-full-access.",
)
# 2. Every skill SKILL.md has valid frontmatter + name matches directory.
skills_dir = root / "skills"
if skills_dir.is_dir():
for skill_md in skills_dir.glob("*/SKILL.md"):
content = skill_md.read_text(encoding="utf-8")
fm, body = parse_frontmatter(content)
if not fm:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="codex",
path=skill_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message="missing or invalid frontmatter",
remediation="SKILL.md must start with `---\\nname: ...\\ndescription: ...\\n---`.",
)
continue
if fm.get("name") != skill_md.parent.name:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="codex",
path=skill_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"frontmatter name {fm.get('name')!r} != directory name {skill_md.parent.name!r}",
remediation="Codex requires the name field to match the skill directory exactly.",
)
if not fm.get("description"):
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="codex",
path=skill_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message="empty description in frontmatter",
remediation="Add a description with a trigger phrase (e.g. 'Use when ...').",
)
# Body cap — entire file ≤ 8 KB (the cap Codex hardcodes).
# Promoted to error: at runtime Codex hard-truncates anything past the cap,
# silently breaking the skill. The fix is mechanical (extract sections to
# references/) and the false-positive risk is zero.
file_bytes = len(content.encode("utf-8"))
if file_bytes > 8 * 1024:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="codex",
path=skill_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"file size {file_bytes} B exceeds Codex 8192 B injection cap",
remediation="Push detail into references/details.md and shorten the SKILL.md body or description.",
)
# 3. AGENTS.md exists and is ≤ 150 lines.
agents_md = WORKTREE / "AGENTS.md"
if not agents_md.is_file():
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="warning",
harness="codex",
path=agents_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message="AGENTS.md not generated",
remediation="Run `make generate HARNESS=codex --all` (or include a global pass).",
)
else:
line_count = len(agents_md.read_text(encoding="utf-8").splitlines())
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
if line_count > 150:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="warning",
harness="codex",
path=agents_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"AGENTS.md is {line_count} lines (cap: 150 — table-of-contents pattern)",
remediation="Move detail into docs/; keep AGENTS.md as a navigation index.",
)
# ── Cursor validators ────────────────────────────────────────────────────────
_ALLOWED_MDC_KEYS = {"description", "globs", "alwaysApply"}
def validate_cursor(report: Report) -> None:
root = WORKTREE / ".cursor-plugin"
if not root.is_dir():
return
# 1. marketplace.json shape
marketplace = root / "marketplace.json"
if marketplace.is_file():
try:
data = json.loads(marketplace.read_text(encoding="utf-8"))
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
except json.JSONDecodeError as e:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="cursor",
path=marketplace,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"JSON parse error: {e}",
remediation="Regenerate via `make generate HARNESS=cursor --all`.",
)
return
if "owner" not in data:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="cursor",
path=marketplace,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message="marketplace.json missing required `owner` field",
remediation="Cursor 2.5+ requires owner.name in marketplace.json.",
)
for entry in data.get("plugins", []):
if "source" not in entry:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="cursor",
path=marketplace,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"plugin entry {entry.get('name', '<unnamed>')} missing `source`",
remediation="Cursor uses `source` (not `path` or `url`) per plugin entry.",
)
# 2. Each per-plugin manifest parses
plugins_dir = root / "plugins"
if plugins_dir.is_dir():
for manifest in plugins_dir.glob("*.json"):
try:
data = json.loads(manifest.read_text(encoding="utf-8"))
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
except json.JSONDecodeError as e:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="cursor",
path=manifest,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"JSON parse error: {e}",
remediation="Regenerate via `make generate HARNESS=cursor`.",
)
continue
if "name" not in data:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="cursor",
path=manifest,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message="missing required `name` field",
remediation="Every Cursor plugin.json needs a name.",
)
# 3. .cursor/rules/*.mdc — only the three allowed frontmatter keys
rules_dir = WORKTREE / ".cursor" / "rules"
if rules_dir.is_dir():
for mdc in rules_dir.glob("*.mdc"):
content = mdc.read_text(encoding="utf-8")
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
fm, _ = parse_frontmatter(content)
if not fm:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="cursor",
path=mdc,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message="missing or invalid frontmatter",
remediation="MDC files need YAML frontmatter with at least `description:`.",
)
continue
invalid = set(fm.keys()) - _ALLOWED_MDC_KEYS
if invalid:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="cursor",
path=mdc,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"invalid MDC keys: {sorted(invalid)}",
remediation=(
"Cursor only supports description/globs/alwaysApply. "
"Keys like agentRequested:, mode:, tags: are folklore."
),
)
# ── OpenCode validators ──────────────────────────────────────────────────────
_OPENCODE_PERMISSION_KEYS = {
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
"read",
"edit",
"write",
"bash",
"grep",
"glob",
"list",
"task",
"skill",
"lsp",
"webfetch",
"websearch",
"external_directory",
"todowrite",
"question",
"doom_loop",
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
}
_OPENCODE_PERMISSION_VALUES = {"allow", "ask", "deny"}
_OPENCODE_MODES = {"primary", "subagent", "all"}
_OPENCODE_SKILL_NAME_RE = re.compile(r"^[a-z0-9]+(-[a-z0-9]+)*$")
_OPENCODE_SKILL_NAME_MAX = 64
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
def _extract_permission_block(raw: str) -> dict | None:
"""Pull the top-level `permission:` block out of raw frontmatter text and return its
key→value mapping. Returns None if no top-level `permission:` block is found.
We hand-roll this because `parse_frontmatter` collapses nested mappings into a list
of strings, losing key→value structure. Validating from raw text catches the
real-shape permission blocks the OpenCode adapter emits.
Only matches `permission:` at column 0 — a nested ` permission:` (e.g. inside a
`metadata:` block) is correctly ignored.
"""
if not raw.startswith("---"):
return None
end = raw.find("\n---", 3)
if end == -1:
return None
fm_text = raw[3:end]
lines = fm_text.splitlines()
block: dict[str, str] = {}
in_perm = False
for line in lines:
# Top-level (column-0) `permission:` is the only one we recognize.
if line.rstrip() == "permission:" and not line.startswith((" ", "\t")):
in_perm = True
continue
if in_perm:
# Continuation if line is indented; otherwise we've left the block.
if line.startswith((" ", "\t")) and ":" in line:
k, _, v = line.strip().partition(":")
block[k.strip()] = v.strip()
elif line.strip() == "":
continue
else:
in_perm = False
return block if block else None
def validate_opencode(report: Report) -> None:
root = WORKTREE / ".opencode"
if not root.is_dir():
return
# 1. opencode.json
cfg = WORKTREE / "opencode.json"
if cfg.is_file():
try:
data = json.loads(cfg.read_text(encoding="utf-8"))
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
except json.JSONDecodeError as e:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="opencode",
path=cfg,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"JSON parse error: {e}",
remediation="Regenerate via `make generate HARNESS=opencode`.",
)
else:
if "$schema" not in data:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="info",
harness="opencode",
path=cfg,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message="missing $schema reference",
remediation='Add "$schema": "https://opencode.ai/config.json" for editor tooling.',
)
# 2. Every agent .md has required frontmatter
agents_dir = root / "agents"
if agents_dir.is_dir():
for agent_md in agents_dir.glob("*.md"):
content = agent_md.read_text(encoding="utf-8")
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
fm, _ = parse_frontmatter(content)
if not fm:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="opencode",
path=agent_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message="missing or invalid frontmatter",
remediation="Regenerate via `make generate HARNESS=opencode`.",
)
continue
if fm.get("mode") not in _OPENCODE_MODES:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="opencode",
path=agent_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"mode {fm.get('mode')!r} not in {sorted(_OPENCODE_MODES)}",
remediation="Set mode: subagent on transpiled agents.",
)
model = fm.get("model", "")
if model and "/" not in model:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="warning",
harness="opencode",
path=agent_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"model {model!r} is not provider-prefixed (e.g. 'anthropic/claude-...')",
remediation="OpenCode requires `provider/model-id`; check MODEL_ALIASES in capabilities.py.",
)
# Validate the permission block by re-parsing raw frontmatter — `fm` from
# parse_frontmatter flattens nested mappings into lists, losing structure.
perm = _extract_permission_block(content)
if perm:
unknown_keys = set(perm.keys()) - _OPENCODE_PERMISSION_KEYS
if unknown_keys:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="opencode",
path=agent_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"unknown permission keys: {sorted(unknown_keys)}",
remediation=(
f"Valid keys: {sorted(_OPENCODE_PERMISSION_KEYS)}. "
"Update _OPENCODE_PERMISSIONS in tools/adapters/opencode.py if a new key was added."
),
)
for k, v in perm.items():
if v not in _OPENCODE_PERMISSION_VALUES:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
harness="opencode",
path=agent_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"permission.{k} = {v!r} not in {sorted(_OPENCODE_PERMISSION_VALUES)}",
remediation="Values must be allow/ask/deny.",
)
# 3. Every skill has OpenCode-valid frontmatter with name matching directory.
skills_dir = root / "skills"
if skills_dir.is_dir():
for skill_md in skills_dir.glob("*/SKILL.md"):
content = skill_md.read_text(encoding="utf-8")
fm, _ = parse_frontmatter(content)
if not fm:
report.add(
severity="error",
harness="opencode",
path=skill_md,
message="missing or invalid frontmatter",
remediation="Regenerate via `make generate HARNESS=opencode`.",
)
continue
name = str(fm.get("name") or "")
if name != skill_md.parent.name:
report.add(
severity="error",
harness="opencode",
path=skill_md,
message=f"frontmatter name {name!r} != directory {skill_md.parent.name!r}",
remediation="OpenCode skill names must match their directory.",
)
if not _OPENCODE_SKILL_NAME_RE.fullmatch(name):
report.add(
severity="error",
harness="opencode",
path=skill_md,
message=f"skill name {name!r} is not OpenCode-safe",
remediation="Use lowercase alphanumeric names with single hyphen separators.",
)
if len(name) > _OPENCODE_SKILL_NAME_MAX:
report.add(
severity="error",
harness="opencode",
path=skill_md,
message=(
f"skill name {name!r} is {len(name)} chars; "
f"limit is {_OPENCODE_SKILL_NAME_MAX}"
),
remediation="Shorten the plugin or skill name before generating OpenCode skills.",
)
if not str(fm.get("description") or "").strip():
report.add(
severity="error",
harness="opencode",
path=skill_md,
message="empty description in frontmatter",
remediation="OpenCode skills require a description.",
)
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
# ── Antigravity validators ───────────────────────────────────────────────────
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
_ANTIGRAVITY_PLUGIN_NAME_RE = re.compile(r"^[a-zA-Z0-9_-]+$")
_ANTIGRAVITY_MODEL_TIERS = {"inherit", "flash", "pro"}
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
def validate_antigravity(report: Report) -> None:
root = WORKTREE / ".antigravity" / "plugins"
if not root.is_dir():
return
for plugin_dir in sorted(p for p in root.iterdir() if p.is_dir()):
# 1. plugin.json parses, has a non-empty agy-safe `name`, and it matches the
# directory (agy discovers plugins by directory; a mismatch is confusing at best).
plugin_json = plugin_dir / "plugin.json"
if not plugin_json.is_file():
report.add(
severity="error",
harness="antigravity",
path=plugin_dir,
message="missing plugin.json",
remediation="Regenerate via `make generate HARNESS=antigravity`.",
)
else:
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
try:
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
data = json.loads(plugin_json.read_text(encoding="utf-8"))
except json.JSONDecodeError as e:
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
harness="antigravity",
path=plugin_json,
message=f"JSON parse error: {e}",
remediation="Regenerate via `make generate HARNESS=antigravity`.",
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
)
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
data = None
if data is not None and not isinstance(data, dict):
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
harness="antigravity",
path=plugin_json,
message=f"plugin.json must be a JSON object, got {type(data).__name__}",
remediation="Regenerate via `make generate HARNESS=antigravity`.",
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
)
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
data = None
if data is not None:
name = data.get("name")
if not isinstance(name, str) or not name:
report.add(
severity="error",
harness="antigravity",
path=plugin_json,
message="missing or empty required `name` field",
remediation="Every agy plugin.json needs a non-empty `name`.",
)
else:
if not _ANTIGRAVITY_PLUGIN_NAME_RE.fullmatch(name):
report.add(
severity="error",
harness="antigravity",
path=plugin_json,
message=f"plugin name {name!r} is not agy-safe (must match ^[a-zA-Z0-9_-]+$)",
remediation="Rename the plugin to letters, digits, hyphens, underscores only.",
)
if name != plugin_dir.name:
report.add(
severity="error",
harness="antigravity",
path=plugin_json,
message=f"plugin.json name {name!r} != directory name {plugin_dir.name!r}",
remediation="agy discovers plugins by directory; name must match.",
)
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
# 2. Every skill's frontmatter name matches its directory.
for skill_md in sorted((plugin_dir / "skills").glob("*/SKILL.md")):
fm, _ = parse_frontmatter(skill_md.read_text(encoding="utf-8"))
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
if fm.get("name") != skill_md.parent.name:
report.add(
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
severity="error",
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
harness="antigravity",
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
path=skill_md,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
message=f"frontmatter name {fm.get('name')!r} != directory {skill_md.parent.name!r}",
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
remediation="agy discovers skills by directory; name must match.",
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
)
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
# 3. Every agent has name + description, and model is a valid tier alias.
for agent_md in sorted((plugin_dir / "agents").glob("*.md")):
fm, _ = parse_frontmatter(agent_md.read_text(encoding="utf-8"))
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
_check_nonempty_str_field(report, fm, "name", "antigravity", agent_md, label="agent")
_check_nonempty_str_field(
report, fm, "description", "antigravity", agent_md, label="agent"
)
model = fm.get("model")
if model not in _ANTIGRAVITY_MODEL_TIERS:
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
report.add(
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
severity="error",
harness="antigravity",
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
path=agent_md,
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
message=f"model {model!r} not in {sorted(_ANTIGRAVITY_MODEL_TIERS)}",
remediation="agy subagent `model` must be a tier alias: inherit, flash, or pro.",
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
)
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
# 4. Every command TOML parses and has description + prompt + {{args}}.
for toml_path in sorted((plugin_dir / "commands").rglob("*.toml")):
try:
data = tomllib.loads(toml_path.read_text(encoding="utf-8"))
except tomllib.TOMLDecodeError as e:
report.add(
severity="error",
harness="antigravity",
path=toml_path,
message=f"TOML parse error: {e}",
remediation="Regenerate via `make generate HARNESS=antigravity`.",
)
continue
_check_nonempty_str_field(
report, data, "description", "antigravity", toml_path, label="command"
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
)
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
has_prompt = _check_nonempty_str_field(
report, data, "prompt", "antigravity", toml_path, label="command"
)
if has_prompt and "{{args}}" not in data["prompt"]:
report.add(
severity="warning",
harness="antigravity",
path=toml_path,
message="prompt does not include {{args}} placeholder",
remediation="Append {{args}} so user input is appended to the prompt.",
)
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
# ── Pi validators ────────────────────────────────────────────────────────────
_PI_SKILL_NAME_RE = re.compile(r"^[a-z0-9]+(-[a-z0-9]+)*$")
_PI_SKILL_NAME_MAX = 64
_PI_DESCRIPTION_MAX = 1024
def validate_pi(report: Report) -> None:
"""Validate the `.pi/` tree: skills (Agent Skills standard), prompt templates,
and agents in the reference subagent-extension format."""
root = WORKTREE / ".pi"
if not root.is_dir():
return
fix = "Regenerate via `make generate HARNESS=pi`."
# 1. Skills: every skill directory has a SKILL.md, and that file names its own
# directory, carries a description, and follows Pi's name rules (warnings).
skills_root = root / "skills"
if skills_root.is_dir():
for plugin_dir in sorted(p for p in skills_root.iterdir() if p.is_dir()):
for skill_dir in sorted(p for p in plugin_dir.iterdir() if p.is_dir()):
if not (skill_dir / "SKILL.md").is_file():
report.add(
severity="error",
harness="pi",
path=skill_dir,
message="skill directory has no SKILL.md",
remediation=fix,
)
for skill_md in sorted((root / "skills").glob("*/*/SKILL.md")):
fm, _ = parse_frontmatter(skill_md.read_text(encoding="utf-8"))
name = fm.get("name")
_check_nonempty_str_field(report, fm, "name", "pi", skill_md, label="skill")
_check_nonempty_str_field(report, fm, "description", "pi", skill_md, label="skill")
if isinstance(name, str) and name:
if name != skill_md.parent.name:
report.add(
severity="error",
harness="pi",
path=skill_md,
message=f"frontmatter name {name!r} != directory {skill_md.parent.name!r}",
remediation=fix,
)
if len(name) > _PI_SKILL_NAME_MAX or not _PI_SKILL_NAME_RE.fullmatch(name):
report.add(
severity="warning",
harness="pi",
path=skill_md,
message=(
f"skill name {name!r} violates Pi's name rules "
f"(lowercase a-z, 0-9, single hyphens, max {_PI_SKILL_NAME_MAX})"
),
remediation="Rename the source skill directory; Pi loads it with a warning.",
)
description = fm.get("description")
if isinstance(description, str) and len(description) > _PI_DESCRIPTION_MAX:
report.add(
severity="warning",
harness="pi",
path=skill_md,
message=f"description is {len(description)} chars (Pi warns above {_PI_DESCRIPTION_MAX})",
remediation="Shorten the source skill description.",
)
# 2. Prompt templates: namespaced filename, frontmatter with a description.
for prompt_md in sorted((root / "prompts").glob("*.md")):
plugin_part, separator, command_part = prompt_md.stem.partition("__")
if not separator or not plugin_part or not command_part:
report.add(
severity="error",
harness="pi",
path=prompt_md,
message="prompt filename is not `<plugin>__<command>.md`",
remediation=fix,
)
fm, _ = parse_frontmatter(prompt_md.read_text(encoding="utf-8"))
_check_nonempty_str_field(report, fm, "description", "pi", prompt_md, label="prompt")
# 3. Agents: namespaced filename, name + description, model is provider/id when present.
for agent_md in sorted((root / "agents").glob("*.md")):
plugin_part, separator, agent_part = agent_md.stem.partition("__")
if not separator or not plugin_part or not agent_part:
report.add(
severity="error",
harness="pi",
path=agent_md,
message="agent filename is not `<plugin>__<agent>.md`",
remediation=fix,
)
fm, _ = parse_frontmatter(agent_md.read_text(encoding="utf-8"))
_check_nonempty_str_field(report, fm, "name", "pi", agent_md, label="agent")
_check_nonempty_str_field(report, fm, "description", "pi", agent_md, label="agent")
model = fm.get("model")
if model is not None and (not isinstance(model, str) or "/" not in model):
report.add(
severity="error",
harness="pi",
path=agent_md,
message=f"model {model!r} is not provider/id",
remediation="Pi models are `provider/id`; check MODEL_ALIASES['pi'].",
)
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
# ── Driver ───────────────────────────────────────────────────────────────────
def validate_copilot(report: Report) -> None:
"""Validate Copilot agent, skill, and command markdown files under WORKTREE/.copilot.
Checks that generated agent, skill, and command markdown files have valid
frontmatter and the minimum metadata Copilot needs to discover them.
"""
candidate_roots = [WORKTREE / ".copilot"]
found_any = False
for root in candidate_roots:
agents_dir = root / "agents"
if not agents_dir.is_dir():
continue
found_any = True
for agent_md in agents_dir.glob("*.agent.md"):
content = agent_md.read_text(encoding="utf-8")
fm, _ = parse_frontmatter(content)
if not fm:
report.add(
severity="error",
harness="copilot",
path=agent_md,
message="missing or invalid frontmatter",
remediation="Regenerate via `make generate HARNESS=copilot`.",
)
continue
_check_nonempty_str_field(report, fm, "name", "copilot", agent_md, label="agent")
_check_nonempty_str_field(report, fm, "description", "copilot", agent_md, label="agent")
# 3. Skills: validate .copilot/skills/*/SKILL.md exists and has valid frontmatter.
skills_dir = WORKTREE / ".copilot" / "skills"
if skills_dir.is_dir():
found_any = True
for skill_md in skills_dir.glob("*/SKILL.md"):
content = skill_md.read_text(encoding="utf-8")
fm, _ = parse_frontmatter(content)
if not fm:
report.add(
severity="error",
harness="copilot",
path=skill_md,
message="missing or invalid frontmatter",
remediation="Regenerate via `make generate HARNESS=copilot`.",
)
continue
_check_nonempty_str_field(report, fm, "name", "copilot", skill_md, label="skill")
_check_nonempty_str_field(report, fm, "description", "copilot", skill_md, label="skill")
commands_dir = WORKTREE / ".copilot" / "commands"
if commands_dir.is_dir():
found_any = True
for command_md in commands_dir.rglob("*.md"):
content = command_md.read_text(encoding="utf-8")
fm, body = parse_frontmatter(content)
if not fm:
report.add(
severity="error",
harness="copilot",
path=command_md,
message="missing or invalid frontmatter",
remediation="Regenerate via `make generate HARNESS=copilot`.",
)
continue
description = fm.get("description")
if not isinstance(description, str) or not description.strip():
report.add(
severity="error",
harness="copilot",
path=command_md,
message="missing required `description` field in frontmatter",
remediation="Copilot command prompt files need a non-empty description.",
)
if command_md.name != "index.md" and not body.strip():
report.add(
severity="warning",
harness="copilot",
path=command_md,
message="command body is empty",
remediation="Keep the prompt body in the source command markdown.",
)
if not found_any:
return
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
_VALIDATORS = {
"codex": validate_codex,
"copilot": validate_copilot,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
"cursor": validate_cursor,
"opencode": validate_opencode,
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
"antigravity": validate_antigravity,
"pi": validate_pi,
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
}
def main() -> int:
parser = argparse.ArgumentParser(description="Validate generated harness artifacts.")
feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542) * feat: AGENTS.md canonical context + OpenAI harness-engineering layout Promote AGENTS.md to the committed cross-harness context file (per the agents.md convention and OpenAI's harness-engineering blog). Harness- specific files become thin redirects: - AGENTS.md — canonical, committed (~74 lines, table-of-contents) - CLAUDE.md — `@AGENTS.md` import + Claude-specific addenda - GEMINI.md — Gemini-specific setup only - .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md - ARCHITECTURE.md — new at root, top-level architectural map - gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md - .gitignore — drops the AGENTS.md entry (file is now committed) Harness support verified: - Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap) - Cursor 2.5+ reads AGENTS.md natively - OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist) - Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's documented interop pattern) - Gemini CLI: `.gemini/settings.json` context.fileName redirect (Gemini doesn't support @-imports) Codex adapter no longer generates AGENTS.md — `emit_global` instead validates the committed file fits Codex's 32 KiB cap and the 150-line table-of-contents convention. Tests updated. Clean-output target no longer touches AGENTS.md. ## Auxiliary files updated for multi-harness reality - `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness + component path; renames "subagent" → "plugin/agent/skill/command" - `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers framework / harness / tooling / docs / CI in addition to components - `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component Proposal" with component-type dropdown (plugin/agent/skill/command/ harness adapter) and cross-harness portability field - `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring guide, per-harness docs; updated Contributing link to root - `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md - `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness checklists, test-plan checklist, portability-notes section - CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md - gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects to AGENTS.md as contextFileName ## Code-quality CI New `.github/workflows/code-quality.yml` with three jobs: - `python-lint` — `ruff check`, `ruff format --check`, `ty check` on the adapter framework + plugin-eval. yt-design-extractor.py legacy code excluded. - `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE, CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`. - `json-lint` — validates every JSON / TOML / YAML in the repo (excluding generated trees). Required ty environment config added to plugin-eval/pyproject.toml so `tools.adapters.*` resolves from outside the package. Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities` (return-type annotation didn't match the `Capability` dataclass returned). Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py. ruff format applied across all in-scope files (formatting-only diffs). ## Tests + verification - 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior) - `make validate STRICT=1` clean - `make garden` 0 errors (10 warnings — remaining oversize source skills) - `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code - Real-CLI round-trip: `opencode agent list` discovers 193 subagents, `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse ## Tag recommendations (separate task — for repo About panel) Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`): automation mcp ai-agents developer-tools claude-code anthropic agentic-ai agents prompt-engineering cursor multi-agent agent-skills orchestration opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules claude-code-plugins * fix(ci): YAML multi-doc + markdownlint scope/rules Two CI failures on PR #542, both fixed: ## JSON/TOML/YAML syntax job (3 false-positive YAML errors) The job's YAML validation used `yaml.safe_load` which only reads the first document in a multi-document YAML stream. Three Kubernetes manifest templates use the standard `---` document separator (valid YAML) and were mis-flagged: plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted. ## Markdown lint job (lots of pre-existing plugin-README violations) Two changes: 1. **Narrow the lint glob** — markdownlint now runs against top-level guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup, CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs (`plugins/*/README.md`) are owned by their plugin authors and not lint-gated as part of this framework PR. Lint enforcement for those belongs at the plugin-author layer, not the framework PR layer. 2. **Tighten `.markdownlint.json`** — disable two rules that produce noise without catching real defects: - MD040 (fenced-code-language) — terminal output / shell command blocks frequently omit a language by convention - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't affect rendering Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks- around-fences), MD032 (blanks-around-lists), MD056 (table-column- count), MD058 (blanks-around-tables). ## Real defects caught and fixed The narrower scope still caught 5 real issues: - `docs/agent-skills.md:397` — code fence inside an ordered list item needed a blank line before the fence - `docs/authoring.md:90` — bulleted list needed a blank line above - `docs/plugin-eval.md:87` — table needed a blank line above - `OPENCODE.md:39` — table-column-count error caused by literal `|` inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5) Rewrote as ``mode:` one of `primary` / `subagent` / `all`` ## Verification - `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors - `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors) - All other CI jobs already passing (Python ruff/ty, multi-harness generate, CLI smoke test, plugin-eval pytest, tools pytest) * fix: address PR #542 bot feedback - Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>` produced a false "missing" warning even when AGENTS.md was committed at the real root, breaking --strict generation outside the repo. Added a constructor arg `repo_root` so tests can stage a fake AGENTS.md without touching the committed file, plus a regression test that proves the two paths are decoupled. - CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions: contents: read` and `persist-credentials: false` on every checkout. Skipped the SHA-pinning recommendation — it's a heavier blanket-policy decision and the workflow has no write scope to abuse. - CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups by moving `ty` into `[project.optional-dependencies].dev` and removing the now-empty `[dependency-groups]` block. Dropped `--group dev` from `code-quality.yml`'s `uv sync` since `--all-extras` now covers it. - Verified gemini-extension.json counts (82/191/155/102) against the actual source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json, 191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
parser.add_argument(
"--harness", choices=supported_harnesses(), help="Only validate one harness."
)
parser.add_argument("--strict", action="store_true", help="Exit nonzero on any warning.")
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
args = parser.parse_args()
targets = [args.harness] if args.harness else supported_harnesses()
report = Report()
for h in targets:
_VALIDATORS[h](report)
if not report.findings:
print(f"OK: no issues across {len(targets)} harness(es).")
return 0
# Sort findings by (severity priority, harness, path) for triage-friendly output.
severity_order = {"error": 0, "warning": 1, "info": 2}
sorted_findings = sorted(
report.findings,
key=lambda f: (severity_order.get(f.severity, 9), f.harness, str(f.path)),
)
errors = report.errors()
warnings = report.warnings()
infos = report.infos()
for f in sorted_findings:
print(f.render())
print()
print(f"Total: {len(errors)} error(s), {len(warnings)} warning(s), {len(infos)} info.")
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
if errors:
return 1
if args.strict and warnings:
return 1
return 0
if __name__ == "__main__":
sys.exit(main())