Files

252 lines
15 KiB
Markdown
Raw Permalink Normal View History

feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
# Cross-harness capability matrix
claude-agents is a multi-harness plugin marketplace. Source-of-truth lives under `plugins/`
as Claude Code markdown. Per-harness artifacts are generated by adapters under `tools/adapters/`.
> This file mirrors the capability matrix in `tools/adapters/capabilities.py`. Edit there;
> regenerate via `make docs`.
## Supported harnesses
| Harness | Status | Generated paths |
|---|---|---|
| **Claude Code** | source-of-truth | `plugins/`, `.claude-plugin/marketplace.json` |
| **OpenAI Codex CLI** | supported | committed: `.agents/plugins/marketplace.json`, `plugins/*/.codex-plugin/plugin.json`; gitignored: `.codex/skills/`, `.codex/agents/` |
| **Cursor** (2.5+) | supported | committed: `.cursor-plugin/`, `.cursor/rules/` (curated) — points at source `plugins/` |
| **OpenCode** (`sst/opencode`) | supported | gitignored: `.opencode/agents/`, `.opencode/commands/`, `.opencode/skills/`, `opencode.json` |
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
| **Google Antigravity CLI** (`agy`) | supported | gitignored: `.antigravity/plugins/<name>/{skills/,agents/,commands/}` |
| **Pi** (`pi`, `@earendil-works/pi-coding-agent`) | supported | gitignored: `.pi/{skills/<plugin>/<skill>/,prompts/,agents/}` |
feat: support gh skill and npx skills installers (#693) * feat: support gh skill and npx skills installers Both Agent Skills installers already discover every skill in this repo through the plugins/<plugin>/skills/<skill>/ layout, so support is documentation plus a real-CLI gate rather than a layout change. - docs: README quick start block; docs/harnesses.md "Skills-only installers" (selectors, install paths, release-freeze rule, local-checkout caveat); docs/agent-skills.md pointer; AGENTS.md bullet; authoring note that a skill's name must equal its directory name - smoke tests: gh skill local discovery, gh skill publish --dry-run spec validation, npx skills discovery, and skill-name uniqueness across plugins (both installers install under the bare skill name) - ci: set up Node for npx and fail loudly if the runner's gh predates gh skill - rename database-design/skills/postgresql to postgresql-table-design so the directory matches the frontmatter name, the one spec error the dry-run found * test: type the smoke helper env as dict[str, str] * fix(database-design): correct three PostgreSQL facts in postgresql-table-design Flagged by review once the rename made the file appear new. Pre-existing content, separate commit so it can be dropped if the rename should stay content-free. - UNIQUE NULLS NOT DISTINCT (...) places the clause before the column list - the RLS example used current_user_id(), which is not a built-in; use current_user or an app-set setting - foreign keys on partitioned tables work from PG11 (from) and PG12 (referencing); triggers are only the pre-11 fallback * refactor(database-design): split postgresql-table-design under the Codex 8 KB cap SKILL.md keeps the rules, decision points, a When to Use section, and the three quick-start DDL examples (7.6 KB body, 126 lines). The full data-type catalog, table types, row-level security, constraint and index notes, partitioning DDL, workload patterns, generated columns, extensions, and JSONB indexing move verbatim to references/details.md. The "avoid these types" list becomes a table. Clears the SKILL_OVER_CODEX_CAP warning this skill carried. Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs * test(smoke): capture opencode agent list through a file, tighten npx parser `opencode agent list` prints every agent's expanded permission array (about 330k lines) and exits before a pipe drains, so pipe capture intermittently lost the alphabetically last agents (measured: 1 in 5 runs short via subprocess pipes, 0 in 5 via a file sink). Both OpenCode tests now write stdout to a file. The npx parser matches only the `│ <name>` line shape. Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs * docs,ci: cover the skills installers everywhere install is documented - docs/usage.md, docs/plugins.md, docs/architecture.md, ARCHITECTURE.md, README multi-harness section, docs/harnesses.md supported table, and CONTRIBUTING's portability checklist now name gh skill and npx skills - docs/round-trip-results.md gains summary rows and a reproduce recipe - docs/authoring.md: skill directory names are identities for installers and adapters alike, so a rename is user-visible - CI installs GitHub CLI from the official apt repository when the runner build predates `gh skill`, so the job no longer depends on the image Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs
2026-09-01 18:28:17 -04:00
| **Agent Skills installers** (`gh skill` 2.90+, `npx skills`) | supported, skills only | nothing generated; both read `plugins/*/skills/` from GitHub directly, see [Skills-only installers](#skills-only-installers) |
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
## Capability matrix
| Capability | Claude Code | Codex | Cursor | OpenCode | Antigravity | Pi |
|---|---|---|---|---|---|---|
| Skills (SKILL.md native) | ✅ | ✅ | ✅ via `.claude/` | ✅ via `.opencode/skills/` | ✅ (native, self-contained per plugin) | ✅ (recursive discovery) |
| Subagents (markdown native) | ✅ | TOML format | ✅ via `.claude/` | ✅ (different frontmatter) | ✅ (`agents/<name>.md` + `invoke_subagent`/`define_subagent`) | via the reference `subagent` extension (`agents/<plugin>__<agent>.md`) |
| Slash commands | ✅ | converted to skills | ✅ | ✅ | TOML at `commands/<p>/<cmd>.toml` (agy reports these as "converted to skills") | prompt templates at `prompts/<plugin>__<cmd>.md` |
| Plugin marketplace | ✅ | — | ✅ (2.5+) | — | ✅ (`agy plugin install <name>@marketplace` / `agy plugin link`) | — (packages via npm, git, or a local path) |
| Parallel subagents | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ (extension) |
| Per-agent tool allowlist | ✅ (`tools:`) | only `sandbox_mode` | only `readonly:` | ✅ (`permission:` block) | ✅ (`tools:`, agy-native names) | ✅ (`tools:`, extension) |
| `TodoWrite` tool | ✅ | — | — | ✅ | — | — |
| `Task`/`Agent` spawn tool | ✅ | name in prose | ✅ | ✅ (`task`) | ✅ (`invoke_subagent`/`define_subagent`) | `subagent` (extension) |
| MCP servers | ✅ | ✅ | ✅ | ✅ | ✅ | via extension |
| Lifecycle hooks | ✅ | — | — | ✅ (TS plugins) | ✅ | ✅ (TypeScript extensions) |
| Context file | `CLAUDE.md` | `AGENTS.md` (32 KiB cap) | `AGENTS.md` | `AGENTS.md` / `~/.claude/CLAUDE.md` | `AGENTS.md` (read natively) | `AGENTS.md` |
| Context file recommended cap | 150 lines / 500 tokens | 150 lines / 500 tokens | 150 lines / 500 tokens | 150 lines / 500 tokens | 150 lines / 500 tokens | 150 lines / 500 tokens |
| Skill body hard cap | none | **8 KB** | none | none | none | none |
| Tool name case | CamelCase (`Read`) | action verbs (no tool vocab) | lowercase | lowercase (strict) | lowercase (agy-native names) | lowercase (`read`, `bash`) |
| Bare model aliases | ✅ (`fable`/`opus`/`sonnet`/`haiku`) | mapped to GPT-5.x family | use `inherit` | full provider/model-id | mapped to tier alias (`pro`/`flash`/`inherit`) | full provider/model-id |
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
## Claude Code native features
Claude Code is the source-of-truth harness. It reads the canonical context file via `CLAUDE.md`,
a symlink to `AGENTS.md`. Features it supports that other harnesses degrade or lack:
- **Per-agent tool allowlist** — `tools:` frontmatter honored verbatim (Cursor / Codex are coarser; the OpenCode adapter translates this into a `permission:` block).
- **`Task` / `Agent` spawn tool** — fan-out parallel subagent execution. (Codex requires naming an agent in prose to delegate.)
- **`TodoWrite`** — native progress tracking. (Not available in Codex / Cursor / Antigravity / Pi.)
- **Slash-command marketplace** — full `/plugin install`, `/plugin marketplace` workflow.
Claude-Code-only paths:
- `.claude-plugin/marketplace.json` — plugin registry (source of truth)
- `plugins/<name>/.claude-plugin/plugin.json` — per-plugin manifest
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
## Graceful degradation
Each adapter handles incompatibilities mechanically — authors don't need to know the per-harness
rules to write portable content.
| Source pattern | Codex | Cursor | OpenCode | Antigravity | Pi |
|---|---|---|---|---|---|
| `tools: Read, Grep` (agent allowlist) | dropped; `sandbox_mode = "read-only"` heuristic | dropped (Cursor doesn't honor) | converted to `permission:` deny block | rewritten to agy-native tool names | rewritten to Pi built-in names; `tools: []` becomes `read, grep, find, ls` |
| `color: blue` (agent) | dropped | dropped | dropped | dropped | dropped |
| `model: opus` (agent) | mapped to `gpt-5.5` | rewritten to `inherit` | rewritten to `anthropic/claude-opus-4-8` | mapped to `pro` | rewritten to `anthropic/claude-opus-4-8` |
| `model: fable` (agent) | mapped to `gpt-5.5` | rewritten to `inherit` | rewritten to `anthropic/claude-fable-5` | mapped to `pro` | rewritten to `anthropic/claude-fable-5` |
| `TodoWrite` in body | no equivalent — leave as-is | no equivalent — leave as-is | works as-is | no equivalent | no equivalent |
| Skill body > 8 KB | split into `references/details.md` | passed through | passed through | passed through | passed through |
| Agent named `worker` | namespaced to `<plugin>__worker` | passed through | passed through | passed through (no `<plugin>__` namespacing — the plugin dir already scopes it) | namespaced to `<plugin>__worker.md` (the agents directory is flat) |
| Slash command (`commands/<x>.md`) | converted to skill | passed through | rewritten to `.opencode/commands/` | TOML at `commands/<plugin>/<x>.toml`, body always inlined (never `@{path}`-injected) | prompt template at `prompts/<plugin>__<x>.md`, the body is copied as written once Claude tool references are rewritten to Pi names, `$ARGUMENTS` is left in place for Pi to substitute, and no wrapper text is added |
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
## Output paths (committed vs gitignored)
Native install is **lean**: only small JSON registries (pointing at the source `plugins/`) are
committed. The large transformed skill/agent trees stay gitignored — regenerate them locally.
**Committed:**
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
```
.claude-plugin/marketplace.json # SOURCE OF TRUTH
plugins/ # SOURCE OF TRUTH
AGENTS.md # canonical context file
.agents/plugins/marketplace.json # Codex marketplace registry (source.path: ./plugins/<name>)
plugins/*/.codex-plugin/plugin.json # per-plugin Codex manifest (skills: ./skills/)
.cursor-plugin/, .cursor/rules/ # Cursor marketplace + curated rules (point at source)
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
```
**Gitignored (regenerate with `make generate`):**
```
.codex/skills/, .codex/agents/ # transformed Codex trees (for ~/.codex/skills symlink recipe)
.opencode/agents/, .opencode/commands/, .opencode/skills/, opencode.json
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
.antigravity/plugins/<name>/ # self-contained agy plugins (skills/, agents/, commands/)
.copilot/agents/, .copilot/skills/, .copilot/commands/
.pi/skills/, .pi/prompts/, .pi/agents/ # transformed Pi trees (skills are nested per plugin)
```
`.pi/` is also Pi's project-local config directory, so you may keep your own files there such
as `.pi/settings.json` or `.pi/extensions/*.ts`. The adapter owns only `.pi/skills`,
`.pi/prompts` and `.pi/agents`. Cleaning and pruning stay inside those three subdirectories and
leave everything else under `.pi/` alone.
## Native install
- **Codex** — `npx codex-marketplace add wshobson/agents` (or it's auto-discovered as a project
marketplace when the repo is the cwd), then install individual plugins. Codex reads `SKILL.md`
straight from `plugins/<name>/skills/`; skills over the 8 KB cap are truncated by Codex at load.
The gitignored `.codex/skills/` copies remain for the `~/.codex/skills` symlink recipe.
- **Cursor** — add the marketplace, then `/plugin install <name>`. Entries point at source
`./plugins/<name>`; Cursor reads `SKILL.md` + `.md` agents from source directly.
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
- **Antigravity** — no one-step-from-URL install (the lean tradeoff). Clone the repo, then
`make generate HARNESS=antigravity` and either `agy plugin install .antigravity/plugins/<name>`
per plugin, or `make install-antigravity` to symlink every generated plugin into
`~/.gemini/antigravity-cli/plugins/` (agy's config dir) at once.
- **OpenCode** — no one-step-from-URL install. Clone the repo, then `make install-opencode`
(runs generate + symlinks `.opencode/``~/.config/opencode/`).
- **Pi** — no one-step-from-URL install. Clone the repo, then `make install-pi` symlinks every
generated skill directory, prompt template, and agent into `~/.pi/agent/` (override with
`PI_CODING_AGENT_DIR`). Skills and prompts can also be installed as a package with
`pi install /path/to/agents/.pi`; agents need the symlink route because Pi packages have no
agents resource, and they only work with the reference `subagent` extension or a compatible
package. Running `pi` inside the clone also works. Pi asks to trust the project and then reads
`.pi/` directly. Pick one route. If you install globally with `make install-pi` and also run
`pi` inside the clone, Pi sees every skill twice and warns on each name.
feat: support gh skill and npx skills installers (#693) * feat: support gh skill and npx skills installers Both Agent Skills installers already discover every skill in this repo through the plugins/<plugin>/skills/<skill>/ layout, so support is documentation plus a real-CLI gate rather than a layout change. - docs: README quick start block; docs/harnesses.md "Skills-only installers" (selectors, install paths, release-freeze rule, local-checkout caveat); docs/agent-skills.md pointer; AGENTS.md bullet; authoring note that a skill's name must equal its directory name - smoke tests: gh skill local discovery, gh skill publish --dry-run spec validation, npx skills discovery, and skill-name uniqueness across plugins (both installers install under the bare skill name) - ci: set up Node for npx and fail loudly if the runner's gh predates gh skill - rename database-design/skills/postgresql to postgresql-table-design so the directory matches the frontmatter name, the one spec error the dry-run found * test: type the smoke helper env as dict[str, str] * fix(database-design): correct three PostgreSQL facts in postgresql-table-design Flagged by review once the rename made the file appear new. Pre-existing content, separate commit so it can be dropped if the rename should stay content-free. - UNIQUE NULLS NOT DISTINCT (...) places the clause before the column list - the RLS example used current_user_id(), which is not a built-in; use current_user or an app-set setting - foreign keys on partitioned tables work from PG11 (from) and PG12 (referencing); triggers are only the pre-11 fallback * refactor(database-design): split postgresql-table-design under the Codex 8 KB cap SKILL.md keeps the rules, decision points, a When to Use section, and the three quick-start DDL examples (7.6 KB body, 126 lines). The full data-type catalog, table types, row-level security, constraint and index notes, partitioning DDL, workload patterns, generated columns, extensions, and JSONB indexing move verbatim to references/details.md. The "avoid these types" list becomes a table. Clears the SKILL_OVER_CODEX_CAP warning this skill carried. Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs * test(smoke): capture opencode agent list through a file, tighten npx parser `opencode agent list` prints every agent's expanded permission array (about 330k lines) and exits before a pipe drains, so pipe capture intermittently lost the alphabetically last agents (measured: 1 in 5 runs short via subprocess pipes, 0 in 5 via a file sink). Both OpenCode tests now write stdout to a file. The npx parser matches only the `│ <name>` line shape. Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs * docs,ci: cover the skills installers everywhere install is documented - docs/usage.md, docs/plugins.md, docs/architecture.md, ARCHITECTURE.md, README multi-harness section, docs/harnesses.md supported table, and CONTRIBUTING's portability checklist now name gh skill and npx skills - docs/round-trip-results.md gains summary rows and a reproduce recipe - docs/authoring.md: skill directory names are identities for installers and adapters alike, so a rename is user-visible - CI installs GitHub CLI from the official apt repository when the runner build predates `gh skill`, so the job no longer depends on the image Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs
2026-09-01 18:28:17 -04:00
## Skills-only installers
`gh skill` (GitHub CLI 2.90+) and `npx skills` ([vercel-labs/skills](https://github.com/vercel-labs/skills))
install Agent Skills into any supported agent straight from GitHub. Both discover every
`plugins/<plugin>/skills/<skill>/` directory in this repo without a clone, a marketplace, or a
generate step. They carry skills only: no agents, commands, or hooks.
```bash
# gh skill: lists as `[plugins] <plugin>/<skill>`, selects by bare skill name or exact path
gh skill install wshobson/agents # interactive browse
gh skill install wshobson/agents python-testing-patterns
gh skill install wshobson/agents plugins/python-development/skills/python-testing-patterns # exact path skips the tree walk
gh skill install wshobson/agents --all --agent claude-code --scope user
gh skill install wshobson/agents python-testing-patterns --pin <sha>
# npx skills: lists and selects by bare skill name
npx skills add wshobson/agents --list
npx skills add wshobson/agents --skill python-testing-patterns -a claude-code
npx skills add wshobson/agents --all -g
```
Gotchas:
- **Both install under the bare skill name** (`<agent>/skills/<skill>/`). The `<plugin>/` prefix
in `gh skill` listings is display only; `python-development/python-testing-patterns` is not a
valid selector, `python-testing-patterns` and the exact `plugins/...` path are. Skill directory
names are unique across plugins and `make smoke-test` keeps them that way; a duplicate would
collide on install.
- **`gh skill` installs from the latest GitHub release when one exists**, and from `main` only
when the repo has none. This repo publishes no releases, so installs track `main`. Creating a
release would freeze `gh skill` installs at that tag until the next one.
- **Local checkouts.** After `make generate-all`, `npx skills add ./agents` also walks the
gitignored `.codex/`, `.opencode/`, `.copilot/`, `.antigravity/` and `.pi/` trees and lists
their copies. Install from the GitHub source instead, or use `gh skill install . --from-local`,
which skips hidden directories.
feat: support gh skill and npx skills installers (#693) * feat: support gh skill and npx skills installers Both Agent Skills installers already discover every skill in this repo through the plugins/<plugin>/skills/<skill>/ layout, so support is documentation plus a real-CLI gate rather than a layout change. - docs: README quick start block; docs/harnesses.md "Skills-only installers" (selectors, install paths, release-freeze rule, local-checkout caveat); docs/agent-skills.md pointer; AGENTS.md bullet; authoring note that a skill's name must equal its directory name - smoke tests: gh skill local discovery, gh skill publish --dry-run spec validation, npx skills discovery, and skill-name uniqueness across plugins (both installers install under the bare skill name) - ci: set up Node for npx and fail loudly if the runner's gh predates gh skill - rename database-design/skills/postgresql to postgresql-table-design so the directory matches the frontmatter name, the one spec error the dry-run found * test: type the smoke helper env as dict[str, str] * fix(database-design): correct three PostgreSQL facts in postgresql-table-design Flagged by review once the rename made the file appear new. Pre-existing content, separate commit so it can be dropped if the rename should stay content-free. - UNIQUE NULLS NOT DISTINCT (...) places the clause before the column list - the RLS example used current_user_id(), which is not a built-in; use current_user or an app-set setting - foreign keys on partitioned tables work from PG11 (from) and PG12 (referencing); triggers are only the pre-11 fallback * refactor(database-design): split postgresql-table-design under the Codex 8 KB cap SKILL.md keeps the rules, decision points, a When to Use section, and the three quick-start DDL examples (7.6 KB body, 126 lines). The full data-type catalog, table types, row-level security, constraint and index notes, partitioning DDL, workload patterns, generated columns, extensions, and JSONB indexing move verbatim to references/details.md. The "avoid these types" list becomes a table. Clears the SKILL_OVER_CODEX_CAP warning this skill carried. Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs * test(smoke): capture opencode agent list through a file, tighten npx parser `opencode agent list` prints every agent's expanded permission array (about 330k lines) and exits before a pipe drains, so pipe capture intermittently lost the alphabetically last agents (measured: 1 in 5 runs short via subprocess pipes, 0 in 5 via a file sink). Both OpenCode tests now write stdout to a file. The npx parser matches only the `│ <name>` line shape. Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs * docs,ci: cover the skills installers everywhere install is documented - docs/usage.md, docs/plugins.md, docs/architecture.md, ARCHITECTURE.md, README multi-harness section, docs/harnesses.md supported table, and CONTRIBUTING's portability checklist now name gh skill and npx skills - docs/round-trip-results.md gains summary rows and a reproduce recipe - docs/authoring.md: skill directory names are identities for installers and adapters alike, so a rename is user-visible - CI installs GitHub CLI from the official apt repository when the runner build predates `gh skill`, so the job no longer depends on the image Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs
2026-09-01 18:28:17 -04:00
- **Spec gate.** `gh skill publish --dry-run` validates every SKILL.md against the
[agentskills.io spec](https://agentskills.io/specification): name pattern, name equal to the
directory name, required frontmatter. `make smoke-test` runs it, plus discovery through both
CLIs, against the real binaries.
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
## Regenerating
The committed registries point at source; the transformed trees are regenerated on demand.
Contributors must run `make generate-all` before committing source changes — CI fails on drift
of the committed registries.
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
```bash
make generate HARNESS=codex
make generate HARNESS=cursor
make generate HARNESS=opencode
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
make generate HARNESS=antigravity
make generate HARNESS=pi
# Or all at once (run before committing source changes):
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
make generate-all
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
# Optional global installs:
make install-opencode
make uninstall-opencode
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
make install-antigravity
make uninstall-antigravity
make install-pi
make uninstall-pi
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
```
## External Pensyve integrations
The Claude Code marketplace includes Pensyve as an external `git-subdir` plugin.
For generated harnesses, use Pensyve's upstream harness-native integration:
| Harness | Upstream integration |
|---|---|
| Claude Code | `https://github.com/major7apps/pensyve.git`, path `integrations/claude-code` |
| Codex CLI | `integrations/codex-plugin` |
| Cursor | `integrations/cursor` |
| OpenCode | `integrations/opencode-plugin` |
| Copilot | `.copilot/` (repo-level) or `~/.copilot/` (global install via `make install-copilot`) |
## External HOL Guard integration
The Claude Code marketplace includes HOL Guard as an external `git-subdir` plugin from
`https://github.com/hashgraph-online/hol-guard-plugin.git`, path `distributions/wshobson-agents`.
The reviewed payload exposes the portable `hol-guard` and `plugin-scanner` skills and
keeps decisioning local by default. Guard Cloud is neither required nor promoted. This
marketplace entry is a Claude Code discovery surface only; it does not add HOL Guard to the
generated Codex, Cursor, OpenCode, Antigravity, Copilot, or Pi outputs.
The reviewed payload is pinned to commit `43b2dda59e9f07057c52e69fd7426188faae1488` and installs the exact local CLI versions
`hol-guard==2.2.119` and `plugin-scanner==2.2.119`, with user approval required before
installation. For a reviewed payload update, advance the marketplace `sha` and matching
marketplace/external manifest versions together.
When the user explicitly requests protection, the local HOL Guard runtime can modify
supported harness hook/settings configuration. Generated harness outputs in this repository
do not vendor or rewrite the external HOL Guard payload.
## Global install
OpenCode, Copilot, Antigravity, and Pi support installing generated artifacts globally for
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
user-level discovery:
```bash
make install-opencode # symlink .opencode/ → ~/.config/opencode/
make uninstall-opencode
make install-copilot # symlink .copilot/ → ~/.copilot/
make uninstall-copilot
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
make install-antigravity # symlink each .antigravity/plugins/<p>/ → ~/.gemini/antigravity-cli/plugins/<p>/
make uninstall-antigravity
make install-pi # symlink each .pi/ skill, prompt, and agent → ~/.pi/agent/
make uninstall-pi
# Force-replace conflicting symlinks:
make install-copilot FORCE=1
feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669) * feat(antigravity): add Google Antigravity CLI harness adapter (#644) * feat(antigravity)!: retire Gemini CLI harness (#644) Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter, validator, and doc-gardener drift pairs, and removes the committed gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local build-only skills/, agents/, commands/ trees they produced. The Google Antigravity CLI (agy), added in the prior commit, is now the harness those users should migrate to: native plugins at .antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file redirect needed), with its own marketplace, tier-based model aliases (pro/flash/inherit), and `make install-antigravity` for global installs. - tools/adapters/gemini.py deleted; capabilities.py/generate.py/ validate_generated.py/doc_gardener.py/Makefile lose their Gemini dispatch, targets, and drift pairs. - Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip, TestGeminiSmoke removed along with now-unused imports. - CI: cli-smoke-test now installs the Antigravity CLI instead of the Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the legacy top-level skills/agents/commands/ output. - Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md, docs/round-trip-results.md, docs/plugin-eval.md, README.md, CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as the fifth harness in place of Gemini. BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated, or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md consumers should switch to `make generate HARNESS=antigravity` and `make install-antigravity`. * fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644) Address CodeRabbit + Codex review feedback on PR #669: - antigravity.py: mirror every skill support file (scripts/, assets/, resources/, examples/), not just references/ — matches OpenCode's pattern. Excludes hidden files. - antigravity.py: translate $ARGUMENTS to {{args}} in place within command bodies; only append a trailing {{args}} block when the source has none. - antigravity.py: serialize frontmatter with YAML-safe scalar quoting and preserve dict-valued fields (e.g. metadata) as nested mappings instead of stringifying the Python repr. - validate_generated.py: guard against non-dict plugin.json and non-string command description/prompt fields so malformed input is reported as a finding instead of crashing with AttributeError/TypeError. - Sync stale plugin/agent/skill/command counts in claude-code-review.yml and ARCHITECTURE.md to the canonical 92/202/181/105. - CONTRIBUTING.md: add the missing Antigravity entry to the six-harness portability checklist. - docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct the TodoWrite/hooks support matrix for Antigravity. - harness_portability.py: fix the bare-model-alias comment — Antigravity maps aliases to tier values, not full model IDs. - .cursor/rules/020-agent-skill-authoring.mdc (source in tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but does support Task-spawn and hooks via native equivalents. - README.md: narrow the Pensyve integration claim to the harnesses it actually covers. - .gitignore: document that Antigravity follows OpenCode's clone+generate install pattern; give .antigravity/ its own comment. - Extend adapter and validator test suites for both fixes. * fix(antigravity): quote comma-containing items in flow-style YAML lists CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't treat ',' or ']' as needing quotes, so a list item containing a comma (e.g. tags: ["foo, bar", baz]) split into two list entries on round-trip since flow sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list items specifically (top-level scalars don't need this — commas are only ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
make install-antigravity FORCE=1
make install-pi FORCE=1
```
> Copilot discovers agents from `.copilot/agents/` and skills from `.copilot/skills/` at the repo level, and from `~/.copilot/agents/` and `~/.copilot/skills/` at the user level. The adapter emits to `.copilot/`; use `make install-copilot` for user-level discovery.
feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541) * feat(adapters): multi-harness framework + harness_portability eval dimension Turn this Claude Code plugin marketplace into a generic agentic-harness marketplace. Adapters under tools/adapters/ emit harness-native artifacts for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source. Source-of-truth stays under plugins/ — Claude Code is unchanged. Framework (tools/adapters/): - base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author - capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS, MODEL_ALIASES, resolve_model() with explicit warnings - codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line table-of-contents). Fence-aware body splitter, _utf8_safe_cut for multibyte safety, _yaml_scalar with reserved-word + special-char quoting. Skill/command name collision detection (and second-order __cmd fallback). - cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars (no false positives on colons in description body). _normalize_author handles dict, npm-style strings, and author lists. - opencode.py — transpiles agents to .opencode/agents/<id>.md with mode:subagent + permission: deny-everything-else block (skill/task always allowed as base capabilities — Claude's implicit defaults). - gemini.py — emits native skills/, agents/, and commands/ at extension root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS. CLI + tooling: - tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`, with --clean (containment-guarded; case-insensitive on Darwin/Win32), --prune (orphan removal across all per-harness output trees), --strict (warnings fail), per-plugin error aggregation, refuses --clean --plugin (would silently wipe other plugins' artifacts). - tools/validate_generated.py — structural validation across all four harness outputs. Codex 8KB cap → error. _extract_permission_block correctly handles nested permission keys (column-0 only). - tools/doc_gardener.py — recurring drift detection per OpenAI harness- engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error), MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped output sorted by severity. plugin-eval (extends existing framework): - New harness_portability dimension (6% weight, rebalanced from existing static sub-scores). Surfaces non-portable patterns with concrete remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS, CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS. - _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool names, case-insensitive on the leading article. - Findings do NOT also feed anti_pattern_penalty (no double-counting). Documentation: - Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each, table-of-contents pattern per OpenAI harness-engineering post) - docs/harnesses.md — capability matrix, graceful-degradation table, generated output paths - docs/authoring.md — portable-content style guide (tools, models, collision rules, fence-respect) - docs/round-trip-results.md — real-CLI verification recipes (OpenCode discovers 193 subagents, Gemini extensions validate passes, Codex TOMLs all parse) - CONTRIBUTING.md — new file pointing at docs/authoring.md - README.md — rewritten for multi-harness (145 lines, was 460) - CLAUDE.md — trimmed to 60-line table-of-contents - GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously) Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip verified for OpenCode, Gemini, and Codex (TOML parses). Replaces tools/generate_gemini_commands.py with the unified CLI. * refactor(skills): extract detail to references/details.md (~75 skills) Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing + quick start), detailed templates and worked examples move to references/details.md (loaded on demand by the agent). Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this change, ~90 skills exceeded that cap and would silently break on Codex. The progressive-disclosure pattern is also Anthropic's documented recommendation for token efficiency — Claude Code reads references/ files on demand when the body navigation says to. What's extracted, by pattern: - Pass 1 (## Templates section): 19 skills — full template libraries moved to references/details.md - Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills - Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills - Conservative re-extraction for 8 skills that got over-reduced — kept ~6-7 KB inline (most of the quick-start tier) plus references/ overflow What stays inline (SKILL.md navigation tier): - description: frontmatter (triggering — unchanged for all skills) - ## When to Use This Skill / ## Core Concepts / ## Quick Start - ## Best Practices / ## Troubleshooting / ## See Also wrap-ups - A pointer note ("see references/details.md") so the agent knows where to look for detail What goes to references/details.md (detail tier, on-demand load): - ## Templates (full code template libraries) - ## Implementation Patterns / ## Advanced Patterns (deep examples) - Mid-skill walkthroughs that exceed the inline budget Also in this commit: - plugins/brand-landingpage description trimmed from 958→543 chars (preserves trigger phrasing, drops verbose example-quote list) Net effect: - SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction) - All triggers unchanged — discovery behavior identical across harnesses - 75 new references/details.md files with the extracted content - Same depth of guidance, loaded progressively Remaining 10 oversized skills are complex multi-section docs (e.g. postgresql, code-review-excellence, evaluation-methodology) that need per-skill manual judgment — flagged by `make garden` for future work. * chore: bump all plugin versions (multi-harness release) Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the multi-harness adapter framework addition. The external git-subdir entry (qa-orchestra) is unaffected — its version is governed by its upstream repo. * fix(opencode): preserve explicit tools:[] + word-boundary subtask match Addresses two Codex review findings on PR #541. ## P1 — `tools: []` silently upgraded to permissive (privilege escalation) Before: `_build_permission_block` returned `{}` for any empty list, which omits the `permission:` block entirely from the emitted agent. An author who explicitly wrote `tools: []` to lock down an advisory-only agent got an UNRESTRICTED agent in OpenCode. Affected agent in this tree: `plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`. Fix: `_build_permission_block` now takes a `has_tools_field` flag so the caller can distinguish "tools: key missing" (Claude default permissive) from "tools: []" (explicit lock-down). The lock-down case emits a deny-everything block that allows ONLY the base capabilities (skill, task) that Claude Code always grants implicitly. Verified against the real arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list: deny, task/skill: allow. ## P2 — `"agent" in cmd.body.lower()` false-positives on substrings Before: a command body containing `PerformanceReviewAgent` (class name in a code snippet) or `useragent` triggered `subtask: true`, changing runtime behavior based on incidental text. Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b` (case-insensitive). Tests confirm the substring `PerformanceReviewAgent` no longer fires, while a real "spawn a subagent" sentence still does. ## Tests 3 new regression tests in tools/tests/test_adapters.py: - `test_explicit_empty_tools_yields_locked_permission_block` (P1) - `test_missing_tools_field_yields_no_permission_block` (P1 boundary) - `test_subtask_inference_word_boundary` (P2) 184 total tests pass (was 181). OpenCode round-trip still discovers all 193 subagents; arm-cortex-expert agent is now properly locked down. * test: behavioral verification + CI gates for multi-harness pipeline Adds three layers of automated verification that pure-Python parser tests miss, plus the CI jobs that turn them into hard gates. Catches the kinds of issues that previously only surfaced when a real user installed the marketplace and tried to use it. ## test_real_world.py — real-source structural tests Runs against the actual `plugins/` tree (not synthetic fixtures). Catches issues that only appear on real content: - every marketplace entry resolves to a plugins/<name>/ dir - every local plugin dir appears in marketplace.json - marketplace.json version == per-plugin plugin.json version (catches drift) - every plugin loads via load_plugin() without error - no plugin name contains `__` (adapter namespace separator) - every agent has name + description; every skill has a trigger phrase (same regex plugin_eval's MISSING_TRIGGER check uses) - no agent name collides with Codex built-ins - every refactored skill (with `references/details.md`) has: - meaningful detail content (>=500 B in details.md) - a pointer to references/ in the SKILL.md body - a navigation-tier heading preserved (When to Use, Overview, etc.) - body >= 600 B (not a stub) - every plugin.json has name + version matching the dir This test pass found and fixed three real defects before commit: - ship-mate/skills/scan: description had no trigger phrase ("Use when…") - reverse-engineering/skills/memory-forensics: nav-tier section lost during extraction - reverse-engineering/skills/binary-analysis-patterns: same All three are now fixed (preserved trigger phrasing, added When-to-Use sections back to the skills my extraction over-trimmed). ## test_round_trip.py — generate→parse→verify CI runs this AFTER `make generate-all`. Catches generation-time regressions: - OpenCode/Codex/Gemini agent counts match source agent count (no skips) - every Codex SKILL.md under 8 KB (the cap that would silently truncate) - every Codex agent TOML has required fields + valid sandbox_mode - every OpenCode agent has mode in {primary,subagent,all} and provider-prefixed model - locked agents (source `tools: []`) emit proper deny-everything permission block with skill/task allow (regression guard for PR-541 P1) - every Gemini @{path} injection resolves to a real source file - every Gemini command TOML has prompt + {{args}} placeholder - every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within 150-line cap - Cursor marketplace + per-plugin manifests cover all local plugins - .cursor/rules/*.mdc only use the 3 documented frontmatter keys ## test_cli_smoke.py — real-CLI subprocess tests Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code) against the generated artifacts. Catches CLI-level issues pure-Python parsing can't see: schema-loader drift, plugin-discovery bugs, version incompatibilities. - `opencode agent list` — must succeed AND discover every source agent (currently 191 + 2 OpenCode built-ins) - `gemini extensions validate <repo>` — must return success - `codex doctor` — must report healthy install - every Codex agent TOML must parse with stdlib `tomllib` - `claude --version` — sanity check the Claude Code CLI loads - marketplace.json must have owner + metadata.version for Claude Code's loader Per-CLI tests skip gracefully when the binary isn't on PATH, so local devs only exercise what they have installed. CI installs OpenCode + Gemini and turns those skips into hard gates. ## Makefile + CI - `make test` — full pytest suite (plugin-eval + tools/tests/) - `make smoke-test` — generates if needed, then runs real-CLI smoke tests - `.github/workflows/validate.yml` extended with: - `tools-tests` job — runs pytest tools/tests/ - `multi-harness-generate` job — `make generate-all && make validate STRICT=1 && make garden`, uploads generated artifacts on every run - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py ## Test counts - Before: 184 tests - After: 386 tests (parameterized real-source tests over all 82 plugins) - All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0 + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
## See also
- [`authoring.md`](authoring.md) — portable-content style guide for plugin authors
- [`architecture.md`](architecture.md) — overall design principles
- [`plugin-eval.md`](plugin-eval.md) — the `harness_portability` scoring dimension