32 Commits

Author SHA1 Message Date
Seth Hobson 4236bb91f8 ci: rebuild the Claude Code review workflow from scratch (#708)
Pins anthropics/claude-code-action to the v1.0.223 release commit (the old pin
was from May), moves the review model to claude-opus-5, adds a concurrency
group so superseded runs stop, uses a sticky summary comment, and rewrites the
review prompt with the current harness list, the generated-versus-committed
tree rules, and no hard-coded component counts. The header explains the two
things that make this check look broken: the action refuses to run when a PR
edits this file, and the Bun directory-mismatch message is noise.

Claude-Session: https://claude.ai/code/session_01DZazzWVyb8MxPCuLC1w5Qo
2026-09-13 10:43:53 -04:00
Seth Hobson 1533e7787b feat: add the Pi harness (pi.dev) as a sixth generated target (#707)
* feat(pi): register the Pi harness in capabilities, model aliases, and tool maps

* feat(pi): add the Pi adapter (skills, prompt templates, subagent-extension agents)

* feat(pi): generate the .pi tree via make generate HARNESS=pi

* feat(pi): validate the generated .pi tree

* feat(pi): gardener reports stale .pi artifacts

* feat(pi): make install-pi symlinks the .pi tree into the Pi config dir

* test(pi): real-CLI smoke tests via token-free slash expansion; install pi in CI

* test(pi): override every Anthropic credential variable in the smoke env

* docs: add the Pi harness everywhere the harness list appears

* fix(pi): close the final-review findings (CI type-check, base URL override, validator gap, docs)

* docs(pi): align harness-count wording and the no-billing explanation

Claude-Session: https://claude.ai/code/session_01DZazzWVyb8MxPCuLC1w5Qo

* fix(pi): scope clean and prune to the adapter-owned .pi subtrees; tighten filename checks

Claude-Session: https://claude.ai/code/session_01DZazzWVyb8MxPCuLC1w5Qo
2026-09-13 09:51:18 -04:00
Seth Hobson 062ca77305 ci: pin the OpenCode installer version in the smoke job (#703)
The unpinned installer asks api.github.com for the latest release and fails on
shared runners when the unauthenticated rate limit is hit. Dependabot PR #698
failed this way on 2026-09-07 while the same job passed seconds earlier. Pin
to 1.18.30, the current release, so the smoke job installs the same CLI every
run.
2026-09-13 07:37:14 -04:00
Seth Hobson c80732de69 feat: support gh skill and npx skills installers (#693)
* feat: support gh skill and npx skills installers

Both Agent Skills installers already discover every skill in this repo through
the plugins/<plugin>/skills/<skill>/ layout, so support is documentation plus a
real-CLI gate rather than a layout change.

- docs: README quick start block; docs/harnesses.md "Skills-only installers"
  (selectors, install paths, release-freeze rule, local-checkout caveat);
  docs/agent-skills.md pointer; AGENTS.md bullet; authoring note that a skill's
  name must equal its directory name
- smoke tests: gh skill local discovery, gh skill publish --dry-run spec
  validation, npx skills discovery, and skill-name uniqueness across plugins
  (both installers install under the bare skill name)
- ci: set up Node for npx and fail loudly if the runner's gh predates gh skill
- rename database-design/skills/postgresql to postgresql-table-design so the
  directory matches the frontmatter name, the one spec error the dry-run found

* test: type the smoke helper env as dict[str, str]

* fix(database-design): correct three PostgreSQL facts in postgresql-table-design

Flagged by review once the rename made the file appear new. Pre-existing
content, separate commit so it can be dropped if the rename should stay
content-free.

- UNIQUE NULLS NOT DISTINCT (...) places the clause before the column list
- the RLS example used current_user_id(), which is not a built-in; use
  current_user or an app-set setting
- foreign keys on partitioned tables work from PG11 (from) and PG12
  (referencing); triggers are only the pre-11 fallback

* refactor(database-design): split postgresql-table-design under the Codex 8 KB cap

SKILL.md keeps the rules, decision points, a When to Use section, and the
three quick-start DDL examples (7.6 KB body, 126 lines). The full data-type
catalog, table types, row-level security, constraint and index notes,
partitioning DDL, workload patterns, generated columns, extensions, and JSONB
indexing move verbatim to references/details.md. The "avoid these types" list
becomes a table. Clears the SKILL_OVER_CODEX_CAP warning this skill carried.

Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs

* test(smoke): capture opencode agent list through a file, tighten npx parser

`opencode agent list` prints every agent's expanded permission array (about
330k lines) and exits before a pipe drains, so pipe capture intermittently
lost the alphabetically last agents (measured: 1 in 5 runs short via
subprocess pipes, 0 in 5 via a file sink). Both OpenCode tests now write
stdout to a file. The npx parser matches only the `│    <name>` line shape.

Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs

* docs,ci: cover the skills installers everywhere install is documented

- docs/usage.md, docs/plugins.md, docs/architecture.md, ARCHITECTURE.md,
  README multi-harness section, docs/harnesses.md supported table, and
  CONTRIBUTING's portability checklist now name gh skill and npx skills
- docs/round-trip-results.md gains summary rows and a reproduce recipe
- docs/authoring.md: skill directory names are identities for installers
  and adapters alike, so a rename is user-visible
- CI installs GitHub CLI from the official apt repository when the runner
  build predates `gh skill`, so the job no longer depends on the image

Claude-Session: https://claude.ai/code/session_01LjJmzuuxXSwGNEYdBvsmFs
2026-09-01 18:28:17 -04:00
Seth Hobson 554237f751 ci: add MLOps lab pipeline (W&B + HF + reusable mlops workflow) (#692)
* ci: add MLOps lab pipeline (W&B + HF + reusable mlops workflow)

- mlops.yml: reusable CI baseline (ruff+ty+pytest) and model-release
  job that pushes model dirs to the Hugging Face major7 org, triggered
  by model/* tags or manual dispatch
- eval-report.yml: add concurrency guard (no overlapping runs on same
  ref) and a log_wandb dispatch input that pushes per-plugin eval
  metrics to W&B m7/major7-lab
- docs/mlops.md: lab reference covering env vars, training, eval,
  release, W&B project layout, offline mode, troubleshooting

* ci: fix mlops.yml — env context unavailable in job-level name expression

* ci: address PR review — injection, dirname, HF API, release gating, pinning

- mlops.yml resolve step: pass inputs.hf_target / inputs.model_path via
  env: instead of interpolating into bash (zizmor template-injection)
- mlops.yml: tag fallback now uses the model dir itself instead of
  dirname() (model/my-model-v1 -> my-model-v1/, not the repo root);
  add an explicit refusal for . / '' / / as release source
- mlops.yml: HfApi.upload_file now uses path_or_fileobj= /
  path_in_repo= (correct huggingface_hub 0.35 signature)
- mlops.yml: model-release now needs: [lint, test] and skips only when
  the gate succeeded or was skipped, so a broken tree cannot be
  published on a model/* tag push
- mlops.yml: summary step reads resolve outputs via env:, not
  expression interpolation
- eval-report.yml: pin actions/checkout + astral-sh/setup-uv +
  actions/upload-artifact to the repo's established SHAs and add
  persist-credentials: false (was @v4/@v5 + default credentials)
- AGENTS.md: add docs/mlops.md to the Map index
2026-09-01 10:48:41 -04:00
Seth Hobson 775cc037a2 fix(ci): accept ref/sha pins on git-subdir marketplace entries (#679)
Claude Code's git-subdir plugin source schema is url, path, ref?, sha?, with
sha as the effective pin. The marketplace validation step rejected sha as an
unsupported key, which blocks any pinned external entry. Allow ref and sha and
check that sha is a 40-character hex commit.
2026-08-23 15:02:29 -04:00
Seth Hobson ca2287c776 chore(tooling): close the ruff scope gap and correct two stale docs (#677)
* chore(tooling): close the ruff scope gap and correct two stale docs

Pre-existing issues found while working on #676.

CI checked ruff against five explicit paths under tools/, so four files outside
that list drifted out of format: check_agent_name_collisions.py, both installers
and yt-design-extractor.py. Ruff now runs over all of tools/, those four are
formatted, and two real B007 findings in yt-design-extractor are fixed. ty keeps
explicit paths because yt-design-extractor imports optional OCR dependencies that
only `make install-ocr` provides.

default_config_dir in install_opencode.py assigned os.environ, a Mapping, back
over a parameter annotated dict[str, str] | None, which produced five ty errors
and left the value possibly-None to the reader. It now binds a separate
Mapping-typed local.

New `make lint` and `make format` targets run the checks from
plugins/plugin-eval/ the way CI does. Running ruff from the repo root silently
falls back to line-length 88 and disagrees with CI, which is easy to hit and
wastes time; the targets remove the choice. I tried a root ruff.toml first and
reverted it: it changes isort's first-party resolution for tools/ and makes CI
fail with I001.

AGENTS.md said generated artifacts are committed. Only the small registries are;
the transformed trees under .codex/, .opencode/, .copilot/ and .antigravity/ are
gitignored. CONTRIBUTING listed `make garden STRICT=1` as a gate while main
carries ten SKILL_OVER_CODEX_CAP warnings that make it always fail.

* fix(make): pin lint tools to the dev extra

Review feedback on PR #677. ruff and ty are declared under the dev optional
dependency, and the new targets called uv run without it. On a clean checkout
that does not fail as expected: uv provisions an unpinned ruff on the fly, so
make lint ran 0.16.4 while CI, which does uv sync --all-extras, ran the locked
0.16.3. A gate whose purpose is matching CI was free to disagree with it about
formatting.

--extra dev on all five invocations. Verified the version now resolves to
0.16.3.

* docs(install-opencode): describe the empty-mapping invariant

Review feedback on PR #677. The comment still described the old
dict[str, str] | None annotation and missed the behavior that actually matters.

The old `env or os.environ` fell back whenever env was falsy, so an explicitly
empty mapping silently read the real environment. The explicit None check keeps
it empty, which is what install_copilot.default_config_dir already did, so the
two now agree. Two tests pin both paths.
2026-08-22 21:55:32 -04:00
Seth Hobson 367cb6a4a1 feat(antigravity)!: migrate from Gemini CLI to Google Antigravity CLI harness (#669)
* feat(antigravity): add Google Antigravity CLI harness adapter (#644)

* feat(antigravity)!: retire Gemini CLI harness (#644)

Google deprecated the Gemini CLI in May 2026. This drops the Gemini adapter,
validator, and doc-gardener drift pairs, and removes the committed
gemini-extension.json / .gemini/ / GEMINI.md artifacts and the local
build-only skills/, agents/, commands/ trees they produced.

The Google Antigravity CLI (agy), added in the prior commit, is now the
harness those users should migrate to: native plugins at
.antigravity/plugins/<name>/, reading AGENTS.md directly (no context-file
redirect needed), with its own marketplace, tier-based model aliases
(pro/flash/inherit), and `make install-antigravity` for global installs.

- tools/adapters/gemini.py deleted; capabilities.py/generate.py/
  validate_generated.py/doc_gardener.py/Makefile lose their Gemini
  dispatch, targets, and drift pairs.
- Tests: TestGeminiAdapter, TestGeminiValidator, TestGeminiRoundTrip,
  TestGeminiSmoke removed along with now-unused imports.
- CI: cli-smoke-test now installs the Antigravity CLI instead of the
  Gemini CLI; multi-harness-generate uploads .antigravity/ instead of the
  legacy top-level skills/agents/commands/ output.
- Docs (AGENTS.md, ARCHITECTURE.md, docs/harnesses.md, docs/authoring.md,
  docs/round-trip-results.md, docs/plugin-eval.md, README.md,
  CONTRIBUTING.md, issue/PR templates) swept to describe Antigravity as
  the fifth harness in place of Gemini.

BREAKING CHANGE: the Gemini CLI harness is no longer generated, validated,
or supported. Existing gemini-extension.json / .gemini/ / GEMINI.md
consumers should switch to `make generate HARNESS=antigravity` and
`make install-antigravity`.

* fix(antigravity): mirror skill support dirs, translate $ARGUMENTS, harden validator (#644)

Address CodeRabbit + Codex review feedback on PR #669:

- antigravity.py: mirror every skill support file (scripts/, assets/,
  resources/, examples/), not just references/ — matches OpenCode's pattern.
  Excludes hidden files.
- antigravity.py: translate $ARGUMENTS to {{args}} in place within command
  bodies; only append a trailing {{args}} block when the source has none.
- antigravity.py: serialize frontmatter with YAML-safe scalar quoting and
  preserve dict-valued fields (e.g. metadata) as nested mappings instead of
  stringifying the Python repr.
- validate_generated.py: guard against non-dict plugin.json and non-string
  command description/prompt fields so malformed input is reported as a
  finding instead of crashing with AttributeError/TypeError.
- Sync stale plugin/agent/skill/command counts in claude-code-review.yml and
  ARCHITECTURE.md to the canonical 92/202/181/105.
- CONTRIBUTING.md: add the missing Antigravity entry to the six-harness
  portability checklist.
- docs/authoring.md: add fable to ARCHITECTURE.md's valid model list; correct
  the TodoWrite/hooks support matrix for Antigravity.
- harness_portability.py: fix the bare-model-alias comment — Antigravity maps
  aliases to tier values, not full model IDs.
- .cursor/rules/020-agent-skill-authoring.mdc (source in
  tools/adapters/cursor_rules/, regenerated): Antigravity lacks TodoWrite but
  does support Task-spawn and hooks via native equivalents.
- README.md: narrow the Pensyve integration claim to the harnesses it
  actually covers.
- .gitignore: document that Antigravity follows OpenCode's clone+generate
  install pattern; give .antigravity/ its own comment.
- Extend adapter and validator test suites for both fixes.

* fix(antigravity): quote comma-containing items in flow-style YAML lists

CodeRabbit follow-up on the frontmatter YAML-safety fix: _yaml_scalar() didn't
treat ',' or ']' as needing quotes, so a list item containing a comma (e.g.
tags: ["foo, bar", baz]) split into two list entries on round-trip since flow
sequences use ',' as the item delimiter. Add _yaml_flow_scalar() for list
items specifically (top-level scalars don't need this — commas are only
ambiguous inside [...]). Regression test added.
2026-08-18 11:42:59 -04:00
Seth Hobson 7f38372210 fix(test): keep real-CLI smoke tests out of the ordinary test target (#664) (#667)
* fix(test): keep real-CLI smoke tests out of the ordinary test target (#664)

* docs(architecture): fix stale make-test composition claim

make test now excludes tools/tests/test_cli_smoke.py (#664); update
the quality-gates description so it no longer claims real-CLI smoke
tests run as part of make test.
2026-08-18 10:31:27 -04:00
Seth Hobson 608c3840ca feat: native plugin-install for Codex/Cursor/Gemini + CLAUDE.md→AGENTS.md symlink
Add lean, native plugin-install entry points so each harness's own plugin
manager can install this marketplace (mirroring obra/superpowers) — committing
only small JSON registries, not duplicated skill/agent content trees.

- Codex: committed marketplace registry (.agents/plugins/marketplace.json) +
  per-plugin manifests (plugins/*/.codex-plugin/plugin.json). Entries point at
  source ./plugins/<name>; Codex reads SKILL.md directly. Transformed
  .codex/skills|agents stay gitignored.
- Cursor: commit the existing .cursor-plugin/ marketplace + .cursor/rules/
  (these already point at source plugins/).
- Gemini: gemini-extension.json already committed (contextFileName: AGENTS.md);
  transformed trees stay gitignored (install via clone + make generate).
- OpenCode: unchanged — install via `make install-opencode` (transformed tree
  stays gitignored).
- CLAUDE.md is now a symlink to AGENTS.md; Claude-Code addenda moved to
  docs/harnesses.md.
- CI: new step fails if `make generate-all` drifts from the committed registries.

Net new committed: ~720 KB of manifests (no skill/agent duplication). Adds
round-trip tests for the new registries + the symlink. Docs updated across
README, docs/harnesses.md, ARCHITECTURE.md, CONTRIBUTING.md, GEMINI.md,
docs/authoring.md, and the PR template.
2026-05-29 13:02:40 -04:00
Seth Hobson f31aba246e Skip Claude review on fork PRs (#559)
Fork PRs don't have access to secrets or OIDC tokens, causing
claude-code-action to fail on OIDC authentication. Skip the job
entirely for cross-repository PRs so CI stays green.
2026-05-25 20:26:53 -04:00
Seth Hobson 8df77ecd46 Fix duplicate agent name collisions (#554)
* fix: namespace duplicate agent names

* fix: format duplicate-name regression test

* fix: align command phase output references

* docs: clarify agent naming formula
2026-05-24 19:55:35 -04:00
Seth Hobson 9834a5f38d Add Pensyve external integration (#552)
* feat: add Pensyve external integration

* docs: fix skill plugin count
2026-05-24 17:35:47 -04:00
Seth Hobson b2b62b2b61 ci: add dependabot for uv-managed Python projects (#545)
* ci: add dependabot for uv-managed Python projects

Configure weekly dependency updates for the two uv projects in the
repo — tools/yt-design-extractor and plugins/plugin-eval. Minor and
patch bumps are grouped per-project so each ecosystem produces at most
one rollup PR per week; majors arrive as individual PRs.

Native uv support is available in Dependabot since 2024 — it updates
pyproject.toml and uv.lock in lock-step.

* deps: drop include:scope to avoid double-parenthesized commit subject

With prefix "deps(yt-design-extractor)" + include:"scope", dependabot
would render commits as "deps(yt-design-extractor)(uv): bump foo from
1 to 2" — double parens. Drop include:scope on both ecosystem entries;
project identity stays in the prefix. Caught by the claude-code-review
workflow on PR #545.
2026-05-22 15:05:08 -04:00
Seth Hobson 1385ee046d Add Claude Code GitHub Workflow (#544)
* "Claude PR Assistant workflow"

* "Claude Code Review workflow"

* ci: tune Claude Code workflows for marketplace review

- Auto-review (claude-code-review.yml): replace generic
  /code-review:code-review slash command with a multi-line prompt that
  carries the repo's invariants (source-of-truth under plugins/, no __
  in plugin names, Codex 8 KB skill cap, AGENTS.md/CLAUDE.md/GEMINI.md
  sync, quality gates). Bump model to claude-opus-4-7. Grant
  pull-requests: write so inline comments and PR summaries can post.
  Add track_progress: true. Skip drafts and dependabot.
- @claude responder (claude.yml): bump model to claude-opus-4-7. Grant
  write scopes for contents/pull-requests/issues so @claude can
  actually fix things. Gate on author_association so only OWNER /
  MEMBER / COLLABORATOR can trigger writes (closes a real auth gap
  flagged by coderabbit, critical now that scopes are write). Curate
  --allowedTools for this repo's uv + ruff + ty + pytest + make + gh +
  git toolchain.
- Both: pin actions/checkout to SHA with persist-credentials: false to
  match .github/workflows/validate.yml convention.

Inspired by major7apps/pensyve and major7apps/maverick-bot review
workflows.

* ci: address coderabbit review on claude workflows

- Pin anthropics/claude-code-action to v1 commit SHA
  (4481e6d3c7bbb88db2a928ca3444c536f589c7c1) — satisfies zizmor's
  blanket-pin policy.
- Drop the implicit "approve with a one-line summary" instruction from
  the review prompt; the workflow's allowedTools only expose inline
  comments and gh pr comment, not a review-submission API. Replace
  with an explicit "post a one-line  summary via gh pr comment and
  stop" so the success path is actually executable.
- Swap issues.types from [opened, assigned] to [opened, edited] on
  claude.yml. The if:-condition only checks issue.body/title for
  @claude, so reassignment-firing produces false triggers while edits
  to existing issues currently fire nothing.

Skipped (with reason): "checkout PR head ref" finding. Anthropic's
canonical claude.yml template (anthropics/claude-code-action examples)
does not set with.ref — the action resolves the PR head internally
via the GitHub API. Forcing a specific ref would diverge from the
upstream contract.

* ci: enumerate git subcommands in @claude allowedTools

Replace the `Bash(git:*)` wildcard with an explicit list of safe
subcommands (status/diff/log/show/rev-parse/branch/tag/ls-files for
inspection; add/commit/checkout/switch/fetch/pull/push/stash for
workflow). Destructive ops like `reset`, `clean`, `rebase`,
`filter-branch`, and `branch -D` are no longer callable — they are
absent from the enumeration.

Addresses coderabbit finding: the wildcard conflicted with the repo's
own security checklist ("no destructive git in scripts") embedded in
claude-code-review.yml.

* ci: further tighten @claude allowedTools to read-only git + scoped gh

Drop `Bash(gh:*)` wildcard and the write-side `git` subcommands.

- `gh` is now enumerated: pr view/diff/list/comment/edit/checks,
  issue view/list/comment/edit, api, run view, workflow view, search.
- `git` is now read-only: status, diff, log, show, rev-parse, ls-files.
  Removed `branch`, `tag`, `add`, `commit`, `checkout`, `switch`,
  `fetch`, `pull`, `push`, `stash` — `git branch -D`, `git tag -d`,
  `git checkout -- .`, `git push --force`, etc. are no longer reachable.

Commits and branch creation are handled by anthropics/claude-code-action
internally via the GitHub API (`branch_prefix` / `branch_name_template`
action inputs), so local write-side git is not required for @claude to
fix things — `Edit`/`Write` produce the changes; the action commits.

Addresses coderabbit follow-up: --allowedTools doesn't understand
destructive flags, so safety comes from precise command patterns.

* ci: scope @claude gh api allowlist to specific endpoints

Drop `Bash(gh api:*)` (which exposed the entire GitHub REST API
surface, including DELETE/PATCH endpoints) in favor of three scoped
patterns sufficient for reading discussion context:

- repos/*/pulls/*/comments  — PR review-thread reads
- repos/*/issues/*/comments — issue-thread reads
- repos/*/issues/*/timeline — issue history

Mirrors the same scoping style used in claude-code-review.yml. The
enumerated `gh pr` / `gh issue` / `gh run` / `gh workflow` / `gh search`
subcommands cover the rest of the responder's typical needs.

Addresses coderabbit finding.
2026-05-22 14:55:49 -04:00
Seth Hobson 2d3f6a8527 chore: move toolchain to uv-native (no pip, no requirements.txt) (#543)
* chore: move toolchain to uv-native (no pip, no requirements.txt)

The repo previously mixed uv (for plugin-eval) with pip + requirements.txt
(for yt-design-extractor) and raw `python3` invocations in the Makefile and
CI workflows. This makes the toolchain uniformly uv-managed.

## yt-design-extractor

Moved `tools/yt-design-extractor.py` into `tools/yt-design-extractor/` with its
own `pyproject.toml` + `uv.lock`. EasyOCR (and its ~2 GB torch dependency)
becomes an optional extra (`uv sync --extra easyocr`). Deleted
`tools/requirements.txt`.

## Makefile

All targets now route through uv:
- `make install` / `make install-easyocr` / `make deps` / `make check` / `make run*`:
  `cd tools/yt-design-extractor && uv run/uv sync`
- `make generate` / `make validate` / `make garden` / `make clean-generated`:
  `uv run --project plugins/plugin-eval python tools/...` (reuses plugin-eval's
  venv — it already has pyyaml and `extra-paths = ["../.."]` for tools/adapters)
- `make test` / `make smoke-test`: `uv run --project plugins/plugin-eval pytest`

## CI

- `.github/workflows/validate.yml`:
  - `multi-harness-generate` swapped from `actions/setup-python@v5` to
    `astral-sh/setup-uv@v5` + `uv sync` of plugin-eval before `make generate-all`
  - Workflow-level `permissions: contents: read` + `persist-credentials: false`
    on every checkout (carrying the security hardening forward consistently)
- `.github/workflows/code-quality.yml`:
  - `json-lint` job swapped from `setup-python` to `setup-uv`
  - YAML validator now uses `uv run --with pyyaml python` (no pre-step install)
  - JSON/TOML validators use `uv run python` (stdlib `json.tool`, `tomllib`)
  - Dropped the standalone `pip install pyyaml --quiet` line

## Inline hints

- `plugins/plugin-eval/src/plugin_eval/layers/judge.py`: error message
  recommends `uv sync --extra llm` instead of `pip install plugin-eval[llm]`
- `tools/yt-design-extractor/yt-design-extractor.py`: usage docstring and
  install hints now reference `make install` / `make install-easyocr` / `uv run`

The pip refs inside plugin-authored commands (deps-audit.md, doc-generate.md,
error-trace.md) describe scanning **user** Python projects — those legitimately
use pip and are out of scope.

Local gates green: make test (385 pass), make garden (0 errors), make validate
clean, ruff + format + ty all pass, markdownlint clean.

* chore: address PR #543 review feedback

- checkmake (Makefile): consolidate the multi-line `.PHONY:` declaration
  onto a single line. The previous backslash-continued form was readable
  but checkmake couldn't parse it, falsely flagging `validate` and
  `clean-generated` as missing from .PHONY. They were already declared —
  just invisible to the linter.

Other CodeRabbit feedback declined:

- pyproject.toml dep version constraints (CodeRabbit self-tagged "Low value"):
  uv.lock pins exact versions for reproducibility; upper bounds on yt-dlp,
  Pillow, etc. would invite stale-pin churn without changing the locked
  installation. Tracking upstream aggressively is the right default for a
  utility tool.

- SHA-pinning for GitHub Actions (CodeRabbit "Major" but defensible):
  Workflow has no write scope (permissions: contents: read on both files),
  no secrets are exposed, and the actions involved are first-party Anthropic
  (astral-sh, actions/*, DavidAnson). Same policy decision as PR #542 —
  blanket SHA pinning is a bigger commitment than this PR's scope warrants.

* chore: pin all GitHub Actions to commit SHAs

Addresses CodeRabbit's Major finding (zizmor `unpinned-uses`). Reverses the
earlier policy decision now that the workflow scope has grown — pins both
workflows consistently in one pass:

- actions/checkout         v4  → 34e114876b0b11c390a56381ad16ebd13914f8d5
- actions/setup-node       v4  → 49933ea5288caeca8642d1e84afbd3f7d6820020
- actions/upload-artifact  v4  → ea165f8d65b6e75b540449e92b4886f43607fa02
- astral-sh/setup-uv       v5  → e58605a9b6da7c637471fab8847a5e5a6b8df081
- DavidAnson/markdownlint-cli2-action v18 → eb5ca3ab411449c66620fe7f1b3c9e10547144b0
- oven-sh/setup-bun        v2  → 0c5077e51419868618aeaa5fe8019c62421857d6

Each pin keeps the version tag in a trailing comment for human readability
and dependabot/renovate compatibility.

Both workflows YAML-validated locally; functionality unchanged.
2026-05-22 12:20:56 -04:00
Seth Hobson 03e6dae50a docs: remove redundant per-harness setup files
CODEX.md, CURSOR.md, and OPENCODE.md were human-readable setup guides that
no harness actually loads as a context file — Codex, Cursor, and OpenCode
all read AGENTS.md natively. The substantive content (install one-liner,
capability deltas, authoring caveats) lives in:

- README.md "Pick your harness" section — install flow
- docs/harnesses.md — full capability matrix + graceful-degradation table
- docs/authoring.md — portable-content style guide

GEMINI.md stays because .gemini/settings.json includes it in
context.fileName — it IS loaded into context every Gemini prompt.

Updated references:
- README.md — badge links + setup-guide line point at docs/harnesses.md
- AGENTS.md, ARCHITECTURE.md — table-of-contents + tree diagram
- docs/authoring.md — context-file cap rule
- .github/PULL_REQUEST_TEMPLATE.md — scope checklist
- tools/doc_gardener.py — CONTEXT_FILES map + dead-link traversal roots
- tools/tests/test_round_trip.py — TestContextFileBudgets parametrize list

Local lints: ruff (CI scope) clean, ruff format clean, ty clean,
markdownlint clean, doc_gardener 0 errors, 31 round-trip + gardener
tests pass.
2026-05-22 11:22:17 -04:00
Seth Hobson 98862b56d8 feat: AGENTS.md canonical context + OpenAI harness-engineering layout (#542)
* feat: AGENTS.md canonical context + OpenAI harness-engineering layout

Promote AGENTS.md to the committed cross-harness context file (per the
agents.md convention and OpenAI's harness-engineering blog). Harness-
specific files become thin redirects:

- AGENTS.md          — canonical, committed (~74 lines, table-of-contents)
- CLAUDE.md          — `@AGENTS.md` import + Claude-specific addenda
- GEMINI.md          — Gemini-specific setup only
- .gemini/settings.json — redirects Gemini CLI's context to read AGENTS.md
- ARCHITECTURE.md    — new at root, top-level architectural map
- gemini-extension.json — bumps version to 1.7.0, sets contextFileName: AGENTS.md
- .gitignore         — drops the AGENTS.md entry (file is now committed)

Harness support verified:
- Codex CLI reads AGENTS.md natively (root → cwd walk, 32 KiB cap)
- Cursor 2.5+ reads AGENTS.md natively
- OpenCode reads AGENTS.md natively (wins over CLAUDE.md if both exist)
- Claude Code: `CLAUDE.md` first line is `@AGENTS.md` (Anthropic's
  documented interop pattern)
- Gemini CLI: `.gemini/settings.json` context.fileName redirect
  (Gemini doesn't support @-imports)

Codex adapter no longer generates AGENTS.md — `emit_global` instead
validates the committed file fits Codex's 32 KiB cap and the 150-line
table-of-contents convention. Tests updated. Clean-output target no
longer touches AGENTS.md.

## Auxiliary files updated for multi-harness reality

- `.github/ISSUE_TEMPLATE/bug_report.yml` — dropdown for harness +
  component path; renames "subagent" → "plugin/agent/skill/command"
- `.github/ISSUE_TEMPLATE/feature_request.yml` — scope dropdown covers
  framework / harness / tooling / docs / CI in addition to components
- `.github/ISSUE_TEMPLATE/new_subagent.yml` — relabeled "New Component
  Proposal" with component-type dropdown (plugin/agent/skill/command/
  harness adapter) and cross-harness portability field
- `.github/ISSUE_TEMPLATE/config.yml` — links to AGENTS.md, authoring
  guide, per-harness docs; updated Contributing link to root
- `.github/CONTRIBUTING.md` — thin pointer to canonical root CONTRIBUTING.md
- `.github/PULL_REQUEST_TEMPLATE.md` — new; scope + affected-harness
  checklists, test-plan checklist, portability-notes section
- CONTRIBUTING.md (root) — updated to reference AGENTS.md / ARCHITECTURE.md
- gemini-extension.json — version 1.6.0 → 1.7.0, count fixes, redirects
  to AGENTS.md as contextFileName

## Code-quality CI

New `.github/workflows/code-quality.yml` with three jobs:
- `python-lint` — `ruff check`, `ruff format --check`, `ty check` on
  the adapter framework + plugin-eval. yt-design-extractor.py legacy
  code excluded.
- `markdown-lint` — markdownlint-cli2 against README, AGENTS, ARCHITECTURE,
  CLAUDE, top-level guides, and docs/. Config in `.markdownlint.json`.
- `json-lint` — validates every JSON / TOML / YAML in the repo (excluding
  generated trees).

Required ty environment config added to plugin-eval/pyproject.toml so
`tools.adapters.*` resolves from outside the package.

Fixed one ty error in `tools/adapters/base.py:HarnessAdapter.capabilities`
(return-type annotation didn't match the `Capability` dataclass returned).
Fixed two ruff SIM108 ternary suggestions in codex.py and doc_gardener.py.
ruff format applied across all in-scope files (formatting-only diffs).

## Tests + verification

- 387 pytest tests pass (1 new test for the AGENTS.md validate-don't-overwrite behavior)
- `make validate STRICT=1` clean
- `make garden` 0 errors (10 warnings — remaining oversize source skills)
- `make smoke-test` clean against locally installed OpenCode/Gemini/Codex/Claude Code
- Real-CLI round-trip: `opencode agent list` discovers 193 subagents,
  `gemini extensions validate .` succeeds, all 191 Codex agent TOMLs parse

## Tag recommendations (separate task — for repo About panel)

Top 20 by reach + relevance (from `gh api search/repositories?q=topic:<tag>`):
automation mcp ai-agents developer-tools claude-code anthropic agentic-ai
agents prompt-engineering cursor multi-agent agent-skills orchestration
opencode workflows gemini-cli codex-cli claude-code-skills cursor-rules
claude-code-plugins

* fix(ci): YAML multi-doc + markdownlint scope/rules

Two CI failures on PR #542, both fixed:

## JSON/TOML/YAML syntax job (3 false-positive YAML errors)

The job's YAML validation used `yaml.safe_load` which only reads the first
document in a multi-document YAML stream. Three Kubernetes manifest
templates use the standard `---` document separator (valid YAML) and were
mis-flagged:

  plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml
  plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/service-template.yaml
  plugins/kubernetes-operations/skills/k8s-security-policies/assets/network-policy-template.yaml

Switched to `list(yaml.safe_load_all(...))` so multi-doc YAML is accepted.

## Markdown lint job (lots of pre-existing plugin-README violations)

Two changes:

1. **Narrow the lint glob** — markdownlint now runs against top-level
   guides (README, AGENTS, ARCHITECTURE, CLAUDE, per-harness setup,
   CONTRIBUTING) and our authored `docs/` only. Per-plugin READMEs
   (`plugins/*/README.md`) are owned by their plugin authors and not
   lint-gated as part of this framework PR. Lint enforcement for those
   belongs at the plugin-author layer, not the framework PR layer.

2. **Tighten `.markdownlint.json`** — disable two rules that produce
   noise without catching real defects:
   - MD040 (fenced-code-language) — terminal output / shell command
     blocks frequently omit a language by convention
   - MD060 (table-column-style) — cosmetic table-pipe spacing; doesn't
     affect rendering

   Genuine formatting rules kept: MD029 (ol-prefix), MD031 (blanks-
   around-fences), MD032 (blanks-around-lists), MD056 (table-column-
   count), MD058 (blanks-around-tables).

## Real defects caught and fixed

The narrower scope still caught 5 real issues:

- `docs/agent-skills.md:397` — code fence inside an ordered list item
  needed a blank line before the fence
- `docs/authoring.md:90` — bulleted list needed a blank line above
- `docs/plugin-eval.md:87` — table needed a blank line above
- `OPENCODE.md:39` — table-column-count error caused by literal `|`
  inside backticks: ``mode: primary|subagent|all`` (3 cells reads as 5)
  Rewrote as ``mode:` one of `primary` / `subagent` / `all``

## Verification

- `npx markdownlint-cli2 "*.md" "docs/*.md"` → 0 errors
- `yaml.safe_load_all` accepts all multi-doc YAMLs (0 errors)
- All other CI jobs already passing (Python ruff/ty, multi-harness
  generate, CLI smoke test, plugin-eval pytest, tools pytest)

* fix: address PR #542 bot feedback

- Codex P2 (chatgpt-codex-connector): emit_global now reads AGENTS.md from
  the repo root (WORKTREE), not output_root. Previously `--output-root <scratch>`
  produced a false "missing" warning even when AGENTS.md was committed at the
  real root, breaking --strict generation outside the repo. Added a constructor
  arg `repo_root` so tests can stage a fake AGENTS.md without touching the
  committed file, plus a regression test that proves the two paths are decoupled.

- CodeRabbit nitpick (code-quality.yml): added workflow-level `permissions:
  contents: read` and `persist-credentials: false` on every checkout. Skipped
  the SHA-pinning recommendation — it's a heavier blanket-policy decision and
  the workflow has no write scope to abuse.

- CodeRabbit nitpick (pyproject.toml): consolidated the duplicate `dev` groups
  by moving `ty` into `[project.optional-dependencies].dev` and removing the
  now-empty `[dependency-groups]` block. Dropped `--group dev` from
  `code-quality.yml`'s `uv sync` since `--all-extras` now covers it.

- Verified gemini-extension.json counts (82/191/155/102) against the actual
  source-of-truth: 81 local plugins + 1 external = 82 in marketplace.json,
  191 agent .md files, 155 SKILL.md files, 102 command .md files. Counts are
  correct as-is — CodeRabbit's quick-win was a regex miscount.
2026-05-22 11:16:39 -04:00
Seth Hobson be57c0b2e3 feat: multi-harness plugin marketplace (Codex, Cursor, OpenCode, Gemini) (#541)
* feat(adapters): multi-harness framework + harness_portability eval dimension

Turn this Claude Code plugin marketplace into a generic agentic-harness
marketplace. Adapters under tools/adapters/ emit harness-native artifacts
for OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single
Markdown source. Source-of-truth stays under plugins/ — Claude Code is
unchanged.

Framework (tools/adapters/):
- base.py — PluginSource parser, HarnessAdapter ABC, write/mirror helpers
  (path-traversal guard, UTF-8-safe), inline-list + block-list + block-scalar
  YAML-ish parser, _utf8_safe_cut, _split_inline_list, _normalize_author
- capabilities.py — per-harness capability matrix, TOOL_NAME_MAPS,
  MODEL_ALIASES, resolve_model() with explicit warnings
- codex.py — emits .codex/{skills,agents}/ + AGENTS.md (≤150-line
  table-of-contents). Fence-aware body splitter, _utf8_safe_cut for
  multibyte safety, _yaml_scalar with reserved-word + special-char quoting.
  Skill/command name collision detection (and second-order __cmd fallback).
- cursor.py — emits .cursor-plugin/{plugin,marketplace}.json + curated
  .cursor/rules/*.mdc. _validate_mdc_frontmatter handles YAML block scalars
  (no false positives on colons in description body). _normalize_author
  handles dict, npm-style strings, and author lists.
- opencode.py — transpiles agents to .opencode/agents/<id>.md with
  mode:subagent + permission: deny-everything-else block (skill/task always
  allowed as base capabilities — Claude's implicit defaults).
- gemini.py — emits native skills/, agents/, and commands/ at extension
  root (April 2026 spec). Tool-allowlist remapped via TOOL_NAME_MAPS.

CLI + tooling:
- tools/generate.py — unified `make generate HARNESS=<x> [PLUGIN=<y>]`,
  with --clean (containment-guarded; case-insensitive on Darwin/Win32),
  --prune (orphan removal across all per-harness output trees), --strict
  (warnings fail), per-plugin error aggregation, refuses --clean --plugin
  (would silently wipe other plugins' artifacts).
- tools/validate_generated.py — structural validation across all four
  harness outputs. Codex 8KB cap → error. _extract_permission_block
  correctly handles nested permission keys (column-0 only).
- tools/doc_gardener.py — recurring drift detection per OpenAI harness-
  engineering principle. STALE_ARTIFACT (info), DEAD_LINK (error),
  MARKETPLACE_ORPHAN (error), SKILL_OVER_CODEX_CAP (warning), grouped
  output sorted by severity.

plugin-eval (extends existing framework):
- New harness_portability dimension (6% weight, rebalanced from existing
  static sub-scores). Surfaces non-portable patterns with concrete
  remediation hints: SKILL_OVER_CODEX_CAP, CLAUDE_TOOL_REFS,
  CLAUDE_TOOL_PROSE, AGENT_NAME_COLLISION, BARE_MODEL_ALIAS.
- _CAMEL_TOOL_PATTERN requires Claude-tool context (no false positives
  on Rust's `Task` etc.). _TOOL_PROSE_PATTERN case-sensitive on tool
  names, case-insensitive on the leading article.
- Findings do NOT also feed anti_pattern_penalty (no double-counting).

Documentation:
- Top-level guides: CODEX.md, CURSOR.md, OPENCODE.md (≤150 lines each,
  table-of-contents pattern per OpenAI harness-engineering post)
- docs/harnesses.md — capability matrix, graceful-degradation table,
  generated output paths
- docs/authoring.md — portable-content style guide (tools, models,
  collision rules, fence-respect)
- docs/round-trip-results.md — real-CLI verification recipes (OpenCode
  discovers 193 subagents, Gemini extensions validate passes, Codex
  TOMLs all parse)
- CONTRIBUTING.md — new file pointing at docs/authoring.md
- README.md — rewritten for multi-harness (145 lines, was 460)
- CLAUDE.md — trimmed to 60-line table-of-contents
- GEMINI.md — trimmed from 1500 to 500 tokens (3× over budget previously)

Tests: 181 passing (103 plugin-eval + 78 tools/tests). Real-CLI round-trip
verified for OpenCode, Gemini, and Codex (TOML parses).

Replaces tools/generate_gemini_commands.py with the unified CLI.

* refactor(skills): extract detail to references/details.md (~75 skills)

Apply Anthropic's canonical SKILL.md progressive-disclosure pattern across
the marketplace: SKILL.md body becomes a navigation tier (trigger phrasing
+ quick start), detailed templates and worked examples move to
references/details.md (loaded on demand by the agent).

Motivation: OpenAI Codex CLI hard-truncates skills at 8 KB. Before this
change, ~90 skills exceeded that cap and would silently break on Codex.
The progressive-disclosure pattern is also Anthropic's documented
recommendation for token efficiency — Claude Code reads references/ files
on demand when the body navigation says to.

What's extracted, by pattern:
- Pass 1 (## Templates section): 19 skills — full template libraries
  moved to references/details.md
- Pass 2 (## Implementation Patterns / ## Advanced Patterns): 13 skills
- Pass 3 (everything between nav-tier and wrap-tier headings): 53 skills
- Conservative re-extraction for 8 skills that got over-reduced — kept
  ~6-7 KB inline (most of the quick-start tier) plus references/ overflow

What stays inline (SKILL.md navigation tier):
- description: frontmatter (triggering — unchanged for all skills)
- ## When to Use This Skill / ## Core Concepts / ## Quick Start
- ## Best Practices / ## Troubleshooting / ## See Also wrap-ups
- A pointer note ("see references/details.md") so the agent knows where
  to look for detail

What goes to references/details.md (detail tier, on-demand load):
- ## Templates (full code template libraries)
- ## Implementation Patterns / ## Advanced Patterns (deep examples)
- Mid-skill walkthroughs that exceed the inline budget

Also in this commit:
- plugins/brand-landingpage description trimmed from 958→543 chars
  (preserves trigger phrasing, drops verbose example-quote list)

Net effect:
- SKILL_OVER_CODEX_CAP findings: 90 → 10 (88% reduction)
- All triggers unchanged — discovery behavior identical across harnesses
- 75 new references/details.md files with the extracted content
- Same depth of guidance, loaded progressively

Remaining 10 oversized skills are complex multi-section docs (e.g.
postgresql, code-review-excellence, evaluation-methodology) that need
per-skill manual judgment — flagged by `make garden` for future work.

* chore: bump all plugin versions (multi-harness release)

Patch-bump every local plugin (81) in both .claude-plugin/marketplace.json
entries and each plugins/<name>/.claude-plugin/plugin.json. Minor-bump the
top-level marketplace metadata.version (1.6.0 → 1.7.0) to signal the
multi-harness adapter framework addition.

The external git-subdir entry (qa-orchestra) is unaffected — its version
is governed by its upstream repo.

* fix(opencode): preserve explicit tools:[] + word-boundary subtask match

Addresses two Codex review findings on PR #541.

## P1 — `tools: []` silently upgraded to permissive (privilege escalation)

Before: `_build_permission_block` returned `{}` for any empty list, which
omits the `permission:` block entirely from the emitted agent. An author
who explicitly wrote `tools: []` to lock down an advisory-only agent got
an UNRESTRICTED agent in OpenCode. Affected agent in this tree:
`plugins/arm-cortex-microcontrollers/agents/arm-cortex-expert.md`.

Fix: `_build_permission_block` now takes a `has_tools_field` flag so the
caller can distinguish "tools: key missing" (Claude default permissive)
from "tools: []" (explicit lock-down). The lock-down case emits a
deny-everything block that allows ONLY the base capabilities (skill, task)
that Claude Code always grants implicitly. Verified against the real
arm-cortex-expert agent — now emits read/edit/write/bash/grep/glob/list:
deny, task/skill: allow.

## P2 — `"agent" in cmd.body.lower()` false-positives on substrings

Before: a command body containing `PerformanceReviewAgent` (class name
in a code snippet) or `useragent` triggered `subtask: true`, changing
runtime behavior based on incidental text.

Fix: switch to a compiled word-boundary regex `\b(agent|subagent)s?\b`
(case-insensitive). Tests confirm the substring `PerformanceReviewAgent`
no longer fires, while a real "spawn a subagent" sentence still does.

## Tests

3 new regression tests in tools/tests/test_adapters.py:
- `test_explicit_empty_tools_yields_locked_permission_block` (P1)
- `test_missing_tools_field_yields_no_permission_block` (P1 boundary)
- `test_subtask_inference_word_boundary` (P2)

184 total tests pass (was 181). OpenCode round-trip still discovers all
193 subagents; arm-cortex-expert agent is now properly locked down.

* test: behavioral verification + CI gates for multi-harness pipeline

Adds three layers of automated verification that pure-Python parser tests
miss, plus the CI jobs that turn them into hard gates. Catches the kinds
of issues that previously only surfaced when a real user installed the
marketplace and tried to use it.

## test_real_world.py — real-source structural tests

Runs against the actual `plugins/` tree (not synthetic fixtures). Catches
issues that only appear on real content:

- every marketplace entry resolves to a plugins/<name>/ dir
- every local plugin dir appears in marketplace.json
- marketplace.json version == per-plugin plugin.json version (catches drift)
- every plugin loads via load_plugin() without error
- no plugin name contains `__` (adapter namespace separator)
- every agent has name + description; every skill has a trigger phrase
  (same regex plugin_eval's MISSING_TRIGGER check uses)
- no agent name collides with Codex built-ins
- every refactored skill (with `references/details.md`) has:
  - meaningful detail content (>=500 B in details.md)
  - a pointer to references/ in the SKILL.md body
  - a navigation-tier heading preserved (When to Use, Overview, etc.)
  - body >= 600 B (not a stub)
- every plugin.json has name + version matching the dir

This test pass found and fixed three real defects before commit:
- ship-mate/skills/scan: description had no trigger phrase ("Use when…")
- reverse-engineering/skills/memory-forensics: nav-tier section lost
  during extraction
- reverse-engineering/skills/binary-analysis-patterns: same

All three are now fixed (preserved trigger phrasing, added When-to-Use
sections back to the skills my extraction over-trimmed).

## test_round_trip.py — generate→parse→verify

CI runs this AFTER `make generate-all`. Catches generation-time regressions:

- OpenCode/Codex/Gemini agent counts match source agent count (no skips)
- every Codex SKILL.md under 8 KB (the cap that would silently truncate)
- every Codex agent TOML has required fields + valid sandbox_mode
- every OpenCode agent has mode in {primary,subagent,all} and
  provider-prefixed model
- locked agents (source `tools: []`) emit proper deny-everything permission
  block with skill/task allow (regression guard for PR-541 P1)
- every Gemini @{path} injection resolves to a real source file
- every Gemini command TOML has prompt + {{args}} placeholder
- every context file (CLAUDE.md, AGENTS.md, GEMINI.md, etc.) within
  150-line cap
- Cursor marketplace + per-plugin manifests cover all local plugins
- .cursor/rules/*.mdc only use the 3 documented frontmatter keys

## test_cli_smoke.py — real-CLI subprocess tests

Invokes the actual harness binaries (OpenCode, Gemini, Codex, Claude Code)
against the generated artifacts. Catches CLI-level issues pure-Python
parsing can't see: schema-loader drift, plugin-discovery bugs, version
incompatibilities.

- `opencode agent list` — must succeed AND discover every source agent
  (currently 191 + 2 OpenCode built-ins)
- `gemini extensions validate <repo>` — must return success
- `codex doctor` — must report healthy install
- every Codex agent TOML must parse with stdlib `tomllib`
- `claude --version` — sanity check the Claude Code CLI loads
- marketplace.json must have owner + metadata.version for Claude Code's loader

Per-CLI tests skip gracefully when the binary isn't on PATH, so local
devs only exercise what they have installed. CI installs OpenCode +
Gemini and turns those skips into hard gates.

## Makefile + CI

- `make test` — full pytest suite (plugin-eval + tools/tests/)
- `make smoke-test` — generates if needed, then runs real-CLI smoke tests
- `.github/workflows/validate.yml` extended with:
  - `tools-tests` job — runs pytest tools/tests/
  - `multi-harness-generate` job — `make generate-all && make validate
    STRICT=1 && make garden`, uploads generated artifacts on every run
  - `cli-smoke-test` job — installs OpenCode + Gemini, runs test_cli_smoke.py

## Test counts

- Before: 184 tests
- After: 386 tests (parameterized real-source tests over all 82 plugins)
- All passing locally on OpenCode 1.15.7 + Gemini 0.42.0 + Codex 0.133.0
  + Claude Code 2.1.148
2026-05-22 08:18:21 -04:00
Seth Hobson 08ded5e7b0 fix: agent teams coordination guardrails (#535)
* fix agent teams coordination guardrails

* address agent teams review feedback
2026-05-16 20:46:39 -04:00
Seth Hobson 0041fc7194 ci: add plugin eval reporting workflow with per-plugin scores
Adds a GitHub Action that runs the full plugin-eval engine across every
local plugin and emits a structured report for review — not a PR gate.

- scripts/eval_all.py batch-runs EvalEngine.evaluate_plugin across
  plugins/*/, writing per-plugin JSON plus summary.md and summary.json
- Report surfaces composite score, 95% CI, badge, confidence label,
  anti-pattern flags, and three weakest dimensions per plugin
- Separate "Issues requiring attention" section for plugins scoring
  under 60 or with anti-pattern flags raised
- Workflow triggers: workflow_dispatch (choose depth + optional
  comma-separated plugin filter) and weekly cron (Mondays 06:00 UTC)
- Depth quick = static only (no credentials); standard/deep require
  ANTHROPIC_API_KEY repo secret for LLM judge / Monte Carlo layers
- Posts summary.md to the job summary, uploads full eval-reports/
  directory as an artifact for 30 days

Smoke-tested locally: 77 plugins evaluated at quick depth in < 5s,
mean 83.1/100, 2 plugins flagged (incident-response with
DEAD_CROSS_REF, plugin-eval with ORPHAN_REFERENCE + DEAD_CROSS_REF).
2026-04-16 13:16:47 -04:00
Seth Hobson 6625d1ead4 docs: refresh counts for protect-mcp + qa-orchestra; add CI validation
- Counts now reflect 79 plugins (77 local + 2 external via git-subdir),
  184 agents, 150 skills, 98 commands across 25 categories
- README/docs: add Governance category (protect-mcp), bump Testing to 2
  (adds qa-orchestra), expand Security table to include reverse-engineering
  and block-no-verify, add Protect MCP skills section, surface pensyve in
  AI & ML
- marketplace.json: bump metadata to 1.6.0, bump pensyve to 1.2.0
- Add .github/workflows/validate.yml to gate PRs on:
  - marketplace.json + every plugin.json + hooks.json parse as JSON
  - every ./plugins/<name> source resolves on disk with a plugin.json
  - plugin-eval pytest suite (would have caught the #482 sdk.stream bug class)
2026-04-16 13:10:31 -04:00
Seth Hobson 47a5dbc3f9 fix(skills): remove phantom resource references and fix CoC links (#447)
Remove references to non-existent resource files (references/, assets/,
scripts/, examples/) from 115 skill SKILL.md files. These sections
pointed to directories and files that were never created, causing
confusion when users install skills.

Also fix broken Code of Conduct links in issue templates to use
absolute GitHub URLs instead of relative paths that 404.
2026-03-07 10:53:17 -05:00
Seth Hobson f662524f9a feat: add Conductor plugin for Context-Driven Development
Add comprehensive Conductor plugin implementing Context-Driven Development
methodology with tracks, specs, and phased implementation plans.

Components:
- 5 commands: setup, new-track, implement, status, revert
- 1 agent: conductor-validator
- 3 skills: context-driven-development, track-management, workflow-patterns
- 18 templates for project artifacts

Documentation updates:
- README.md: Updated counts (68 plugins, 100 agents, 110 skills, 76 tools)
- docs/plugins.md: Added Conductor to Workflows section
- docs/agents.md: Added conductor-validator agent
- docs/agent-skills.md: Added Conductor skills section

Also includes Prettier formatting across all project files.
2026-01-15 17:38:21 -05:00
Seth Hobson f95810b340 chore: remove Claude Code GitHub workflows
The workflows were failing for fork PRs due to GitHub security
restrictions on secrets access. Removing until a better solution
is implemented.
2025-12-30 16:25:52 -05:00
Seth Hobson cf82055296 docs(conduct): add automated abuse and LLM spam policies
Add explicit language prohibiting bulk automated issue/PR creation
and LLM-generated spam, with zero-tolerance enforcement policy.
2025-12-30 15:20:47 -05:00
Seth Hobson 27a246a8c6 Add Claude Code GitHub Workflow (#140)
* "Claude PR Assistant workflow"

* "Claude Code Review workflow"
2025-12-10 15:12:53 -05:00
Seth Hobson cc31a8f777 remove github workflows 2025-08-19 19:39:52 -04:00
Seth Hobson e88a6ee87d Add GitHub Sponsors button
Add FUNDING.yml to enable sponsor button on repository
2025-08-10 10:51:19 -04:00
Seth Hobson 4dad1937e9 Format all YAML files to pass yamllint validation
- Added document start markers (---) to all YAML files
- Fixed line length issues by breaking long lines appropriately
- Removed trailing whitespace throughout all files
- Added proper newlines at end of files
- Fixed truthy value format ('on' -> 'on')
- Standardized YAML formatting across workflows and issue templates
- Used multi-line strings (>) for long descriptions
- Maintained readability while adhering to 80-character line limit

All YAML files now pass yamllint validation with only minor warnings remaining.
2025-08-01 17:33:42 -04:00
Seth Hobson fe236a76b5 Fix GitHub Actions welcome workflow using best practices
Problem Analysis:
- Original workflow used complex GitHub search API causing rate limiting issues
- Custom first-time contributor detection was unreliable and fragile
- Used pull_request instead of pull_request_target for PRs (security issue)
- Complex github-script logic prone to failures

Solution Implemented:
- Replaced custom logic with GitHub's official actions/first-interaction@v1
- Changed to pull_request_target for PR security and reliability
- Eliminated API rate limiting issues by removing search calls
- Simplified permissions and workflow structure
- Added comprehensive welcome messages with community guidelines

Benefits:
- More reliable first-time contributor detection
- No rate limiting issues
- Better security with pull_request_target
- Easier to maintain using official GitHub action
- Consistent messaging across issues and PRs

Also included alternative implementation example using garg3133/welcome-new-contributors@v1.2
2025-08-01 17:19:23 -04:00
Seth Hobson ea5644ab9b Implement comprehensive content moderation and community protection
- Added Code of Conduct with clear behavioral standards
- Created Contributing guidelines with submission requirements
- Implemented structured issue templates (bug reports, features, new agents, moderation)
- Disabled blank issues to enforce template usage
- Added automated content moderation via GitHub Actions:
  * Real-time scanning for hate speech, threats, and profanity
  * Automatic closure/locking of critical violations
  * Moderation alerts for maintainer review
- Set up welcome system for new contributors with community guidelines
- Enabled GitHub Discussions as alternative to issues for general questions
- Closed and locked existing hate speech issue #30
- Blocked offending user account

This creates a multi-layered defense against inappropriate content while
maintaining an open, welcoming environment for legitimate contributors.
2025-08-01 16:44:10 -04:00