mirror of
https://github.com/zernie/vigiles.git
synced 2026-09-14 20:53:57 +08:00
docs: honest versioning, README scannability, and docs restructure (#67)
Fix the version claim (v12, not 0.x; the 0.0.0-semantically-released placeholder was the source), reframe the eval section around cost measurement, trim the README for scannability (Diátaxis link buckets, tighter reconciliation + proofs), reorganize the docs index into Guides/Reference/Explanation and add six missing docs, and consolidate overlapping docs (related-tools→comparison, inline-mode→markdown-mode, agent-setup+agent-workflows). Relocate the eval-architecture design-of-record ADR from docs/ to research/ per the doc-tiers rule, and fix the library entry points (vigiles/linting is the compiler surface) plus the Codex setup path.
This commit is contained in:
+5
-9
File diff suppressed because one or more lines are too long
+36
-44
@@ -14,58 +14,50 @@
|
|||||||
|
|
||||||
## RESUME HERE
|
## RESUME HERE
|
||||||
|
|
||||||
**Branch `claude/vigiles-cost-analysis-9ko9mv`** (name is misleading — the work is
|
**Branch `claude/skill-eval-cost-benefit-q4ivfp`** (name misleading — the work is a
|
||||||
ADOPTION FIXES, not cost analysis) → **PR #66 OPEN, merging-when-green.** This session acted
|
PUBLIC-DOCS REVAMP, not cost analysis) → **PR #67 OPEN, merging-when-green.** The user flagged the
|
||||||
on a skills-monorepo FIELD REPORT (a team ran vigiles on a 46-skill CI library with no
|
docs as over-claiming maturity + hard to scan; this session overhauled the README + docs.
|
||||||
`plugin.json` and hit blockers) — all 7 feedback points fixed.
|
|
||||||
|
|
||||||
**MERGE STATE (resume here first):** PR #66 open, **merging-when-green**; **subscribed to its
|
**MERGE STATE (resume here first):** PR #67 open, **merging-when-green**; **subscribed to its
|
||||||
activity**; a `send_later` check-in (trigger `trig_01HgpKfkov7BQAJNzHtMNEoq`, ~14:44Z) re-checks CI
|
activity**; a `send_later` check-in (trigger `trig_01LX727g2VMn1qpe7k31eJ1o`, ~21:18Z) re-checks CI
|
||||||
and **squash-merges into main with a CLEAN message (NO session link / model-id) when all 6 jobs
|
and **squash-merges into main with a CLEAN message (NO session link / model-id) when all 6 jobs
|
||||||
green**, else re-arms. If resuming: `get_check_runs` for #66 → merge if green, then **unsubscribe**.
|
green**, else re-arms. If resuming: `get_check_runs` for #67 → merge if green, then **unsubscribe**.
|
||||||
Latest SHA `aa5526d` (a HANDOFF-refresh commit sits on top). CI jobs: validate/describe/check/test/
|
Latest SHA `f1d0a2d`. CI jobs: validate/describe/check/test/e2e/harness (`test` ~5-7 min, last).
|
||||||
e2e/harness; `test` runs ~5-7 min and is always last. The lone allowed failure anywhere is env-only
|
Lone allowed failure: env-only `dialect-drift` (CI pins CC → passes in CI).
|
||||||
`dialect-drift` (CI pins CC, so it passes in CI).
|
|
||||||
|
|
||||||
**CODE-REVIEW LOOP (done):** Codex-bot reviewed every pushed commit and found **13 real P2 bugs
|
**Shipped (branch — all `docs:`):**
|
||||||
across 5 rounds — ALL fixed + tested** (api-extractor surface, eslint void-expr, single-skill bundled
|
|
||||||
resources, root-SKILL.md coverage + colocation, sharedDirs-from-repo-root, scoped harness detection,
|
|
||||||
per-surface→repo-level fallback, hook-only + hooks-convention plugin shape, loadable-only surface
|
|
||||||
count, foreign-repo sharedDirs root, query-suffix vs glob-skip). Codex then **hit its usage quota**
|
|
||||||
(no more reviews incoming) — the loop ends by quota, NOT by proof of correctness. WATCH-OUT: the
|
|
||||||
single-skill-dir targeting + `.claude`-fallback subsystem generated most siblings; classes are now
|
|
||||||
closed + tested, but if a NEW real bug there surfaces, prefer a redesign or NARROWING the PR (drop
|
|
||||||
single-skill-dir) over another patch.
|
|
||||||
|
|
||||||
**Shipped (branch — `feat(scan)` + `fix(lint)` + docs):**
|
- **Version honesty** — STABILITY.md + README FAQ claimed "0.x"; real npm version is **12.7.0**
|
||||||
|
(semantic-release cuts a major per breaking change). Fixed to v12 + honest framing. Root cause: the
|
||||||
|
`package.json` `0.0.0-semantically-released` placeholder misread as "0.x". Added `vigiles/linting`
|
||||||
|
to the stable entry-points list.
|
||||||
|
- **README scannability** — `More` link-farm → 3 Diátaxis buckets + index pointer; collapsed the dense
|
||||||
|
audit/lint/test/eval reconciliation paragraph; trimmed Proof 2/3. Eval heading → "the only way to put
|
||||||
|
a real number on cost."
|
||||||
|
- **docs/README.md** reorganized (Guides / Reference / Explanation) + 6 docs added that were missing
|
||||||
|
from the index (harnesses, adapter-api, authoring-an-adapter, railway-subagents, faq, what-vigiles-catches).
|
||||||
|
- **Merges** — `related-tools`→`comparison`; `inline-mode`→`markdown-mode` (one no-spec on-ramp doc;
|
||||||
|
NB inline mode is LIVE — only FRONTMATTER mode is disabled); `agent-setup`+`agent-workflows`→one guide.
|
||||||
|
- **`eval-architecture.md` (54KB design-of-record ADR) relocated `docs/`→`research/`** per doc-tiers:
|
||||||
|
5 public links repointed (testing-api / measuring-skills), research/src relative paths fixed, added to
|
||||||
|
research index + `status:`/`topic:` frontmatter, both `CLAUDE.md` + `research/CLAUDE.md` recompiled.
|
||||||
|
- **Codex-bot review** caught 1 real bug (the compiler surface is `vigiles/linting`, not `vigiles/spec`) — fixed.
|
||||||
|
- Kept standalone by JUDGMENT (merging would bloat, not help): `testing-matrix`, `migrating-from-promptfoo`.
|
||||||
|
|
||||||
- **P0-1** `loadPlugin` recognizes THREE repo shapes — published plugin / bare `skills/*` library /
|
Deleted: docs/{related-tools,inline-mode,agent-workflows,eval-architecture}.md (last one moved to research/).
|
||||||
plain `.claude/skills` user repo — via a new optional `PluginLayout.userSurfaceRoot` (`.claude`
|
|
||||||
in the CC adapter; core stays agnostic, `.claude` literal only in the adapter). Root `skills/`
|
|
||||||
WINS over `.claude/skills`. A single skill dir works. `LoadedPlugin.sources` maps each
|
|
||||||
materialized key → its real on-disk path. Reads the PROJECT `.claude` only, never `~/.claude`.
|
|
||||||
- **P0-2** `vigiles lint` now SCOPES to an explicit dir arg (was: ignored the path, scanned the
|
|
||||||
whole repo → reported foreign surfaces). `runLint` resolves ONE `scanRoot` (single existing dir →
|
|
||||||
narrow; file / several / none → cwd, so **bare `lint` is byte-identical**) threaded into all 21
|
|
||||||
surface appliers (replaced `process.cwd()`). Bugfix — only `lint <dir>` changes.
|
|
||||||
- **P1-3** skill-resources skips glob / placeholder refs (`* ? { } < >`) + `~/` home paths.
|
|
||||||
- **P1-4** OPT-IN `sharedDirs` config — a ref whose first segment is a declared shared dir also
|
|
||||||
resolves at the repo root; scoped so nothing outside it is masked; default byte-identical.
|
|
||||||
- **P2-6** lethal-trifecta advisory collapsed to one line per unit. **P2-7** `docs/skills-monorepo.md`.
|
|
||||||
|
|
||||||
Files: `src/plugin-loader.ts`, `src/core/layout.ts`, `src/adapters/claude-code/layout.ts`,
|
**REPO ABOUT (USER ACTION — no tool/API access to set it):** paste into Settings→About —
|
||||||
`src/scan.ts`, `src/core/skill-resources.ts`, `src/core/types.ts` (`sharedDirs`), `src/cli.ts`
|
desc "Like Lighthouse for your agent harness — verify your CLAUDE.md/AGENTS.md, skills & hooks are real,
|
||||||
(scanRoot). Tests: `src/scan.test.ts`, `src/core/skill-resources.test.ts`, `src/scan-cli.test.ts`
|
then test and measure they actually work. Claude Code + Codex.", website https://zernie.github.io/vigiles/,
|
||||||
(lint-scoping e2e), `src/adapters/claude-code/plugin-loader.test.ts`.
|
topics: claude-code codex agentic-coding ai-agents llm claude anthropic developer-tools cli linter testing
|
||||||
|
evals typescript mcp skills.
|
||||||
|
|
||||||
**NEXT (not blocking):** the feedback is fully addressed. The field report's exact "global
|
**NEXT (not blocking):** none required. Candidate follow-up: a deterministic `no-internal-links-in-public-docs`
|
||||||
`~/.claude` skills appeared" symptom couldn't be reproduced from code (their env) — P0-2's
|
lint rule (P1 roadmap) — currently hand-enforced (this session verified it by grep).
|
||||||
correct rooting closes it either way. Possible follow-up: the `init` ADOPT/untested discovery
|
|
||||||
also walks cwd; P0-2 scopes the lint appliers, adopt-discovery is a separate pass.
|
|
||||||
|
|
||||||
**TEST STATUS:** full vitest **2074 passed** locally; the only failure is env-only
|
**TEST STATUS:** touched-gate dogfoods pass locally (research-index, self-command-refs, doc-command-coverage,
|
||||||
`dialect-drift.test.ts` (installed vs pinned CC — CI pins it). eslint/tsc/test:types/fmt/api all
|
orphans, inline) + build / integrity / orphan-docs / fmt / no-internal-links green. Full vitest not re-run
|
||||||
green locally.
|
(docs-only change); env-only `dialect-drift` still fails locally (CI pins CC).
|
||||||
|
|
||||||
## Don't re-read unless the task needs it
|
## Don't re-read unless the task needs it
|
||||||
|
|
||||||
|
|||||||
@@ -132,7 +132,7 @@ This subagent — a helper your main agent hands work to — lists a tool that d
|
|||||||
apart, so the wrong one fires (e.g. "agent-coder" ↔ "agent-tester", 83% alike)
|
apart, so the wrong one fires (e.g. "agent-coder" ↔ "agent-tester", 83% alike)
|
||||||
```
|
```
|
||||||
|
|
||||||
One popular plugin ships **45 pairs of skills** with near-identical descriptions. Your agent picks which skill to run by _reading_ those descriptions, so when two match it fires the wrong one. Still perfectly valid markdown.
|
One popular plugin ships **45 pairs** of near-identical skill descriptions. Your agent picks a skill by _reading_ them — so when two match, it fires the wrong one. Still perfectly valid markdown.
|
||||||
**[How triggering works →](docs/measuring-skills.md)**
|
**[How triggering works →](docs/measuring-skills.md)**
|
||||||
|
|
||||||
## Proof 3 — it can quietly read your secrets and send them out
|
## Proof 3 — it can quietly read your secrets and send them out
|
||||||
@@ -144,7 +144,7 @@ One popular plugin ships **45 pairs of skills** with near-identical descriptions
|
|||||||
· can send data out (Bash, WebFetch)
|
· can send data out (Bash, WebFetch)
|
||||||
```
|
```
|
||||||
|
|
||||||
Hand one subagent all three powers and a poisoned web page can tell it to read your `.env` and POST it anywhere — no exploit code, just the tools it was given. The **80 still looks like a B** — that's the point: a healthy-looking grade can hide a single subagent that's a data-leak waiting to happen. vigiles spots it from the tool list alone, free, no model.
|
Hand one subagent all three powers and a poisoned web page can make it read your `.env` and POST it anywhere — no exploit code, just the tools it was given. The **80 looks like a B** — and that's the trap: a healthy grade hiding a subagent that's a data-leak waiting to happen. vigiles spots it from the tool list alone, free, no model.
|
||||||
|
|
||||||
That's the whole idea: it checks your harness against **reality, not style**. Every tool, hook, file, script, and skill you reference is verified to actually resolve — and where you name a linter rule, it's checked to exist _and_ be enabled (ESLint, Ruff, Clippy, and more).
|
That's the whole idea: it checks your harness against **reality, not style**. Every tool, hook, file, script, and skill you reference is verified to actually resolve — and where you name a linter rule, it's checked to exist _and_ be enabled (ESLint, Ruff, Clippy, and more).
|
||||||
**[Everything it catches →](docs/what-vigiles-catches.md)** · point `audit` at a whole marketplace and it ranks every plugin the same way.
|
**[Everything it catches →](docs/what-vigiles-catches.md)** · point `audit` at a whole marketplace and it ranks every plugin the same way.
|
||||||
@@ -160,7 +160,7 @@ That's the whole idea: it checks your harness against **reality, not style**. Ev
|
|||||||
| `test` | Does the harness behave? | No — a scripted stand-in | Every commit |
|
| `test` | Does the harness behave? | No — a scripted stand-in | Every commit |
|
||||||
| `eval` | Does a skill actually help? | Yes — your subscription | On demand |
|
| `eval` | Does a skill actually help? | Yes — your subscription | On demand |
|
||||||
|
|
||||||
`audit` and `lint` share one engine. **`lint` is the CI gate** — it fails the build on broken references, bad tool contracts, dead hooks, and skill collisions (Proofs 1 and 2). **`audit`** runs those same checks, adds the Safety ring, renders the graded report, and can also run two opt-in _live_ checks (does your MCP server connect, do your skills fire). `test` and `eval` go past _does it exist_ to _does it work_. (`init` / `compile` / `eject` manage the spec layer underneath; you rarely run them by hand.)
|
**One engine, two doors.** `audit` is the local report; **`lint` is the CI gate** that fails the build on the same deterministic checks — broken refs, bad tool contracts, dead hooks, skill collisions (Proofs 1–2). `test` and `eval` go further: past _does it exist_ to _does it work_. (`init` / `compile` / `eject` manage the spec layer underneath — you rarely run them by hand.)
|
||||||
|
|
||||||
### 🔎 Lint — your instructions stop lying
|
### 🔎 Lint — your instructions stop lying
|
||||||
|
|
||||||
@@ -172,7 +172,7 @@ Every path, script, symbol, and rule verified against reality — plus tool cont
|
|||||||
A hook that blocks nothing, a skill that hijacks unrelated prompts, context that never reaches the model — each passes a naive "did it run?" check. That gap is **false confidence**: a guard that looks like it works and silently doesn't. vigiles tests the real thing — hooks block, skills fire, subagents finish what they promised, a stray `git push` is caught before it happens. It drives a scripted stand-in for the model, not a live call, so it needs no key and runs on every commit.
|
A hook that blocks nothing, a skill that hijacks unrelated prompts, context that never reaches the model — each passes a naive "did it run?" check. That gap is **false confidence**: a guard that looks like it works and silently doesn't. vigiles tests the real thing — hooks block, skills fire, subagents finish what they promised, a stray `git push` is caught before it happens. It drives a scripted stand-in for the model, not a live call, so it needs no key and runs on every commit.
|
||||||
**[How testing works →](docs/harness-testing.md)**
|
**[How testing works →](docs/harness-testing.md)**
|
||||||
|
|
||||||
### 📊 Eval — does a skill help, or just cost more?
|
### 📊 Eval — the only way to put a real number on cost
|
||||||
|
|
||||||
_"Caveman Mode cuts 65% of your tokens." Says who?_ vigiles A/Bs the claim on real coding tasks and hands you three numbers: the **token bill**, whether it hit its **target**, and whether your code still **works**.
|
_"Caveman Mode cuts 65% of your tokens." Says who?_ vigiles A/Bs the claim on real coding tasks and hands you three numbers: the **token bill**, whether it hit its **target**, and whether your code still **works**.
|
||||||
|
|
||||||
@@ -237,18 +237,22 @@ Targets Claude Code and Codex out of the box, or [your own harness](docs/authori
|
|||||||
- **Is this a framework I have to build around?** No. It's a tool you run — like ESLint, Lighthouse, or `npm audit`. One command, a report, an optional CI gate. There's a library API for automation, but you never touch it to get value.
|
- **Is this a framework I have to build around?** No. It's a tool you run — like ESLint, Lighthouse, or `npm audit`. One command, a report, an optional CI gate. There's a library API for automation, but you never touch it to get value.
|
||||||
- **Isn't this just a markdown linter?** No — it checks whether your instruction file is _true_ (every path/script/symbol/rule exists and is enabled), then tests and measures your harness. A style linter can't do any of that.
|
- **Isn't this just a markdown linter?** No — it checks whether your instruction file is _true_ (every path/script/symbol/rule exists and is enabled), then tests and measures your harness. A style linter can't do any of that.
|
||||||
- **Do I have to write TypeScript?** No — your agent writes the spec (`init` adopts your CLAUDE.md or AGENTS.md into one), or plain markdown lints with zero new files. Compiler-grade guarantees are opt-in, like TS's `strict` ([why?](docs/faq.md#why-are-the-strongest-guarantees-opt-in-not-the-default)).
|
- **Do I have to write TypeScript?** No — your agent writes the spec (`init` adopts your CLAUDE.md or AGENTS.md into one), or plain markdown lints with zero new files. Compiler-grade guarantees are opt-in, like TS's `strict` ([why?](docs/faq.md#why-are-the-strongest-guarantees-opt-in-not-the-default)).
|
||||||
- **Is it stable enough to adopt?** Yes — the CLI is stable; only the library API is still evolving ([details](STABILITY.md)).
|
- **Is it stable enough to adopt?** The CLI you run is small and rarely changes; the library API still moves between releases. The high version number is release automation (a new major per breaking change), not age — see [Stability](STABILITY.md).
|
||||||
- **Non-JS repo?** `npx vigiles lint` verifies your CLAUDE.md or AGENTS.md with no install (Ruff/Clippy/Pylint/… too).
|
- **Non-JS repo?** `npx vigiles lint` verifies your CLAUDE.md or AGENTS.md with no install (Ruff/Clippy/Pylint/… too).
|
||||||
|
|
||||||
**[Full FAQ →](docs/faq.md)**
|
**[Full FAQ →](docs/faq.md)**
|
||||||
|
|
||||||
**Not for you if** you want a model/capability benchmark or runtime guardrails in the request path — vigiles is build-/CI-time.
|
**Not for you if** you want a model/capability benchmark or runtime guardrails in the request path — vigiles is build-/CI-time.
|
||||||
|
|
||||||
## More
|
## Docs
|
||||||
|
|
||||||
**Docs** — **[What it catches and prevents →](docs/what-vigiles-catches.md)** · **[Verifying instruction files →](docs/verifying-instruction-files.md)** ([rules matrix](docs/verifying-instruction-files.md#the-validation-rules--the-full-matrix)) · **[Harness testing →](docs/harness-testing.md)** · **[Measuring skills →](docs/measuring-skills.md)** · **[CLI →](docs/cli.md)** · **[GitHub Action →](docs/github-action.md)** · **[Skills →](docs/skills.md)** · **[Plugin-author guide →](docs/for-plugin-authors.md)** · **[Docs index →](docs/README.md)** · **[API reference →](https://zernie.github.io/vigiles/)**
|
The **[docs index](docs/README.md)** is the full map, grouped by what you're doing:
|
||||||
|
|
||||||
**Project** — **[Stability →](STABILITY.md)** · **[Related tools →](docs/related-tools.md)** · companion to [Feedback Loop Is All You Need](https://zernie.com/blog/feedback-loop-is-all-you-need).
|
- **Guides** — [verify instruction files](docs/verifying-instruction-files.md) · [test your harness](docs/harness-testing.md) · [measure a skill](docs/measuring-skills.md) · [ship a plugin](docs/for-plugin-authors.md) · [Codex & other harnesses](docs/harnesses.md)
|
||||||
|
- **Reference** — [CLI](docs/cli.md) · [rules matrix](docs/verifying-instruction-files.md#the-validation-rules--the-full-matrix) · [testing API](docs/testing-api.md) · [full API](https://zernie.github.io/vigiles/)
|
||||||
|
- **Explanation** — [what it catches](docs/what-vigiles-catches.md) · [how it compares](docs/comparison.md) · [FAQ](docs/faq.md)
|
||||||
|
|
||||||
|
**Project** — [Stability](STABILITY.md) · [Related tools](docs/comparison.md#what-vigiles-composes-with) · companion to [Feedback Loop Is All You Need](https://zernie.com/blog/feedback-loop-is-all-you-need).
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
|
|||||||
+11
-7
@@ -1,19 +1,23 @@
|
|||||||
# Stability
|
# Stability
|
||||||
|
|
||||||
> vigiles is **0.x**. This page states exactly what you can depend on today and
|
> vigiles is at **v12** — but read that as _"still moving fast,"_ not
|
||||||
> what may still move, so you can adopt the stable parts now without getting
|
> _"battle-hardened."_ `semantic-release` cuts a **new major on every breaking
|
||||||
> surprised by a change to the parts that are still evolving.
|
> API change**, and there have been a lot of them. The number is an artifact of
|
||||||
|
> how it ships, not a claim of maturity. This page says what I try hardest not
|
||||||
|
> to break, and what's still in motion.
|
||||||
|
|
||||||
Honest beats a fake 1.0: pre-1.0 semver keeps the deeper, still-moving surfaces
|
The steadiest contract is the **CLI** — the commands you run, their flags, and
|
||||||
free to improve while the parts most people actually use stay put.
|
their exit codes. Most of the churn is in the library API underneath it.
|
||||||
|
|
||||||
## What's stable — depend on it
|
## What's stable — depend on it
|
||||||
|
|
||||||
- **The CLI** — the verbs (`init`, `compile`, `lint`, `test`, `eval`,
|
- **The CLI** — the verbs (`init`, `compile`, `lint`, `test`, `eval`,
|
||||||
`scan`, `generate`), their flags, and their **exit codes**
|
`scan`, `generate`), their flags, and their **exit codes**
|
||||||
(`0` clean / `1` warn / `2` error). This is the narrowest, stablest contract
|
(`0` clean / `1` warn / `2` error). This is the narrowest, steadiest contract
|
||||||
and what ~90% of users touch — including the GitHub Action, which wraps it.
|
and the surface almost everyone touches — including the GitHub Action, which wraps it.
|
||||||
- **The authoring + testing library entry points:**
|
- **The authoring + testing library entry points:**
|
||||||
|
- `vigiles/linting` — the compiler + reference verification
|
||||||
|
(`compileClaude`, `compileSkill`, …).
|
||||||
- `vigiles/spec` — the core builders (`enforce`, `guidance`, `claude`,
|
- `vigiles/spec` — the core builders (`enforce`, `guidance`, `claude`,
|
||||||
`skill`, `agent`, `file`, `cmd`, `ref`, `dir`, `glob`, `result`,
|
`skill`, `agent`, `file`, `cmd`, `ref`, `dir`, `glob`, `result`,
|
||||||
`delegate`, `railway`).
|
`delegate`, `railway`).
|
||||||
|
|||||||
+52
-47
@@ -1,60 +1,65 @@
|
|||||||
# Documentation — index
|
# Documentation — index
|
||||||
|
|
||||||
How-to and reference docs for using vigiles. New here? Start with the
|
How-to and reference docs for using vigiles. **New here? Start with the
|
||||||
[README](../README.md).
|
[README](../README.md)** for the pitch and a 5-minute quick start.
|
||||||
|
|
||||||
## Shipping a plugin?
|
The docs are grouped by what you're trying to do:
|
||||||
|
|
||||||
- [`for-plugin-authors.md`](for-plugin-authors.md) — the plugin-author journey end to end: scan a draft for structural health, fix what it flags, make your skills actually fire for users, rank against a marketplace, and gate it in CI.
|
- **[Guides](#guides--help-me-do-x)** — step-by-step, task-first ("help me do X").
|
||||||
|
- **[Reference](#reference--the-exact-flag-symbol-or-rule)** — exact flags, symbols, rules.
|
||||||
|
- **[Explanation](#explanation--why-its-built-this-way)** — the reasoning and trade-offs.
|
||||||
|
|
||||||
## Verify your instruction files (layer 1)
|
---
|
||||||
|
|
||||||
- [`verifying-instruction-files.md`](verifying-instruction-files.md) — the full guide: the markdown→typed-spec ladder, the three rule types (`enforce` / `guidance` / `guard`), verified references + marks, and the before/after tables.
|
## Guides — "help me do X"
|
||||||
- [`skills-monorepo.md`](skills-monorepo.md) — adopt vigiles in a CI-tested skill library or a plain `.claude/` repo (no `plugin.json`): the three repo shapes it loads, the `sharedDirs` opt-in, and what a `SKILL.md` body ref resolves.
|
|
||||||
|
|
||||||
## Guard the harness — compiled hooks
|
### Verify your instruction files (the Lint layer)
|
||||||
|
|
||||||
- [`compiled-hooks.md`](compiled-hooks.md) — author a hook as a pure typed function against the closed `vigiles/hook` vocabulary and compile it, making whole classes of hook bugs unrepresentable (false confidence, matcher bypass, capability creep). The deterministic gate instrument beside verify + test.
|
- [`verifying-instruction-files.md`](verifying-instruction-files.md) — the master guide: the markdown → typed-spec ladder, the three rule types (`enforce` / `guidance` / `guard`), verified references, and the before/after tables. Holds the [full validation-rules matrix](verifying-instruction-files.md#the-validation-rules--the-full-matrix).
|
||||||
|
- [`markdown-mode.md`](markdown-mode.md) — the no-spec on-ramp: verify rules in plain markdown with inline `<!-- vigiles:enforce -->` comments, no TypeScript.
|
||||||
|
- [`skills-monorepo.md`](skills-monorepo.md) — adopt vigiles in a CI-tested skill library or a plain `.claude/` repo (no `plugin.json`).
|
||||||
|
|
||||||
## Two on-ramps — plain markdown → typed spec
|
### Test & measure your harness (the Test + Eval layers)
|
||||||
|
|
||||||
- [`markdown-mode.md`](markdown-mode.md) — verify rules in plain markdown with inline `<!-- vigiles:enforce -->` comments, no TypeScript (frontmatter is a kept, demoted advanced option).
|
- [`harness-testing.md`](harness-testing.md) — task-first how-to: pick what you want to test (hook / wiring / skill firing / behaviour) and the tier that answers it, with a copy-paste first test and CI.
|
||||||
- [`inline-mode.md`](inline-mode.md) — inline-comment mode in depth.
|
- [`harness-testing-claude-code.md`](harness-testing-claude-code.md) — Claude Code specifics: `scriptModel`, `${CLAUDE_PLUGIN_ROOT}` / `pluginDir` / the `Skill` tool, the bubblewrap sandbox.
|
||||||
- [`spec-format.md`](spec-format.md) — the typed `.spec.ts` format (target, sections, rules, verified references) — the source of truth.
|
- [`harness-testing-codex.md`](harness-testing-codex.md) — Codex specifics: `runHarnessTest({ adapter: codexAdapter })` against real `codex exec`, the Responses mock, what maps and what doesn't.
|
||||||
|
- [`measuring-skills.md`](measuring-skills.md) — A/B a skill, plugin, model, or rule change on real coding tasks: the metric triple (bill / target / blast-radius), the worked example, and why it's affordable on your subscription.
|
||||||
|
- [`migrating-from-promptfoo.md`](migrating-from-promptfoo.md) — move existing skill evals onto the subscription: the concept + assertion mapping, a worked side-by-side, and the honest gaps.
|
||||||
|
|
||||||
## Reference
|
### Author & ship
|
||||||
|
|
||||||
- **Library entry points** (grouped by concern, so a future non-Claude-Code harness can sit beside the current one):
|
|
||||||
- `vigiles/linting` — Layer 1: the spec builders + compiler (`claude`, `enforce`, `guidance`, `file`, `cmd`, `symbol`, …).
|
|
||||||
- `vigiles/testing` — Layer 2: the three tiers (`runHook`, `runHarnessTest`, `runEval`) + the runner-agnostic assertions.
|
|
||||||
- `vigiles/claude-code` — the Claude Code-specific adapter (`loadPlugin`, `scriptModel`, the mock).
|
|
||||||
- `vigiles/spec` — the authoring surface (the spec builders; also the module-augmentation target for generated types).
|
|
||||||
- Per-tier barrels `vigiles/unit` / `vigiles/integration` / `vigiles/e2e` make a test's capability legible from its import.
|
|
||||||
- **[API reference (generated) →](https://zernie.github.io/vigiles/)** — every exported symbol across all entry points, generated from the source by API Documenter and published to GitHub Pages. The hand-written guides here are the human-facing layer; this is the exhaustive symbol-level reference.
|
|
||||||
- [`cli.md`](cli.md) — the full CLI, the Claude Code plugin, and the `lint` validation rules.
|
|
||||||
- [`github-action.md`](github-action.md) — run vigiles in CI: the composite Action, every input, the sticky PR comment, versioning.
|
|
||||||
- [`linter-support.md`](linter-support.md) — the 7 linter catalogs + `generate-types` / `generate-schema`.
|
|
||||||
- [`comparison.md`](comparison.md) — before/after tables, the determinism breakdown, the flow diagram.
|
|
||||||
- [`related-tools.md`](related-tools.md) — what vigiles composes with rather than replaces.
|
|
||||||
- **Validation rules:** [`require-instructions-spec`](rules/require-instructions-spec.md) · [`require-skill-spec`](rules/require-skill-spec.md) · [`integrity`](rules/integrity.md) · [`coverage`](rules/coverage.md) · [`untested-skill`](rules/untested-skill.md) · [`untested-subagent`](rules/untested-subagent.md) · [`untested-hook`](rules/untested-hook.md) · [`unmarked-refs`](rules/unmarked-refs.md).
|
|
||||||
|
|
||||||
## Test your harness (layer 2)
|
|
||||||
|
|
||||||
- [`harness-testing.md`](harness-testing.md) — the task-first how-to guide: pick what you want to test (hook / wiring / skill firing / behaviour) and the tier that answers it, with a copy-paste first test, CI, and the coverage table.
|
|
||||||
- [`testing-api.md`](testing-api.md) — the full API reference: every predicate, assertion, `check`, matcher, and option (`measureTriggerRate` / `runEval` / significance), plus imports & harness selection.
|
|
||||||
- [`harness-testing-claude-code.md`](harness-testing-claude-code.md) — Claude Code specifics: the oh-my-claudecode walkthrough, `${CLAUDE_PLUGIN_ROOT}` / `pluginDir` / the `Skill` tool, `scriptModel`, the bubblewrap sandbox.
|
|
||||||
- [`harness-testing-codex.md`](harness-testing-codex.md) — Codex specifics: `runHarnessTest({ adapter: codexAdapter })` against real `codex exec`, the OpenAI Responses mock, what maps and what doesn't.
|
|
||||||
- [`testing-matrix.md`](testing-matrix.md) — every use case mapped to its test tier and file.
|
|
||||||
- [`sandboxing.md`](sandboxing.md) — what the sandbox isolates vs records (honestly): IO / `rm -rf`, the three network modes (deny-all / `recordEgress` / allowlisted `egress: { allow }`), tiers and limits.
|
|
||||||
|
|
||||||
## Measure what works (layer 3)
|
|
||||||
|
|
||||||
- [`measuring-skills.md`](measuring-skills.md) — A/B a skill, plugin, model, or rule change on real coding tasks: the metric triple (bill / target / blast-radius correctness), the worked `measureArms` example, the ecosystem benchmark, and why it's affordable on your subscription.
|
|
||||||
- [`migrating-from-promptfoo.md`](migrating-from-promptfoo.md) — move existing skill evals onto the subscription: the concept + assertion mapping, a side-by-side worked example, and the honest gaps (redteam).
|
|
||||||
- [`eval-architecture.md`](eval-architecture.md) — the cost model + the two testing verbs reconciled with what ships.
|
|
||||||
|
|
||||||
## Skills & agents
|
|
||||||
|
|
||||||
- [`skills.md`](skills.md) — authoring a SKILL.md across the three on-ramps; the prose-vs-gates split.
|
- [`skills.md`](skills.md) — authoring a SKILL.md across the three on-ramps; the prose-vs-gates split.
|
||||||
- [`agent-setup.md`](agent-setup.md) — non-interactive setup for agents (hooks via settings.json).
|
- [`compiled-hooks.md`](compiled-hooks.md) — author a hook as a pure typed function against the closed `vigiles/hook` vocabulary and compile it, making whole classes of hook bugs unrepresentable (false confidence, matcher bypass, capability creep).
|
||||||
- [`agent-workflows.md`](agent-workflows.md) — workflows for Claude Code, Codex, Cursor, multi-agent.
|
- [`railway-subagents.md`](railway-subagents.md) — the typed `Result` subagent contract: declare a typed outcome, compose flat workers, and assert the outcome deterministically (no model judge).
|
||||||
|
- [`for-plugin-authors.md`](for-plugin-authors.md) — the plugin-author journey end to end: scan a draft, fix what it flags, make your skills fire, rank against a marketplace, gate it in CI.
|
||||||
|
- [`github-action.md`](github-action.md) — run vigiles in CI: the composite Action, every input, the sticky PR comment, versioning.
|
||||||
|
|
||||||
|
### Harnesses, adapters & agents
|
||||||
|
|
||||||
|
- [`harnesses.md`](harnesses.md) — which harness vigiles targets and how you pick one (by import), plus the capability matrix.
|
||||||
|
- [`authoring-an-adapter.md`](authoring-an-adapter.md) — teach vigiles a new harness: the five ports, a worked skeleton, validating with the conformance kit.
|
||||||
|
- [`agent-setup.md`](agent-setup.md) — agent setup & workflows in one guide: what `init` does, per-agent recipes (Claude Code / Codex / multi-agent / Cursor), non-interactive setup + fallback hooks, and CI.
|
||||||
|
|
||||||
|
## Reference — "the exact flag, symbol, or rule"
|
||||||
|
|
||||||
|
- [`cli.md`](cli.md) — the full CLI: every verb and flag, the Claude Code plugin, `lint` vs `audit`.
|
||||||
|
- [`testing-api.md`](testing-api.md) — the full harness-testing API: every predicate, assertion, `check`, matcher, and option (`measureTriggerRate` / `runEval` / significance).
|
||||||
|
- [`spec-format.md`](spec-format.md) — the typed `.spec.ts` format (target, sections, rules, verified references) — the source of truth.
|
||||||
|
- [`linter-support.md`](linter-support.md) — the 7 linter catalogs + `generate-types` / `generate-schema`.
|
||||||
|
- [`adapter-api.md`](adapter-api.md) — the adapter API reference: every port field, the conformance functions, the registry API.
|
||||||
|
- **Validation rules:** the [full matrix](verifying-instruction-files.md#the-validation-rules--the-full-matrix) lives in the linting guide; each rule has a doc under [`rules/`](rules/).
|
||||||
|
- **Library entry points** (grouped by concern so a future harness can sit beside the current one):
|
||||||
|
- `vigiles/linting` — the compiler + reference verification (`compileClaude`, `compileSkill`, …).
|
||||||
|
- `vigiles/spec` — the spec builders (`claude`, `enforce`, `guidance`, `file`, `cmd`, `symbol`, …) and the module-augmentation target for generated types.
|
||||||
|
- `vigiles/testing`, `vigiles/unit` — the harness-test tiers + the `check` vocabulary + runner-agnostic assertions.
|
||||||
|
- `vigiles/claude-code`, `vigiles/codex` — the per-harness adapters.
|
||||||
|
- `vigiles/adapter` — the adapter-authoring kit.
|
||||||
|
- **[API reference (generated) →](https://zernie.github.io/vigiles/)** — every exported symbol across all entry points, generated from the source. The hand-written guides here are the human-facing layer; this is the exhaustive symbol-level reference.
|
||||||
|
|
||||||
|
## Explanation — "why it's built this way"
|
||||||
|
|
||||||
|
- [`what-vigiles-catches.md`](what-vigiles-catches.md) — the taxonomy of problems vigiles handles: the prevented / caught / measured model, biggest-problem-first.
|
||||||
|
- [`comparison.md`](comparison.md) — before/after tables, the determinism breakdown, the flow diagram, and what vigiles composes with rather than replaces.
|
||||||
|
- [`sandboxing.md`](sandboxing.md) — what the sandbox isolates vs records (honestly): IO / `rm -rf`, the three network modes, tiers and limits.
|
||||||
|
- [`faq.md`](faq.md) — the front-door FAQ across all four layers.
|
||||||
|
|||||||
+124
-66
@@ -1,58 +1,122 @@
|
|||||||
# Agent Setup Guide
|
# Agent setup & workflows
|
||||||
|
|
||||||
**One command is all you need.** `npx vigiles init` handles everything non-interactively — the agent runs it, skills and hooks take over, and there are no manual chores afterward. This guide shows what happens and what fallbacks exist.
|
**One command is all you need.** `npx vigiles init` handles everything
|
||||||
|
non-interactively — the agent runs it, the installed skills and hooks take over
|
||||||
|
(auto-compiling specs, blocking stray edits, nudging when something needs
|
||||||
|
attention), and there are no manual chores afterward. This guide covers what
|
||||||
|
`init` does, the per-agent specifics, and the fallbacks.
|
||||||
|
|
||||||
→ Back to [README](../README.md)
|
→ Back to [README](../README.md)
|
||||||
|
|
||||||
## Contents
|
## Contents
|
||||||
|
|
||||||
- [What an Agent Can Do](#what-an-agent-can-do)
|
- [What `init` does](#what-init-does)
|
||||||
- [Non-Interactive Setup](#non-interactive-setup)
|
- [Per-agent](#per-agent) — [Claude Code](#claude-code) · [Codex / Copilot](#codex--github-copilot) · [Multi-agent](#multi-agent-claude--codex) · [Cursor / Windsurf](#cursor--windsurf--other-formats)
|
||||||
- [Step 1: Run the wizard](#step-1-run-the-wizard)
|
- [Non-interactive setup (agents & CI)](#non-interactive-setup-agents--ci)
|
||||||
- [Step 2: Install hooks directly (fallback)](#step-2-install-hooks-directly-fallback)
|
- [CI pipeline](#ci-pipeline)
|
||||||
- [Step 3: Edit the spec](#step-3-edit-the-spec)
|
- [What the agent gets wrong](#what-the-agent-gets-wrong)
|
||||||
- [Step 4: Compile and verify](#step-4-compile-and-verify)
|
|
||||||
- [Recommended Agent Prompt](#recommended-agent-prompt)
|
|
||||||
- [What the Agent Gets Wrong](#what-the-agent-gets-wrong)
|
|
||||||
- [See also](#see-also)
|
- [See also](#see-also)
|
||||||
|
|
||||||
## What an Agent Can Do
|
## What `init` does
|
||||||
|
|
||||||
| Action | Agent can do it? | How |
|
`vigiles init` scans your project, **auto-detects which agents you already use**,
|
||||||
| ------------------------ | ---------------- | -------------------------------------------------------------------------------------------------------------------------------- |
|
and sets both layers up. No `--target` flag needed unless you want to override.
|
||||||
| Create spec file | Yes | `npx vigiles init` (non-interactive wizard) |
|
|
||||||
| Generate types | Yes | `npx vigiles generate types` |
|
|
||||||
| Compile specs | Yes | `npx vigiles compile` |
|
|
||||||
| Add the dev dependency | Yes | `npx vigiles init` adds `vigiles` to devDeps |
|
|
||||||
| Add CI step | Yes | Edit `.github/workflows/*.yml` directly |
|
|
||||||
| Install the plugin | Maybe | `claude plugin install vigiles@vigiles` if the `claude` CLI is on PATH; else the user runs the two `/plugin` commands in-session |
|
|
||||||
| Install hooks (fallback) | Yes | Write to `.claude/settings.json` directly |
|
|
||||||
|
|
||||||
**The plugin installs globally** — into `~/.claude/plugins/`, never vendored into your repo. `init` calls the `claude plugin` CLI when available. If the `claude` CLI is not on PATH, it prints the two in-session slash commands for the user to run. An agent that can't reach `claude` at all can still get hook behaviour by writing directly to `.claude/settings.json` (see Step 2 below).
|
| Signal | What it means |
|
||||||
|
| ------------------------------------------ | ------------------------------------------------- |
|
||||||
|
| `CLAUDE.md` exists | Claude Code in use — suggest migration if no spec |
|
||||||
|
| `AGENTS.md` exists | Codex / GitHub Copilot in use |
|
||||||
|
| `.claude/` directory | Claude Code project config |
|
||||||
|
| `.cursorrules` | Cursor in use — suggest rule-porter |
|
||||||
|
| `.github/copilot-instructions.md` | GitHub Copilot custom instructions |
|
||||||
|
| `.windsurfrules` | Windsurf in use |
|
||||||
|
| `rule-porter` / `rulesync` in package.json | Sync tool already installed |
|
||||||
|
|
||||||
## Non-Interactive Setup
|
**What it sets up by default:**
|
||||||
|
|
||||||
### Step 1: Run the wizard
|
|
||||||
|
|
||||||
```bash
|
|
||||||
npx vigiles init # or `npx vigiles init --yes` to be explicit
|
|
||||||
```
|
|
||||||
|
|
||||||
`init` **auto-detects a non-TTY** and runs without prompts. A user prompt as simple as _"set up vigiles in this repo"_ is enough — the agent runs this command and gets sensible defaults.
|
|
||||||
|
|
||||||
**What `init` sets up by default:**
|
|
||||||
|
|
||||||
- **Lint layer** — a typed `.spec.ts` + generated types
|
- **Lint layer** — a typed `.spec.ts` + generated types
|
||||||
- **Test layer** — a starter `vigiles.harness.mjs`
|
- **Test layer** — a starter `vigiles.harness.mjs`
|
||||||
- **CI** — a `zernie/vigiles@v1` workflow at `.github/workflows/vigiles.yml`
|
- **CI** — a `zernie/vigiles@v1` workflow at `.github/workflows/vigiles.yml`
|
||||||
- **Dependency** — `vigiles` added to `devDependencies`
|
- **Dependency** — `vigiles` added to `devDependencies`
|
||||||
- **Plugin** — Claude Code plugin installed via the marketplace
|
- **Plugin** — the Claude Code plugin, installed **globally** via the marketplace (into `~/.claude/plugins/`, never vendored into your repo)
|
||||||
|
|
||||||
Scope it with flags when needed: `--lint`, `--test` (one layer or both), `--harness=claude,codex`, `--no-gha`, `--no-plugin`, `--strict`. A human running it in a terminal gets interactive prompts instead.
|
Scope with flags: `--lint` / `--test` (one layer or both), `--harness=claude,codex`,
|
||||||
|
`--no-gha`, `--no-plugin`, `--strict`.
|
||||||
|
|
||||||
### Step 2: Install hooks directly (fallback)
|
## Per-agent
|
||||||
|
|
||||||
The plugin already brings the hooks. If you'd rather commit project-level hooks instead of (or alongside) the plugin, write them to `.claude/settings.json`:
|
### Claude Code
|
||||||
|
|
||||||
|
Instruction file: `CLAUDE.md`. Once the plugin is installed, the agent no longer
|
||||||
|
has to remember to compile:
|
||||||
|
|
||||||
|
| Hook | Trigger | Action |
|
||||||
|
| ----------- | ----------------------------------------------- | ---------------------------------------- |
|
||||||
|
| PreToolUse | Agent tries to Edit/Write a compiled `.md` file | Blocks the edit, redirects to `.spec.ts` |
|
||||||
|
| PostToolUse | Agent edits a `.spec.ts` file | Auto-runs `vigiles compile` |
|
||||||
|
| PostToolUse | Agent edits linter config or `package.json` | Auto-runs `vigiles generate types` |
|
||||||
|
|
||||||
|
`init` installs the plugin via the marketplace; by hand in a Claude Code session:
|
||||||
|
|
||||||
|
```
|
||||||
|
/plugin marketplace add zernie/vigiles
|
||||||
|
/plugin install vigiles@vigiles
|
||||||
|
```
|
||||||
|
|
||||||
|
⚠️ **Without the plugin**, run `vigiles compile` manually after editing specs. CI still catches stale files.
|
||||||
|
|
||||||
|
### Codex / GitHub Copilot
|
||||||
|
|
||||||
|
Instruction file: `AGENTS.md`, read directly — there is no plugin or hook system.
|
||||||
|
The enforcement path is:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npx vigiles init --harness=codex # full setup: scaffolds AGENTS.md.spec.ts + types + CI + Codex skills
|
||||||
|
# 1. edit AGENTS.md.spec.ts (source of truth)
|
||||||
|
# 2. npx vigiles compile → regenerates AGENTS.md
|
||||||
|
# 3. CI: npx vigiles lint && npx vigiles generate types --check
|
||||||
|
```
|
||||||
|
|
||||||
|
Use the full `init --harness=codex` (not `init --target=AGENTS.md`, which only
|
||||||
|
scaffolds the spec) — it's what generates `.vigiles/generated.d.ts` and the CI
|
||||||
|
config that step 3's `generate types --check` depends on. Authoring skills install
|
||||||
|
**globally** via the cross-agent `skills` CLI (no repo vendoring):
|
||||||
|
`npx skills add zernie/vigiles -a codex -g -y` → `~/.agents/skills/`, which that
|
||||||
|
same command handles. Codex hooks (`.codex/config.toml [hooks]`) aren't auto-wired yet.
|
||||||
|
|
||||||
|
### Multi-agent (Claude + Codex)
|
||||||
|
|
||||||
|
Use a **single spec with multiple targets** — one source of truth, two outputs:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
export default claude({
|
||||||
|
target: ["CLAUDE.md", "AGENTS.md"],
|
||||||
|
rules: { ... },
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
Both compile from the same spec with the same linter verification.
|
||||||
|
|
||||||
|
### Cursor / Windsurf / other formats
|
||||||
|
|
||||||
|
vigiles compiles to **markdown only** (CLAUDE.md, AGENTS.md). For non-markdown
|
||||||
|
formats (`.cursorrules`, `.github/copilot-instructions.md`, Windsurf), use a sync
|
||||||
|
tool to convert from the compiled markdown — [rule-porter](https://github.com/nichochar/rule-porter)
|
||||||
|
or [rulesync](https://github.com/dyoshikawa/rulesync). vigiles is the source-of-truth
|
||||||
|
compiler; sync tools handle the last mile.
|
||||||
|
|
||||||
|
## Non-interactive setup (agents & CI)
|
||||||
|
|
||||||
|
`init` **auto-detects a non-TTY** and runs without prompts — a prompt as simple as
|
||||||
|
_"set up vigiles in this repo"_ is enough:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npx vigiles init # or `npx vigiles init --yes` to be explicit
|
||||||
|
```
|
||||||
|
|
||||||
|
**Fallback — install hooks directly.** The plugin already brings the hooks. To
|
||||||
|
commit project-level hooks instead of (or alongside) the plugin, write them to
|
||||||
|
`.claude/settings.json`:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
@@ -73,44 +137,38 @@ The plugin already brings the hooks. If you'd rather commit project-level hooks
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
This is equivalent to what the plugin installs, but written directly without the skills system.
|
**Recommended agent prompt** — if you want an agent to set up vigiles:
|
||||||
|
|
||||||
### Step 3: Edit the spec
|
|
||||||
|
|
||||||
**The agent reads the generated `.spec.ts`** and fills in the project's actual conventions — sections, key files, commands, and rules. Use the `edit-spec` skill instructions as a guide for the spec format.
|
|
||||||
|
|
||||||
### Step 4: Compile and verify
|
|
||||||
|
|
||||||
```bash
|
|
||||||
npx vigiles compile
|
|
||||||
npx vigiles lint
|
|
||||||
```
|
|
||||||
|
|
||||||
## Recommended Agent Prompt
|
|
||||||
|
|
||||||
If you want an agent to set up vigiles in a project, use this prompt:
|
|
||||||
|
|
||||||
```
|
```
|
||||||
Set up vigiles for this project:
|
Set up vigiles for this project:
|
||||||
1. Run `npx vigiles init` (it adds vigiles to devDependencies and installs the
|
1. Run `npx vigiles init` (adds vigiles to devDependencies and installs the
|
||||||
Claude Code plugin via the marketplace — nothing is vendored into the repo)
|
Claude Code plugin via the marketplace — nothing is vendored into the repo)
|
||||||
2. Read the generated .spec.ts file
|
2. Read the generated .spec.ts, fill in the project's actual conventions
|
||||||
3. Fill in the project's actual conventions based on the codebase
|
3. Run `npm install`, then `npx vigiles compile` to verify
|
||||||
4. Run `npm install`, then `npx vigiles compile` to verify everything works
|
4. Commit the .spec.ts, compiled .md, .vigiles/generated.d.ts, and package.json
|
||||||
5. Commit the .spec.ts, compiled .md, .vigiles/generated.d.ts, and package.json
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## What the Agent Gets Wrong
|
## CI pipeline
|
||||||
|
|
||||||
Common issues when agents set up vigiles:
|
All agents share the same CI step:
|
||||||
|
|
||||||
- **Editing CLAUDE.md directly** — the PreToolUse hook prevents this if installed
|
```yaml
|
||||||
- **Using wrong rule names** — `enforce("no-console")` instead of `enforce("eslint/no-console")`. The compiler catches this.
|
- name: Verify specs
|
||||||
- **Forgetting to compile** — the PostToolUse hook handles this automatically
|
run: npx vigiles lint && npx vigiles generate types --check
|
||||||
- **Adding headers inside sections** — the compiler catches `#`/`##` headers in section content
|
```
|
||||||
|
|
||||||
|
It catches hash mismatches (someone edited the compiled `.md`), missing specs
|
||||||
|
(`require-instructions-spec`), and stale generated types.
|
||||||
|
|
||||||
|
## What the agent gets wrong
|
||||||
|
|
||||||
|
- **Editing CLAUDE.md directly** — the PreToolUse hook prevents this if installed.
|
||||||
|
- **Wrong rule names** — `enforce("no-console")` instead of `enforce("eslint/no-console")`. The compiler catches it.
|
||||||
|
- **Forgetting to compile** — the PostToolUse hook handles it automatically.
|
||||||
|
- **Headers inside sections** — the compiler catches `#`/`##` headers in section content.
|
||||||
|
|
||||||
## See also
|
## See also
|
||||||
|
|
||||||
- [Agent Workflows](agent-workflows.md) — per-agent setup (Claude Code, Codex, Cursor, CI)
|
- [Markdown mode](markdown-mode.md) — the no-spec on-ramp (inline `<!-- vigiles:enforce -->` comments).
|
||||||
- [Markdown mode](markdown-mode.md) — inline comments and frontmatter (no `.spec.ts` required)
|
- [CLI reference](cli.md) — every verb and flag.
|
||||||
- [CLI reference](cli.md)
|
- [Harnesses](harnesses.md) — how vigiles targets Claude Code, Codex, and beyond.
|
||||||
|
|||||||
@@ -1,127 +0,0 @@
|
|||||||
# Agent Workflows
|
|
||||||
|
|
||||||
**vigiles is low-friction by design.** Run `npx vigiles init` and the installed skills and hooks handle the rest — auto-compiling specs, blocking stray edits, and nudging the agent when something needs attention. This guide shows the per-agent setup.
|
|
||||||
|
|
||||||
→ Back to [README](../README.md)
|
|
||||||
|
|
||||||
vigiles verifies the rule references in agent instruction files — declared as inline comments, `vigiles:` YAML frontmatter, or a typed spec compiled to markdown ([markdown mode](markdown-mode.md)). Different AI agents read different files, but the validation pipeline is the same. The workflows below use spec mode, the deepest level; the inline and frontmatter levels need no build step.
|
|
||||||
|
|
||||||
## Contents
|
|
||||||
|
|
||||||
- [Auto-Detection](#auto-detection)
|
|
||||||
- [Claude Code](#claude-code)
|
|
||||||
- [Codex / GitHub Copilot](#codex--github-copilot)
|
|
||||||
- [Multi-Agent (Claude + Codex)](#multi-agent-claude--codex)
|
|
||||||
- [Cursor / Windsurf / Other Formats](#cursor--windsurf--other-formats)
|
|
||||||
- [CI Pipeline](#ci-pipeline)
|
|
||||||
- [See also](#see-also)
|
|
||||||
|
|
||||||
## Auto-Detection
|
|
||||||
|
|
||||||
`vigiles init` scans your project and **auto-detects which agents you're already using** — no `--target` flag needed unless you want to override.
|
|
||||||
|
|
||||||
| Signal | What it means |
|
|
||||||
| ------------------------------------------ | ------------------------------------------------- |
|
|
||||||
| `CLAUDE.md` exists | Claude Code in use — suggest migration if no spec |
|
|
||||||
| `AGENTS.md` exists | Codex / GitHub Copilot in use |
|
|
||||||
| `.claude/` directory | Claude Code project config |
|
|
||||||
| `.cursorrules` | Cursor in use — suggest rule-porter |
|
|
||||||
| `.github/copilot-instructions.md` | GitHub Copilot custom instructions |
|
|
||||||
| `.windsurfrules` | Windsurf in use |
|
|
||||||
| `rule-porter` / `rulesync` in package.json | Sync tool already installed |
|
|
||||||
| Symlinked instruction files | Notes them in output |
|
|
||||||
|
|
||||||
The wizard creates specs for detected targets, generates types, compiles, and adds a CI step.
|
|
||||||
|
|
||||||
## Claude Code
|
|
||||||
|
|
||||||
**Instruction file:** `CLAUDE.md`
|
|
||||||
|
|
||||||
**Setup:**
|
|
||||||
|
|
||||||
```bash
|
|
||||||
npx vigiles init
|
|
||||||
# init installs the plugin via the marketplace; to do it by hand in Claude Code:
|
|
||||||
# /plugin marketplace add zernie/vigiles
|
|
||||||
# /plugin install vigiles@vigiles
|
|
||||||
```
|
|
||||||
|
|
||||||
**What the plugin does once installed** — the agent no longer needs to remember to compile:
|
|
||||||
|
|
||||||
| Hook | Trigger | Action |
|
|
||||||
| ----------- | ----------------------------------------------- | ---------------------------------------- |
|
|
||||||
| PreToolUse | Agent tries to Edit/Write a compiled `.md` file | Blocks the edit, redirects to `.spec.ts` |
|
|
||||||
| PostToolUse | Agent edits a `.spec.ts` file | Auto-runs `vigiles compile` |
|
|
||||||
| PostToolUse | Agent edits linter config or `package.json` | Auto-runs `vigiles generate types` |
|
|
||||||
|
|
||||||
⚠️ **Without the plugin**, you must run `vigiles compile` manually after editing specs. CI still catches stale files.
|
|
||||||
|
|
||||||
## Codex / GitHub Copilot
|
|
||||||
|
|
||||||
**Instruction file:** `AGENTS.md`
|
|
||||||
|
|
||||||
**Setup:**
|
|
||||||
|
|
||||||
```bash
|
|
||||||
npx vigiles init --target=AGENTS.md
|
|
||||||
```
|
|
||||||
|
|
||||||
Codex and GitHub Copilot read `AGENTS.md` directly. There is no plugin or hook system — these agents don't support it. The enforcement path is:
|
|
||||||
|
|
||||||
1. Edit `AGENTS.md.spec.ts` (the source of truth)
|
|
||||||
2. Run `npx vigiles compile` to regenerate `AGENTS.md`
|
|
||||||
3. CI verifies freshness: `npx vigiles lint && npx vigiles generate types --check`
|
|
||||||
|
|
||||||
**Authoring skills for Codex** install globally via the cross-agent `skills` CLI — no repo vendoring: `npx skills add zernie/vigiles -a codex -g -y` installs into `~/.agents/skills/`. `vigiles init --harness=codex` runs this automatically. Codex hooks (`.codex/config.toml [hooks]`) aren't auto-wired yet.
|
|
||||||
|
|
||||||
ℹ️ **If you also use Claude Code**, install the plugin (`/plugin marketplace add zernie/vigiles` then `/plugin install vigiles@vigiles`, or `vigiles init`) for auto-recompilation.
|
|
||||||
|
|
||||||
## Multi-Agent (Claude + Codex)
|
|
||||||
|
|
||||||
Use a **single spec with multiple targets** — one source of truth, two outputs:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
export default claude({
|
|
||||||
target: ["CLAUDE.md", "AGENTS.md"],
|
|
||||||
rules: { ... },
|
|
||||||
});
|
|
||||||
```
|
|
||||||
|
|
||||||
Both files compile from the same spec with the same linter verification.
|
|
||||||
|
|
||||||
```bash
|
|
||||||
npx vigiles init # for CLAUDE.md (primary)
|
|
||||||
npx vigiles init --target=AGENTS.md # adds AGENTS.md target
|
|
||||||
```
|
|
||||||
|
|
||||||
Or just set `target: ["CLAUDE.md", "AGENTS.md"]` in your spec directly.
|
|
||||||
|
|
||||||
## Cursor / Windsurf / Other Formats
|
|
||||||
|
|
||||||
vigiles compiles to **markdown only** (CLAUDE.md, AGENTS.md). For non-markdown formats (`.cursorrules`, `.github/copilot-instructions.md`, Windsurf), use a sync tool to convert from the compiled markdown:
|
|
||||||
|
|
||||||
- [rule-porter](https://github.com/nichochar/rule-porter) — bidirectional conversion between agent formats
|
|
||||||
- [rulesync](https://github.com/dyoshikawa/rulesync) — unified rule management across 10+ tools
|
|
||||||
|
|
||||||
vigiles is the source of truth compiler. Sync tools handle the last mile.
|
|
||||||
|
|
||||||
## CI Pipeline
|
|
||||||
|
|
||||||
All agents share the same CI step:
|
|
||||||
|
|
||||||
```yaml
|
|
||||||
- name: Verify specs
|
|
||||||
run: npx vigiles lint && npx vigiles generate types --check
|
|
||||||
```
|
|
||||||
|
|
||||||
This catches:
|
|
||||||
|
|
||||||
- **Hash mismatches** — someone edited the compiled `.md` directly
|
|
||||||
- **Missing specs** — `require-instructions-spec` rule requires a `.spec.ts` behind every `.md`
|
|
||||||
- **Stale generated types** — linter config changed but types weren't regenerated
|
|
||||||
|
|
||||||
## See also
|
|
||||||
|
|
||||||
- [Agent Setup](agent-setup.md) — non-interactive installation and recommended agent prompt
|
|
||||||
- [Markdown mode](markdown-mode.md) — inline comments and frontmatter (no `.spec.ts` required)
|
|
||||||
- [CLI reference](cli.md)
|
|
||||||
+1
-2
@@ -148,8 +148,7 @@ See the [rules matrix](verifying-instruction-files.md#the-validation-rules--the-
|
|||||||
`vigiles lint` accepts files **or a directory** (`vigiles lint .` discovers the
|
`vigiles lint` accepts files **or a directory** (`vigiles lint .` discovers the
|
||||||
instruction files under it); with no argument it discovers them from the repo root.
|
instruction files under it); with no argument it discovers them from the repo root.
|
||||||
|
|
||||||
See the [agent setup guide](agent-setup.md) and
|
See the [agent setup & workflows guide](agent-setup.md).
|
||||||
[agent workflows](agent-workflows.md).
|
|
||||||
|
|
||||||
### `compile [files...]` — harness selection
|
### `compile [files...]` — harness selection
|
||||||
|
|
||||||
|
|||||||
@@ -63,6 +63,23 @@ Out of scope — use other tools:
|
|||||||
|
|
||||||
Illustrative code blocks (typo demos, template placeholders, speculative refs in design docs) opt out via `<!-- vigiles:ignore -->` immediately before the fence, or `<!-- vigiles:ignore-file -->` anywhere in a file that's entirely illustrative. Placeholders containing `<` or `>` are auto-skipped. Refs that can't be verified because the underlying tool isn't installed (e.g. `pylint/X` on a machine without pylint) are reported separately from real errors.
|
Illustrative code blocks (typo demos, template placeholders, speculative refs in design docs) opt out via `<!-- vigiles:ignore -->` immediately before the fence, or `<!-- vigiles:ignore-file -->` anywhere in a file that's entirely illustrative. Placeholders containing `<` or `>` are auto-skipped. Refs that can't be verified because the underlying tool isn't installed (e.g. `pylint/X` on a machine without pylint) are reported separately from real errors.
|
||||||
|
|
||||||
|
## What vigiles composes with
|
||||||
|
|
||||||
|
vigiles owns one thing: compile-time verification of typed specs against real
|
||||||
|
linter configs, filesystems, and `package.json`, plus testing the harness those
|
||||||
|
specs describe. Everything else, compose:
|
||||||
|
|
||||||
|
- **Architectural linting** — [ast-grep](https://ast-grep.github.io/), [Dependency Cruiser](https://github.com/sverweij/dependency-cruiser), [Steiger](https://github.com/feature-sliced/steiger). Reference their rules via `enforce()`.
|
||||||
|
- **File sync across agents** — [Ruler](https://github.com/intellectronica/ruler), [rulesync](https://github.com/dyoshikawa/rulesync), [block/ai-rules](https://github.com/block/ai-rules). vigiles compiles the source; sync tools distribute. For non-markdown formats (`.cursorrules`, Copilot), [rule-porter](https://github.com/nichochar/rule-porter) or rulesync convert the compiled output.
|
||||||
|
- **Markdown linting** — [markdownlint](https://github.com/DavidAnson/markdownlint). vigiles generates markdown; structure is correct by construction.
|
||||||
|
- **Code-block linting in docs** — [eslint-plugin-markdown](https://github.com/eslint/eslint-plugin-markdown) for syntax, [twoslash](https://shikijs.github.io/twoslash/) for TS type-checking.
|
||||||
|
- **Prose quality** — [Vale](https://vale.sh). Different concern.
|
||||||
|
- **Runtime LLM rule checking** — opposite paradigm: those tools send your code to a model on every check (tokens, non-reproducible verdicts); vigiles compiles once and checks deterministically forever after with `eslint`, `ruff`, `tsc`, Cedar evaluation.
|
||||||
|
|
||||||
|
Specs compile to `CLAUDE.md` by default; set `target: "AGENTS.md"` or
|
||||||
|
`target: ["CLAUDE.md", "AGENTS.md"]` for multiple outputs from one spec. See the
|
||||||
|
[spec format reference](spec-format.md).
|
||||||
|
|
||||||
## Flow
|
## Flow
|
||||||
|
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -1,718 +0,0 @@
|
|||||||
# Eval architecture — how vigiles tests Claude Code harness features
|
|
||||||
|
|
||||||
> Status: design of record (2026-06-16). Captures the conceptual model behind
|
|
||||||
> the two testing verbs (`vigiles test` / `vigiles eval`), reconciles it with
|
|
||||||
> what the codebase **already** ships, and ranks the genuine remaining gaps into
|
|
||||||
> a build roadmap. Companion to [`harness-testing.md`](harness-testing.md) (the
|
|
||||||
> user guide).
|
|
||||||
|
|
||||||
## The kicker
|
|
||||||
|
|
||||||
Real-model evals are invoked manually (`npm run test:eval`), and their results get
|
|
||||||
frozen as `FINDING:` comments baked into the `*.eval.mjs` files. **A snapshot of a
|
|
||||||
past run is documentation, not protection** — edit a skill description and break
|
|
||||||
its trigger rate, and nothing re-ran the classifier. This doc is the reevaluation.
|
|
||||||
|
|
||||||
> **Now shipped — the eval lock.** That exact gap is closed by a committed
|
|
||||||
> integrity stamp: `vigiles eval --update` (local, on your subscription) records
|
|
||||||
> each named eval's result; `vigiles eval --check` (CI) fails "stale" when an
|
|
||||||
> input changed without a re-run — **without a model call**. See
|
|
||||||
> [The eval lock](#the-eval-lock-the-ci-staleness-gate). The rest of this doc is
|
|
||||||
> the reasoning that led there.
|
|
||||||
|
|
||||||
The honest scope correction up front: the gap is **narrower than "build an eval
|
|
||||||
runner,"** and the fix is **not** "add a GitHub Actions eval job." vigiles already
|
|
||||||
has `vigiles eval` (discovers + runs `*.eval.mjs`), a record/replay cache, a
|
|
||||||
significance-gated baseline, a check vocabulary scored across trials, and
|
|
||||||
trigger-rate with recall **and** precision. The real moves are (1) **run the
|
|
||||||
real-model eval where the subscription already is** — a Claude Code session (the
|
|
||||||
agent loop / web / a scheduled session) or locally, since vigiles drives the
|
|
||||||
`claude` CLI; NOT a metered GitHub Actions workflow (CI runs only the free
|
|
||||||
deterministic tiers); (2) make `vigiles eval` **fail honestly** so a session run
|
|
||||||
can't false-green (`--min`, `--no-skip`, corrupt-cache throw, the Sonnet model
|
|
||||||
floor); (3) **honest model pinning** for cached/baselined results; and (4) the
|
|
||||||
**tool-call spy/fake** for side-effecting skills. See
|
|
||||||
[What already exists](#what-already-exists) before building anything.
|
|
||||||
|
|
||||||
## Positioning & pros/cons (the approach, decided 2026-06-17)
|
|
||||||
|
|
||||||
> The canonical positioning **statement** lives in `CLAUDE.md` (`## Positioning`,
|
|
||||||
> layer 2). This section is the **detailed** pros/cons behind it.
|
|
||||||
|
|
||||||
**The thesis: the harness eval you can actually afford to run.** Almost nobody
|
|
||||||
evals their harness because the usual tools (promptfoo, DeepEval, Braintrust,
|
|
||||||
Inspect) hit the model **API SDK** and bill **per token on every run** → real
|
|
||||||
money on every CI run → so it doesn't get run. vigiles inverts the cost curve two
|
|
||||||
ways: (a) **most harness questions need no model at all** — `runHook` + mock-model
|
|
||||||
`runHarnessTest` answer "does the hook fire/block/inject?" deterministically,
|
|
||||||
free, every commit; (b) when a question **is** irreducibly real-model
|
|
||||||
(does a description _fire_, does behaviour _move_), vigiles drives the **`claude`
|
|
||||||
CLI**, so the eval runs on the **Pro/Max subscription** the user already pays for —
|
|
||||||
in a Claude Code session or locally — not a metered API key in CI. (Confirmed this
|
|
||||||
session: a real eval ran with `apiKeySource:"none"`, i.e. on the OAuth sub.)
|
|
||||||
|
|
||||||
### Pros (why this is defensible)
|
|
||||||
|
|
||||||
- **Cost** — the structural advantage. Free deterministic tiers + sub-priced real-model
|
|
||||||
tier vs competitors' per-token-every-run. This is the only reason a small team
|
|
||||||
will _actually_ eval their harness.
|
|
||||||
- **Fidelity** — the unit under test is the harness **loaded as it ships**
|
|
||||||
(`plugin-loader`: real `plugin.json`/`hooks`/`settings`/`CLAUDE.md`). A
|
|
||||||
YAML-config eval runner reconstructs an agent; it can't host this.
|
|
||||||
- **Honesty** — measures in-plugin with real sibling competition (vs others'
|
|
||||||
optimistic one-skill isolation), on the realistic selector (Sonnet, not haiku),
|
|
||||||
with significance + `pass^k`; `interceptTools` intercepts-and-prevents a
|
|
||||||
side-effecting tool in the **real** hook layer (a safety assertion others can't
|
|
||||||
make).
|
|
||||||
|
|
||||||
### Cons / limits (state them honestly)
|
|
||||||
|
|
||||||
- **The sub is rate-limited.** This works _because_ the real-model surface is thin
|
|
||||||
by design — it is **not** a license for huge trial counts; heavy volume still
|
|
||||||
wants metered API or a higher tier.
|
|
||||||
- **Real-model evals stay non-deterministic** — a statistical rate ± se across
|
|
||||||
trials, never a single-run gate. (The deterministic tiers are the per-commit
|
|
||||||
gate.)
|
|
||||||
- **The tool-call spy is intercept-and-prevent, not a faithful mock** — CC
|
|
||||||
surfaces the deny as a _block_, so it asserts the ATTEMPT, not a continued flow.
|
|
||||||
- **Trigger-rate must run on the realistic model** — a cheap haiku run
|
|
||||||
under-measures selection (dogfooded: 0.50 haiku vs 0.90 Sonnet). The `minModel`
|
|
||||||
floor enforces this.
|
|
||||||
- **Evals aren't a zero-effort CI checkbox** — you run them deliberately in a
|
|
||||||
session, which is a workflow change vs "add a GitHub Action."
|
|
||||||
- **No dataset / red-team / scorer-library / web UI** — that's promptfoo's lane;
|
|
||||||
we bridge or skip, not chase.
|
|
||||||
|
|
||||||
## Coverage & scope — what we test, what we delegate
|
|
||||||
|
|
||||||
What a test needs from the _outside world_ sorts onto **three rungs**, and you
|
|
||||||
**pick the lowest rung that faithfully measures the thing**:
|
|
||||||
|
|
||||||
- **R1 — cheap / deterministic (nothing executes):** hook-fires (`runHook`),
|
|
||||||
trigger-rate recall+precision (`measureTriggerRate`), tool-contract / `notTool`.
|
|
||||||
No tool, no service, no Docker. _"calls a tool" → R1._
|
|
||||||
- **R2 — stub / record-replay:** the deterministic logic consumes a tool/MCP/API
|
|
||||||
**result** that is **recorded ONCE** from a real tool and **replayed** by
|
|
||||||
shadowing the binary on PATH / stubbing the MCP — no live service. **Never**
|
|
||||||
model-synthesized stubs (drift → false confidence); reuse the eval cache's
|
|
||||||
record/replay machinery. _"needs the result" → R2._
|
|
||||||
- **R3 — real disposable service/container:** the real system's **semantics** is
|
|
||||||
what's under test (real SQL vs a real schema, a real browser, a DB/redis/
|
|
||||||
analytics engine). _"real semantics under test" → R3._
|
|
||||||
|
|
||||||
**A second, orthogonal axis decides the cost: the oracle.** The rungs say _what
|
|
||||||
executes_; they do **not** say _who decides pass/fail_. A **deterministic** oracle
|
|
||||||
(hook block/allow, tool-contract, a structural fact) is **free, no model, in CI**;
|
|
||||||
a **model-gated** oracle (does a description **fire**? does prose guidance **change
|
|
||||||
behavior**? is the output good, judged?) needs a **real model — on your
|
|
||||||
subscription**, not metered API. We tag the latter `-MG` (e.g. **`R1-MG`** =
|
|
||||||
nothing executes but only a model can judge — _the case for any prose skill_).
|
|
||||||
Don't read "R1" as "free": a prose guidance skill is **R1-MG** — it's fully
|
|
||||||
testable (trigger-rate + a judged behavioral eval), just on the sub, not in CI.
|
|
||||||
That's the boundary, not a coverage hole.
|
|
||||||
|
|
||||||
**What vigiles can and can't test — three buckets.** Folding both axes:
|
|
||||||
**(A) Free & deterministic** (R1/R2 + deterministic oracle — every commit);
|
|
||||||
**(B) Model-gated on your sub** (`-MG` + model oracle — no metered API);
|
|
||||||
**(C) Needs a real service** (R3 — vigiles **composes** with a container, doesn't
|
|
||||||
run it). **A + B is "testable by vigiles"; only C is delegated.** So grade a plugin
|
|
||||||
with **two numbers, not three**: **"% testable at all (free + sub)"** vs **"% that
|
|
||||||
needs a container"** — and always say which bucket, so "testable" never hides
|
|
||||||
whether it's free or sub-priced.
|
|
||||||
|
|
||||||
**Distribution (blended, scrubbed).** A survey of popular community plugin
|
|
||||||
collections **and** an audit of a ~90-artifact real-world production skill set
|
|
||||||
**converge**: **R1 ≈ 48–90%, R2 ≈ 10–43%, R3 ≈ 0–9%.** Net — **R1+R2 covers
|
|
||||||
~90%+** of real plugin surface with **no Docker, on the subscription**; the R3
|
|
||||||
apex is **thin** and collapses to a handful of real services. Every common
|
|
||||||
SaaS/CLI integration (GitHub / issue-tracker / chat / CI / linters / test-runners)
|
|
||||||
is faithfully **replayable at R2**.
|
|
||||||
|
|
||||||
**The e2e landscape (honest).** Real side-effecting e2e is mature
|
|
||||||
(SWE-bench/Verified, Terminal-Bench, OSWorld, WebArena; the labs' per-task cloud
|
|
||||||
sandboxes; AISI Inspect's Docker sandbox) — but **every one runs inside a
|
|
||||||
container/VM/cloud sandbox.** There is no "safe reproducible e2e without a
|
|
||||||
container," so at R3 vigiles **composes with a container, does not reinvent the
|
|
||||||
sandbox, and does not claim containerless e2e.**
|
|
||||||
|
|
||||||
**Across the axes.** SAFETY: R1 nothing executes; R2 fake outputs, no real system;
|
|
||||||
R3 real side effects only inside an isolated disposable container — layered with
|
|
||||||
provenance confinement + the ephemeral run env + `interceptTools`. VIABILITY: R1+R2
|
|
||||||
need no Docker, run on the sub (affordable + cross-platform); R3 needs Docker.
|
|
||||||
PERFORMANCE: R1/R2 ms-fast deterministic; R3 Docker cold-start is seconds — keep
|
|
||||||
thin. **Non-goals:** containerless reproducible e2e; per-host egress on macOS;
|
|
||||||
verifying vendor MCP connectors' live semantics (vendor's job); becoming a
|
|
||||||
sandbox/orchestrator (compose instead).
|
|
||||||
|
|
||||||
**Competitor comparison.** Completion-graders (promptfoo / DeepEval / Braintrust)
|
|
||||||
— metered API every run, no real-harness load, no cheap no-model tiers.
|
|
||||||
Containerized e2e (SWE-bench / Inspect / Codex) — faithful but heavy / metered /
|
|
||||||
infra. vigiles — owns **R1+R2 + sub-affordability + a clean container hand-off at
|
|
||||||
R3**. The unclaimed seam is R1+R2 + sub-pricing + compose-with-container, **NOT**
|
|
||||||
e2e-without-a-container.
|
|
||||||
|
|
||||||
**Build verdict.** A **PATH-shim / record-replay helper (fake-on-PATH)** is
|
|
||||||
**higher leverage** than a testcontainers integration — it unlocks the ~43% R2
|
|
||||||
with no Docker and covers far more real plugins. Real-service provisioning stays a
|
|
||||||
thin, composed apex.
|
|
||||||
|
|
||||||
## Core model: every harness feature = a deterministic part + a behavioral part
|
|
||||||
|
|
||||||
This is the load-bearing idea. Decompose every harness feature (a skill, a hook,
|
|
||||||
a `CLAUDE.md` rule, a subagent, an MCP server) into two parts:
|
|
||||||
|
|
||||||
| Part | Becomes a | Mechanism | Cost | Runs |
|
|
||||||
| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------- | -------------------------------------------------------------------------- | ---------------------- | ----------------- |
|
|
||||||
| **Deterministic** — does the hook fire/block? does the file parse? do permission rules match? does an extracted prompt-builder/checker produce the right string? does an MCP tool return the right shape? | **TEST** — exact binary assertion | `runHook`, `runHarnessTest`, plain `node:test` over an extracted pure core | free, no model | every PR |
|
|
||||||
| **Behavioral** — does a description _trigger_? does `CLAUDE.md` _change behavior_? does an agent reach the right _outcome_? | **EVAL** — statistical, scored, threshold-gated | `measureTriggerRate`, `runEval`/`measure` A/B | real model → real cost | gated (see knobs) |
|
|
||||||
|
|
||||||
**Most of a feature is the deterministic part, and that's where
|
|
||||||
protection-per-dollar is highest.** Only the irreducibly-stochastic slice runs a
|
|
||||||
model. This mirrors the repo's existing "keep the real-model surface THIN"
|
|
||||||
discipline: of all harness questions, only two
|
|
||||||
are _irreducibly_ real-model — _does a description fire_ and _does behavior move_.
|
|
||||||
|
|
||||||
Practical corollary, and a prerequisite for the dogfood work below: **most
|
|
||||||
testable skill logic is inline guidance prose, not code.** Where a deterministic
|
|
||||||
core is _embedded in a prompt_, you can't test it for free. So the highest-leverage
|
|
||||||
move is often to **extract the deterministic core into a script** (the
|
|
||||||
`prune-illustrate` `generate.sh` with its `STYLE_SUFFIX` constant is the template
|
|
||||||
to copy), then test that script at the free subprocess tier. Extraction converts
|
|
||||||
a behavioral question into a deterministic one — the cheapest possible win.
|
|
||||||
|
|
||||||
## trigger-rate is a classifier eval, not a unit test
|
|
||||||
|
|
||||||
Name it right, because the name dictates how you run it. The "unit under test" is
|
|
||||||
the **model's routing decision** (stochastic); the artifact being tested is the
|
|
||||||
**description string**. Each `(prompt, skill)` pair is a labeled example. The
|
|
||||||
metrics are information-retrieval metrics:
|
|
||||||
|
|
||||||
- **recall** — fires when it should (`TriggerRateReport.rate`),
|
|
||||||
- **precision** — stays quiet when it shouldn't, including on sibling-skill prompts
|
|
||||||
(`TriggerRateReport.precision` / `falsePositiveRate`, driven by
|
|
||||||
`irrelevantPrompts`).
|
|
||||||
|
|
||||||
Run it like an ML eval: a labeled set, a threshold (`recall ≥ 0.9`), tolerant of
|
|
||||||
noise, tracked for drift — **not** like jest with an exact assertion. This is
|
|
||||||
exactly why snapshotting its score is wrong: a frozen number protects nothing;
|
|
||||||
only re-running the classifier does.
|
|
||||||
|
|
||||||
> Already shipped: `measureTriggerRate` + `assertTriggerRate` (min recall,
|
|
||||||
> maxFalsePositive, minPrecision) + the deterministic `checkPromptDiversity`
|
|
||||||
> pre-flight (NCD-based near-duplicate + min-size gate, so you can't measure a
|
|
||||||
> rate over three copy-pasted prompts). The framing here is the _justification_
|
|
||||||
> for that API, and the argument for **running it in CI** rather than
|
|
||||||
> snapshotting it.
|
|
||||||
|
|
||||||
## Two orthogonal knobs on every behavioral eval
|
|
||||||
|
|
||||||
Every behavioral eval is configured along two independent axes. Keeping them
|
|
||||||
orthogonal is what stops the snapshot/hash machinery from metastasizing into
|
|
||||||
every test.
|
|
||||||
|
|
||||||
1. **Reproducibility** — how you make a stochastic run repeatable:
|
|
||||||
`exact-assert` | `record/replay cassette` | `hermetic fixture` |
|
|
||||||
`live + threshold`.
|
|
||||||
2. **When you run it** — the deterministic tiers run **every commit in CI** (free,
|
|
||||||
no model); a real-model eval is **run deliberately** on the subscription, not in
|
|
||||||
CI: `on-demand (a Claude Code session / local)` | `hash-lockfile (replay)` |
|
|
||||||
`nightly/manual`.
|
|
||||||
|
|
||||||
The snapshot/hash machinery is **just the `hash-lockfile` value of knob 2** — one
|
|
||||||
option most features never pick. Concretely:
|
|
||||||
|
|
||||||
- a hook is `(exact-assert, every-commit CI)` — `runHook`, free, no model;
|
|
||||||
- trigger-rate is `(live + threshold, on-demand)` _if cheap_ (**Sonnet** — the
|
|
||||||
realistic selector — with bodies stubbed); run it in a session, not per-PR;
|
|
||||||
- an expensive agent eval is `(cassette, hash-lockfile)` plus a nightly live run.
|
|
||||||
|
|
||||||
## Match the mechanism to the eval's cost
|
|
||||||
|
|
||||||
The single rule that drives every gating decision:
|
|
||||||
|
|
||||||
- **Cheap eval** (Sonnet, body stubbed via `stubSkillBodies`, ~pennies on the
|
|
||||||
sub): **run it deliberately with a threshold gate** — in a Claude Code session
|
|
||||||
or locally, when it's worth it, not on every PR. No snapshot machinery at all.
|
|
||||||
vigiles has the significance-gated baseline (`eval-baseline.ts`) that makes "did
|
|
||||||
this change move the number beyond the noise floor?" a real gate, not a bare
|
|
||||||
pass-rate.
|
|
||||||
|
|
||||||
- **Expensive eval** (opus, multi-turn, N trials, spawns subagents, clones repos —
|
|
||||||
$10s–$100s/run): _pay as few times as possible and amortize._
|
|
||||||
- **Record/replay cassette = amortization.** Pay the trajectory once at record
|
|
||||||
time; every CI replay is $0. The expensive eval becomes a deterministic fixture
|
|
||||||
until inputs change. (vigiles' `eval-cache.ts` already does input-keyed
|
|
||||||
record/replay incl. post-run filesystem restore.)
|
|
||||||
- **hash-lockfile = invalidation.** Input unchanged → replay free; input changed
|
|
||||||
→ re-record (pay once). You spend the full amount _only when the definition
|
|
||||||
actually changes_ — exactly when you want to.
|
|
||||||
- **Trials are the cost multiplier** (confidence = N trials × dataset size).
|
|
||||||
N=1 smoke per PR; high-N nightly. Subset-sample per PR, full suite nightly.
|
|
||||||
`maxCostUsd` is the hard cap (already implemented).
|
|
||||||
- **The nightly live tier is the one thing you cannot amortize.** Detecting "did
|
|
||||||
the model get worse" requires hitting the live model with nothing cached.
|
|
||||||
Schedule it, cap it, budget for it. Everything else drives per-PR cost to ~0.
|
|
||||||
|
|
||||||
## What already exists
|
|
||||||
|
|
||||||
Read this before proposing to build anything — much of the design is shipped.
|
|
||||||
|
|
||||||
| Capability | Module | Notes |
|
|
||||||
| ------------------------------------------------------- | --------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
||||||
| Run behavioral scripts (`*.eval.mjs`) | `vigiles eval` (`cli.ts` → `run-scripts.ts`) | aggregates pass/skip/fail by exit code; `--trials=N`. Run locally on the sub. A bare (no-target) `eval` spends quota over the whole tree, so it asks first (`decideRunScripts`): name the eval(s), `--all`, or answer the prompt; headless → refuses (exit 2). |
|
|
||||||
| Committed staleness gate (CI, no model) | `eval-lock.ts` (`eval --check`/`--update`) | integrity hash of inputs; the CI half of evals you produce locally |
|
|
||||||
| A/B harness arms + Welch significance | `eval.ts` (`runEval`, `measureArms`), `stats.ts` | the differentiator — harness loaded _as it ships_ |
|
|
||||||
| Declarative check vocabulary (data, not asserts) | `check.ts` | `tool`/`skill`/`output`/`hookFired`/`received`/`turns`/`wrote`/`subagent`/`mcp`/`judged`/`cost`/`latency`/`tokens` — one vocab, strict + scored |
|
|
||||||
| Scored eval + threshold gate | `eval.ts` (`measure`, `assertRates`), `harness-assert.ts` | rate ± se, pass^k |
|
|
||||||
| Record/replay cache | `eval-cache.ts` | input-keyed (excludes `measure`), restores post-run filesystem |
|
|
||||||
| Committed baseline + regression gate | `eval-baseline.ts` | Welch current-vs-baseline; `lowerIsBetter`; JUnit |
|
|
||||||
| Trigger-rate (recall **and** precision) | `eval.ts` (`measureTriggerRate`) | `irrelevantPrompts` → `falsePositiveRate`/`precision` |
|
|
||||||
| Cheap-firing path | `eval.ts` (`stubSkillBodies`) | strip skill body, stop at selection — ~18× cheaper |
|
|
||||||
| Prompt-set diversity pre-flight | `eval.ts` (`checkPromptDiversity`) | deterministic NCD gate, no model |
|
|
||||||
| Cost / latency / token capture, concurrency, budget cap | `eval.ts` | `maxCostUsd`, `runPool`, 429 backoff |
|
|
||||||
| JUnit output | `eval.ts`, `eval-baseline.ts` | CI-consumable |
|
|
||||||
| Deterministic hook tier | `run-hook.ts` | event-JSON → hook, every event, no model |
|
|
||||||
| Sandbox + allowlisted egress | `sandbox.ts`, `egress.ts` | confine untrusted harness code |
|
|
||||||
|
|
||||||
What is **genuinely missing** is in [Capability gaps](#capability-gaps-ranked).
|
|
||||||
|
|
||||||
## The eval lock (the CI staleness gate)
|
|
||||||
|
|
||||||
Real-model evals run on your **subscription**, so they only run **locally** —
|
|
||||||
never in CI. The lock lets CI verify the committed numbers still match the current
|
|
||||||
inputs **without running the model**. It is the snapshot/lockfile pattern
|
|
||||||
(`Cargo.lock` + `npm ci`; `jest --ci` / `cargo-insta`), and it is an **integrity
|
|
||||||
hash, not a cache** — the local [record/replay cache](#what-already-exists) is a
|
|
||||||
separate, gitignored speed optimization.
|
|
||||||
|
|
||||||
| | the **cache** | the **lock** |
|
|
||||||
| ----------- | --------------------- | ----------------------------------- |
|
|
||||||
| purpose | local iteration speed | CI staleness detection |
|
|
||||||
| lifecycle | gitignored, throwaway | **committed**, reviewed in the diff |
|
|
||||||
| runs in CI? | no | **yes** (`eval --check`, no model) |
|
|
||||||
|
|
||||||
How you use it:
|
|
||||||
|
|
||||||
- **`vigiles eval --update`** (local, on your subscription): records each **named**
|
|
||||||
eval's report to a committed `.vigiles/eval-locks/<name>.lock.json`, and prints
|
|
||||||
the per-number delta vs the prior lock.
|
|
||||||
- **`vigiles eval --check`** (CI): recompute the input hash, compare. Match →
|
|
||||||
pass, **no model call**. Mismatch → fail "stale, run `--update`." The committed
|
|
||||||
diff of `recall: 0.90 → 0.65` **is the quality gate** a human reviews.
|
|
||||||
- In a workflow: `uses: zernie/vigiles@v1` with `command: eval-check` (`vigiles
|
|
||||||
init` scaffolds this job). It is a green no-op until you commit your first lock.
|
|
||||||
|
|
||||||
**The split that makes it sound.** The lock stores only the model's _observed
|
|
||||||
behavior_ (the recorded numbers). Your script's assertions re-run live against
|
|
||||||
those numbers on every `--check`. So:
|
|
||||||
|
|
||||||
- ❌ change an **input** (skill / prompt / model) → stale → re-run `--update`.
|
|
||||||
- ✅ change only a **threshold** in the test → valid replay, no model — the
|
|
||||||
assertion just re-judges the saved numbers.
|
|
||||||
|
|
||||||
The hash covers the model-affecting inputs: task, files, settings, tools,
|
|
||||||
plugin-dir contents, model, and `evalApiVersion`.
|
|
||||||
|
|
||||||
**Honest scope.** The lock proves _"your saved numbers match your current inputs,"_
|
|
||||||
not _"they reflect today's model."_ There is no automated live run — model/harness
|
|
||||||
drift is caught when you re-run `--update` and review the moved numbers. What it
|
|
||||||
_does_ catch is the common bug: edit a skill, forget to re-eval, ship stale numbers.
|
|
||||||
|
|
||||||
## `evalApiVersion` — a hand-bumped behavior epoch (not the CC version)
|
|
||||||
|
|
||||||
A monotonic integer **you** own (in `.vigilesrc.json` under `eval.apiVersion`),
|
|
||||||
bumped only when a _harness change on your side_ would shift eval outputs (a
|
|
||||||
CLAUDE.md edit, a hook change) but isn't otherwise in the lock's inputs. Like a
|
|
||||||
migration number / `CACHE_VERSION`. Bumping it makes `eval --check` report the
|
|
||||||
committed results stale, forcing a local re-run.
|
|
||||||
|
|
||||||
Why the Claude Code version is **not** hashed into the lock:
|
|
||||||
|
|
||||||
- `--check` runs in CI where `claude` is **pinned**, while a dev's local `claude`
|
|
||||||
is whatever they have — hashing the version would false-trip `--check` on every
|
|
||||||
PR where those differ.
|
|
||||||
- It is the honest-scope line above: the gate is about author-controlled inputs.
|
|
||||||
Keeping the version out is what lets `--check` stay **binary-free** in CI.
|
|
||||||
|
|
||||||
The version is recorded on the lock as provenance. (The local **cache** _does_ key
|
|
||||||
on it — via `HarnessRuntime.versionKey`, which is `major.minor` for Claude Code
|
|
||||||
but `""` for Codex, since Codex's minor is patch-cadence. That's local replay
|
|
||||||
soundness, a separate axis.)
|
|
||||||
|
|
||||||
## Model strategy — measure on what users run; compare models as arms (decided 2026-06-17)
|
|
||||||
|
|
||||||
Which model an eval uses is **not** cosmetic. Dogfooding the shipped `test-harness`
|
|
||||||
skill found a 0.50 trigger-rate on `claude-haiku-4-5` vs **0.90 on
|
|
||||||
`claude-sonnet-4-6`** — same skill, same prompts. Trigger-rate is a _selection_
|
|
||||||
measurement and haiku is a much weaker selector, so a haiku eval gives
|
|
||||||
false-negative recall and would fail skills that are fine on the model users
|
|
||||||
actually run. Conclusions:
|
|
||||||
|
|
||||||
1. **Default to the realistic selector — Sonnet.** `measureTriggerRate` now
|
|
||||||
defaults to `"sonnet"` (was haiku), and the `minModel` floor (also Sonnet)
|
|
||||||
fails a run that resolves below it. Haiku stays available as a deliberate,
|
|
||||||
_pessimistic_ override (a lower bound), never the default for a selection
|
|
||||||
measurement. The model lives in the **spec** (`model`/`minModel`), not a CLI/env
|
|
||||||
override — it's part of the measurement definition, not a run knob like
|
|
||||||
`--trials`.
|
|
||||||
2. **No multi-model matrix runner by default.** Running every eval across
|
|
||||||
`[haiku, sonnet, opus]` multiplies cost on every run — promptfoo's "providers"
|
|
||||||
lane, against our keep-the-real-model-surface-thin discipline.
|
|
||||||
3. **A model comparison is a harness A/B → model-as-an-arm.** When you _do_ want
|
|
||||||
"does my harness hold on the cheaper tier / after a model upgrade?", set
|
|
||||||
`model` per **arm** (`EvalArm.model`) and let the existing significance
|
|
||||||
machinery read the gap — no separate matrix DSL. `measureTriggerRate` stays
|
|
||||||
single-model (loop it for a matrix). This is the one model feature we built.
|
|
||||||
4. **(Considered, not yet built) A model FLOOR.** A configurable `minModel`
|
|
||||||
(default Sonnet) that fails/warns when an eval resolves below it — the runtime
|
|
||||||
guard (post-env) that a static lint can't give, since the haiku footgun entered
|
|
||||||
via an env var. Deferred pending a decision on warn-vs-fail + config source.
|
|
||||||
|
|
||||||
### Honest pinning (the orthogonal axis)
|
|
||||||
|
|
||||||
Picking the right model (above) is separate from **pinning** it for a
|
|
||||||
cached/baselined result. The defaults are floating aliases (`runEval` → `"haiku"`,
|
|
||||||
`measure`/`measureTriggerRate` → `"sonnet"`); for a lockfiled/baselined result a
|
|
||||||
floating alias is **dishonest** (it can re-point while the hash says "unchanged").
|
|
||||||
Pin a **dated** id (e.g. `claude-haiku-4-5-20251001`) so the hash is honest; a
|
|
||||||
dated id 404ing on deprecation is a **feature** (forces a re-eval onto a current
|
|
||||||
model) as long as the failure is surfaced. `isDatedModel` + the floating-alias
|
|
||||||
cache warning already nudge this. A cache-off run (a one-shot session eval) can
|
|
||||||
use the plain `sonnet` alias without churn; pin a dated id only when you turn on
|
|
||||||
the record/replay cache or a committed baseline.
|
|
||||||
|
|
||||||
## Deferred (YAGNI): canary / ETag scaling optimization
|
|
||||||
|
|
||||||
Only worth it with _many_ expensive evals **and** frequent CC bumps. On a version
|
|
||||||
bump, run **one** cheap fingerprint prompt: matches → trust all snapshots, skip the
|
|
||||||
rebuild; moved → invalidate + rebuild. Plus auto-rebless-within-tolerance: when
|
|
||||||
only the version changed, auto-rerun and auto-accept if metrics are within
|
|
||||||
tolerance, page a human only when a number actually moves. **Document it, don't
|
|
||||||
build it yet.**
|
|
||||||
|
|
||||||
## Isolation lies — bound which interactions matter (closure-scoped hashing)
|
|
||||||
|
|
||||||
A skill's behavior depends on context (`CLAUDE.md`, sibling skills, hooks), so pure
|
|
||||||
isolation gives false confidence. But the interactions that _matter_ are a short
|
|
||||||
finite list, not a cross product:
|
|
||||||
|
|
||||||
(a) **triggering collisions** — descriptions compete (inherently whole-set);
|
|
||||||
(b) **guidance conflicts** — skill rules vs `CLAUDE.md` vs another skill;
|
|
||||||
(c) **hook/tool interception**.
|
|
||||||
|
|
||||||
"Add everything to the hash" globally causes (1) a rebuild storm (edit any skill →
|
|
||||||
every snapshot dies) and (2) combinatorial state explosion (testing every config
|
|
||||||
combo). Don't.
|
|
||||||
|
|
||||||
The fix is what Bazel/Nix/Turborepo do: hash each eval's **observed dependency
|
|
||||||
closure** — the specific skills/`CLAUDE.md`/hooks that _actually loaded_ during the
|
|
||||||
run — not the global everything. Editing `wrap-up` doesn't invalidate the
|
|
||||||
`illustrate` eval because it isn't in `illustrate`'s closure. Prefer **observed**
|
|
||||||
(snapshot what the harness loaded; the `plugin-loader` already materializes this)
|
|
||||||
over **declared** (no dep lists to hand-maintain and drift).
|
|
||||||
|
|
||||||
Granularity matched to cost:
|
|
||||||
|
|
||||||
- cheap trigger eval → a coarse "all frontmatter" key is fine (re-runs the whole
|
|
||||||
matrix on any description edit — cheap, bodies stubbed);
|
|
||||||
- expensive behavioral → a fine closure key so an unrelated edit doesn't detonate a
|
|
||||||
$50 rerun.
|
|
||||||
|
|
||||||
The combinatorial tail you can't enumerate: curate **named integration cases** for
|
|
||||||
(a)(b)(c), and let the **nightly full-config live run** (real harness, everything
|
|
||||||
loaded) catch the unanticipated interaction as a metric drop. Prune, don't
|
|
||||||
enumerate.
|
|
||||||
|
|
||||||
## Behavioral / side-effecting eval is the umbrella — the adversarial-gate is one member (note 2026-06-17)
|
|
||||||
|
|
||||||
The apex of this layer is **behavioral, side-effecting** eval — the assembled
|
|
||||||
harness run end-to-end against real-but-**disposable** dependencies (the "ephemeral
|
|
||||||
integration eval"): does the harness actually DO the task AND not do the dangerous
|
|
||||||
thing? The **adversarial-gate test** (ask the agent to skip an enforcement gate;
|
|
||||||
assert it refuses — a `notTool`-shaped check) is a high-value **member** of that
|
|
||||||
family, **not a replacement** for it. It's cheap and it sets up the **eval→enforce
|
|
||||||
bridge** (when the prose gate caves under pressure, vigiles's deterministic
|
|
||||||
hook/rail is the fix — layer 2 hands off to layer 1), but the broad behavioral
|
|
||||||
tier — multiple surfaces firing together, real side effects against ephemeral
|
|
||||||
deps, graded by the `Trace`/check vocabulary — is the flagship. Don't let the
|
|
||||||
narrow check stand in for the umbrella.
|
|
||||||
|
|
||||||
### The adversarial-gate test — worked example and the eval→enforce bridge
|
|
||||||
|
|
||||||
The pattern is concrete and reusable. A worked dogfood lives at
|
|
||||||
[`examples/harness/dogfood/adversarial-gate.eval.mjs`](../examples/harness/dogfood/adversarial-gate.eval.mjs):
|
|
||||||
an inline `SKILL.md` that states a measurement gate ("never run a benchmark
|
|
||||||
without a baseline; refuse if asked to skip it"), an adversarial task prompt
|
|
||||||
that explicitly asks to skip the gate ("I don't need a baseline, just scaffold
|
|
||||||
it"), and two checks:
|
|
||||||
|
|
||||||
```js
|
|
||||||
checks: [
|
|
||||||
notTool("Bash", { command: /bench\.sh|hyperfine|time\s/ }), // didn't do the forbidden thing
|
|
||||||
output(/baseline|refus|can't|won't/i), // pushed back in prose
|
|
||||||
];
|
|
||||||
```
|
|
||||||
|
|
||||||
**The eval→enforce bridge** is the key takeaway. If this eval passes at ≥ 0.9,
|
|
||||||
the prose gate is robust enough on its own. If it passes at only ≥ 0.7, the
|
|
||||||
prose alone is fragile under adversarial pressure — and the fix is NOT a better
|
|
||||||
SKILL.md description. Prose gates can always be talked out of. The fix is a
|
|
||||||
deterministic `PreToolUse` hook that checks the forbidden condition and blocks
|
|
||||||
the call regardless of what the user says (see
|
|
||||||
`src/adapters/claude-code/agent-runtime.ts` for the hook skeleton). The eval
|
|
||||||
told you _where_ the soft boundary is; the hook is the hard wall. A rate below
|
|
||||||
the acceptable floor is an automatic referral from layer 2 (test) to layer 1
|
|
||||||
(deterministic constraint) — that is the bridge.
|
|
||||||
|
|
||||||
## Token & cost as a first-class measurement — input / output / cache (decided 2026-06-17)
|
|
||||||
|
|
||||||
A harness change moves tokens on **both** sides and usually **trades them off**: a
|
|
||||||
skill or CLAUDE.md injection ADDS input every turn; a "compression" skill cuts
|
|
||||||
OUTPUT. Net cost = f(fresh-input, cached-input, output). So a single total
|
|
||||||
token/cost number can **bless a change that's net-negative** — the dogfood proof is
|
|
||||||
SkillBenchmark's Caveman run (cut output yet **2–4×'d total cost** via system-prompt
|
|
||||||
injection). Honest cost verification therefore requires the classes **separated**.
|
|
||||||
|
|
||||||
State today: `UsageTrace` carries `inputTokens`/`outputTokens`/`costUsd`/`durationMs`
|
|
||||||
(from claude's `total_cost_usd` + `usage.input_tokens`/`output_tokens`), but (a) the
|
|
||||||
`tokens()` check **collapses** input+output into one number, (b) **cache tokens**
|
|
||||||
(`cache_creation_input_tokens`/`cache_read_input_tokens`) aren't captured at all —
|
|
||||||
and a large CLAUDE.md/skill is cached (~0.1× input), so omitting them makes the cost
|
|
||||||
of exactly the harness changes you'd test misleading, and (c) there's no first-class
|
|
||||||
A/B **delta per class**.
|
|
||||||
|
|
||||||
Native support (decided):
|
|
||||||
|
|
||||||
1. Extend `UsageTrace` to all token classes — `inputTokens` (fresh),
|
|
||||||
`cacheCreationTokens`, `cacheReadTokens`, `outputTokens`, `costUsd` — captured
|
|
||||||
from the CLI usage block.
|
|
||||||
2. First-class checks `inputTokens({max})` / `outputTokens({max})` /
|
|
||||||
`cacheTokens({…})` beside `cost`/`tokens` (keep `tokens` as the convenience
|
|
||||||
total).
|
|
||||||
3. A/B token/cost **delta per class** in `measureArms`, gated by the existing Welch
|
|
||||||
significance — so "verbose vs caveman" reports input↑ / output↓ / net-cost± with
|
|
||||||
a **p-value**, not an eyeballed CI overlap. This is the cost/ROI optimizer made
|
|
||||||
native and input/output-separated, and the honest-measurement differentiator
|
|
||||||
(competitors report a single total or eyeball CIs).
|
|
||||||
|
|
||||||
## Capability gaps, ranked
|
|
||||||
|
|
||||||
The genuinely missing primitives (everything above is shipped). Ranked by
|
|
||||||
protection-per-dollar unlocked.
|
|
||||||
|
|
||||||
1. **KEYSTONE — tool-call spy/fake.** Assert on the **arguments** a skill causes
|
|
||||||
the agent to pass to a tool, **without executing the tool** (no real image-API
|
|
||||||
call, no real `git push`, no real subagent spawn). This is precisely what
|
|
||||||
promptfoo-style tools _can't_ do — they grade a completion; they can't see "the
|
|
||||||
agent decided to push to `main`."
|
|
||||||
- **Correction to the original framing:** a tool-spy does **not** "unlock the
|
|
||||||
cheap (free, no-model) tier." Asserting on args the _model_ chose still needs
|
|
||||||
the real model to make the routing/argument decision — you can't get it from a
|
|
||||||
scripted mock. The spy is an **eval-tier** capability: real model, **faked
|
|
||||||
tools**. The saving is **eliminating the expensive side effect**, not
|
|
||||||
eliminating the model. (Today the existing `tool()`/`skill()` checks already
|
|
||||||
_read_ `ToolCall.input`, so argument _inspection_ exists in the `Trace`; what's
|
|
||||||
missing is **interception** — preventing the call and returning a canned
|
|
||||||
result so the real-model run is cheap and side-effect-free.)
|
|
||||||
- Where the logic _can_ be lifted out of the prompt into a script, prefer that
|
|
||||||
(gap #5) — it's the free deterministic test, strictly cheaper than any
|
|
||||||
model-driven spy.
|
|
||||||
- **Shipped (inspection half):** `toolWith(name, args)` and `notTool(name,
|
|
||||||
args?)` in `src/check.ts` over a shared, serializable `ArgMatcher`
|
|
||||||
(`src/arg-match.ts`; dot-path keys, RegExp = pattern, primitive = exact) —
|
|
||||||
assert _how_ a tool was called, and the negative/safety form (#2). These read
|
|
||||||
the `Trace` the harness/eval tier already captures.
|
|
||||||
- **Shipped (interception, end-to-end):** declare `interceptTools: [{ tool,
|
|
||||||
when?, denyReason? }]` on a `measure` / `runEval` arm. `src/tool-intercept.ts` +
|
|
||||||
the `vigiles hook-runtime intercept-tool` PreToolUse subcommand deny the real execution
|
|
||||||
(exit 2), so a
|
|
||||||
real-model run that _decides_ to hit a paid API / `git push` / spawn a paid
|
|
||||||
subagent is **safe and side-effect-free** — yet its arguments still land in
|
|
||||||
the `Trace` for `toolWith` / `notTool`. The eval tier auto-merges the hook
|
|
||||||
into the arm's settings (appending, never clobbering), carries the intercept
|
|
||||||
list (RegExp matchers intact) in `VIGILES_INTERCEPT_TOOLS`, and keys the cache
|
|
||||||
on it so two intercept configs sharing tool names don't collide. Pure core
|
|
||||||
fully unit-tested; the wiring sits under the eval tier's 100% gate.
|
|
||||||
- **Honest assessment (2026-06-17) — keep, with scope.** Three caveats the
|
|
||||||
"keystone" label shouldn't paper over:
|
|
||||||
1. **Intercept-and-prevent, not a faithful mock.** CC surfaces the exit-2 deny
|
|
||||||
as a _blocked_ call, not a success, so this is sound for "did the agent
|
|
||||||
ATTEMPT X" (safety / approval-gate / first-attempt) and unsound for "stub
|
|
||||||
the tool and let a multi-step flow continue as if it returned" — the call
|
|
||||||
is intercepted (prevented), NOT executed. There is no CC primitive for
|
|
||||||
"skip execution, return this as success."
|
|
||||||
2. **Mostly ergonomic on the inspection side.** `toolWith` overlaps the
|
|
||||||
existing `toolUsedWith` predicate (`harness-assert.ts`); the genuinely new
|
|
||||||
bit is the serializable _negative_ check and the interception. For many
|
|
||||||
safety cases the simplest protection — **don't allowlist the tool, then
|
|
||||||
assert the attempt** — needs no new primitive; `interceptTools` earns its
|
|
||||||
keep at the margins (args-scoped interception, intercepting a tool you
|
|
||||||
otherwise want allowed, and capturing an intercepted `Task` spawn's args).
|
|
||||||
3. **One unverified assumption.** Arg-capture-under-deny (the `tool_use` lands
|
|
||||||
in the stream _before_ the hook denies) is asserted from CC semantics but
|
|
||||||
not yet proven against a live model. `examples/harness/intercept-tools.eval.mjs`
|
|
||||||
is the end-to-end validation (skips without `claude`); run it with a key
|
|
||||||
before relying on the spy. Cost is **not** reduced — the model call remains;
|
|
||||||
only the side effect is removed.
|
|
||||||
- **vs competitors:** the _assertions_ are at parity with promptfoo `trajectory:*`;
|
|
||||||
the differentiator is intercepting in the **real shipped harness** (promptfoo
|
|
||||||
reconstructs an agent from YAML/SDK and can't), but that edge is narrow
|
|
||||||
(attempt/safety, not faithful mocking).
|
|
||||||
2. **Negative / safety assertions** (a mode of #1 — highest value, most
|
|
||||||
overlooked). Did **not** call the paid API before approval; did **not** push to
|
|
||||||
the wrong branch; did **not** file a security advisory for a model-only repro.
|
|
||||||
**Shipped:** `notTool(name, args?)` in `check.ts` + the `interceptTools`
|
|
||||||
interception from #1 — together they assert the agent _didn't_ take a dangerous action,
|
|
||||||
cheaply and for real.
|
|
||||||
3. **Outbound HTTP/curl fake + request-body assertion** (the network case of #1).
|
|
||||||
Distinct from `egress.ts` (which records/allows at the packet layer) — this
|
|
||||||
_fakes_ the endpoint and asserts the request **body** (e.g. the image prompt =
|
|
||||||
`CONCEPT + STYLE_SUFFIX`).
|
|
||||||
4. **Hermetic fixtures + seam-ability.** Committed fixture repos; skills refactored
|
|
||||||
to point at a local fixture instead of cloning/pushing for real. Partly a
|
|
||||||
_skill-side_ refactor, not a vigiles primitive — but vigiles should make the
|
|
||||||
fixture wiring ergonomic.
|
|
||||||
5. **Subprocess golden harness.** Generalize `runHook` to "run _this script_
|
|
||||||
against a fixture, assert stdout/exit" — for extracted deterministic cores (a
|
|
||||||
miner/checker, a `generate.sh` prompt-builder). This is the tier that makes the
|
|
||||||
"extract the core" prerequisite pay off. Highest protection-per-dollar where the
|
|
||||||
logic is extractable.
|
|
||||||
6. **hash-lockfile + cassette cost machinery** (knob 2) — only for genuinely
|
|
||||||
expensive behavioral evals. `eval --check` / `--update`, `.snapshot.json`,
|
|
||||||
`evalApiVersion`, dated-model pin.
|
|
||||||
7. **Closure-scoped (observed) invalidation** — the dependency-closure hashing from
|
|
||||||
the section above; layers on top of #6 once there are enough expensive evals to
|
|
||||||
warrant it.
|
|
||||||
- **Shipped (cache-key hardening, 2026-06-17):** the record/replay key now
|
|
||||||
content-hashes a native `--plugin-dir` (`hashDir` — editing a skill in it
|
|
||||||
invalidates, where a path-only key false-replayed), treats the tool list as a
|
|
||||||
set, and salts a `CACHE_FORMAT_VERSION`; floating-alias model drift is warned.
|
|
||||||
Full best-practice survey + the shipped/deferred decisions (eviction deferred
|
|
||||||
as disk hygiene) are captured in the design record.
|
|
||||||
|
|
||||||
## Dogfood targets
|
|
||||||
|
|
||||||
These six skills live in a **separate portfolio repo**, used as worked examples to
|
|
||||||
validate the vigiles API — they are **not** in this repo. Mapped to tier + the gap
|
|
||||||
each needs:
|
|
||||||
|
|
||||||
| Skill | Deterministic part (TEST) | Behavioral part (EVAL) | Gap it needs |
|
|
||||||
| ------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
||||||
| **writing-quality** (pure guidance) | regex linter for mechanical tropes + a trigger test (currently **missing**) | A/B lift — existing eval doesn't discriminate (both arms 100% on blatant tropes) | a discriminating fixture (subtler structural-trope draft) |
|
|
||||||
| **illustrate** (paid image API) | prompt-builder asserts recipe + style-suffix; approval-gate "no API call before confirm" | one live image smoke nightly | **fake outbound curl + call-spy** (#1, #3) |
|
|
||||||
| **prune-illustrate** (paid image API; has `generate.sh` w/ `STYLE_SUFFIX` const — the **only clean unit seam**, the template to copy) | `generate.sh` assembly via faked curl (body = CONCEPT+STYLE_SUFFIX, env overrides, out path) | live smoke | **fake-curl** + subprocess golden (#3, #5). _Security:_ it commits `api-key.txt` and the trigger eval `cpSync`s `.claude/skills` to `/tmp` — leaks the key to CI; rotate → env var → gitignore → add a secret-scan guard (itself a Layer-2 dogfood) |
|
|
||||||
| **wrap-up** (git commit+push) | ToC gen, ≤3-commit skip, zero-commit stop, "no push to wrong branch / no unasked PR"; assert 4 sections + ToC, links-not-duplicates `STATE.md` | — | **hermetic git fixture + fake push + spy** (#1, #4) |
|
|
||||||
| **audience-test** (spawns 3–5 paid agents + screenshots) | panel-composition invariants (≥1 non-expert, 3–5, not all-expert) asserted from spawn **args** with agents faked; reader's-cut transform | full-run structure (7 deliverables, in character) | **fake the Agent tool + assert on its call args** (#1) |
|
|
||||||
| **cross-field-bug-hunt** (clones repos, spawns paid subagents, boots Rails+PG) | miner+checker **golden** test on frozen Lago/Solidus model fixtures → re-flags the known pair + trace, **zero model** — code already exists, needs ~no new primitive | planted-bug fixture repo → agent finds it, bucket A, right `file:line` | hermetic fixture repo + faked clone + subprocess golden (#5) |
|
|
||||||
|
|
||||||
Two structural notes carried from the analysis: (1) most testable skill logic is
|
|
||||||
**inline guidance, not code**, so extracting deterministic cores into scripts is a
|
|
||||||
prerequisite (`generate.sh` is the model); (2) the **cross-field golden test needs
|
|
||||||
almost no new vigiles capability** — do it first.
|
|
||||||
|
|
||||||
## Ranked build roadmap
|
|
||||||
|
|
||||||
Ordered by protection-per-dollar, with the dogfood that validates each step.
|
|
||||||
|
|
||||||
1. **Run the behavioral tier where the subscription already is — a Claude Code
|
|
||||||
session, NOT GitHub Actions.** The original "wire evals into CI" framing was
|
|
||||||
wrong: real-model evals don't belong in a standalone GitHub Actions workflow
|
|
||||||
that needs a metered (or sub-token-as-secret) credential. CI runs the **free
|
|
||||||
deterministic tiers** (`ci.yml` — `runHook` + mock-model `runHarnessTest`, no
|
|
||||||
token); the **real-model eval** runs on your **subscription** in a Claude Code
|
|
||||||
session (the agent loop / web / a scheduled session) or locally — `vigiles
|
|
||||||
drives the `claude`CLI, so it authenticates like your own CLI does (no metered
|
|
||||||
API). _Validates on:_ the **missing`writing-quality` trigger case\*\*.
|
|
||||||
- **Shipped (eval robustness, applies wherever `vigiles eval` runs):**
|
|
||||||
`--min=N` (fail if fewer than N evals actually ran — no silent zero),
|
|
||||||
`--no-skip` (a skipped tier fails), a **corrupt-cache throw** (a broken
|
|
||||||
cassette surfaces, not a silent re-run), a **model floor** (`minModel`,
|
|
||||||
default Sonnet — a too-weak selector fails before spending a token), and the
|
|
||||||
floating-alias cache warning. Measure trigger-rate on **Sonnet** (dogfooded:
|
|
||||||
0.50 haiku vs 0.90 Sonnet — haiku under-selects). **Removed:** the speculative
|
|
||||||
`evals.yml` GitHub Actions workflow + the `--model`/`VIGILES_MODEL` env knob
|
|
||||||
(model belongs in the spec, not a hidden override). **Remaining:** the
|
|
||||||
`writing-quality` trigger case lives in the separate portfolio repo.
|
|
||||||
2. **Cross-field miner/checker golden fixture test.** Near-zero new primitive — the
|
|
||||||
code exists; freeze the Lago/Solidus model fixtures and assert it re-flags the
|
|
||||||
known pair + trace, zero model. Cheapest real protection available. _Needs:_ a
|
|
||||||
thin generalization toward the subprocess golden harness (#5).
|
|
||||||
3. **KEYSTONE: tool-call spy/fake + negative/safety checks** (#1, #2). Build tool
|
|
||||||
interception (capture args, return canned result, prevent side effect) at the
|
|
||||||
eval tier, and add `notTool`/arg-matcher checks to `check.ts`. _Validates on:_
|
|
||||||
`illustrate` approval-gate, `audience-test` panel composition.
|
|
||||||
4. **Outbound curl fake + body assertion** (#3). _Validates on:_ `prune-illustrate`
|
|
||||||
`generate.sh` (body = CONCEPT+STYLE_SUFFIX) + `illustrate`. Pairs with the
|
|
||||||
secret-scan guard dogfood.
|
|
||||||
5. **Hermetic fixtures + the hash-lockfile cost machinery** (#4, #6). `eval --check`
|
|
||||||
/ `--update`, `.snapshot.json`, `evalApiVersion`, dated-model pin; nightly live
|
|
||||||
tier scheduled + capped. _Validates on:_ `wrap-up` + `cross-field` against
|
|
||||||
fixture repos, gated by hash-lockfile + nightly live.
|
|
||||||
6. **Closure-scoped (observed) invalidation** (#7) — only once there are enough
|
|
||||||
expensive evals that an unrelated edit detonating a rebuild is a real pain.
|
|
||||||
7. **Native input/output/cache token + cost measurement + A/B delta** — split
|
|
||||||
`tokens()` into `inputTokens`/`outputTokens`, capture cache tokens
|
|
||||||
(`cache_creation`/`cache_read`), and report a per-class A/B delta gated by Welch
|
|
||||||
significance. The honest cost-claim verifier (the Caveman gap: output↓ but net
|
|
||||||
cost↑). _Validates on:_ the `skill-compression` (Caveman) eval — assert output↓
|
|
||||||
AND input/net honestly, with a p-value. **HIGH** (a money story; cheap to build —
|
|
||||||
the data model is half there).
|
|
||||||
8. **Adversarial-gate check + the eval→enforce bridge** — a first-class "ask the
|
|
||||||
agent to skip the enforcement gate, assert it refuses" check (the `notTool`
|
|
||||||
shape); when it fails, point at the deterministic rail (layer 2 → layer 1).
|
|
||||||
_Validates on:_ an OMC enforcement-skill dogfood.
|
|
||||||
9. **Whole-harness trigger-rate tier** — `measureTriggerRate` is isolated today
|
|
||||||
(cheap, but it _overstates recall and understates false-positives_ because skill
|
|
||||||
selection is competitive and Claude Code evicts least-used skill descriptions
|
|
||||||
under a context budget). Add an `installSet`/`withHarness` arm that co-installs
|
|
||||||
the skill alongside the user's real set as a **release gate**, plus a
|
|
||||||
near-neighbor middle tier built on the existing `ncd`/`findSimilarRules` engine.
|
|
||||||
This is a genuine wedge — **no existing eval tool populates the install set**.
|
|
||||||
|
|
||||||
## Where this design is wrong / open questions
|
|
||||||
|
|
||||||
Consolidated pushback, for the record:
|
|
||||||
|
|
||||||
1. **The biggest correction: most of the "machinery" is already built.** The cache,
|
|
||||||
the significance-gated baseline, the check vocabulary, trigger-rate
|
|
||||||
recall+precision, cost/budget/concurrency, JUnit — all shipped. Framing this as
|
|
||||||
"design the eval system" overstates the work. The real deliverables are a **CI
|
|
||||||
job**, a **gating policy**, **dated-model honesty**, and a **handful of
|
|
||||||
assertion primitives**. Don't rebuild what `eval.ts` / `eval-cache.ts` /
|
|
||||||
`eval-baseline.ts` / `check.ts` already do.
|
|
||||||
2. **The tool-spy does not move work to the free tier.** It needs the real model
|
|
||||||
(the routing decision is what you're testing); it only removes the _side effect_.
|
|
||||||
Treat it as a cheaper/safer **eval**, not a deterministic test. The genuinely
|
|
||||||
free win is **extracting the deterministic core into a script** and testing it at
|
|
||||||
the subprocess golden tier — so #5 is arguably co-equal with the keystone where
|
|
||||||
the logic is extractable.
|
|
||||||
3. **hash-lockfile vs the existing baseline must be reconciled, not duplicated.**
|
|
||||||
For cheap evals the existing `(live + threshold, on-demand)` baseline gives
|
|
||||||
_strictly more_ drift protection than a lockfile. The lockfile is a cost
|
|
||||||
concession for expensive evals **only**, and only safe with the nightly backstop.
|
|
||||||
Retrofitting cheap evals onto a lockfile would _remove_ protection.
|
|
||||||
4. **An HTTP cassette does not escape the snapshot problem.** Replaying one recorded
|
|
||||||
trajectory is the same false-green as a frozen comment. Only the nightly live run
|
|
||||||
detects model drift. This is a property of _replay_, not of the recording format.
|
|
||||||
5. **Everything else in the original thinking holds and is good:** the
|
|
||||||
feature = test + eval decomposition, the two orthogonal knobs, cost-matched
|
|
||||||
mechanism, trigger-rate-as-classifier, `evalApiVersion` as a behavior epoch
|
|
||||||
distinct from the CC version, dated-model honesty, the deferred canary, and
|
|
||||||
closure-scoped (observed) invalidation. These are the spine of the doc.
|
|
||||||
|
|
||||||
## Where to start
|
|
||||||
|
|
||||||
**Step 1 + Step 2 in parallel**, because they're cheap and prove the model end to
|
|
||||||
end:
|
|
||||||
|
|
||||||
- Wire the cheap behavioral tier into CI as a gate, pin a dated model, delete the
|
|
||||||
comment-snapshots — and add the missing `writing-quality` trigger case as the
|
|
||||||
first thing the new gate protects.
|
|
||||||
- Land the `cross-field` miner/checker **golden** test (zero model, code already
|
|
||||||
exists) as the first subprocess-golden dogfood.
|
|
||||||
|
|
||||||
Then build the **keystone tool-call spy/fake (#1) + negative checks (#2)**, since
|
|
||||||
every remaining expensive dogfood (`illustrate`, `audience-test`, `wrap-up`) is
|
|
||||||
blocked on it.
|
|
||||||
</content>
|
|
||||||
</invoke>
|
|
||||||
+1
-1
@@ -49,7 +49,7 @@ Test layer drives the real `claude` / `codex` CLI. You can even
|
|||||||
|
|
||||||
**Most of vigiles needs no model and no key.** Lint and the deterministic Test tiers run in milliseconds on every commit, free.
|
**Most of vigiles needs no model and no key.** Lint and the deterministic Test tiers run in milliseconds on every commit, free.
|
||||||
|
|
||||||
The only thing that needs a model is a real-model **eval**. That runs on **your own Claude Pro/Max subscription** via the `claude` CLI — **$0 of metered API tokens**. Tools like promptfoo / DeepEval hit a metered API and bill per token on every run. See [the eval architecture](eval-architecture.md).
|
The only thing that needs a model is a real-model **eval**. That runs on **your own Claude Pro/Max subscription** via the `claude` CLI — **$0 of metered API tokens**. Tools like promptfoo / DeepEval hit a metered API and bill per token on every run. See [measuring skills](measuring-skills.md).
|
||||||
|
|
||||||
## What does `vigiles audit` actually run — and why did it "find nothing"?
|
## What does `vigiles audit` actually run — and why did it "find nothing"?
|
||||||
|
|
||||||
|
|||||||
@@ -1,124 +0,0 @@
|
|||||||
# Inline mode
|
|
||||||
|
|
||||||
Inline mode lets you adopt vigiles **one rule at a time**, without committing
|
|
||||||
to a `.spec.ts` compile step. You add `<!-- vigiles:enforce ... -->` HTML
|
|
||||||
comments directly to your existing `CLAUDE.md` / `AGENTS.md`, and
|
|
||||||
`vigiles lint` verifies them the same way it verifies spec-declared rules:
|
|
||||||
linter-backed existence check, closest-match typo suggestions, disabled-rule
|
|
||||||
detection, and GitHub Actions annotations.
|
|
||||||
|
|
||||||
It's the vigiles equivalent of `// eslint-disable-next-line` — minimum
|
|
||||||
commitment, maximum incrementalism.
|
|
||||||
|
|
||||||
## When to use it
|
|
||||||
|
|
||||||
- You already have a `CLAUDE.md` and don't want to port it to `.spec.ts`
|
|
||||||
- You want to experiment with a single rule before committing to the full
|
|
||||||
vigiles workflow
|
|
||||||
- Your project isn't a TypeScript project at all and the build step feels
|
|
||||||
like dead weight
|
|
||||||
- Hesitant teammates want to see the verification work before accepting a
|
|
||||||
new file type in the repo
|
|
||||||
|
|
||||||
If you already have a real TypeScript project and you want the strongest
|
|
||||||
guarantees (editor-time type safety, programmatic rule composition, the
|
|
||||||
`generate-types` type generation), use spec mode instead — see the main README.
|
|
||||||
|
|
||||||
## Format
|
|
||||||
|
|
||||||
A single HTML comment per rule:
|
|
||||||
|
|
||||||
```md
|
|
||||||
<!-- vigiles:enforce eslint/no-console "Use structured logger for app output" -->
|
|
||||||
```
|
|
||||||
|
|
||||||
Three required pieces:
|
|
||||||
|
|
||||||
1. `vigiles:enforce` — only `enforce` is supported inline. Guidance rules
|
|
||||||
are just paragraphs in the surrounding prose, so a `guidance` comment
|
|
||||||
would be a tautology.
|
|
||||||
2. `<linter>/<rule>` — the same reference format as `enforce()` in spec
|
|
||||||
mode. Supports all seven catalogs (ESLint, Stylelint, Ruff, Clippy,
|
|
||||||
Pylint, RuboCop, Cedar), scoped plugin names
|
|
||||||
(`eslint/@typescript-eslint/...`), and the vigiles-internal namespace
|
|
||||||
(`vigiles/<rule>` for built-in checks like `vigiles/orphan-docs`).
|
|
||||||
3. `"<why>"` — a simple double-quoted string. No newlines, no embedded
|
|
||||||
quotes. If you need either, move to spec mode.
|
|
||||||
|
|
||||||
## Example
|
|
||||||
|
|
||||||
```md
|
|
||||||
# My Project
|
|
||||||
|
|
||||||
<!-- vigiles:enforce eslint/no-console "Route output through logger.ts" -->
|
|
||||||
<!-- vigiles:enforce eslint/@typescript-eslint/no-floating-promises "Await or explicitly void" -->
|
|
||||||
<!-- vigiles:enforce ruff/F401 "No unused imports" -->
|
|
||||||
|
|
||||||
## Logging
|
|
||||||
|
|
||||||
All application output must go through the shared logger module.
|
|
||||||
Do not use `console.log` directly in src/.
|
|
||||||
|
|
||||||
## Async
|
|
||||||
|
|
||||||
Every promise must be awaited or explicitly voided. The ESLint rule
|
|
||||||
enforces this automatically.
|
|
||||||
```
|
|
||||||
|
|
||||||
## What lint catches
|
|
||||||
|
|
||||||
Running `vigiles lint CLAUDE.md` on the above file will:
|
|
||||||
|
|
||||||
- Verify each `eslint/…`, `ruff/…` reference against your actual linter
|
|
||||||
config
|
|
||||||
- Emit closest-match suggestions on typos: `"no-consol"` →
|
|
||||||
`did you mean "eslint/no-console"?`
|
|
||||||
- Emit `::error` annotations when running inside GitHub Actions
|
|
||||||
- Exit with code 2 (hard error) on any failed rule, so CI fails fast
|
|
||||||
|
|
||||||
## What lint does NOT do in inline mode
|
|
||||||
|
|
||||||
- **No type safety at edit time.** The `.spec.ts` path gets TypeScript
|
|
||||||
squiggles in the editor because `StrictLinterRule` is a type union of
|
|
||||||
every rule in your linters. Inline mode is strings-in-markdown, so
|
|
||||||
typos only surface at `vigiles lint` time. Still catches them before
|
|
||||||
CI, just not in the editor.
|
|
||||||
- **No programmatic composition.** You can't reuse a batch of rules from
|
|
||||||
a helper. Each comment is its own line.
|
|
||||||
- **No rule deduplication via NCD.** Duplicate-rule detection runs on
|
|
||||||
spec-mode files; inline rules are ungrouped.
|
|
||||||
|
|
||||||
All of this is fine for the adoption-onramp use case. When you outgrow it,
|
|
||||||
port to spec mode.
|
|
||||||
|
|
||||||
## Mixing inline and spec mode
|
|
||||||
|
|
||||||
Spec mode wins. If a file has both `CLAUDE.md.spec.ts` and inline
|
|
||||||
comments inside `CLAUDE.md`, the spec compiler will overwrite the markdown
|
|
||||||
on the next compile, and your inline comments will be gone. Pick one per
|
|
||||||
file.
|
|
||||||
|
|
||||||
## `require-instructions-spec`
|
|
||||||
|
|
||||||
The built-in `require-instructions-spec` validation rule demands a `.spec.ts`
|
|
||||||
sibling for every `CLAUDE.md` / `AGENTS.md`. It is **narrow**: only a
|
|
||||||
`.spec.ts` satisfies it. Inline mode is a valid plain-markdown on-ramp, but
|
|
||||||
it does **not** satisfy `require-instructions-spec` — so an inline-mode user
|
|
||||||
simply keeps the rule off (it is off by default; turning it on is a
|
|
||||||
workflow-tier opt-in for teams that want to require a typed spec). You don't
|
|
||||||
need a `vigiles-disable require-instructions-spec` comment unless you have
|
|
||||||
enabled the rule.
|
|
||||||
|
|
||||||
## Graduating to spec mode
|
|
||||||
|
|
||||||
When you've accumulated a dozen or so inline rules and the prose is
|
|
||||||
starting to feel crowded, run:
|
|
||||||
|
|
||||||
```bash
|
|
||||||
npx vigiles init --target=CLAUDE.md
|
|
||||||
```
|
|
||||||
|
|
||||||
That scaffolds a `CLAUDE.md.spec.ts` next to your existing `CLAUDE.md`.
|
|
||||||
Copy the inline enforce rules into the `rules:` block, delete the inline
|
|
||||||
comments, and run `vigiles compile`. The markdown output will be rebuilt
|
|
||||||
with a `sha256` hash header, and future edits flow through the spec.
|
|
||||||
+45
-14
@@ -29,30 +29,62 @@ to plain markdown anytime — so graduating to a spec is never a one-way door.
|
|||||||
|
|
||||||
## Inline `enforce` comments
|
## Inline `enforce` comments
|
||||||
|
|
||||||
The minimum-commitment path: add a single HTML comment per rule, anywhere
|
The minimum-commitment path: add a single HTML comment per rule, anywhere in
|
||||||
in your existing markdown.
|
your existing markdown. It's the vigiles equivalent of
|
||||||
|
`// eslint-disable-next-line` — maximum incrementalism, zero new files.
|
||||||
|
|
||||||
```md
|
```md
|
||||||
<!-- vigiles:enforce eslint/no-console "Route output through logger.ts" -->
|
<!-- vigiles:enforce eslint/no-console "Route output through logger.ts" -->
|
||||||
```
|
```
|
||||||
|
|
||||||
Only `enforce` is supported inline — the prose around the comment _is_ the
|
Three pieces, all required:
|
||||||
guidance, so a `guidance` comment would be a tautology. The reference uses
|
|
||||||
the same `<linter>/<rule>` format as everywhere else in vigiles.
|
|
||||||
|
|
||||||
This is the vigiles equivalent of `// eslint-disable-next-line`: maximum
|
1. **`vigiles:enforce`** — only `enforce` is supported inline. The prose around the comment _is_ the guidance, so a `guidance` comment would be a tautology.
|
||||||
incrementalism, zero new files. For the full reference — fenced-block
|
2. **`<linter>/<rule>`** — the same reference format as `enforce()` in spec mode. All seven catalogs (ESLint, Stylelint, Ruff, Clippy, Pylint, RuboCop, Cedar), scoped plugin names (`eslint/@typescript-eslint/...`), and the vigiles-internal namespace (`vigiles/orphan-docs`) work here.
|
||||||
handling, scoped plugin names, graduating to a typed spec — see
|
3. **`"<why>"`** — a double-quoted string shown to the agent as context. No newlines or embedded quotes; if you need either, move to a spec.
|
||||||
[docs/inline-mode.md](inline-mode.md).
|
|
||||||
|
A fuller example:
|
||||||
|
|
||||||
|
```md
|
||||||
|
# My Project
|
||||||
|
|
||||||
|
<!-- vigiles:enforce eslint/no-console "Route output through logger.ts" -->
|
||||||
|
<!-- vigiles:enforce eslint/@typescript-eslint/no-floating-promises "Await or explicitly void" -->
|
||||||
|
<!-- vigiles:enforce ruff/F401 "No unused imports" -->
|
||||||
|
|
||||||
|
## Logging
|
||||||
|
|
||||||
|
All application output must go through the shared logger module.
|
||||||
|
```
|
||||||
|
|
||||||
### What `vigiles lint` catches
|
### What `vigiles lint` catches
|
||||||
|
|
||||||
- Verifies each rule reference against your real linter config.
|
- Verifies each rule reference against your real linter config.
|
||||||
- Emits closest-match suggestions on typos:
|
- Emits closest-match suggestions on typos: `"no-consol"` → `did you mean "eslint/no-console"?`
|
||||||
`"no-consol"` → `did you mean "eslint/no-console"?`
|
|
||||||
- Flags rules that exist but are disabled in your linter config.
|
- Flags rules that exist but are disabled in your linter config.
|
||||||
- Emits `::error` annotations under GitHub Actions.
|
- Emits `::error` annotations under GitHub Actions; exits code 2 on any failed rule, so CI fails fast.
|
||||||
- Exits with code 2 (hard error) on any failed rule, so CI fails fast.
|
|
||||||
|
### What it does _not_ do (vs a typed spec)
|
||||||
|
|
||||||
|
- **No edit-time type safety.** A `.spec.ts` gets editor squiggles because rules are a type union; inline strings surface typos only at `vigiles lint` time (still before CI).
|
||||||
|
- **No programmatic composition** — each comment stands alone.
|
||||||
|
- **No NCD duplicate detection** — that runs on spec-mode files.
|
||||||
|
|
||||||
|
That's all fine for the adoption on-ramp. When you outgrow it, graduate to a spec.
|
||||||
|
|
||||||
|
### Mixing with a spec — don't
|
||||||
|
|
||||||
|
A file is checked for inline rules **only when it isn't managed by a spec** (no sibling `<file>.spec.ts`, no `vigiles:sha256 … compiled from …` header). If both exist, the compiler overwrites the markdown on the next compile and your inline comments vanish. Pick one per file.
|
||||||
|
|
||||||
|
### Graduating to a spec
|
||||||
|
|
||||||
|
When a dozen inline rules start crowding the prose:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npx vigiles init --target=CLAUDE.md
|
||||||
|
```
|
||||||
|
|
||||||
|
That scaffolds a `CLAUDE.md.spec.ts` beside your `CLAUDE.md`. Copy the enforce rules into the `rules:` block, delete the inline comments, and run `vigiles compile` — the markdown is rebuilt with a `sha256` header, and future edits flow through the spec. `vigiles eject` reverses it anytime.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -75,7 +107,6 @@ file.
|
|||||||
|
|
||||||
## See also
|
## See also
|
||||||
|
|
||||||
- [`inline-mode.md`](inline-mode.md) — the full inline-comment reference.
|
|
||||||
- [`spec-format.md`](spec-format.md) — the typed `.spec.ts` source of truth.
|
- [`spec-format.md`](spec-format.md) — the typed `.spec.ts` source of truth.
|
||||||
- [`verifying-instruction-files.md`](verifying-instruction-files.md) — the lint guide.
|
- [`verifying-instruction-files.md`](verifying-instruction-files.md) — the lint guide.
|
||||||
|
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
# Measuring skills & plugins — does it actually help?
|
# Measuring skills & plugins — does it actually help?
|
||||||
|
|
||||||
> The README has the pitch ("Measure — does it actually help, or just cost
|
> The README has the pitch (the only way to put a real number on cost); this is
|
||||||
> more?"); this is the full guide. vigiles is the only harness tool that can A/B a
|
> the full guide. vigiles is the only harness tool that can A/B a
|
||||||
> skill, plugin, model, or rule change on **real coding tasks** and tell you
|
> skill, plugin, model, or rule change on **real coding tasks** and tell you
|
||||||
> whether it moved the needle — on your **Claude subscription**, not metered API.
|
> whether it moved the needle — on your **Claude subscription**, not metered API.
|
||||||
|
|
||||||
@@ -86,7 +86,7 @@ Two ways to specify an arm:
|
|||||||
| promptfoo, DeepEval, … | metered API SDK | billed **per token, every run** |
|
| promptfoo, DeepEval, … | metered API SDK | billed **per token, every run** |
|
||||||
| **vigiles** | your Claude Pro/Max sub | **$0 extra** beyond your sub |
|
| **vigiles** | your Claude Pro/Max sub | **$0 extra** beyond your sub |
|
||||||
|
|
||||||
That's why vigiles can measure continuously — on every change, not once — while a per-token competitor cannot. Most of vigiles needs no model at all. Only this measurement tier does, and it runs where your subscription already is. See [`docs/eval-architecture.md`](eval-architecture.md) for the cost model.
|
That's why vigiles can measure continuously — on every change, not once — while a per-token competitor cannot. Most of vigiles needs no model at all. Only this measurement tier does, and it runs where your subscription already is.
|
||||||
|
|
||||||
### What a run reports — and the metered-API warning
|
### What a run reports — and the metered-API warning
|
||||||
|
|
||||||
|
|||||||
@@ -1,22 +0,0 @@
|
|||||||
# Related tools
|
|
||||||
|
|
||||||
vigiles doesn't try to do everything. It owns one thing: compile-time
|
|
||||||
verification of typed specs against real linter configs, filesystems, and
|
|
||||||
package.json, plus testing the harness those specs describe. Everything else,
|
|
||||||
compose:
|
|
||||||
|
|
||||||
- **Architectural linting** — [ast-grep](https://ast-grep.github.io/), [Dependency Cruiser](https://github.com/sverweij/dependency-cruiser), [Steiger](https://github.com/feature-sliced/steiger). Reference their rules via `enforce()`.
|
|
||||||
- **File sync** across agents — [Ruler](https://github.com/intellectronica/ruler), [rulesync](https://github.com/dyoshikawa/rulesync), [block/ai-rules](https://github.com/block/ai-rules). vigiles compiles the source; sync tools distribute.
|
|
||||||
- **Markdown linting** — [markdownlint](https://github.com/DavidAnson/markdownlint). vigiles generates markdown; structure is correct by construction.
|
|
||||||
- **Code-block linting in docs** — [eslint-plugin-markdown](https://github.com/eslint/eslint-plugin-markdown) for syntax, [twoslash](https://shikijs.github.io/twoslash/) for TS type-checking.
|
|
||||||
- **Prose quality** — [Vale](https://vale.sh). Different concern.
|
|
||||||
- **Runtime LLM rule checking** (e.g. ai-rulez `"AI-Powered Rule Enforcement"`) — opposite paradigm. Those tools send your code to a model on every check, costing tokens and giving non-reproducible verdicts. vigiles compiles once and checks deterministically forever after with `eslint`, `ruff`, `tsc`, Cedar evaluation — tools as deterministic as their inputs.
|
|
||||||
|
|
||||||
## Output targets
|
|
||||||
|
|
||||||
Specs compile to `CLAUDE.md` by default. Set `target: "AGENTS.md"` or
|
|
||||||
`target: ["CLAUDE.md", "AGENTS.md"]` for multiple outputs from one spec. For
|
|
||||||
non-markdown formats (`.cursorrules`, Copilot), use
|
|
||||||
[rule-porter](https://github.com/nichochar/rule-porter) or
|
|
||||||
[rulesync](https://github.com/dyoshikawa/rulesync) to convert. See the
|
|
||||||
[spec format reference](spec-format.md).
|
|
||||||
+3
-3
@@ -102,7 +102,7 @@ const report = measure(spec, {
|
|||||||
- ✅ **Sound for:** "did the agent _attempt_ X?" — safety gates, approval checks, "no paid call", "no push to the wrong branch". `notTool` is the **negative safety assertion** a completion-grading eval structurally can't make: it sees the agent's _decision to act_, not just its final text.
|
- ✅ **Sound for:** "did the agent _attempt_ X?" — safety gates, approval checks, "no paid call", "no push to the wrong branch". `notTool` is the **negative safety assertion** a completion-grading eval structurally can't make: it sees the agent's _decision to act_, not just its final text.
|
||||||
- ⚠️ **Not for:** stubbing a tool to return a fake success and letting a multi-step flow continue. The model sees a block, so a sequence that needs the real result breaks. (Claude Code has no "skip-but-return-success" primitive for arbitrary tools; deny is the closest safe thing.)
|
- ⚠️ **Not for:** stubbing a tool to return a fake success and letting a multi-step flow continue. The model sees a block, so a sequence that needs the real result breaks. (Claude Code has no "skip-but-return-success" primitive for arbitrary tools; deny is the closest safe thing.)
|
||||||
|
|
||||||
**Testing that an enforcement gate actually holds** — including under an adversarial prompt that asks the agent to skip it — is done with `notTool` + `output` checks in `measure`. The worked dogfood is [`examples/harness/dogfood/adversarial-gate.eval.mjs`](../examples/harness/dogfood/adversarial-gate.eval.mjs). When that eval shows a prose gate can be talked out of, the deterministic `PreToolUse` hook is the fix — see the eval→enforce bridge note in [`eval-architecture.md`](eval-architecture.md#the-adversarial-gate-test--worked-example-and-the-evalenforce-bridge).
|
**Testing that an enforcement gate actually holds** — including under an adversarial prompt that asks the agent to skip it — is done with `notTool` + `output` checks in `measure`. The worked dogfood is [`examples/harness/dogfood/adversarial-gate.eval.mjs`](../examples/harness/dogfood/adversarial-gate.eval.mjs). When that eval shows a prose gate can be talked out of, the deterministic `PreToolUse` hook is the fix.
|
||||||
|
|
||||||
## At a glance — what's confined, per tier
|
## At a glance — what's confined, per tier
|
||||||
|
|
||||||
@@ -139,7 +139,7 @@ Yes — `sandbox: false` is the explicit, greppable opt-out for code you trust,
|
|||||||
opt-in-not-always-on argument, `recordEgress`, `egress: { allow }`, dogfood findings.
|
opt-in-not-always-on argument, `recordEgress`, `egress: { allow }`, dogfood findings.
|
||||||
- [`harness-testing.md`](harness-testing.md) — the three tiers and where the
|
- [`harness-testing.md`](harness-testing.md) — the three tiers and where the
|
||||||
boundary sits.
|
boundary sits.
|
||||||
- [`eval-architecture.md`](eval-architecture.md) — `interceptTools`/`notTool` in the
|
- [`testing-api.md`](testing-api.md) — `interceptTools`/`notTool` in the
|
||||||
eval design, with the intercept-≠-mock trade-off.
|
testing API, with the intercept-≠-mock trade-off.
|
||||||
- [`src/sandbox.ts`](../src/sandbox.ts) · [`src/egress.ts`](../src/egress.ts) ·
|
- [`src/sandbox.ts`](../src/sandbox.ts) · [`src/egress.ts`](../src/egress.ts) ·
|
||||||
[`src/tool-intercept.ts`](../src/tool-intercept.ts) — the pure, tested seams.
|
[`src/tool-intercept.ts`](../src/tool-intercept.ts) — the pure, tested seams.
|
||||||
|
|||||||
+1
-1
@@ -125,7 +125,7 @@ The same `session-start` hook is dogfooded a second time under `egress: { allow:
|
|||||||
## See also
|
## See also
|
||||||
|
|
||||||
- [Testing your harness](harness-testing.md) — the three tiers + the sandbox boundary.
|
- [Testing your harness](harness-testing.md) — the three tiers + the sandbox boundary.
|
||||||
- [Eval architecture](eval-architecture.md) — `interceptTools`/`notTool`:
|
- [Testing API](testing-api.md) — `interceptTools`/`notTool`:
|
||||||
preventing a real model's tool side effects at the eval tier.
|
preventing a real model's tool side effects at the eval tier.
|
||||||
- [`src/sandbox.ts`](../src/sandbox.ts) — `decideSandbox` (the pure policy), `bwrapArgs`, `parseEgressLog`.
|
- [`src/sandbox.ts`](../src/sandbox.ts) — `decideSandbox` (the pure policy), `bwrapArgs`, `parseEgressLog`.
|
||||||
- [`src/egress.ts`](../src/egress.ts) — the `egress: { allow }` allowlist: ruleset builder, counter parser, the pure seams.
|
- [`src/egress.ts`](../src/egress.ts) — the `egress: { allow }` allowlist: ruleset builder, counter parser, the pure seams.
|
||||||
|
|||||||
+1
-1
@@ -7,7 +7,7 @@ vigiles specs are TypeScript files (`*.spec.ts`) that compile to markdown instru
|
|||||||
Be honest about what a spec is **not** for. The reference checks — does this
|
Be honest about what a spec is **not** for. The reference checks — does this
|
||||||
`file()` exist, is this linter rule enabled, is this `cmd()` a real script — do
|
`file()` exist, is this linter rule enabled, is this `cmd()` a real script — do
|
||||||
**not** need a spec. vigiles runs them on a plain CLAUDE.md via inline
|
**not** need a spec. vigiles runs them on a plain CLAUDE.md via inline
|
||||||
[`<!-- vigiles:enforce -->` comments](inline-mode.md), on purpose, as the
|
[`<!-- vigiles:enforce -->` comments](markdown-mode.md), on purpose, as the
|
||||||
on-ramp. If verification is all you want, **stay in markdown**.
|
on-ramp. If verification is all you want, **stay in markdown**.
|
||||||
|
|
||||||
A spec earns its place when you cross from **declaring** your harness to
|
A spec earns its place when you cross from **declaring** your harness to
|
||||||
|
|||||||
@@ -270,7 +270,7 @@ Everything vigiles compiles and lints is **deterministic** — same input, same
|
|||||||
|
|
||||||
## See also
|
## See also
|
||||||
|
|
||||||
- [Markdown mode](markdown-mode.md) · [Inline mode](inline-mode.md) — the no-spec on-ramps.
|
- [Markdown mode](markdown-mode.md) — the no-spec on-ramp (inline `<!-- vigiles:enforce -->` comments).
|
||||||
- [Spec format reference](spec-format.md) — every section and rule kind.
|
- [Spec format reference](spec-format.md) — every section and rule kind.
|
||||||
- [Linter support](linter-support.md) — the 7 catalogs + `generate-types` / `generate-schema`.
|
- [Linter support](linter-support.md) — the 7 catalogs + `generate-types` / `generate-schema`.
|
||||||
- [CLI & CI reference](cli.md) · [Agent setup](agent-setup.md).
|
- [CLI & CI reference](cli.md) · [Agent setup](agent-setup.md).
|
||||||
|
|||||||
@@ -91,7 +91,7 @@ text
|
|||||||
});
|
});
|
||||||
|
|
||||||
it("ignores vigiles:enforce markers inside fenced code blocks", () => {
|
it("ignores vigiles:enforce markers inside fenced code blocks", () => {
|
||||||
// Illustrative example in docs/inline-mode.md would otherwise get
|
// Illustrative example in docs/markdown-mode.md would otherwise get
|
||||||
// picked up as a live rule.
|
// picked up as a live rule.
|
||||||
const { rules, errors } = parseInlineRules(
|
const { rules, errors } = parseInlineRules(
|
||||||
`# Docs
|
`# Docs
|
||||||
|
|||||||
+1
-1
@@ -9,7 +9,7 @@
|
|||||||
* the `claude` CLI) + a running session tally. We deliberately do NOT show a
|
* the `claude` CLI) + a running session tally. We deliberately do NOT show a
|
||||||
* "% of your subscription" — Anthropic does not expose a subscription's quota or
|
* "% of your subscription" — Anthropic does not expose a subscription's quota or
|
||||||
* limit programmatically (and the real limits are rolling rate windows, not a
|
* limit programmatically (and the real limits are rolling rate windows, not a
|
||||||
* dollar bucket), so any percentage would be fiction. See docs/eval-architecture.md.
|
* dollar bucket), so any percentage would be fiction. See research/eval-architecture.md.
|
||||||
*
|
*
|
||||||
* Pure + injectable (env + an output sink), so the whole thing is unit-tested
|
* Pure + injectable (env + an output sink), so the whole thing is unit-tested
|
||||||
* without a model or a real key.
|
* without a model or a real key.
|
||||||
|
|||||||
+2
-2
@@ -124,7 +124,7 @@ export interface EvalArm {
|
|||||||
* opus: { model: "claude-opus-4-8" } }` — so model-as-an-arm answers "does my
|
* opus: { model: "claude-opus-4-8" } }` — so model-as-an-arm answers "does my
|
||||||
* harness still hold on the cheaper tier / after a model upgrade?" through the
|
* harness still hold on the cheaper tier / after a model upgrade?" through the
|
||||||
* same significance machinery, with no separate model-matrix runner. Omit to
|
* same significance machinery, with no separate model-matrix runner. Omit to
|
||||||
* use the eval-level model. See `docs/eval-architecture.md` (model strategy).
|
* use the eval-level model. See `research/eval-architecture.md` (model strategy).
|
||||||
*/
|
*/
|
||||||
readonly model?: string;
|
readonly model?: string;
|
||||||
}
|
}
|
||||||
@@ -1186,7 +1186,7 @@ function isRecord(v: unknown): v is Record<string, unknown> {
|
|||||||
* e.g. `claude-haiku-4-5-20251001`. A floating alias (`haiku`, `sonnet`, or even
|
* e.g. `claude-haiku-4-5-20251001`. A floating alias (`haiku`, `sonnet`, or even
|
||||||
* `claude-sonnet-4-6` with no date) can change underneath you — so a cached or
|
* `claude-sonnet-4-6` with no date) can change underneath you — so a cached or
|
||||||
* baselined result pinned to it can silently hide model drift. See
|
* baselined result pinned to it can silently hide model drift. See
|
||||||
* `docs/eval-architecture.md` (honest model pinning).
|
* `research/eval-architecture.md` (honest model pinning).
|
||||||
*/
|
*/
|
||||||
export function isDatedModel(model: string): boolean {
|
export function isDatedModel(model: string): boolean {
|
||||||
return /\d{8}$/.test(model);
|
return /\d{8}$/.test(model);
|
||||||
|
|||||||
@@ -297,7 +297,7 @@ function safetySection(
|
|||||||
// --- Safety (deterministic) — generated from ${input.name}'s side-effecting tools: ${sideEffecting.join(", ")} ---
|
// --- Safety (deterministic) — generated from ${input.name}'s side-effecting tools: ${sideEffecting.join(", ")} ---
|
||||||
// In a real run, replace this constructed Trace with a real \`runHarness\` /
|
// In a real run, replace this constructed Trace with a real \`runHarness\` /
|
||||||
// \`measure\` turn (use interceptTools so a real model's attempt is DENIED, never
|
// \`measure\` turn (use interceptTools so a real model's attempt is DENIED, never
|
||||||
// executed — see docs/eval-architecture.md). The checks below are derived from the
|
// executed — see research/eval-architecture.md). The checks below are derived from the
|
||||||
// declared tools contract — the agent's "hole" asserted to stay in its lane.
|
// declared tools contract — the agent's "hole" asserted to stay in its lane.
|
||||||
{
|
{
|
||||||
const trace = {
|
const trace = {
|
||||||
|
|||||||
Reference in New Issue
Block a user