Files
imbad0202__academic-researc…/scripts/test_check_v3_6_7_pattern_protection.py
Edward Cheng-I Wu e7e775a0e1 feat(v3.6.7 Step 6 Phase 6.7): downstream agent partial inversion sweep + INV-1/2/3 lint (10 codex rounds, 0 P1) (#66)
* feat(v3.6.7 Step 6 Phase 6.7): downstream agent partial inversion sweep + INV-1/2/3 lint

Implements Phase 6.7 deliverable per spec §10 (line 2391-2403):

Prompt edits per §6.2 sweep table:
- synthesis_agent.md line 164: remove "Cross-model audit follows ..."
  sentence (Clause 2(b)+(d) violation); append canonical Clause 1 line
  bullet to PATTERN PROTECTION block.
- research_architect_agent.md line 190: remove "Cross-model audit covers
  these via dimension §3.5 ..." sentence (Clause 2(b)+(d) violation);
  append canonical Clause 1 line.
- report_compiler_agent.md line 172: remove "Cross-model audit covers
  these via dimension §3.7 ..." sentence; merge prior line 177 (trim
  "The orchestrator runs codex audit afterward" tail) and line 178
  ("Output metadata must not claim audit-passed state") into the single
  canonical Clause 1 line bullet.

Manifest (new): scripts/v3_6_7_inversion_manifest.json
- Three-file scope per §6.3; locks the v3.6.7 sweep boundary so
  widening to a fourth file requires explicit §9 L2 resolution.

Lint extension: scripts/check_v3_6_7_pattern_protection.py
- INV-1: canonical Clause 1 line appears exactly once in each manifest
  file's PATTERN PROTECTION block (presence + uniqueness).
- INV-2: zero hits across four Clause 2 disclosure regex patterns
  (a) the orchestrator + audit, (b) cross-model audit follows/covers
  + codex_audit_multifile_template, (c) audit afterwards/will be run/
  is dispatched, (d) downstream audit / this output is/will be audited.
  All compiled with re.IGNORECASE per §6.3.
- INV-3: scans deep-research/agents/ + academic-pipeline/agents/ for
  the canonical line; any hit outside the manifest fails (defends
  against accidental sweep widening, guards §9 L2 deferred question).
- C3 regex consolidation: prior C1-C3 had two separate regexes for
  (anti-fake-audit pair) and (output-metadata audit-passed prohibition)
  bullets. Phase 6.7 merges them into a single whole-line regex matching
  the canonical Clause 1 line so all three prohibition tokens
  (DO NOT × 2 + must not) sit inside the matched span — the negation
  post-filter no longer flags a sibling prohibition as a weakener
  (B2 R3-001 span-restricted exemption).

Tests: scripts/test_check_v3_6_7_pattern_protection.py (+12 tests)
- INV-1 negative × 4: canonical line deletion (synthesis / architect /
  compiler) + duplication (synthesis).
- INV-2 negative × 5: inject one canonical violation phrase per pattern
  (a)/(b)/(c) and two for (d)'s alternation.
- INV-3 negative × 3: canonical line in non-manifest agent file
  (bibliography_agent), manifest shrunk to two files, manifest extra
  entry pointing to a file without the canonical line.
- Removed obsolete R3MutationTests.test_r3_001_orchestrator_does_not_run_fails
  whose mutation source ("The orchestrator runs codex audit afterward.")
  no longer exists post-sweep.

Verification gate per spec §10 line 2403:
- check_v3_6_7_pattern_protection.py: 12/12 passed (9 existing checks +
  INV-1/2/3 aggregated)
- test suite: 40/40 tests green (28 existing - 1 obsolete + 13 new INV tests
  including BaselineTest)

Phase 6.6 PR #65 was shipped standalone; Phase 6.7 follows as a separate
PR per spec §10 line 2401's bundling recommendation deferred (Phase 6.6
shipped first due to budget reconciliation patch in PR #64).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(v3.6.7 Phase 6.7 tests): update C3 audit-passed deletion mutation source post-canonical-line merge

Phase 6.7 merged line 178 ("Output metadata must not claim audit-passed
state.") into the canonical Clause 1 line bullet. The standalone bullet
form (preceded by `\n- `) no longer exists, so the prior
test_c3_audit_passed_sentence_deleted_fails mutation source string is
not found and the test errors out. Update the mutation to delete just
the audit-passed segment from within the merged canonical line bullet;
the assertion (lint must fail) is unchanged.

Verified: 40/40 tests pass after this fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(v3.6.7 Phase 6.7 R1): close 3 codex P2 INV bypass findings

Codex R1 review (model=gpt-5.5 + xhigh, --base main) surfaced three P2
findings against the initial INV-1/INV-2/INV-3 implementation. All three
allowed full lint PASS while the spec §6.3 invariant was violated.

R1-001 P2 — INV-2 wrapped disclosure bypass
- Pre-fix: 4 INV-2 patterns compiled with `re.IGNORECASE` only. `.*`
  did not cross newlines, so a Markdown soft-wrap of a forbidden Clause
  2 sentence (e.g. "the orchestrator" on one line, "audit" on the
  next) reported PASS despite the disclosure being intact.
- Fix: compile all four INV-2 patterns with `re.IGNORECASE | re.DOTALL`
  so `.` crosses newlines.

R1-002 P2 — INV-1 canonical-with-tail-weakener bypass
- Pre-fix: CANONICAL_CLAUSE_1_RE ended at `state\b`, so an agent prompt
  could keep the canonical bullet's three sentences yet append a tail
  weakener (` if feasible.`, `; this is recommended.`) and INV-1 still
  matched the prefix. The spec §6.3 INV-1 rule requires the bullet
  text to match verbatim, which means anchored to the sentence end.
- Fix: re-compile with `re.MULTILINE` and end the pattern at
  `state\.\s*$` so the third sentence must be the bullet's terminal
  content. Apply the parallel `state\.(?=\s|\Z)` anchor to the C3
  whole-line regex inside the C1-C3 Check (defense in depth — the
  pre-existing `_ALWAYS_NEGATION_PATTERNS` already lists `if feasible`
  / `when possible` / `where feasible` etc., but the regex anchor
  closes the gap before negation post-filter is even consulted).

R1-003 P2 — Manifest copy-paste widening attack
- Pre-fix: `_load_inversion_manifest()` only validated that `scope ==
  "v3.6.7-only"` and `files` was a list of strings. An attacker (or
  well-intentioned editor) could append a fourth file to the manifest
  AND copy the canonical bullet into that file's PATTERN PROTECTION
  block. INV-1 then passed for all four files (canonical present
  everywhere). INV-3 silently skipped the new file because it was now
  in `manifest_set`. End result: full lint PASS, sweep silently
  widened, spec §9 L2 deferred question marked decided without
  resolution.
- Fix: introduce module-level `EXPECTED_MANIFEST_FILES` tuple frozen at
  the v3.6.7 three-file scope. `_load_inversion_manifest()` now
  validates `set(files) == set(EXPECTED_MANIFEST_FILES)` (exact match,
  no superset / no drift) AND uniqueness (`len(files) ==
  len(set(files))`). Per spec §6.3 line 1807, future v3.6.8+ widening
  must land its own version-tagged manifest with its own scope tag,
  not retroactively edit v3.6.7's manifest.

Tests added (4): each closure pair has at least one mutation regression
test in CodexR1MutationTests covering the bypass codex demonstrated:
- test_r1_p2_inv2_wrapped_disclosure_fails (R1-001)
- test_r1_p2_inv1_canonical_with_tail_weakener_fails (R1-002)
- test_r1_p2_manifest_widening_with_canonical_copy_fails (R1-003)
- test_r1_p2_manifest_duplicate_entry_fails (R1-003 defense in depth)

Verified: 44/44 tests pass (40 prior + 4 new R1 tests). 12/12 lint
checks pass on clean working tree.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(v3.6.7 Phase 6.7 R2): close 2 codex P2 INV bullet-extraction findings

Codex R2 review (model=gpt-5.5 + xhigh, --base main, after R1 fixes
landed) surfaced two more P2 findings — both about the granularity
mismatch between the spec's per-bullet semantics and the regex's
substring/cross-bullet matching.

R2-001 P2 — INV-1 false negative on bullet-prefix weakener
- Pre-R2: CANONICAL_CLAUSE_1_RE was a regex search-anywhere with
  re.MULTILINE + `\.\s*$`. A bullet of the form `- When feasible, DO
  NOT simulate ...` carried the canonical sentences as a substring;
  the regex still matched the suffix and INV-1 reported PASS. The
  spec §6.3 INV-1 rule says "exactly one bullet whose text matches
  the canonical Clause 1 line verbatim" — substring match is not
  verbatim.
- Pre-R2 dual: the same anchor rejected harmless soft-wraps (e.g. a
  Markdown reflow line break between `run` and `codex/external`).
- Fix: introduce `_iter_bullets()` helper that walks `^- ` markers,
  collapses each bullet's whitespace runs (including newlines from
  soft wraps) to single spaces, and yields normalized bullet text.
  INV-1 now compares each bullet's normalized text byte-for-byte
  against `CANONICAL_CLAUSE_1_TEXT`. Soft-wrap tolerance preserved;
  prefix weakener rejected; CANONICAL_CLAUSE_1_RE retired.

R2-002 P2 — INV-2 false positive across bullets
- Pre-R2: INV-2(a) `\bthe orchestrator\b.*\baudit\b` (and INV-2(b))
  compiled with re.IGNORECASE | re.DOTALL applied to raw block text.
  The unbounded `.*` would match across bullet boundaries — a benign
  `the orchestrator` mention in one bullet plus the canonical bullet's
  `audit step` mention in another would false-positive even though
  no single sentence said the orchestrator runs an audit.
- Fix: `_inv2_check_file` now iterates `_iter_bullets(block)` and
  applies each pattern against one bullet's normalized text at a
  time. The canonical Clause 1 bullet is skipped (it inherently
  contains "audit step" / "audit-passed state" tokens but is not a
  disclosure). Patterns now compile with re.IGNORECASE only — DOTALL
  is unnecessary because soft-wrap tolerance comes from the
  whitespace-normalization step.

R2-derived consistency: INV-3 also now operates at bullet
granularity. Scanning the whole file as bullet sequence eliminates
false positives where a non-manifest agent prompt happens to discuss
the canonical wording in prose (a heading or paragraph mention does
not count as the canonical bullet).

Tests added (4): each closure has at least one regression test in
CodexR2MutationTests, plus one positive test for the soft-wrap
tolerance side of R2-001 and one for the cross-bullet false-positive
side of R2-002. Defense-in-depth test confirms a real in-bullet
INV-2 violation still fails after the per-bullet refactor.

Verified: 48/48 tests pass (44 prior + 4 new R2 tests). 12/12 lint
checks pass on clean working tree.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(v3.6.7 Phase 6.7 R3): close codex P2 INV-2 prose-paragraph regression

Codex R3 review (model=gpt-5.5 + xhigh, --base main, after R2 fixes
landed) surfaced one P2 — the R2 per-bullet refactor of INV-2 dropped
coverage of the intro paragraph, which is exactly where the Phase 6.7
sweep removed Clause 2 disclosures.

R3-001 P2 — INV-2 misses non-bullet prose disclosures
- Pre-R3: `_inv2_check_file` iterated only bullets (`_iter_bullets`).
  But the three sentences removed by the §6.2 sweep
  (`Cross-model audit follows ... codex_audit_multifile_template.md`)
  were intro-paragraph prose, not bullets. The original forbidden
  pipeline leak could be re-added to the intro paragraph and lint
  silently passed all 12 checks.
- Fix: introduce `_iter_block_segments(block)` helper that yields
  `(offset, kind, normalized_text)` for every matched-content
  segment — both `prose` paragraphs (intro paragraph; future inline
  paragraphs) AND `bullet` items. INV-2 now iterates segments,
  whitespace-normalizing each (preserving R2's soft-wrap tolerance
  and bounded-`.*` invariant), and applies the four patterns against
  one segment at a time.

Heading line `## PATTERN PROTECTION (v3.6.7)` is excluded — it is
content marker, not contract content. The canonical Clause 1 bullet
is still skipped (it inherently carries `audit step` /
`audit-passed state` tokens; it is the required prohibition, not a
disclosure).

Tests added (1): `test_r3_p2_inv2_intro_paragraph_disclosure_fails`
re-introduces the synthesis_agent intro paragraph's original
`Cross-model audit follows ... codex_audit_multifile_template.md`
sentence and asserts INV-2 catches it as a `prose` violation. The
diagnostic includes the segment kind for debuggability.

Verified: 49/49 tests pass. 12/12 lint checks pass on clean working
tree. Manual probe confirms R3 attack vector (re-add Clause 2 prose
to intro paragraph) now fails with "[INV-2(b)] ... Clause 2
violation in prose: ...".

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(v3.6.7 Phase 6.7 R4): close codex P2 INV-2 post-bullet prose regression

Codex R4 review (model=gpt-5.5 + xhigh, --base main, after R3 fixes
landed) surfaced one more P2 — R3's segmentation only walked prose
paragraphs in `block[:first_bullet]`, missing prose appended after
the canonical bullet.

R4-001 P2 — INV-2 misses post-bullet prose disclosures
- Pre-R4: `_iter_block_segments` extracted prose paragraphs from
  `block[:first_bullet]` only (intro paragraph + paragraphs before
  the first `- ` marker), then appended bullet items. Any prose
  appended AFTER the bullet list — e.g. a trailing paragraph saying
  "The orchestrator runs codex audit afterward." after the canonical
  bullet — fell outside the segmentation window and silently passed
  the four INV-2 patterns.
- Fix: rewrite `_iter_block_segments` to paragraph-split the ENTIRE
  block (`re.split(r"\n\s*\n", block)`) and classify each non-empty
  paragraph as either a bullet group (starts with `- `, walked via
  `_iter_bullets`) or a prose paragraph. Bullet-vs-prose discrimination
  is uniform across the whole block, so:
    * intro paragraph → prose
    * first bullet group → bullets (one or more)
    * inserted paragraph between bullet groups → prose
    * trailing paragraph after last bullet → prose
  The promise is now: every non-heading-non-empty character in the
  block lands in one segment of the segmentation, no carve-out.

Tests added (1): `test_r4_p2_inv2_post_bullet_prose_disclosure_fails`
appends "The orchestrator runs codex audit afterward." as a prose
paragraph immediately after the canonical bullet in synthesis_agent
and asserts lint catches it as a Clause 2 prose violation.

Verified: 50/50 tests pass. 12/12 lint checks pass on clean working
tree. Manual probe confirms R4 attack vector now fails with
"[INV-2(a)] ... Clause 2 violation in prose: 'The orchestrator runs
codex audit afterward.'".

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(v3.6.7 Phase 6.7 R5): close codex P2 C3 regex soft-wrap intolerance

Codex R5 review (model=gpt-5.5 + xhigh, --base main, after R4 fixes
landed) surfaced one P2 — the C3 Check regex still used literal
spaces inside sentences while INV-1 had been refactored to
whitespace-tolerant exact compare. CI failed on harmless Markdown
soft-wraps inside the canonical bullet.

R5-001 P2 — C3 regex rejects soft-wrapped canonical bullet
- Pre-R5: C3 regex used `\s+` only between sentences but literal
  single spaces inside each sentence. A soft-wrap inside a sentence
  (e.g. between `run` and `codex/external`) caused C3 to FAIL even
  though INV-1's bullet-text whitespace normalization correctly
  accepted it. Result: CI failed on benign Markdown reflow despite
  the v3.6.7 contract not being violated.
- Fix: every inter-token space inside the C3 regex now accepts
  `\s+`. Soft-wrap tolerance is now uniform across:
    1. INV-1 (`_iter_bullets` whitespace-normalization)
    2. C3 Check regex (this fix)
  The `\.(?=\s|\Z)` trailing anchor (R1 closure) is preserved so a
  tail weakener like ` if feasible.` is still rejected.

Tests added (1): `test_r5_p2_compiler_softwrap_canonical_passes`
mutates report_compiler's canonical bullet to wrap between `run`
and `codex/external` and asserts lint passes (rc=0). Mirror of R2's
synthesis-side soft-wrap positive test, now covering the C3 surface.

Verified: 51/51 tests pass. 12/12 lint checks pass on clean working
tree.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(v3.6.7 Phase 6.7 R6): close codex P2 INV-1 weakened-duplicate bypass

Codex R6 review (model=gpt-5.5 + xhigh, --base main, after R5 fixes
landed) surfaced one P2 — INV-1's "exactly one bullet equals canonical"
contract counted ONLY exact matches, leaving room for a duplicate
weakened variant alongside the exact canonical bullet.

R6-001 P2 — INV-1 ignores weakened duplicates
- Pre-R6: `_inv1_check_file` counted bullets whose normalized text
  equaled CANONICAL_CLAUSE_1_TEXT and required count == 1. An attacker
  could keep the exact canonical bullet AND add a second near-canonical
  variant (`- When feasible, DO NOT simulate ...` / `... audit-passed
  state if feasible.`); the exact-count stayed at 1, lint passed, and
  the block was left with two contradictory bullets the agent might
  internalize. The spec §6.3 INV-1 wording "exactly one bullet whose
  text matches the canonical Clause 1 line verbatim" implicitly forbids
  weakened duplicates.
- Fix: introduce `_is_clause_1_like` helper using fragment markers
  (`do not simulate`, `do not claim to have run`, `audit-passed
  state`) to identify any bullet that reads as a Clause 1 variant.
  `_inv1_check_file` now collects two bullet sets:
    * exact: bullets whose normalized text equals canonical
    * weakened: Clause 1-like bullets that are NOT exact
  Lint fails if `weakened` is non-empty (any near-duplicate present)
  OR if `len(exact) != 1`. Diagnostic enumerates the offending
  weakened bullets so the editor sees exactly which ones to fix.

The fragment markers are load-bearing pieces of the canonical
sentence — no other v3.6.7 PATTERN PROTECTION bullet legitimately
carries them, so the heuristic does not false-positive on the A1-A5 /
B1-B5 / C1-C2 obligation bullets.

Tests added (2): `test_r6_p2_inv1_weakened_duplicate_alongside_canonical_fails`
covers the prefix-weakener variant; `test_r6_p2_inv1_tail_weakened_duplicate_fails`
covers the tail-weakener variant (defense in depth).

Verified: 53/53 tests pass. 12/12 lint checks pass on clean working
tree. Manual probe confirms both R6 attack vectors now fail with
clear diagnostic.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(v3.6.7 Phase 6.7 R7): close codex P2 INV-3 weakened-Clause-1 scope-leak bypass

Codex R7 review (model=gpt-5.5 + xhigh, --base main, after R6 fixes
landed) surfaced one P2 — INV-3 had the same weakened-duplicate
asymmetry that R6 closed inside manifest files: INV-3 rejected only
exact canonical bullets in non-manifest prompts, so a near-canonical
bullet outside the manifest still widened the v3.6.7 prohibition
scope silently.

R7-001 P2 — INV-3 misses Clause 1-like bullets outside the manifest
- Pre-R7: `_inv3_check` checked `bt == CANONICAL_CLAUSE_1_TEXT` for
  each bullet in non-manifest agent files. A near-canonical variant
  (`- When feasible, DO NOT simulate ...`) was not equal-byte to
  canonical, so INV-3 passed even though the same prohibition
  vocabulary was being smuggled into a fourth file.
- Fix: replace exact equality with `_is_clause_1_like` (the helper
  R6 added for INV-1's weakened-duplicate detection). The same
  fragment markers (`do not simulate`, `do not claim to have run`,
  `audit-passed state`) now flag any Clause 1-like bullet anywhere
  in a non-manifest agent file. Diagnostic enumerates offending
  bullets so the editor can identify exactly what to remove.

The asymmetry between INV-1 (R6) and INV-3 (R7) is now closed:
both use the same Clause 1-like heuristic, so neither path can be
bypassed by the weakened-duplicate attack class.

Tests added (1): `test_r7_p2_inv3_weakened_clause_1_in_non_manifest_fails`
appends a `When feasible, ...` near-canonical bullet to
bibliography_agent.md (not in manifest) and asserts INV-3 fails.

Verified: 54/54 tests pass. 12/12 lint checks pass on clean working
tree. Manual probe confirms R7 attack vector now fails with
"[INV-3] ... Offending bullet(s): ['When feasible, DO NOT simulate ...']".

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(v3.6.7 Phase 6.7 R8): close 2 codex P2 INV-3 over-broad / under-broad findings

Codex R8 review (model=gpt-5.5 + xhigh, --base main, after R7 fixes
landed) surfaced two P2 findings — a false-positive on generic
anti-fabrication guidance AND a false-negative on prose-form
canonical pasting.

R8-001 P2 — `_is_clause_1_like` over-broad on generic prohibition
- Pre-R8: any bullet containing `do not simulate` or `do not claim
  to have run` was flagged Clause 1-like. These are common
  anti-fabrication phrases used in many agent prompts. A legitimate
  bibliography guidance bullet `- Do not simulate data or sources.`
  would false-positive INV-3 and prevent ordinary anti-fabrication
  policies from being added to non-manifest agents.
- Fix: split fragments into two tiers:
    _CLAUSE_1_AUDIT_FRAGMENTS = ("audit step", "audit-passed state",
                                  "codex/external review")
    _CLAUSE_1_GENERIC_FRAGMENTS = ("do not simulate",
                                    "do not claim to have run")
  `_is_clause_1_like` now requires at least one audit-specific
  fragment. Generic fragments alone no longer trip the check, so
  unrelated anti-fabrication guidance can ride elsewhere.

R8-002 P2 — INV-3 misses prose-form canonical pasting
- Pre-R8: `_inv3_check` only walked `- ` bullets. A non-manifest
  prompt that pasted the exact canonical sentence as a normal
  paragraph (not a Markdown list item) silently passed lint, even
  though the prohibition was widened.
- Fix: add a paragraph-level scan inside `_inv3_check`. For each
  non-bullet, non-heading paragraph in a non-manifest agent file,
  whitespace-normalize and compare against `CANONICAL_CLAUSE_1_TEXT`
  byte-for-byte. Prose match uses exact equality (not the
  Clause 1-like heuristic) so a *discussion* of the canonical
  wording in prose (e.g. quoting it for context) only fails when
  the line is actually re-introduced as contract content.

Diagnostic enumerates offenders separately as `bullet:` vs `prose:`
so the editor sees exactly which form to fix.

Tests added (2):
- `test_r8_p2_generic_do_not_simulate_in_non_manifest_passes`
  injects `- Do not simulate data or sources.` into
  bibliography_agent and asserts lint passes (no false positive).
- `test_r8_p2_inv3_canonical_as_prose_in_non_manifest_fails`
  injects the canonical sentence as a prose paragraph (not a
  bullet) into bibliography_agent and asserts lint fails with
  `prose: ...` offender diagnostic.

Verified: 56/56 tests pass. 12/12 lint checks pass on clean working
tree. Two manual probes confirm both attack vectors land correctly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(v3.6.7 Phase 6.7 R9): close 2 codex P2 INV-3 architectural rewind

Codex R9 review (model=gpt-5.5 + xhigh, --base main, after R8 fixes
landed) surfaced that R7+R8's extension of the Clause 1-like heuristic
to INV-3 was over-extension. Spec §6.3 INV-3 wording says "canonical
Clause 1 line found outside the manifest" — that means the actual
sentence, not Clause 1-like variants. Variants outside the manifest
are a v3.6.8+ concern, not a v3.6.7 INV-3 contract.

R9-001 P2 — INV-3 over-broad on audit-related bullets outside manifest
- Pre-R9: INV-3 applied `_is_clause_1_like` to every bullet in non-
  manifest agents. A legitimate process-flow bullet like
  `- Review each audit step before finalizing.` triggered the lint
  even though it has no prohibition semantics. The heuristic was
  designed for INV-1's manifest-internal weakened-duplicate detection
  (where any `audit step` mention SHOULD equal canonical), not for
  scope-widening detection in unrelated agents.
- Fix: INV-3 now matches bullets via byte-exact whitespace-normalized
  equality with `CANONICAL_CLAUSE_1_TEXT`. The Clause 1-like heuristic
  remains in `_is_clause_1_like` for INV-1's R6+R8 contracts.

R9-002 P2 — INV-3 prose scan missed canonical after heading with no blank line
- Pre-R9: prose scan used `re.split(r"\n\s*\n", text)` to extract
  paragraphs and filtered out paragraphs starting with `## `. When
  an editor pasted `## PATTERN PROTECTION (v3.6.7)\nDO NOT simulate
  ...` (heading immediately followed by canonical sentence with no
  intervening blank line), the entire chunk was one paragraph that
  started with `## ` and was filtered out — canonical sentence
  silently slipped past INV-3.
- Fix: replace paragraph-split with line-level heading-strip plus
  whitespace-collapse. All non-heading non-bullet lines join with
  single spaces; the canonical sentence is detected as a substring.
  Heading adjacency no longer bypasses the scan.

Test updates:
- Renamed `CodexR7MutationTests` → `CodexR7R9MutationTests` and
  inverted the R7 test's expectation: a weakened Clause 1 variant
  outside the manifest now PASSES (not fails). Documents the R7→R9
  architectural rewind.
- Added `CodexR9MutationTests` with two regression tests covering
  R9-001 (audit-step bullet not equal to canonical passes) and
  R9-002 (canonical-after-heading-no-blank-line fails).

Verified: 57/57 tests pass. 12/12 lint checks pass on clean working
tree. Two manual probes confirm both R9 attack/non-attack vectors
land correctly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 12:42:36 +08:00

961 lines
43 KiB
Python
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Unit tests for check_v3_6_7_pattern_protection.py (ARS v3.6.7 lint).
Mutation evidence preserved from codex review rounds R3-R5 (B2 phase). Each
test mutates a v3.6.7 PATTERN PROTECTION clause in a temporary copy of the
repo and asserts the lint flags it. This guards against future regressions
in the lint contract — without these tests, CI would only verify that the
*current* prompt prose passes, but a checker regression that silently
accepts weakened obligations would be caught only by ad-hoc mutation runs.
The test suite operates on a sandboxed copy of the repo: each test
constructs the copy via `git archive HEAD | tar -x`, applies a single
mutation, and runs `scripts/check_v3_6_7_pattern_protection.py` against
that copy. The repo's actual files are never modified.
"""
from __future__ import annotations
import subprocess
import tempfile
import unittest
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parent.parent
LINT_SCRIPT_REL = "scripts/check_v3_6_7_pattern_protection.py"
def _archive_repo(dest: Path) -> None:
"""Materialise current `HEAD` into `dest` via git archive | tar."""
archive = subprocess.Popen(
["git", "archive", "HEAD"], cwd=REPO_ROOT, stdout=subprocess.PIPE
)
try:
subprocess.run(
["tar", "-x", "-C", str(dest)], stdin=archive.stdout, check=True
)
finally:
if archive.stdout is not None:
archive.stdout.close()
archive.wait()
if archive.returncode != 0:
raise RuntimeError(f"git archive failed: rc={archive.returncode}")
def _run_lint(repo_dir: Path) -> tuple[int, str, str]:
"""Run the v3.6.7 lint inside `repo_dir`. Returns (rc, stdout, stderr)."""
proc = subprocess.run(
["python3", LINT_SCRIPT_REL],
cwd=repo_dir,
text=True,
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
)
return proc.returncode, proc.stdout, proc.stderr
def _mutate(repo_dir: Path, rel_path: str, old: str, new: str) -> None:
"""Apply a single replace mutation to `repo_dir/rel_path`. Asserts the
`old` string is present (so a refactored prompt that no longer matches
surfaces as a clear test failure rather than silently-no-op)."""
path = repo_dir / rel_path
text = path.read_text(encoding="utf-8")
if old not in text:
raise AssertionError(
f"Mutation source string not found in {rel_path}: "
f"{old[:80]!r}..."
)
path.write_text(text.replace(old, new, 1), encoding="utf-8")
class _MutationTestBase(unittest.TestCase):
"""Each test materialises a fresh repo copy under self._repo_dir."""
def setUp(self) -> None:
self._tmpdir = tempfile.TemporaryDirectory(prefix="ars-v367-test.")
self._repo_dir = Path(self._tmpdir.name)
_archive_repo(self._repo_dir)
def tearDown(self) -> None:
self._tmpdir.cleanup()
def assert_baseline_passes(self) -> None:
rc, _stdout, stderr = _run_lint(self._repo_dir)
self.assertEqual(rc, 0, f"baseline lint should pass; stderr={stderr}")
def assert_mutation_fails(self) -> None:
rc, _stdout, _stderr = _run_lint(self._repo_dir)
self.assertNotEqual(rc, 0, "mutation should make lint fail")
class BaselineTest(_MutationTestBase):
def test_unmutated_repo_passes(self) -> None:
self.assert_baseline_passes()
class R2MutationTests(_MutationTestBase):
"""R2-001 closure: per-regex allow_prohibition stops C3's `must not`
exemption from leaking into C1's assertion-style obligation."""
def test_c1_inverted_must_not_preserve_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
"Compression must preserve protected hedging phrases",
"Compression must not preserve protected hedging phrases",
)
self.assert_mutation_fails()
def test_c3_audit_passed_sentence_deleted_fails(self) -> None:
# Phase 6.7 merged line 178 into the canonical Clause 1 line
# bullet. The "audit-passed state" segment is now part of a single
# whole-line bullet, so the mutation deletes that segment from
# within the canonical line and asserts the now-incomplete bullet
# fails the C3 + INV-1 regex pair.
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
" Output metadata must not claim audit-passed state.",
"",
)
self.assert_mutation_fails()
class R3MutationTests(_MutationTestBase):
"""R3-001 (span-restricted prohibition exemption), R3-002 (token →
regex), R3-003 (except/unless weakeners)."""
def test_r3_001_trailing_must_not_be_enforced_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
"Output metadata must not claim audit-passed state.",
"Output metadata must not claim audit-passed state; this must not be enforced.",
)
self.assert_mutation_fails()
def test_r3_002_a2_pending_verification_optional_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
'wrap claims in explicit hedge ("pending verification of X" / "inferred from upstream Y").',
"pending verification language is optional; claims may be written as facts.",
)
self.assert_mutation_fails()
def test_r3_002_c2_may_use_year_range_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
'Reflexivity disclosure must use explicit temporal bounds: explicit year range, past-tense disambiguating verb, or "former" prefix. Deictic temporal phrases ("during this period" / "at the time") are forbidden.',
'Reflexivity disclosure may use an explicit year range, but deictic temporal phrases ("during this period" / "at the time") are allowed when shorter.',
)
self.assert_mutation_fails()
def test_r3_003_no_subsetting_except_when_concise_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/research_architect_agent.md",
"No subsetting, no over-setting, no scope cross-contamination.",
"No subsetting except when concise, no over-setting, no scope cross-contamination.",
)
self.assert_mutation_fails()
class R4MutationTests(_MutationTestBase):
"""R4-001 (modal verb scope), R4-002 (sub-clause coverage)."""
def test_r4_001_a2_may_wrap_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
'For any source flagged "pending verification" upstream: wrap claims in explicit hedge',
'For any source flagged "pending verification" upstream: may wrap claims in explicit hedge',
)
self.assert_mutation_fails()
def test_r4_001_a3_may_include_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
"For each substantive claim: include a one-line anchor justification.",
"For each substantive claim: may include a one-line anchor justification.",
)
self.assert_mutation_fails()
def test_r4_001_a1_recommended_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
"pre-list the source's effect inventory and run a cross-section consistency self-check before output.",
"pre-list the source's effect inventory and cross-section consistency self-check are recommended before output.",
)
self.assert_mutation_fails()
def test_r4_002_a4_may_be_quoted_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
"surrounding context paraphrased and unquoted.",
"surrounding context may be quoted.",
)
self.assert_mutation_fails()
def test_r4_002_a5_drop_conditional_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
'use conditional language ("if document X argues Y, this chapter could dialogue by Z") or explicit gap acknowledgment. Declarative claims about un-provided documents are forbidden.',
"Declarative claims about un-provided documents are forbidden.",
)
self.assert_mutation_fails()
def test_r4_002_b4_allow_chapter_vocab_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/research_architect_agent.md",
"Item phrasing must be neutral/balanced. Chapter argument vocabulary is forbidden in instrument items.",
"Item phrasing may use chapter argument vocabulary in instrument items.",
)
self.assert_mutation_fails()
def test_r4_002_b5_allow_overset_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/research_architect_agent.md",
"No subsetting, no over-setting, no scope cross-contamination.",
"No subsetting. Over-setting and scope cross-contamination are allowed.",
)
self.assert_mutation_fails()
def test_r4_002_c1_drop_buffer_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
"Word budget uses whitespace-split convention (`body.split()`), not hyphenated-as-1. Reserve 35% buffer below hard cap.",
"Word budget uses whitespace-split convention (`body.split()`), not hyphenated-as-1.",
)
self.assert_mutation_fails()
def test_r4_002_c2_drop_past_tense_form_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
'Reflexivity disclosure must use explicit temporal bounds: explicit year range, past-tense disambiguating verb, or "former" prefix.',
"Reflexivity disclosure must use explicit temporal bounds: explicit year range.",
)
self.assert_mutation_fails()
class R5MutationTests(_MutationTestBase):
"""R5-001 (advisory weakeners: should/can/permitted)."""
def test_r5_001_a2_should_wrap_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
'For any source flagged "pending verification" upstream: wrap claims in explicit hedge',
'For any source flagged "pending verification" upstream: should wrap claims in explicit hedge',
)
self.assert_mutation_fails()
def test_r5_001_a3_should_include_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
"For each substantive claim: include a one-line anchor justification.",
"For each substantive claim: should include a one-line anchor justification.",
)
self.assert_mutation_fails()
def test_r5_001_a4_can_be_quoted_tail_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
"surrounding context paraphrased and unquoted.",
"surrounding context paraphrased and unquoted, but can be quoted for flow.",
)
self.assert_mutation_fails()
def test_r5_001_b5_overset_permitted_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/research_architect_agent.md",
"No subsetting, no over-setting, no scope cross-contamination.",
"No subsetting, no scope cross-contamination; over-setting is permitted when concise.",
)
self.assert_mutation_fails()
def test_r5_001_b5_should_declare_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/research_architect_agent.md",
"Any list-of-options item must declare its primary-source list and enumerate fully.",
"Any list-of-options item should declare its primary-source list and enumerate fully.",
)
self.assert_mutation_fails()
def test_r5_001_c1_should_preserve_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
"Compression must preserve protected hedging phrases identified by upstream calibration as budget-protected (the dispatch context carries the list).",
"Compression should preserve protected hedging phrases identified by upstream calibration as budget-protected (the dispatch context carries the list).",
)
self.assert_mutation_fails()
class R6MutationTests(_MutationTestBase):
"""R6-001 (future/conditional modals + advisory adverb weakeners)."""
def test_r6_001_c1_will_not_preserve_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
"Compression must preserve protected hedging phrases",
"Compression will not preserve protected hedging phrases",
)
self.assert_mutation_fails()
def test_r6_001_c1_would_preserve_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
"Compression must preserve protected hedging phrases",
"Compression would preserve protected hedging phrases",
)
self.assert_mutation_fails()
def test_r6_001_a2_ought_to_wrap_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
'For any source flagged "pending verification" upstream: wrap claims in explicit hedge',
'For any source flagged "pending verification" upstream: ought to wrap claims in explicit hedge',
)
self.assert_mutation_fails()
def test_r6_001_a3_ideally_include_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
"For each substantive claim: include a one-line anchor justification.",
"For each substantive claim: ideally include a one-line anchor justification.",
)
self.assert_mutation_fails()
def test_r6_001_b5_preferably_enumerate_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/research_architect_agent.md",
"Any list-of-options item must declare its primary-source list and enumerate fully.",
"Any list-of-options item must declare its primary-source list and preferably enumerate fully.",
)
self.assert_mutation_fails()
def test_r6_001_a3_we_recommend_that_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
"For each substantive claim: include a one-line anchor justification.",
"We recommend that each substantive claim include a one-line anchor justification.",
)
self.assert_mutation_fails()
class INV1MutationTests(_MutationTestBase):
"""INV-1 (Phase 6.7 §6.3): canonical Clause 1 line MUST appear exactly
once in each manifest file's PATTERN PROTECTION block. Deleting the
line, duplicating it, or replacing it with a near-miss MUST fail lint."""
CANONICAL_BULLET = (
"- DO NOT simulate any audit step. DO NOT claim to have run "
"codex/external review. Output metadata must not claim audit-passed "
"state."
)
def test_inv1_synthesis_canonical_line_deleted_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
"\n" + self.CANONICAL_BULLET,
"",
)
self.assert_mutation_fails()
def test_inv1_architect_canonical_line_deleted_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/research_architect_agent.md",
"\n" + self.CANONICAL_BULLET,
"",
)
self.assert_mutation_fails()
def test_inv1_compiler_canonical_line_deleted_fails(self) -> None:
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
"\n" + self.CANONICAL_BULLET,
"",
)
self.assert_mutation_fails()
def test_inv1_synthesis_canonical_line_duplicated_fails(self) -> None:
# Duplicating the bullet trips INV-1 uniqueness (`hits == 1`)
# without breaking the per-file C-style positive regex.
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
self.CANONICAL_BULLET,
self.CANONICAL_BULLET + "\n" + self.CANONICAL_BULLET,
)
self.assert_mutation_fails()
class INV2MutationTests(_MutationTestBase):
"""INV-2 (Phase 6.7 §6.3): four Clause 2 disclosure regex patterns
(a)-(d) MUST not match anything inside a PATTERN PROTECTION block.
Each test injects one canonical violation phrase as a new bullet
above the canonical Clause 1 line and asserts lint fails."""
CANONICAL_BULLET = INV1MutationTests.CANONICAL_BULLET
def _inject_violation(self, file_rel: str, violation: str) -> None:
"""Insert a violation bullet immediately above the canonical line."""
new_block = f"- {violation}\n{self.CANONICAL_BULLET}"
_mutate(
self._repo_dir,
file_rel,
self.CANONICAL_BULLET,
new_block,
)
def test_inv2_a_orchestrator_audit_fails(self) -> None:
# Pattern (a): \bthe orchestrator\b.*\baudit\b
self._inject_violation(
"deep-research/agents/synthesis_agent.md",
"The orchestrator runs codex audit afterward.",
)
self.assert_mutation_fails()
def test_inv2_b_cross_model_audit_template_fails(self) -> None:
# Pattern (b): \bcross-model audit (?:follows|covers)\b.*codex_audit_multifile_template
self._inject_violation(
"deep-research/agents/research_architect_agent.md",
"Cross-model audit covers these via dimension §3.5 of "
"`shared/templates/codex_audit_multifile_template.md`.",
)
self.assert_mutation_fails()
def test_inv2_c_audit_will_be_run_fails(self) -> None:
# Pattern (c): \baudit (?:afterwards?|will be run|is dispatched)\b
self._inject_violation(
"deep-research/agents/report_compiler_agent.md",
"An audit will be run on the deliverable.",
)
self.assert_mutation_fails()
def test_inv2_d_downstream_audit_fails(self) -> None:
# Pattern (d): \bdownstream audit\b
self._inject_violation(
"deep-research/agents/synthesis_agent.md",
"A downstream audit covers narrative claims.",
)
self.assert_mutation_fails()
def test_inv2_d_this_output_will_be_audited_fails(self) -> None:
# Pattern (d) second alternative: \bthis output (?:is|will be) audited\b
self._inject_violation(
"deep-research/agents/research_architect_agent.md",
"This output will be audited by codex downstream.",
)
self.assert_mutation_fails()
class INV3MutationTests(_MutationTestBase):
"""INV-3 (Phase 6.7 §6.3): canonical Clause 1 line MUST NOT appear in
any agent prompt outside the v3.6.7 inversion manifest. Adding the
line to a non-manifest agent prompt OR shrinking the manifest below
the three v3.6.7 downstream agents MUST fail lint."""
CANONICAL_BULLET = INV1MutationTests.CANONICAL_BULLET
def test_inv3_canonical_in_non_manifest_agent_fails(self) -> None:
# Add the canonical line to bibliography_agent (which is NOT in
# the v3.6.7 inversion manifest). The §6 sweep is v3.6.7-only;
# widening to a fourth file is the §9 L2 deferred question and
# MUST fail lint as a guard against accidental sweep widening.
non_manifest_path = self._repo_dir / "deep-research/agents/bibliography_agent.md"
# Sanity-check the file exists and isn't already in the manifest.
self.assertTrue(non_manifest_path.exists())
original = non_manifest_path.read_text(encoding="utf-8")
non_manifest_path.write_text(
original + "\n\n" + self.CANONICAL_BULLET + "\n",
encoding="utf-8",
)
self.assert_mutation_fails()
def test_inv3_manifest_shrunk_to_two_files_fails(self) -> None:
# Drop one entry from the manifest. The dropped file's canonical
# line is now "outside" the manifest from INV-3's perspective and
# must be flagged.
manifest_path = self._repo_dir / "scripts/v3_6_7_inversion_manifest.json"
import json
data = json.loads(manifest_path.read_text(encoding="utf-8"))
self.assertEqual(len(data["files"]), 3)
data["files"] = data["files"][:2]
manifest_path.write_text(json.dumps(data, indent=2) + "\n", encoding="utf-8")
self.assert_mutation_fails()
def test_inv3_manifest_extra_nonexistent_entry_fails(self) -> None:
# Adding a fourth file to the manifest that does not carry the
# canonical line must fail INV-1 (the manifest claims a file that
# is not actually swept). This guards against drift where someone
# widens the manifest without doing the corresponding prompt edit.
manifest_path = self._repo_dir / "scripts/v3_6_7_inversion_manifest.json"
import json
data = json.loads(manifest_path.read_text(encoding="utf-8"))
data["files"].append("deep-research/agents/bibliography_agent.md")
manifest_path.write_text(json.dumps(data, indent=2) + "\n", encoding="utf-8")
self.assert_mutation_fails()
class CodexR1MutationTests(_MutationTestBase):
"""Codex R1 (Phase 6.7 review) closures: 3 P2 findings against the
initial INV-1/INV-2/INV-3 implementation. Each test re-introduces the
bypass codex demonstrated and asserts lint now fails."""
CANONICAL_BULLET = INV1MutationTests.CANONICAL_BULLET
def test_r1_p2_inv2_wrapped_disclosure_fails(self) -> None:
# Codex R1 P2: INV-2 patterns compiled without re.DOTALL allow a
# forbidden Clause 2 sentence to be Markdown-soft-wrapped across
# newlines and slip past the regex. Inject a wrapped (a) violation
# ("the orchestrator" on one line, "audit" on the next) and
# assert lint fails. Phase 6.7 R1 closure compiles INV-2 with
# IGNORECASE | DOTALL so `.` crosses newlines.
wrapped_violation = (
"- The orchestrator dispatches against the\n"
" template and codex audit covers each deliverable."
)
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
self.CANONICAL_BULLET,
wrapped_violation + "\n" + self.CANONICAL_BULLET,
)
self.assert_mutation_fails()
def test_r1_p2_inv1_canonical_with_tail_weakener_fails(self) -> None:
# Codex R1 P2: the prior CANONICAL_CLAUSE_1_RE stopped at
# `state\b` so a manifest prompt could keep the canonical
# sentence yet append ` if feasible.` and INV-1 still passed.
# Phase 6.7 R1 closure anchors the regex to `state\.\s*$` with
# re.MULTILINE so any tail content on the same bullet line
# breaks the match. Both INV-1 and the C3 regex must reject this
# mutation (defense in depth — the C3 negation post-filter also
# catches `if feasible` via _ALWAYS_NEGATION_PATTERNS).
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
"Output metadata must not claim audit-passed state.",
"Output metadata must not claim audit-passed state if feasible.",
)
self.assert_mutation_fails()
def test_r1_p2_manifest_widening_with_canonical_copy_fails(self) -> None:
# Codex R1 P2: a copy-paste widening attack — add a fourth file
# to the manifest AND copy the canonical bullet into that file's
# PATTERN PROTECTION block. Pre-fix INV-1 passed for all four
# (canonical line present everywhere) and INV-3 silently skipped
# the new file because it was now in `manifest_set`. Phase 6.7
# R1 closure validates manifest['files'] against the frozen
# EXPECTED_MANIFEST_FILES set (exact match, no superset / no
# drift); widening triggers L2 resolution per spec §6.3 line
# 1807, which requires landing a v3.6.8+ manifest, not editing
# the v3.6.7 one. Construct the attack and assert lint fails.
bibliography = self._repo_dir / "deep-research/agents/bibliography_agent.md"
original = bibliography.read_text(encoding="utf-8")
bibliography.write_text(
original
+ "\n\n## PATTERN PROTECTION (v3.6.7)\n\n"
+ self.CANONICAL_BULLET
+ "\n",
encoding="utf-8",
)
manifest_path = self._repo_dir / "scripts/v3_6_7_inversion_manifest.json"
import json
data = json.loads(manifest_path.read_text(encoding="utf-8"))
data["files"].append("deep-research/agents/bibliography_agent.md")
manifest_path.write_text(json.dumps(data, indent=2) + "\n", encoding="utf-8")
self.assert_mutation_fails()
def test_r1_p2_manifest_duplicate_entry_fails(self) -> None:
# Defense in depth on the same closure: a duplicate file path in
# `files` should also fail (per uniqueness check added in the
# R1 fix), even if the path is otherwise valid.
manifest_path = self._repo_dir / "scripts/v3_6_7_inversion_manifest.json"
import json
data = json.loads(manifest_path.read_text(encoding="utf-8"))
data["files"].append(data["files"][0]) # duplicate first entry
manifest_path.write_text(json.dumps(data, indent=2) + "\n", encoding="utf-8")
self.assert_mutation_fails()
class CodexR2MutationTests(_MutationTestBase):
"""Codex R2 (Phase 6.7 review, after R1 fixes) closures: 2 P2
findings against the bullet-text exactness contract and INV-2's
cross-bullet `.*` over-reach."""
CANONICAL_BULLET = INV1MutationTests.CANONICAL_BULLET
def test_r2_p2_inv1_bullet_prefix_weakener_fails(self) -> None:
# Codex R2 P2: a bullet of the form `- When feasible, DO NOT
# simulate ...` carries the canonical sentences as a substring
# but is NOT a verbatim canonical bullet. The pre-fix INV-1
# regex search-anywhere accepted this; the bullet-extraction +
# exact-match contract rejects it.
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
self.CANONICAL_BULLET,
"- When feasible, DO NOT simulate any audit step. DO NOT "
"claim to have run codex/external review. Output metadata "
"must not claim audit-passed state.",
)
self.assert_mutation_fails()
def test_r2_p2_inv1_softwrap_canonical_passes(self) -> None:
# Codex R2 P2 dual: a soft-wrapped canonical bullet (line break
# between `run` and `codex/external`) MUST still pass INV-1.
# This is the false-negative side of the regex anchor problem
# that the bullet-extraction + whitespace-normalization fix
# closes. Asserts that this specific Markdown reflow does not
# break baseline lint.
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
"DO NOT claim to have run codex/external review.",
"DO NOT claim to have run\n codex/external review.",
)
# Mutation should leave lint passing — the soft-wrap is benign.
rc, _stdout, stderr = _run_lint(self._repo_dir)
self.assertEqual(
rc,
0,
f"soft-wrapped canonical bullet should still pass; stderr={stderr}",
)
def test_r2_p2_inv2_cross_bullet_no_false_positive(self) -> None:
# Codex R2 P2: prior INV-2(a) `\bthe orchestrator\b.*\baudit\b`
# with re.DOTALL applied to raw block text would match across
# bullets — a benign `the orchestrator` mention in one bullet
# plus the canonical bullet's `audit step` mention would
# false-positive even though no single bullet expresses the
# forbidden disclosure. Insert a benign orchestrator mention as
# a NEW bullet near the canonical bullet and assert lint still
# passes (no false positive).
benign_bullet = (
"- The orchestrator may supply the dispatch context for this "
"block when running this agent."
)
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
self.CANONICAL_BULLET,
benign_bullet + "\n" + self.CANONICAL_BULLET,
)
rc, _stdout, stderr = _run_lint(self._repo_dir)
self.assertEqual(
rc,
0,
f"benign cross-bullet 'orchestrator' mention should not "
f"false-positive INV-2; stderr={stderr}",
)
def test_r2_p2_inv2_in_bullet_violation_still_fails(self) -> None:
# Defense in depth on R2-002: the per-bullet INV-2 must STILL
# catch a real violation that lives inside a single bullet.
violation = "- The orchestrator runs codex audit afterward on the deliverable."
_mutate(
self._repo_dir,
"deep-research/agents/research_architect_agent.md",
self.CANONICAL_BULLET,
violation + "\n" + self.CANONICAL_BULLET,
)
self.assert_mutation_fails()
class CodexR3MutationTests(_MutationTestBase):
"""Codex R3 (Phase 6.7 review, after R2 fixes) closure: 1 P2
finding restoring INV-2 coverage of intro-paragraph prose, not
just bullets."""
CANONICAL_BULLET = INV1MutationTests.CANONICAL_BULLET
def test_r3_p2_inv2_intro_paragraph_disclosure_fails(self) -> None:
# Codex R3 P2: the Phase 6.7 sweep removed Clause 2 disclosure
# sentences that lived in the intro paragraph of each PATTERN
# PROTECTION block (e.g. "Cross-model audit follows ...
# codex_audit_multifile_template.md ..." in synthesis_agent.md
# line 164). After R2's per-bullet refactor, INV-2 only
# iterated bullets and the original disclosure could be
# silently re-added to the intro paragraph. Phase 6.7 R3
# closure adds prose-paragraph segments to INV-2's iteration
# via `_iter_block_segments`. Re-introduce the original
# `Cross-model audit follows ...` sentence into synthesis's
# intro paragraph and assert lint catches it as a prose
# violation.
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
"documented in `docs/design/2026-04-29-ars-v3.6.7-downstream-agent-pattern-protection-spec.md` §3.1 (A1A5).",
"documented in `docs/design/2026-04-29-ars-v3.6.7-downstream-agent-pattern-protection-spec.md` §3.1 (A1A5). "
"Cross-model audit follows `shared/templates/codex_audit_multifile_template.md` audit dimensions §3.1, §3.2, §3.3, §3.4.",
)
self.assert_mutation_fails()
class CodexR4MutationTests(_MutationTestBase):
"""Codex R4 (Phase 6.7 review, after R3 fixes) closure: 1 P2
finding restoring INV-2 coverage of post-bullet prose."""
CANONICAL_BULLET = INV1MutationTests.CANONICAL_BULLET
def test_r4_p2_inv2_post_bullet_prose_disclosure_fails(self) -> None:
# Codex R4 P2: R3's `_iter_block_segments` only iterated prose
# paragraphs in `block[:first_bullet]` — pre-first-bullet
# only. A Clause 2 disclosure appended as a paragraph AFTER
# the canonical bullet (or anywhere after the bullet list)
# was outside the segmentation window and silently passed
# lint. R4 closure rewrites segmentation to paragraph-split
# the entire block, distinguishing bullet groups (`- ...`
# paragraphs) from prose paragraphs uniformly. Append a
# forbidden disclosure as a trailing prose paragraph and
# assert lint catches it.
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
"Output metadata must not claim audit-passed state.\n",
"Output metadata must not claim audit-passed state.\n"
"\nThe orchestrator runs codex audit afterward.\n",
)
self.assert_mutation_fails()
class CodexR5MutationTests(_MutationTestBase):
"""Codex R5 (Phase 6.7 review, after R4 fixes) closure: 1 P2
finding — C3 regex was not whitespace-tolerant inside sentences,
causing CI to fail on harmless Markdown soft-wraps that INV-1
correctly accepts."""
def test_r5_p2_compiler_softwrap_canonical_passes(self) -> None:
# Codex R5 P2: C3 regex used literal spaces inside sentences,
# so a soft-wrap between `run` and `codex/external` (the same
# benign Markdown reflow R2's INV-1 test verifies for synthesis)
# caused CI to fail on report_compiler. Phase 6.7 R5 closure
# rewrites every inter-token space inside the canonical regex
# to `\s+` so soft-wrap tolerance is uniform across INV-1 and
# the C3 Check regex. Mutation should leave lint passing.
_mutate(
self._repo_dir,
"deep-research/agents/report_compiler_agent.md",
"DO NOT claim to have run codex/external review.",
"DO NOT claim to have run\n codex/external review.",
)
rc, _stdout, stderr = _run_lint(self._repo_dir)
self.assertEqual(
rc,
0,
f"compiler soft-wrap canonical bullet should pass; stderr={stderr}",
)
class CodexR6MutationTests(_MutationTestBase):
"""Codex R6 (Phase 6.7 review, after R5 fixes) closure: 1 P2
finding — INV-1 only counted exact canonical matches and ignored
weakened near-duplicates that preserved an exact canonical bullet
while adding a confusing variant alongside."""
CANONICAL_BULLET = INV1MutationTests.CANONICAL_BULLET
def test_r6_p2_inv1_weakened_duplicate_alongside_canonical_fails(self) -> None:
# Codex R6 P2: an attacker keeps the exact canonical bullet
# AND adds a second, weakened variant
# (`- When feasible, DO NOT simulate ...`). Pre-fix INV-1 used
# `count(exact) == 1` which stayed at 1; lint passed and the
# block was left in a state where the agent reads two
# contradictory bullets. Phase 6.7 R6 closure introduces
# `_is_clause_1_like` that flags any bullet carrying canonical
# fragment markers (`do not simulate`, `do not claim to have
# run`, `audit-passed state`); each flagged bullet must equal
# the canonical text byte-for-byte or lint fails.
weakened = (
"- When feasible, DO NOT simulate any audit step. DO NOT "
"claim to have run codex/external review. Output metadata "
"must not claim audit-passed state."
)
_mutate(
self._repo_dir,
"deep-research/agents/synthesis_agent.md",
self.CANONICAL_BULLET,
self.CANONICAL_BULLET + "\n" + weakened,
)
self.assert_mutation_fails()
def test_r6_p2_inv1_tail_weakened_duplicate_fails(self) -> None:
# Defense in depth on R6: tail-weakener variant of the same
# attack (`... audit-passed state if feasible.`) — preserves
# canonical, adds a near-duplicate with a tail weakener.
weakened = (
"- DO NOT simulate any audit step. DO NOT claim to have "
"run codex/external review. Output metadata must not "
"claim audit-passed state if feasible."
)
_mutate(
self._repo_dir,
"deep-research/agents/research_architect_agent.md",
self.CANONICAL_BULLET,
self.CANONICAL_BULLET + "\n" + weakened,
)
self.assert_mutation_fails()
class CodexR7R9MutationTests(_MutationTestBase):
"""Codex R7 → R9 architectural rewind: R7 originally extended INV-3
with `_is_clause_1_like` to catch weakened variants outside the
manifest. R9 surfaced that this over-extends: spec §6.3 INV-3
detects "the canonical Clause 1 line found outside the manifest"
(the actual sentence), not Clause 1-like variants. v3.6.7 does not
lint variants outside the manifest. Test updated to reflect the
correct contract: a weakened variant in a non-manifest agent
passes lint."""
def test_r9_inv3_weakened_clause_1_in_non_manifest_passes(self) -> None:
# Codex R9 P2 architectural rewind: a weakened canonical
# variant outside the manifest is NOT a v3.6.7 INV-3 violation.
# Only the exact canonical sentence (whitespace-normalized
# equal to CANONICAL_CLAUSE_1_TEXT) widens the scope. Variants
# are a v3.6.8+ concern if/when L2 is reopened.
bibliography = self._repo_dir / "deep-research/agents/bibliography_agent.md"
original = bibliography.read_text(encoding="utf-8")
weakened = (
"\n\n- When feasible, DO NOT simulate any audit step. DO NOT "
"claim to have run codex/external review. Output metadata "
"must not claim audit-passed state.\n"
)
bibliography.write_text(original + weakened, encoding="utf-8")
rc, _stdout, stderr = _run_lint(self._repo_dir)
self.assertEqual(
rc,
0,
f"weakened Clause 1 variant outside manifest is not a "
f"v3.6.7 INV-3 violation per R9 architectural rewind; "
f"stderr={stderr}",
)
class CodexR8MutationTests(_MutationTestBase):
"""Codex R8 (Phase 6.7 review, after R7 fixes) closures: 2 P2
findings — `_is_clause_1_like` was over-broad (rejected legitimate
anti-fabrication guidance) AND INV-3 missed prose-form Clause 1
copies."""
CANONICAL_BULLET = INV1MutationTests.CANONICAL_BULLET
def test_r8_p2_generic_do_not_simulate_in_non_manifest_passes(self) -> None:
# Codex R8 P2 (a): the prior `_is_clause_1_like` flagged any
# bullet containing `do not simulate`, which is common
# anti-fabrication language. A legitimate non-manifest bullet
# like `- Do not simulate data or sources.` would
# false-positive INV-3 even though it has no audit-prohibition
# semantics. Phase 6.7 R8 closure narrows the heuristic to
# require an audit-specific fragment (`audit step`,
# `audit-passed state`, or `codex/external review`). Inject
# the generic bullet into bibliography_agent and assert lint
# passes (no false positive).
bibliography = self._repo_dir / "deep-research/agents/bibliography_agent.md"
original = bibliography.read_text(encoding="utf-8")
bibliography.write_text(
original + "\n\n- Do not simulate data or sources.\n",
encoding="utf-8",
)
rc, _stdout, stderr = _run_lint(self._repo_dir)
self.assertEqual(
rc,
0,
f"generic anti-fabrication bullet should not false-positive "
f"INV-3 after R8 audit-specific tightening; stderr={stderr}",
)
def test_r8_p2_inv3_canonical_as_prose_in_non_manifest_fails(self) -> None:
# Codex R8 P2 (b): the prior INV-3 only walked Markdown bullets.
# A non-manifest agent prompt that pasted the exact canonical
# sentence as a normal paragraph (not a `- ` bullet) silently
# passed lint, even though the v3.6.7 prohibition was now
# widened beyond the frozen manifest. Phase 6.7 R8 closure adds
# a paragraph-level scan that compares whitespace-normalized
# prose paragraphs against `CANONICAL_CLAUSE_1_TEXT` exactly.
# Inject the canonical sentence as a prose paragraph in
# bibliography_agent and assert lint fails.
bibliography = self._repo_dir / "deep-research/agents/bibliography_agent.md"
original = bibliography.read_text(encoding="utf-8")
canon_prose = (
"\n\nDO NOT simulate any audit step. DO NOT claim to have "
"run codex/external review. Output metadata must not claim "
"audit-passed state.\n"
)
bibliography.write_text(original + canon_prose, encoding="utf-8")
self.assert_mutation_fails()
class CodexR9MutationTests(_MutationTestBase):
"""Codex R9 (Phase 6.7 review, after R8 fixes) architectural
rewind: R7+R8 had over-extended INV-3 with the
Clause 1-like heuristic (false-positives on legitimate audit-step
mentions in non-manifest agents) AND under-extended prose scan
(missed canonical-after-heading-no-blank-line). R9 closes both."""
def test_r9_p2_inv3_audit_fragment_unrelated_bullet_passes(self) -> None:
# Codex R9 P2 (a): a non-manifest agent prompt may legitimately
# discuss audit steps in process-flow guidance, e.g.
# `- Review each audit step before finalizing.`. R8's
# heuristic flagged this as Clause 1-like even though the
# bullet has no prohibition semantics. R9 closure scopes INV-3
# to exact canonical sentence only.
bibliography = self._repo_dir / "deep-research/agents/bibliography_agent.md"
original = bibliography.read_text(encoding="utf-8")
bibliography.write_text(
original + "\n\n- Review each audit step before finalizing.\n",
encoding="utf-8",
)
rc, _stdout, stderr = _run_lint(self._repo_dir)
self.assertEqual(
rc,
0,
f"unrelated audit-step bullet should not trip INV-3 after "
f"R9 architectural rewind; stderr={stderr}",
)
def test_r9_p2_inv3_canonical_after_heading_no_blank_line_fails(self) -> None:
# Codex R9 P2 (b): when the canonical sentence is pasted as
# prose immediately after a Markdown heading with no blank
# line, R8's `re.split(r"\n\s*\n", text)` kept the heading and
# sentence in one paragraph, and the `startswith("## ")`
# filter skipped the whole thing. R9 closure switches to
# line-level heading-strip + whitespace-collapse so heading
# adjacency does not bypass the prose scan.
bibliography = self._repo_dir / "deep-research/agents/bibliography_agent.md"
original = bibliography.read_text(encoding="utf-8")
injection = (
"\n\n## PATTERN PROTECTION (v3.6.7)\n"
"DO NOT simulate any audit step. DO NOT claim to have run "
"codex/external review. Output metadata must not claim "
"audit-passed state.\n"
)
bibliography.write_text(original + injection, encoding="utf-8")
self.assert_mutation_fails()
if __name__ == "__main__":
unittest.main()