Files
Edward Cheng-I Wu d4523a9343 v3.9.2: Phase scope inflation hot-fix (#133) (#137)
* spec: v3.9.2 38-agent 4-bucket phase classification (Phase 0.2)

Phase 0.2 of #133 hot-fix. Per design doc v4 opus HIGH-1 (4-bucket model)
+ out-of-scope inflation risk column.

Bucket breakdown:
- A single-phase (get hard fence): 23 agents
- B multi-phase (NO fence, honest framing): 4 agents
- C phase-orthogonal (NO fence): 8 agents (incl. 2x socratic_mentor)
- D cross-phase/meta (NO fence): 4 agents

Audit-flag (within Bucket A, existing partial language): synthesis,
literature_strategist, draft_writer — Phase 1 sweep harmonizes.

Refs #133 (hot-fix) #134 (v3.10 conductor carry-over)

* feat: v3.9.2 Phase L1 — routing clarification gate (#133)

Phase L1 of #133 hot-fix. Adds:

L1.1 — Routing Discipline section in .claude/CLAUDE.md (before Routing
Rules 1-5). 3 routing classes (explicit / cross-phase materials clarify /
no-materials clarify) + [direct-mode] escape hatch Step 0 + anti-pattern
naming + v3.10 forward note (#134).

L1.2 — New shared/references/intent_clarification_protocol.md: trigger
condition table, pipeline phase reference, clarification message template,
[direct-mode] mechanism spec (byte-0 case-insensitive stripped before
pass-through), 5 worked examples, v3.10 carry-over notes.

L1.3 — One-line backpointer block added to all 4 SKILL.md files
(deep-research, academic-paper, academic-paper-reviewer, academic-pipeline).
Source-of-truth lives in CLAUDE.md + protocol doc; SKILL.md files only
reference (no rule duplication per opus codex review).

L1.4 — 8 behavioral smoke test fixtures in tests/fixtures/issue_133_routing/:
01 cross-phase abstract+lit (the #133 root case)
02 single-phase explicit lit-review
03 no-materials ambiguous
04 explicit /ars-* slash command
05 [direct-mode] honored at byte-0
06 [direct-mode] mid-message NOT honored
07 [Direct-Mode] case-insensitive
08 cross-phase draft+abstract+lit+reviews
Honestly framed as behavioral smoke tests, not deterministic unit tests.
Acceptance: 100% pass on Opus 4.7; ≥75% on Sonnet 4.6 + GPT-5.5.

Phase 5 lint compat audit confirmed:
- check_spec_consistency.py PASSES on this diff
- check_version_consistency.py fails on pre-existing v3.9.1 ship gap
  (Suite version '3.9.0' vs CHANGELOG '3.9.1') — NOT introduced by L1;
  will be fixed in v3.9.2 atomic version bump at ship time
- pytest 1463 passed + 3 skipped + 111 subtests (no regression vs v3.9.1)

SKILL.md frontmatter version + last_updated intentionally NOT bumped in
this commit — deferred to ship-time atomic bump per Phase 5 audit decision.

Refs #133 (hot-fix)
Refs #134 (v3.10 active conductor — carries lookup, multi-phase schema,
provenance, orchestrator intake)

* feat: v3.9.2 Phase 1 — prompt hard fence on 22 Bucket A agents (#133)

Phase 1 of #133 hot-fix. Adds 'Phase Boundary (v3.9.2)' block to all
22 single-phase agents per docs/design/2026-05-18-ars-v3.9.2-agent-phase-
classification.md.

Coverage (22 of 38 total agents):
- deep-research (9): research_question (P1), research_architect (P1),
  bibliography (P2), source_verification (P2), synthesis (P3, coexists
  w/ v3.6.7 PATTERN PROTECTION), editor_in_chief (P5), ethics_review (P5),
  risk_of_bias (SR-P2), meta_analysis (SR-P3)
- academic-paper (7): literature_strategist (P1, coexists w/ v3.6.5 corpus
  protocol), structure_architect (P2), draft_writer (P4/P6 per invocation,
  coexists w/ v3.6.6 generator-evaluator contract), citation_compliance
  (P5a), abstract_bilingual (P5b), peer_reviewer (P6 coexists w/ v3.6.6),
  formatter (P7 coexists w/ v3.7.1 hard-gate)
- academic-paper-reviewer (6): eic, methodology_reviewer, domain_reviewer,
  perspective_reviewer, devils_advocate_reviewer, editorial_synthesizer
  (all coexist w/ v3.6.2 Sprint Contract Protocol)

NOT fenced (16 of 38, per classification doc):
- 4 Bucket B (multi-phase): devils_advocate (DR P1/3/5), report_compiler
  (DR P4/6), argument_builder (AP P3/Plan), visualization (AP P4/7)
- 8 Bucket C (phase-orthogonal): socratic_mentor x2, monitoring,
  revision_coach, integrity_verification, collaboration_depth,
  claim_ref_alignment_audit, compliance
- 4 Bucket D (cross-phase meta): intake, pipeline_orchestrator,
  state_tracker, field_analyst

Per honest framing (HIGH-2 from v3 review): multi-phase agents receive NO
fence. Placebo prose would create false-enforcement illusion. v3.10
envelope is the deterministic fix.

Each fence block is customized per agent:
- Explicit phase number and deliverable description
- MUST-NOT list of cross-phase writes and downstream deliverable types
- MAY-READ list for upstream context (e.g. Phase 5 reviewers explicitly
  granted READ on Phase 1-4 since review requires upstream context)
- Explicit coexistence notes where the agent already carries other
  protocol blocks (v3.6.2 / v3.6.5 / v3.6.6 / v3.6.7 / v3.7.1)
- Forward note to v3.10 conductor (#134) for deterministic enforcement

Also fixes classification doc typo: A=22 (not 23 as v0 draft said).
9+7+6=22, miscount surfaced during sweep.

Verification:
- check_spec_consistency.py PASS
- check_v3_6_7_pattern_protection.py PASS (existing PATTERN PROTECTION
  block on synthesis_agent unaffected by Phase Boundary block addition)
- check_v3_6_8_pattern_protection.py PASS
- pytest 1463 passed + 3 skipped + 111 subtests (0 regression vs v3.9.1)
- 22 files with 'Phase Boundary (v3.9.2)' marker confirmed via grep
- 16 files without marker (matches Bucket B/C/D classification)

Refs #133 (hot-fix)
Refs #134 (v3.10 active conductor — Bucket B/C/D enforcement)

* feat: v3.9.2 Phase 3 + Phase 4 — SKILL.md contract + advisory verifier (#133)

Phase 3 (4 files modified):
Added 'Phase-by-phase Invocation Contract (v3.9.2)' section to all 4
SKILL.md files. Each describes Mode A (orchestrator-driven, default) vs
Mode B (phase-by-phase across sessions), enumerates Bucket A single-phase
agents specific to that skill, notes coexistence with skill-specific
protocols (v3.6.6 generator-evaluator for academic-paper, v3.6.2 Sprint
Contract for reviewer), and points to v3.10 active conductor (#134) as
the deterministic enforcement layer.

Phase 4 (2 files added):
scripts/check_pipeline_integrity.py — advisory verifier per design v4 §3
Phase 4. Lite version (v3.9.2): no provenance, heuristic-based. Two rules:

1. STRUCTURAL — phase5_missing_independent_reviewer
   Detects #133 pattern: phase5_*/ dir exists but missing files matching
   independent reviewer name conventions (DA / EIC / Ethics / panel
   reviewers). High-value structural detection; catches the exact case
   reported in #133.

2. HEURISTIC — adjacent_phase_same_window (--strict only, default OFF)
   Adjacent phase{N} + phase{N+1} files within 5-min mtime window.
   Acknowledged FP-prone (legitimate fast orchestrator runs trigger);
   opt-in via --strict flag, default disabled.

Cross-platform Python (no provenance dependency), user-invokable post-hoc.
Advisory output (exit 0 on findings) — explicitly NOT a CI gate.
JSON output via --json flag.

scripts/test_check_pipeline_integrity.py — 12 unit tests covering:
- empty workdir / no phase dirs (PASS)
- #133 inflation case (STRUCTURAL flagged)
- legitimate Phase 5 with 3 reviewers (PASS)
- alternative attribution (DA + EIC + 1 panel reviewer) PASSES
- empty phase5_*/ flags phase5_empty
- editorial_synthesizer alone misses DA/ethics categories
- --strict flag enables/disables heuristic
- JSON output format
- invalid workdir exits 1
- non-phase dirs ignored

Verification:
- pytest: 1475 passed (was 1463 + 12 new) + 3 skipped + 111 subtests
- check_spec_consistency.py PASS
- 0 regression on v3.9.1 baseline

v3.10 carry-over: verifier becomes deterministic via task envelope
author provenance (#134). The 5-minute window heuristic disappears.

Refs #133 (hot-fix)
Refs #134 (v3.10 conductor — provenance + deterministic verification)

* fix: v3.9.2 Phase 6 review absorption — [direct-mode] step ref + verifier polish (#133)

Mid-impl checkpoint dual-track review (codex broken corner per memory; 49
files / 1529 lines puts diff firmly in broken zone — inline opus subagent
was the substantive reviewer):

LOW-1 fix (LOGICAL BUG — promoted to ship-blocking):
.claude/CLAUDE.md Step 0 said 'skip directly to Step 3' but Step 3 is
'No materials, ambiguous → clarify'. Should be Step 1 (Explicit clear
intent → Route directly). Reading literally, [direct-mode] users would
have been routed to clarification — defeating the escape hatch entirely.
Fixed wording, also clarified Step 1→3 fallback chain when stripped
message has no clear skill named.

LOW-6 fix (consistency):
shared/references/intent_clarification_protocol.md [direct-mode] spec
was internally inconsistent ('byte-0' vs 'whitespace stripped on parse').
Reframed as 'first non-whitespace token, leading whitespace stripped on
parse'. Added explicit bracket-form restrictions (only [direct-mode],
not (direct-mode) or [direct mode]). Aligned fallback prose with
CLAUDE.md Step 0 wording.

MED-1 absorption (verifier filename convention):
Added NORMATIVE FILENAME CONVENTION block to check_pipeline_integrity.py
docstring. Verifier is filename-regex based; orchestrator runs with terse
filenames (e.g. 'round1_ethics.md' without _review suffix) WILL produce
false positives. Documented trade-off and v3.10 envelope provenance as
the fix.

LOW-2 fix (dotfile noise):
check_pipeline_integrity.py now filters .DS_Store / .gitkeep / Thumbs.db
from phase5 file enumeration. Hidden files are OS/git noise, not
authored content.

LOW-3 test coverage additions (+4 tests):
- test_dotfiles_in_phase5_ignored
- test_multiple_phase5_dirs_each_independently_checked (multi-round
  review fixture)
- test_unicode_filenames_with_canonical_stem_match (Chinese suffix on
  canonical ASCII stem still matches)
- test_nested_files_in_phase5_count (rglob recursion confirmed)

Verification:
- 16 verifier tests PASS (was 12; +4)
- pytest 1479 passed (was 1475; +4) + 3 skipped + 111 subtests
- 0 regression vs v3.9.1 baseline

Carry to ship:
- LOW-4 (SKILL.md contract redundancy) — opus confirmed NOT redundant,
  no fix needed
- LOW-5 (synthesis_agent.md pre-existing v3.7.1 markdown spacing issue)
  — pre-existing, not v3.9.2 scope
- LOW-7 (_test_helpers.py) — pre-existing helper, already committed

Refs #133 (hot-fix)
Refs #134 (v3.10 conductor — provenance replaces filename matching)

* ship: v3.9.2 — Phase 5 lint + Phase 7 audit + Phase 8 version bump (#133)

Phase 5 (new):
- scripts/check_v3_9_2_phase_boundary.py: enforces 22 Bucket A agents have
  '## Phase Boundary (v3.9.2)' block, 16 Bucket B/C/D agents excluded,
  each Bucket A block contains 4 load-bearing phrases (Phase Boundary
  v3.9.2, MUST NOT, MAY READ, Enforcement v3.9.2). Falsifiability
  discipline: phrases scoped to H2 block, not file-wide.
- scripts/test_check_v3_9_2_phase_boundary.py: 3 tests (baseline pass,
  module invariants, REQUIRED_PHRASES constant).
- .github/workflows/spec-consistency.yml: wired both v3.9.2 lints to CI.

Phase 7 (audit pass):
- ~/.claude/personal-boundary/check_boundary.py --root . PASSED
  (648 files / 0 violations).
- Per feedback_ars_public_repo_boundary.md grep: only matches are
  public-known author (Cheng-I Wu) and GitHub handle (Imbad0202),
  both explicitly allowlisted in deny_list.yaml.
- No hei-platform / HEEACT / 評鑑中心 / Springer / school-name leakage.

Phase 8 (atomic version bump, ship):
- .claude-plugin/plugin.json: 3.8.2 → 3.9.2 (also catches v3.9.0 + v3.9.1
  deferrals); description updated for 38-agent ensemble + v3.9.2 fence.
- .claude/CLAUDE.md Skills Overview table:
  deep-research 2.9.3→2.9.4, academic-paper 3.1.1→3.1.2,
  academic-paper-reviewer 1.9.0→1.9.1, academic-pipeline 3.9.0→3.9.2.
- .claude/CLAUDE.md Suite version: 3.9.0 → 3.9.2 (also fixes pre-existing
  v3.9.1 latent bug — that ship missed bumping Suite version).
- 4 SKILL.md frontmatter version + last_updated bumped to match.
- academic-paper-reviewer/SKILL.md body Version Info table synced (lint
  enforces frontmatter + body alignment via check_reviewer_version_block).
- MODE_REGISTRY.md 'Last updated' bumped to v3.9.2 (2026-05-18).
- README.md + README.zh-TW.md: badge bumped + v3.9.2 + v3.9.1 changelog
  entries added in both EN + zh-TW.
- CHANGELOG.md [3.9.2] entry: full release notes covering Phase L1 / 1 /
  3 / 4 / 5 / 6 (review absorption) + carry-over to v3.10 (#134) +
  migration notes for in-flight users.
- docs/design/2026-05-18-ars-v3.9.2-phase-boundary-spec.md: copied from
  .local-plans/ as final spec doc.
- scripts/check_spec_consistency.py: hardcoded version refs updated to
  v3.9.2 (also expects new v3.9.1 + v3.9.2 README sections).

Verification:
- check_spec_consistency.py PASS
- check_version_consistency.py PASS (was failing on v3.9.1 latent bug)
- check_v3_9_2_phase_boundary.py PASS (22 + 16 = 38 agents validated)
- check_v3_6_7_pattern_protection.py PASS (existing PROTECTION blocks
  unaffected by new Phase Boundary additions)
- pytest 1482 passed + 3 skipped + 111 subtests passed
- 0 regression vs v3.9.1 baseline (was 1463; +19 = 1482, all new tests
  from this PR)
- personal-boundary PASSED (648 files, 0 violations)

Closes #133

v3.10 conductor architecture (#134) carries:
- PreToolUse hook (Phase 0.1 verified CC payload has agent_type field)
- Multi-phase ars_phase_writes/reads envelope schema
- Deterministic verifier with author provenance
- Orchestrator cross-phase intake capability

Refs #134
2026-05-18 14:58:35 +08:00

341 lines
13 KiB
Python

#!/usr/bin/env python3
"""Advisory verifier for ARS pipeline phase scope (v3.9.2).
Spec: docs/design/2026-05-18-ars-v3.9.2-phase-boundary-spec.md Phase 4
Issue: #133 (phase scope inflation hot-fix)
Forward note: v3.10 active conductor (#134) replaces this with deterministic
provenance-based verification via task envelope; this v3.9.2 advisory is
heuristic and FP-prone by design.
Usage:
python scripts/check_pipeline_integrity.py [workdir]
python scripts/check_pipeline_integrity.py --json [workdir]
Default workdir: current directory.
This script SCANS a working directory for `phaseN_*/` subdirectories (where
N is 1-6 per the ARS pipeline phase convention) and flags advisory signals
that suggest #133-class scope inflation. Output is ADVISORY ONLY — it does
NOT block any workflow. User reviews findings and decides.
Detection rules (v3.9.2 lite):
1. **Phase 5 missing independent reviewer attribution** (HIGH-VALUE structural rule)
When `phase5_*/` exists but contains no separate files matching the
independent-reviewer naming convention (devils_advocate / editor_in_chief /
ethics_review / eic / methodology / domain / perspective / editorial_synthesizer),
flag INTEGRITY-ADVISORY. This catches the exact #133 reported pattern:
a single agent producing phase5_*/ output that looks like a review but was
never actually crosschecked by independent reviewer agents.
**NORMATIVE FILENAME CONVENTION:** This rule is filename-based regex matching
over a closed set of agent stem names. For the verifier to recognize a
reviewer report, the filename MUST contain one of the canonical agent stems:
- `devils_advocate` (or `devil-s_advocate` / `devils-advocate`)
- `editor_in_chief` or `eic`
- `ethics_review` (the full stem; bare `ethics` does NOT match)
- `methodology_review` / `domain_review` / `perspective_review`
- `editorial_synth` (matches editorial_synthesizer family)
Orchestrator runs that emit terse filenames like `round1_ethics.md` (without
`_review` suffix) WILL produce a false-positive STRUCTURAL finding here.
Resolution options: (a) rename file to match convention, or (b) ignore the
advisory finding (verifier exits 0; advisory does not block). The verifier
trades filename-convention discipline for zero-dependency cross-platform
simplicity in v3.9.2. v3.10 conductor (#134) replaces filename matching
with task envelope author provenance.
2. **Multi-phase same-call generation heuristic** (LOWER-VALUE, FP-prone)
When files in phaseN_*/ and phase{N+1}_*/ share creation timestamps within
a configurable window (default 5 minutes), flag as POSSIBLE same-call
inflation. Acknowledged FP risk: legitimate fast orchestrator runs trigger
this. Use --strict to enable, default OFF.
Exit codes:
0 No findings, or only --strict findings without --strict flag
0 Findings produced (advisory output, NOT a hard gate); printed to stdout
1 Script error (invalid workdir, IO failure)
The verifier intentionally fails open (exit 0) on findings — it is advisory,
not a CI gate. CI should not rely on this for v3.9.2.
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from dataclasses import dataclass, field
from pathlib import Path
from typing import Iterable
PHASE_DIR_RE = re.compile(r"^phase([1-6])(?:_.*)?$")
# Phase 5 reviewer file-name heuristics. A phase5_*/ directory should contain
# files whose names match at least one of these per independent crosscheck
# (Devil's Advocate, Editor-in-Chief, Ethics Review). Missing all three →
# advisory flag per #133 root pattern.
PHASE5_REVIEWER_PATTERNS = {
"devils_advocate": re.compile(r"devils?[_-]?advocate", re.IGNORECASE),
"editor_in_chief": re.compile(r"(?:editor[_-]?in[_-]?chief|^eic|[_-]eic[_-])", re.IGNORECASE),
"ethics_review": re.compile(r"ethics?[_-]?review", re.IGNORECASE),
"methodology_reviewer": re.compile(r"methodology[_-]?review", re.IGNORECASE),
"domain_reviewer": re.compile(r"domain[_-]?review", re.IGNORECASE),
"perspective_reviewer": re.compile(r"perspective[_-]?review", re.IGNORECASE),
"editorial_synthesizer": re.compile(r"editorial[_-]?synth", re.IGNORECASE),
}
# Required minimum reviewer attributions for a valid Phase 5 output per
# academic-pipeline Stage 3 review protocol. At least one from each category
# must appear in phase5_*/ filenames.
PHASE5_REQUIRED_CATEGORIES = [
("devil's advocate", ["devils_advocate"]),
("editorial/EIC", ["editor_in_chief", "editorial_synthesizer"]),
("ethics or panel reviewer", ["ethics_review", "methodology_reviewer", "domain_reviewer", "perspective_reviewer"]),
]
DEFAULT_SAME_CALL_WINDOW_SECONDS = 300 # 5 minutes
@dataclass
class Finding:
rule: str
severity: str # ADVISORY (default), STRUCTURAL (high-value), HEURISTIC (FP-prone)
phase: int | None
path: str
message: str
@dataclass
class Report:
workdir: str
phase_dirs: dict[int, list[str]] = field(default_factory=dict) # phase → list of dir paths
findings: list[Finding] = field(default_factory=list)
def scan_workdir(workdir: Path) -> Report:
"""Locate all phase{1-6}_*/ subdirectories under workdir."""
report = Report(workdir=str(workdir))
if not workdir.is_dir():
return report
for entry in workdir.iterdir():
if not entry.is_dir():
continue
match = PHASE_DIR_RE.match(entry.name)
if not match:
continue
phase = int(match.group(1))
report.phase_dirs.setdefault(phase, []).append(str(entry))
return report
def check_phase5_attribution(report: Report) -> None:
"""Rule 1 — STRUCTURAL: phase5_*/ must contain independent reviewer files."""
phase5_dirs = report.phase_dirs.get(5, [])
if not phase5_dirs:
return
for dir_path in phase5_dirs:
path = Path(dir_path)
try:
# Skip hidden files (.DS_Store, Thumbs.db, .gitkeep, etc.) — they're
# never reviewer reports and would clutter advisory output noise.
files = [
p.name for p in path.rglob("*")
if p.is_file() and not p.name.startswith(".")
]
except OSError as exc:
report.findings.append(Finding(
rule="phase5_attribution_io_error",
severity="ADVISORY",
phase=5,
path=dir_path,
message=f"Could not read phase5 directory: {exc}",
))
continue
if not files:
report.findings.append(Finding(
rule="phase5_empty",
severity="ADVISORY",
phase=5,
path=dir_path,
message="phase5_*/ directory is empty — Phase 5 should produce review reports",
))
continue
# Tally which reviewer categories are present
matched_agents: set[str] = set()
for fname in files:
for agent, pattern in PHASE5_REVIEWER_PATTERNS.items():
if pattern.search(fname):
matched_agents.add(agent)
missing_categories: list[str] = []
for category_label, agent_list in PHASE5_REQUIRED_CATEGORIES:
if not any(agent in matched_agents for agent in agent_list):
missing_categories.append(category_label)
if missing_categories:
report.findings.append(Finding(
rule="phase5_missing_independent_reviewer",
severity="STRUCTURAL",
phase=5,
path=dir_path,
message=(
f"phase5_*/ missing independent reviewer attribution for categories: "
f"{', '.join(missing_categories)}. "
f"Files found: {files[:5]}{'...' if len(files) > 5 else ''}. "
f"#133 pattern: Phase 5 deliverable was likely produced by a single agent "
f"that inflated past its scope, skipping mandatory independent crosschecks "
f"(DA / EIC / Ethics). Re-run via orchestrator-driven Mode A with "
f"`/ars-full` or invoke each reviewer agent separately."
),
))
def _file_mtime(path: Path) -> float | None:
try:
return path.stat().st_mtime
except OSError:
return None
def check_same_call_heuristic(report: Report, window_seconds: int) -> None:
"""Rule 2 — HEURISTIC (--strict only): adjacent-phase same-window timestamps."""
phases_present = sorted(report.phase_dirs.keys())
if len(phases_present) < 2:
return
# Collect (phase, file_path, mtime) tuples
file_records: dict[int, list[tuple[Path, float]]] = {}
for phase, dirs in report.phase_dirs.items():
records: list[tuple[Path, float]] = []
for dir_path in dirs:
path = Path(dir_path)
for file_path in path.rglob("*"):
if not file_path.is_file():
continue
mtime = _file_mtime(file_path)
if mtime is None:
continue
records.append((file_path, mtime))
if records:
file_records[phase] = records
# For each adjacent (N, N+1) pair, find files with mtime within window
for phase in phases_present:
next_phase = phase + 1
if next_phase not in file_records:
continue
for file_a, mtime_a in file_records.get(phase, []):
for file_b, mtime_b in file_records[next_phase]:
delta = abs(mtime_a - mtime_b)
if delta <= window_seconds:
report.findings.append(Finding(
rule="adjacent_phase_same_window",
severity="HEURISTIC",
phase=phase,
path=str(file_a),
message=(
f"phase{phase} file {file_a.name} and phase{next_phase} file "
f"{file_b.name} share mtime within {int(delta)}s "
f"(window={window_seconds}s). POSSIBLE same-call inflation. "
f"Note: legitimate fast orchestrator runs also trigger this "
f"heuristic — verify against orchestrator state ledger before "
f"treating as #133-class violation."
),
))
def format_text(report: Report) -> str:
lines = [f"ARS pipeline integrity check (v3.9.2 advisory)"]
lines.append(f"Workdir: {report.workdir}")
lines.append(f"Phase dirs found: {dict(sorted(report.phase_dirs.items()))}")
lines.append("")
if not report.findings:
lines.append("No advisory findings.")
return "\n".join(lines)
lines.append(f"Findings ({len(report.findings)}):")
for i, finding in enumerate(report.findings, 1):
lines.append("")
lines.append(f" [{i}] {finding.severity}{finding.rule}")
if finding.phase is not None:
lines.append(f" Phase: {finding.phase}")
lines.append(f" Path: {finding.path}")
lines.append(f" {finding.message}")
lines.append("")
lines.append("Reminder: this output is ADVISORY. Findings do NOT block any workflow.")
lines.append("See docs/design/2026-05-18-ars-v3.9.2-phase-boundary-spec.md for rationale.")
return "\n".join(lines)
def format_json(report: Report) -> str:
payload = {
"workdir": report.workdir,
"phase_dirs": {str(k): v for k, v in sorted(report.phase_dirs.items())},
"findings": [
{
"rule": f.rule,
"severity": f.severity,
"phase": f.phase,
"path": f.path,
"message": f.message,
}
for f in report.findings
],
}
return json.dumps(payload, indent=2, ensure_ascii=False)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="Advisory check for ARS pipeline phase scope inflation (v3.9.2)"
)
parser.add_argument(
"workdir",
nargs="?",
default=".",
help="Working directory to scan (default: current directory)",
)
parser.add_argument(
"--strict",
action="store_true",
help="Enable heuristic same-call timestamp check (FP-prone, default OFF)",
)
parser.add_argument(
"--window-seconds",
type=int,
default=DEFAULT_SAME_CALL_WINDOW_SECONDS,
help=f"Same-call window for --strict heuristic (default: {DEFAULT_SAME_CALL_WINDOW_SECONDS}s)",
)
parser.add_argument(
"--json",
action="store_true",
help="Output JSON instead of human-readable text",
)
args = parser.parse_args(argv)
workdir = Path(args.workdir).resolve()
if not workdir.is_dir():
print(f"ERROR: workdir not found or not a directory: {workdir}", file=sys.stderr)
return 1
report = scan_workdir(workdir)
check_phase5_attribution(report)
if args.strict:
check_same_call_heuristic(report, args.window_seconds)
if args.json:
print(format_json(report))
else:
print(format_text(report))
return 0
if __name__ == "__main__":
sys.exit(main())