mirror of
https://github.com/Imbad0202/academic-research-skills.git
synced 2026-09-14 13:51:17 +08:00
d4523a9343
* spec: v3.9.2 38-agent 4-bucket phase classification (Phase 0.2) Phase 0.2 of #133 hot-fix. Per design doc v4 opus HIGH-1 (4-bucket model) + out-of-scope inflation risk column. Bucket breakdown: - A single-phase (get hard fence): 23 agents - B multi-phase (NO fence, honest framing): 4 agents - C phase-orthogonal (NO fence): 8 agents (incl. 2x socratic_mentor) - D cross-phase/meta (NO fence): 4 agents Audit-flag (within Bucket A, existing partial language): synthesis, literature_strategist, draft_writer — Phase 1 sweep harmonizes. Refs #133 (hot-fix) #134 (v3.10 conductor carry-over) * feat: v3.9.2 Phase L1 — routing clarification gate (#133) Phase L1 of #133 hot-fix. Adds: L1.1 — Routing Discipline section in .claude/CLAUDE.md (before Routing Rules 1-5). 3 routing classes (explicit / cross-phase materials clarify / no-materials clarify) + [direct-mode] escape hatch Step 0 + anti-pattern naming + v3.10 forward note (#134). L1.2 — New shared/references/intent_clarification_protocol.md: trigger condition table, pipeline phase reference, clarification message template, [direct-mode] mechanism spec (byte-0 case-insensitive stripped before pass-through), 5 worked examples, v3.10 carry-over notes. L1.3 — One-line backpointer block added to all 4 SKILL.md files (deep-research, academic-paper, academic-paper-reviewer, academic-pipeline). Source-of-truth lives in CLAUDE.md + protocol doc; SKILL.md files only reference (no rule duplication per opus codex review). L1.4 — 8 behavioral smoke test fixtures in tests/fixtures/issue_133_routing/: 01 cross-phase abstract+lit (the #133 root case) 02 single-phase explicit lit-review 03 no-materials ambiguous 04 explicit /ars-* slash command 05 [direct-mode] honored at byte-0 06 [direct-mode] mid-message NOT honored 07 [Direct-Mode] case-insensitive 08 cross-phase draft+abstract+lit+reviews Honestly framed as behavioral smoke tests, not deterministic unit tests. Acceptance: 100% pass on Opus 4.7; ≥75% on Sonnet 4.6 + GPT-5.5. Phase 5 lint compat audit confirmed: - check_spec_consistency.py PASSES on this diff - check_version_consistency.py fails on pre-existing v3.9.1 ship gap (Suite version '3.9.0' vs CHANGELOG '3.9.1') — NOT introduced by L1; will be fixed in v3.9.2 atomic version bump at ship time - pytest 1463 passed + 3 skipped + 111 subtests (no regression vs v3.9.1) SKILL.md frontmatter version + last_updated intentionally NOT bumped in this commit — deferred to ship-time atomic bump per Phase 5 audit decision. Refs #133 (hot-fix) Refs #134 (v3.10 active conductor — carries lookup, multi-phase schema, provenance, orchestrator intake) * feat: v3.9.2 Phase 1 — prompt hard fence on 22 Bucket A agents (#133) Phase 1 of #133 hot-fix. Adds 'Phase Boundary (v3.9.2)' block to all 22 single-phase agents per docs/design/2026-05-18-ars-v3.9.2-agent-phase- classification.md. Coverage (22 of 38 total agents): - deep-research (9): research_question (P1), research_architect (P1), bibliography (P2), source_verification (P2), synthesis (P3, coexists w/ v3.6.7 PATTERN PROTECTION), editor_in_chief (P5), ethics_review (P5), risk_of_bias (SR-P2), meta_analysis (SR-P3) - academic-paper (7): literature_strategist (P1, coexists w/ v3.6.5 corpus protocol), structure_architect (P2), draft_writer (P4/P6 per invocation, coexists w/ v3.6.6 generator-evaluator contract), citation_compliance (P5a), abstract_bilingual (P5b), peer_reviewer (P6 coexists w/ v3.6.6), formatter (P7 coexists w/ v3.7.1 hard-gate) - academic-paper-reviewer (6): eic, methodology_reviewer, domain_reviewer, perspective_reviewer, devils_advocate_reviewer, editorial_synthesizer (all coexist w/ v3.6.2 Sprint Contract Protocol) NOT fenced (16 of 38, per classification doc): - 4 Bucket B (multi-phase): devils_advocate (DR P1/3/5), report_compiler (DR P4/6), argument_builder (AP P3/Plan), visualization (AP P4/7) - 8 Bucket C (phase-orthogonal): socratic_mentor x2, monitoring, revision_coach, integrity_verification, collaboration_depth, claim_ref_alignment_audit, compliance - 4 Bucket D (cross-phase meta): intake, pipeline_orchestrator, state_tracker, field_analyst Per honest framing (HIGH-2 from v3 review): multi-phase agents receive NO fence. Placebo prose would create false-enforcement illusion. v3.10 envelope is the deterministic fix. Each fence block is customized per agent: - Explicit phase number and deliverable description - MUST-NOT list of cross-phase writes and downstream deliverable types - MAY-READ list for upstream context (e.g. Phase 5 reviewers explicitly granted READ on Phase 1-4 since review requires upstream context) - Explicit coexistence notes where the agent already carries other protocol blocks (v3.6.2 / v3.6.5 / v3.6.6 / v3.6.7 / v3.7.1) - Forward note to v3.10 conductor (#134) for deterministic enforcement Also fixes classification doc typo: A=22 (not 23 as v0 draft said). 9+7+6=22, miscount surfaced during sweep. Verification: - check_spec_consistency.py PASS - check_v3_6_7_pattern_protection.py PASS (existing PATTERN PROTECTION block on synthesis_agent unaffected by Phase Boundary block addition) - check_v3_6_8_pattern_protection.py PASS - pytest 1463 passed + 3 skipped + 111 subtests (0 regression vs v3.9.1) - 22 files with 'Phase Boundary (v3.9.2)' marker confirmed via grep - 16 files without marker (matches Bucket B/C/D classification) Refs #133 (hot-fix) Refs #134 (v3.10 active conductor — Bucket B/C/D enforcement) * feat: v3.9.2 Phase 3 + Phase 4 — SKILL.md contract + advisory verifier (#133) Phase 3 (4 files modified): Added 'Phase-by-phase Invocation Contract (v3.9.2)' section to all 4 SKILL.md files. Each describes Mode A (orchestrator-driven, default) vs Mode B (phase-by-phase across sessions), enumerates Bucket A single-phase agents specific to that skill, notes coexistence with skill-specific protocols (v3.6.6 generator-evaluator for academic-paper, v3.6.2 Sprint Contract for reviewer), and points to v3.10 active conductor (#134) as the deterministic enforcement layer. Phase 4 (2 files added): scripts/check_pipeline_integrity.py — advisory verifier per design v4 §3 Phase 4. Lite version (v3.9.2): no provenance, heuristic-based. Two rules: 1. STRUCTURAL — phase5_missing_independent_reviewer Detects #133 pattern: phase5_*/ dir exists but missing files matching independent reviewer name conventions (DA / EIC / Ethics / panel reviewers). High-value structural detection; catches the exact case reported in #133. 2. HEURISTIC — adjacent_phase_same_window (--strict only, default OFF) Adjacent phase{N} + phase{N+1} files within 5-min mtime window. Acknowledged FP-prone (legitimate fast orchestrator runs trigger); opt-in via --strict flag, default disabled. Cross-platform Python (no provenance dependency), user-invokable post-hoc. Advisory output (exit 0 on findings) — explicitly NOT a CI gate. JSON output via --json flag. scripts/test_check_pipeline_integrity.py — 12 unit tests covering: - empty workdir / no phase dirs (PASS) - #133 inflation case (STRUCTURAL flagged) - legitimate Phase 5 with 3 reviewers (PASS) - alternative attribution (DA + EIC + 1 panel reviewer) PASSES - empty phase5_*/ flags phase5_empty - editorial_synthesizer alone misses DA/ethics categories - --strict flag enables/disables heuristic - JSON output format - invalid workdir exits 1 - non-phase dirs ignored Verification: - pytest: 1475 passed (was 1463 + 12 new) + 3 skipped + 111 subtests - check_spec_consistency.py PASS - 0 regression on v3.9.1 baseline v3.10 carry-over: verifier becomes deterministic via task envelope author provenance (#134). The 5-minute window heuristic disappears. Refs #133 (hot-fix) Refs #134 (v3.10 conductor — provenance + deterministic verification) * fix: v3.9.2 Phase 6 review absorption — [direct-mode] step ref + verifier polish (#133) Mid-impl checkpoint dual-track review (codex broken corner per memory; 49 files / 1529 lines puts diff firmly in broken zone — inline opus subagent was the substantive reviewer): LOW-1 fix (LOGICAL BUG — promoted to ship-blocking): .claude/CLAUDE.md Step 0 said 'skip directly to Step 3' but Step 3 is 'No materials, ambiguous → clarify'. Should be Step 1 (Explicit clear intent → Route directly). Reading literally, [direct-mode] users would have been routed to clarification — defeating the escape hatch entirely. Fixed wording, also clarified Step 1→3 fallback chain when stripped message has no clear skill named. LOW-6 fix (consistency): shared/references/intent_clarification_protocol.md [direct-mode] spec was internally inconsistent ('byte-0' vs 'whitespace stripped on parse'). Reframed as 'first non-whitespace token, leading whitespace stripped on parse'. Added explicit bracket-form restrictions (only [direct-mode], not (direct-mode) or [direct mode]). Aligned fallback prose with CLAUDE.md Step 0 wording. MED-1 absorption (verifier filename convention): Added NORMATIVE FILENAME CONVENTION block to check_pipeline_integrity.py docstring. Verifier is filename-regex based; orchestrator runs with terse filenames (e.g. 'round1_ethics.md' without _review suffix) WILL produce false positives. Documented trade-off and v3.10 envelope provenance as the fix. LOW-2 fix (dotfile noise): check_pipeline_integrity.py now filters .DS_Store / .gitkeep / Thumbs.db from phase5 file enumeration. Hidden files are OS/git noise, not authored content. LOW-3 test coverage additions (+4 tests): - test_dotfiles_in_phase5_ignored - test_multiple_phase5_dirs_each_independently_checked (multi-round review fixture) - test_unicode_filenames_with_canonical_stem_match (Chinese suffix on canonical ASCII stem still matches) - test_nested_files_in_phase5_count (rglob recursion confirmed) Verification: - 16 verifier tests PASS (was 12; +4) - pytest 1479 passed (was 1475; +4) + 3 skipped + 111 subtests - 0 regression vs v3.9.1 baseline Carry to ship: - LOW-4 (SKILL.md contract redundancy) — opus confirmed NOT redundant, no fix needed - LOW-5 (synthesis_agent.md pre-existing v3.7.1 markdown spacing issue) — pre-existing, not v3.9.2 scope - LOW-7 (_test_helpers.py) — pre-existing helper, already committed Refs #133 (hot-fix) Refs #134 (v3.10 conductor — provenance replaces filename matching) * ship: v3.9.2 — Phase 5 lint + Phase 7 audit + Phase 8 version bump (#133) Phase 5 (new): - scripts/check_v3_9_2_phase_boundary.py: enforces 22 Bucket A agents have '## Phase Boundary (v3.9.2)' block, 16 Bucket B/C/D agents excluded, each Bucket A block contains 4 load-bearing phrases (Phase Boundary v3.9.2, MUST NOT, MAY READ, Enforcement v3.9.2). Falsifiability discipline: phrases scoped to H2 block, not file-wide. - scripts/test_check_v3_9_2_phase_boundary.py: 3 tests (baseline pass, module invariants, REQUIRED_PHRASES constant). - .github/workflows/spec-consistency.yml: wired both v3.9.2 lints to CI. Phase 7 (audit pass): - ~/.claude/personal-boundary/check_boundary.py --root . PASSED (648 files / 0 violations). - Per feedback_ars_public_repo_boundary.md grep: only matches are public-known author (Cheng-I Wu) and GitHub handle (Imbad0202), both explicitly allowlisted in deny_list.yaml. - No hei-platform / HEEACT / 評鑑中心 / Springer / school-name leakage. Phase 8 (atomic version bump, ship): - .claude-plugin/plugin.json: 3.8.2 → 3.9.2 (also catches v3.9.0 + v3.9.1 deferrals); description updated for 38-agent ensemble + v3.9.2 fence. - .claude/CLAUDE.md Skills Overview table: deep-research 2.9.3→2.9.4, academic-paper 3.1.1→3.1.2, academic-paper-reviewer 1.9.0→1.9.1, academic-pipeline 3.9.0→3.9.2. - .claude/CLAUDE.md Suite version: 3.9.0 → 3.9.2 (also fixes pre-existing v3.9.1 latent bug — that ship missed bumping Suite version). - 4 SKILL.md frontmatter version + last_updated bumped to match. - academic-paper-reviewer/SKILL.md body Version Info table synced (lint enforces frontmatter + body alignment via check_reviewer_version_block). - MODE_REGISTRY.md 'Last updated' bumped to v3.9.2 (2026-05-18). - README.md + README.zh-TW.md: badge bumped + v3.9.2 + v3.9.1 changelog entries added in both EN + zh-TW. - CHANGELOG.md [3.9.2] entry: full release notes covering Phase L1 / 1 / 3 / 4 / 5 / 6 (review absorption) + carry-over to v3.10 (#134) + migration notes for in-flight users. - docs/design/2026-05-18-ars-v3.9.2-phase-boundary-spec.md: copied from .local-plans/ as final spec doc. - scripts/check_spec_consistency.py: hardcoded version refs updated to v3.9.2 (also expects new v3.9.1 + v3.9.2 README sections). Verification: - check_spec_consistency.py PASS - check_version_consistency.py PASS (was failing on v3.9.1 latent bug) - check_v3_9_2_phase_boundary.py PASS (22 + 16 = 38 agents validated) - check_v3_6_7_pattern_protection.py PASS (existing PROTECTION blocks unaffected by new Phase Boundary additions) - pytest 1482 passed + 3 skipped + 111 subtests passed - 0 regression vs v3.9.1 baseline (was 1463; +19 = 1482, all new tests from this PR) - personal-boundary PASSED (648 files, 0 violations) Closes #133 v3.10 conductor architecture (#134) carries: - PreToolUse hook (Phase 0.1 verified CC payload has agent_type field) - Multi-phase ars_phase_writes/reads envelope schema - Deterministic verifier with author provenance - Orchestrator cross-phase intake capability Refs #134
341 lines
13 KiB
Python
341 lines
13 KiB
Python
#!/usr/bin/env python3
|
|
"""Advisory verifier for ARS pipeline phase scope (v3.9.2).
|
|
|
|
Spec: docs/design/2026-05-18-ars-v3.9.2-phase-boundary-spec.md Phase 4
|
|
Issue: #133 (phase scope inflation hot-fix)
|
|
Forward note: v3.10 active conductor (#134) replaces this with deterministic
|
|
provenance-based verification via task envelope; this v3.9.2 advisory is
|
|
heuristic and FP-prone by design.
|
|
|
|
Usage:
|
|
python scripts/check_pipeline_integrity.py [workdir]
|
|
python scripts/check_pipeline_integrity.py --json [workdir]
|
|
|
|
Default workdir: current directory.
|
|
|
|
This script SCANS a working directory for `phaseN_*/` subdirectories (where
|
|
N is 1-6 per the ARS pipeline phase convention) and flags advisory signals
|
|
that suggest #133-class scope inflation. Output is ADVISORY ONLY — it does
|
|
NOT block any workflow. User reviews findings and decides.
|
|
|
|
Detection rules (v3.9.2 lite):
|
|
|
|
1. **Phase 5 missing independent reviewer attribution** (HIGH-VALUE structural rule)
|
|
When `phase5_*/` exists but contains no separate files matching the
|
|
independent-reviewer naming convention (devils_advocate / editor_in_chief /
|
|
ethics_review / eic / methodology / domain / perspective / editorial_synthesizer),
|
|
flag INTEGRITY-ADVISORY. This catches the exact #133 reported pattern:
|
|
a single agent producing phase5_*/ output that looks like a review but was
|
|
never actually crosschecked by independent reviewer agents.
|
|
|
|
**NORMATIVE FILENAME CONVENTION:** This rule is filename-based regex matching
|
|
over a closed set of agent stem names. For the verifier to recognize a
|
|
reviewer report, the filename MUST contain one of the canonical agent stems:
|
|
- `devils_advocate` (or `devil-s_advocate` / `devils-advocate`)
|
|
- `editor_in_chief` or `eic`
|
|
- `ethics_review` (the full stem; bare `ethics` does NOT match)
|
|
- `methodology_review` / `domain_review` / `perspective_review`
|
|
- `editorial_synth` (matches editorial_synthesizer family)
|
|
|
|
Orchestrator runs that emit terse filenames like `round1_ethics.md` (without
|
|
`_review` suffix) WILL produce a false-positive STRUCTURAL finding here.
|
|
Resolution options: (a) rename file to match convention, or (b) ignore the
|
|
advisory finding (verifier exits 0; advisory does not block). The verifier
|
|
trades filename-convention discipline for zero-dependency cross-platform
|
|
simplicity in v3.9.2. v3.10 conductor (#134) replaces filename matching
|
|
with task envelope author provenance.
|
|
|
|
2. **Multi-phase same-call generation heuristic** (LOWER-VALUE, FP-prone)
|
|
When files in phaseN_*/ and phase{N+1}_*/ share creation timestamps within
|
|
a configurable window (default 5 minutes), flag as POSSIBLE same-call
|
|
inflation. Acknowledged FP risk: legitimate fast orchestrator runs trigger
|
|
this. Use --strict to enable, default OFF.
|
|
|
|
Exit codes:
|
|
0 No findings, or only --strict findings without --strict flag
|
|
0 Findings produced (advisory output, NOT a hard gate); printed to stdout
|
|
1 Script error (invalid workdir, IO failure)
|
|
|
|
The verifier intentionally fails open (exit 0) on findings — it is advisory,
|
|
not a CI gate. CI should not rely on this for v3.9.2.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import json
|
|
import re
|
|
import sys
|
|
from dataclasses import dataclass, field
|
|
from pathlib import Path
|
|
from typing import Iterable
|
|
|
|
PHASE_DIR_RE = re.compile(r"^phase([1-6])(?:_.*)?$")
|
|
|
|
# Phase 5 reviewer file-name heuristics. A phase5_*/ directory should contain
|
|
# files whose names match at least one of these per independent crosscheck
|
|
# (Devil's Advocate, Editor-in-Chief, Ethics Review). Missing all three →
|
|
# advisory flag per #133 root pattern.
|
|
PHASE5_REVIEWER_PATTERNS = {
|
|
"devils_advocate": re.compile(r"devils?[_-]?advocate", re.IGNORECASE),
|
|
"editor_in_chief": re.compile(r"(?:editor[_-]?in[_-]?chief|^eic|[_-]eic[_-])", re.IGNORECASE),
|
|
"ethics_review": re.compile(r"ethics?[_-]?review", re.IGNORECASE),
|
|
"methodology_reviewer": re.compile(r"methodology[_-]?review", re.IGNORECASE),
|
|
"domain_reviewer": re.compile(r"domain[_-]?review", re.IGNORECASE),
|
|
"perspective_reviewer": re.compile(r"perspective[_-]?review", re.IGNORECASE),
|
|
"editorial_synthesizer": re.compile(r"editorial[_-]?synth", re.IGNORECASE),
|
|
}
|
|
|
|
# Required minimum reviewer attributions for a valid Phase 5 output per
|
|
# academic-pipeline Stage 3 review protocol. At least one from each category
|
|
# must appear in phase5_*/ filenames.
|
|
PHASE5_REQUIRED_CATEGORIES = [
|
|
("devil's advocate", ["devils_advocate"]),
|
|
("editorial/EIC", ["editor_in_chief", "editorial_synthesizer"]),
|
|
("ethics or panel reviewer", ["ethics_review", "methodology_reviewer", "domain_reviewer", "perspective_reviewer"]),
|
|
]
|
|
|
|
DEFAULT_SAME_CALL_WINDOW_SECONDS = 300 # 5 minutes
|
|
|
|
|
|
@dataclass
|
|
class Finding:
|
|
rule: str
|
|
severity: str # ADVISORY (default), STRUCTURAL (high-value), HEURISTIC (FP-prone)
|
|
phase: int | None
|
|
path: str
|
|
message: str
|
|
|
|
|
|
@dataclass
|
|
class Report:
|
|
workdir: str
|
|
phase_dirs: dict[int, list[str]] = field(default_factory=dict) # phase → list of dir paths
|
|
findings: list[Finding] = field(default_factory=list)
|
|
|
|
|
|
def scan_workdir(workdir: Path) -> Report:
|
|
"""Locate all phase{1-6}_*/ subdirectories under workdir."""
|
|
report = Report(workdir=str(workdir))
|
|
if not workdir.is_dir():
|
|
return report
|
|
|
|
for entry in workdir.iterdir():
|
|
if not entry.is_dir():
|
|
continue
|
|
match = PHASE_DIR_RE.match(entry.name)
|
|
if not match:
|
|
continue
|
|
phase = int(match.group(1))
|
|
report.phase_dirs.setdefault(phase, []).append(str(entry))
|
|
return report
|
|
|
|
|
|
def check_phase5_attribution(report: Report) -> None:
|
|
"""Rule 1 — STRUCTURAL: phase5_*/ must contain independent reviewer files."""
|
|
phase5_dirs = report.phase_dirs.get(5, [])
|
|
if not phase5_dirs:
|
|
return
|
|
|
|
for dir_path in phase5_dirs:
|
|
path = Path(dir_path)
|
|
try:
|
|
# Skip hidden files (.DS_Store, Thumbs.db, .gitkeep, etc.) — they're
|
|
# never reviewer reports and would clutter advisory output noise.
|
|
files = [
|
|
p.name for p in path.rglob("*")
|
|
if p.is_file() and not p.name.startswith(".")
|
|
]
|
|
except OSError as exc:
|
|
report.findings.append(Finding(
|
|
rule="phase5_attribution_io_error",
|
|
severity="ADVISORY",
|
|
phase=5,
|
|
path=dir_path,
|
|
message=f"Could not read phase5 directory: {exc}",
|
|
))
|
|
continue
|
|
|
|
if not files:
|
|
report.findings.append(Finding(
|
|
rule="phase5_empty",
|
|
severity="ADVISORY",
|
|
phase=5,
|
|
path=dir_path,
|
|
message="phase5_*/ directory is empty — Phase 5 should produce review reports",
|
|
))
|
|
continue
|
|
|
|
# Tally which reviewer categories are present
|
|
matched_agents: set[str] = set()
|
|
for fname in files:
|
|
for agent, pattern in PHASE5_REVIEWER_PATTERNS.items():
|
|
if pattern.search(fname):
|
|
matched_agents.add(agent)
|
|
|
|
missing_categories: list[str] = []
|
|
for category_label, agent_list in PHASE5_REQUIRED_CATEGORIES:
|
|
if not any(agent in matched_agents for agent in agent_list):
|
|
missing_categories.append(category_label)
|
|
|
|
if missing_categories:
|
|
report.findings.append(Finding(
|
|
rule="phase5_missing_independent_reviewer",
|
|
severity="STRUCTURAL",
|
|
phase=5,
|
|
path=dir_path,
|
|
message=(
|
|
f"phase5_*/ missing independent reviewer attribution for categories: "
|
|
f"{', '.join(missing_categories)}. "
|
|
f"Files found: {files[:5]}{'...' if len(files) > 5 else ''}. "
|
|
f"#133 pattern: Phase 5 deliverable was likely produced by a single agent "
|
|
f"that inflated past its scope, skipping mandatory independent crosschecks "
|
|
f"(DA / EIC / Ethics). Re-run via orchestrator-driven Mode A with "
|
|
f"`/ars-full` or invoke each reviewer agent separately."
|
|
),
|
|
))
|
|
|
|
|
|
def _file_mtime(path: Path) -> float | None:
|
|
try:
|
|
return path.stat().st_mtime
|
|
except OSError:
|
|
return None
|
|
|
|
|
|
def check_same_call_heuristic(report: Report, window_seconds: int) -> None:
|
|
"""Rule 2 — HEURISTIC (--strict only): adjacent-phase same-window timestamps."""
|
|
phases_present = sorted(report.phase_dirs.keys())
|
|
if len(phases_present) < 2:
|
|
return
|
|
|
|
# Collect (phase, file_path, mtime) tuples
|
|
file_records: dict[int, list[tuple[Path, float]]] = {}
|
|
for phase, dirs in report.phase_dirs.items():
|
|
records: list[tuple[Path, float]] = []
|
|
for dir_path in dirs:
|
|
path = Path(dir_path)
|
|
for file_path in path.rglob("*"):
|
|
if not file_path.is_file():
|
|
continue
|
|
mtime = _file_mtime(file_path)
|
|
if mtime is None:
|
|
continue
|
|
records.append((file_path, mtime))
|
|
if records:
|
|
file_records[phase] = records
|
|
|
|
# For each adjacent (N, N+1) pair, find files with mtime within window
|
|
for phase in phases_present:
|
|
next_phase = phase + 1
|
|
if next_phase not in file_records:
|
|
continue
|
|
for file_a, mtime_a in file_records.get(phase, []):
|
|
for file_b, mtime_b in file_records[next_phase]:
|
|
delta = abs(mtime_a - mtime_b)
|
|
if delta <= window_seconds:
|
|
report.findings.append(Finding(
|
|
rule="adjacent_phase_same_window",
|
|
severity="HEURISTIC",
|
|
phase=phase,
|
|
path=str(file_a),
|
|
message=(
|
|
f"phase{phase} file {file_a.name} and phase{next_phase} file "
|
|
f"{file_b.name} share mtime within {int(delta)}s "
|
|
f"(window={window_seconds}s). POSSIBLE same-call inflation. "
|
|
f"Note: legitimate fast orchestrator runs also trigger this "
|
|
f"heuristic — verify against orchestrator state ledger before "
|
|
f"treating as #133-class violation."
|
|
),
|
|
))
|
|
|
|
|
|
def format_text(report: Report) -> str:
|
|
lines = [f"ARS pipeline integrity check (v3.9.2 advisory)"]
|
|
lines.append(f"Workdir: {report.workdir}")
|
|
lines.append(f"Phase dirs found: {dict(sorted(report.phase_dirs.items()))}")
|
|
lines.append("")
|
|
if not report.findings:
|
|
lines.append("No advisory findings.")
|
|
return "\n".join(lines)
|
|
|
|
lines.append(f"Findings ({len(report.findings)}):")
|
|
for i, finding in enumerate(report.findings, 1):
|
|
lines.append("")
|
|
lines.append(f" [{i}] {finding.severity} — {finding.rule}")
|
|
if finding.phase is not None:
|
|
lines.append(f" Phase: {finding.phase}")
|
|
lines.append(f" Path: {finding.path}")
|
|
lines.append(f" {finding.message}")
|
|
|
|
lines.append("")
|
|
lines.append("Reminder: this output is ADVISORY. Findings do NOT block any workflow.")
|
|
lines.append("See docs/design/2026-05-18-ars-v3.9.2-phase-boundary-spec.md for rationale.")
|
|
return "\n".join(lines)
|
|
|
|
|
|
def format_json(report: Report) -> str:
|
|
payload = {
|
|
"workdir": report.workdir,
|
|
"phase_dirs": {str(k): v for k, v in sorted(report.phase_dirs.items())},
|
|
"findings": [
|
|
{
|
|
"rule": f.rule,
|
|
"severity": f.severity,
|
|
"phase": f.phase,
|
|
"path": f.path,
|
|
"message": f.message,
|
|
}
|
|
for f in report.findings
|
|
],
|
|
}
|
|
return json.dumps(payload, indent=2, ensure_ascii=False)
|
|
|
|
|
|
def main(argv: list[str] | None = None) -> int:
|
|
parser = argparse.ArgumentParser(
|
|
description="Advisory check for ARS pipeline phase scope inflation (v3.9.2)"
|
|
)
|
|
parser.add_argument(
|
|
"workdir",
|
|
nargs="?",
|
|
default=".",
|
|
help="Working directory to scan (default: current directory)",
|
|
)
|
|
parser.add_argument(
|
|
"--strict",
|
|
action="store_true",
|
|
help="Enable heuristic same-call timestamp check (FP-prone, default OFF)",
|
|
)
|
|
parser.add_argument(
|
|
"--window-seconds",
|
|
type=int,
|
|
default=DEFAULT_SAME_CALL_WINDOW_SECONDS,
|
|
help=f"Same-call window for --strict heuristic (default: {DEFAULT_SAME_CALL_WINDOW_SECONDS}s)",
|
|
)
|
|
parser.add_argument(
|
|
"--json",
|
|
action="store_true",
|
|
help="Output JSON instead of human-readable text",
|
|
)
|
|
args = parser.parse_args(argv)
|
|
|
|
workdir = Path(args.workdir).resolve()
|
|
if not workdir.is_dir():
|
|
print(f"ERROR: workdir not found or not a directory: {workdir}", file=sys.stderr)
|
|
return 1
|
|
|
|
report = scan_workdir(workdir)
|
|
check_phase5_attribution(report)
|
|
if args.strict:
|
|
check_same_call_heuristic(report, args.window_seconds)
|
|
|
|
if args.json:
|
|
print(format_json(report))
|
|
else:
|
|
print(format_text(report))
|
|
return 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
sys.exit(main())
|