mirror of
https://github.com/Imbad0202/academic-research-skills.git
synced 2026-09-14 13:51:17 +08:00
main
771 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d5accd6b1f |
docs(changelog): record the #856 es-ES trigger phrases under [Unreleased] (#866)
Extends the #855 entry with the merged companion trigger change (
|
||
|
|
a366e39e2e |
Add conservative es-ES trigger keywords to the four skills (#856)
* add conservative es-ES trigger keywords to the four skills * address review: re-pin pipeline content lock, add verificar citas, fix reseñas typo * address review items 4-6: trim deep-research es subset under 1,024, drop standalone investigación, add es-ES routing fixtures |
||
|
|
91fc74d37e |
docs(contributing): provisional single-owner locale applications (#862) (#863)
Add a provisional-application route under the "Two named owners" locale-pack condition: a single-owner application is recorded in a dedicated issue, is not a supported pack, and gets a 14-day backup window opening with the first minor release after both the locale mechanism and the primary owner's recorded acceptance. An unfilled window lapses the application; a supported pack that loses either owner leaves the supported list. Point the interim activation-layer sentence at #862 (Phase 1) alongside #850. Record the #861 locale-pack policy and this amendment together under [Unreleased] Changed so the pre-tag changelog-covers-merges gate sees both. Claude-Session: https://claude.ai/code/session_01HwTR5WFp4E4gNU8RF2CjNy Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
7cf67a657d |
docs(contributing): locale packs are community-maintained; translation fast-merge excludes operative text (#861)
States the ownership model for non-default output locales ahead of the #850 mechanism: two named owners, recorded currency with a 14-day window per minor release, visible staleness in CI that never delays a core release, configuration-and-presentation scope only, and #509 trigger discipline including the Agent Skills 1,024-character description cap. Narrows the translation fast-merge category so operative instructions follow the review tier of the file they touch. Adds README.es-ES.md to the README sync list. Claude-Session: https://claude.ai/code/session_01EEbmXsHyQ3dzYoep74zSmH Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
4be3d11ca9 |
docs(changelog): record the #855 es-ES README under [Unreleased] (#860)
Claude-Session: https://claude.ai/code/session_01EEbmXsHyQ3dzYoep74zSmH Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
665b3ca0a0 | add es-ES README with language nav links and CI registration (#855) | ||
|
|
90f2176cd5 |
evals: add claude plugin eval suite for the academic-paper citation-check flow (#859)
* evals: add `claude plugin eval` suite for the academic-paper citation-check flow Eight cases (six fire, two negative) under plugin-evals-citation-check/, graded as a with/without-plugin ablation. Fire cases carry a synthetic source pack so the four author-defined citation failures (no source on file, wrong authors, hedged finding cited as established, retracted or concern-flagged paper cited as live) are detectable offline. Styles: APA 7 (en / zh-TW mixed / es), IEEE, Vancouver (style unnamed), Chicago NB. Cases pin model: sonnet; run with --judge-model opus. Calibration (two pilots, 2026-09-13) and caveats are in the suite README. The with-plugin arm cannot load the mode prompt in the eval sandbox because the command stub references plugin files by relative path (#857), and plain-language prompts fired the skill in 3 of 6 cases (#858; the Spanish case is one data point for #850). No uplift figure is claimed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EEbmXsHyQ3dzYoep74zSmH * evals(citation-check): apply cross-model review findings and re-calibrate Twelve of thirteen review findings applied: replace the disputable four-author "et al." planting in 04 with a year mismatch; make the clean citations in 02 and 06 supported by their abstracts; replace real Taiwan journal names in 02 with fictional ones; tie the 08 presence regexes to an "unused" statement; require metadata preservation and reject audit content on the 07 conversion negative; add no-overreach to 05; turn the 02 language check into an llm grader; exempt unchanged entries in a complete corrected list from no-false-positive and drop its DOI claim; drop the 04 style-name regex (both arms fixed the year without naming the style). The one rejected finding (08 skill check "display-only") was wrong: it carries arm: both and is scored; README says so. Pilot 3 on the revised suite: $4.65, max 142 s / 7 turns / $0.45 per run; 04 and 08 re-run clean after the last two grader fixes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EEbmXsHyQ3dzYoep74zSmH --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
88725b8a55 |
fix(academic-paper): advertise revision-coach rebuttal triggers in the SKILL.md description (#851) (#853)
The revision-coach phrases "I got reviewer comments", "revision roadmap", "should we push back", "conference rebuttal", "grant panel response" lived only in the SKILL.md body, which the model reads after deciding to load the skill. Add them (plus zh-TW/ko equivalents) to the frontmatter description (699 chars, under the 1,024 Claude Code allowance) and add the three missing English phrases to the body Trigger Keywords line. Verification: plugin-evals/03-iclr-rebuttal-en skill-fired 0/2 -> 7/7. The case's two llm rubrics are rewritten in enumerate-then-quote style; the earlier claim-list phrasing drew 3-vote FAILs from the runner judge on outputs a reasoning judge passed. README caveats updated. Closes #851 Claude-Session: https://claude.ai/code/session_013fbc5qpXkAac1o4HLMGinE Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
cfbd2c6f63 |
evals: add claude plugin eval suite for the academic-paper revision-coach flow (#852)
Seven cases (five fire, two should-not-fire), twenty graders, run as a with/without-plugin ablation. Primary quality axis is "no unauthorised rewriting" (no manuscript prose drafted, nothing changed that no reviewer asked for, no results or changes asserted that have not happened). Calibrated against five pilots on 2026-09-12; README records the run command, ceilings, pilot cost and caveats. plugin-evals/results/ is gitignored. The ICLR case surfaced the trigger gap filed as #851 and is its acceptance check. Claude-Session: https://claude.ai/code/session_013fbc5qpXkAac1o4HLMGinE Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
f1a57bbcab |
fix: shared file-lock helper with msvcrt backend for the remaining fcntl sites (#845) (#847)
* fix: shared file-lock helper with msvcrt backend for the six fcntl sites (#845) scripts/file_lock.py owns the backend choice (fcntl.flock on POSIX, msvcrt.locking on byte 0 on Windows) and routes adjudication_activity, inquiry_branch_ledger, review_criteria_binding, and ars_mark_read through acquire()/release(). POSIX lock sequences are unchanged. Per-site Windows decisions: adjudication reads degrade to exclusive with a 5 s bounded wait; the review-criteria manifest lock is capped at 30 s on Windows only; the inquiry ledger alpha keeps refusing non-POSIX hosts. Two finally blocks that released an unacquired lock now release only what they acquired. SETUP docs state the best-effort Windows posture; no Windows CI job is added. Refs #845, #843, #844. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2 * fix(file_lock): interrupted attempts honour the deadline; pin adjudication wait policy (#845) Cross-model review round 1 (gpt-6-astra, xhigh): a persistent InterruptedError could retry past the bound; the Windows-shape test did not exercise adjudication's reader-waits / writer-does-not-wait policy; the adjudication contention message now names LockTimeout instead of BlockingIOError, recorded in the CHANGELOG rather than masked. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2 * refactor(file_lock): held() context manager, single BACKEND source, one fake msvcrt (#845) /simplify pass (four cleanup reviewers): the release-only-if-acquired invariant moves into file_lock.held() and review_criteria_binding / inquiry_branch_ledger use it; runtime branches key off BACKEND and SHARED_LOCKS_SUPPORTED is dropped; EINTR joins the retryable errno set and the unreachable EDEADLK entry goes; backend calls are deduplicated; all four consumers try the sibling import first so one module instance is shared; the Windows fake lives once in tests/fake_msvcrt.py; test scaffolding is folded into a lock_pair fixture and a parametrized wait test. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2 * fix(file_lock): keep lock acquisition and the guarded body in separate try blocks (#845) Cross-model review round 3 (gpt-6-astra, xhigh): wrapping the body in the same handler that translates LockTimeout meant a contended inner lock inside the body was reported as the outer manifest/passport lock failing. Both consumers now acquire in their own try block and release only after a successful acquire; held() is dropped from the helper. The subprocess test pins that a LockTimeout raised inside the binding body surfaces as itself. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2 * test(file_lock): let the body LockTimeout leave _locked() so the attribution check bites (#845) Cross-model review round 4: the inner LockTimeout was caught inside the binding body, so the erroneous outer translation would still have passed. Verified by mutation: restoring the outer translation fails this test. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2 * ci(673): whitelist scripts/test_file_lock.py as a non-consumer importer of the activity runtime (#845) The shared file-lock test imports adjudication_activity in a subprocess to exercise its lock backend under a fake msvcrt; it never reads or writes an activity store. The exact-owner whitelist is the lint's route for that. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2 --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
c7af8b9017 |
docs(changelog): record the #844 Windows msvcrt lock backend for /ars-mark-read under [Unreleased] (#846)
Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2 Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
32f754aae5 |
fix: make /ars-mark-read work on Windows via msvcrt lock backend (#843) (#844)
scripts/ars_mark_read.py imported POSIX-only fcntl at module load, so the documented /ars-mark-read CLI failed on Windows before parsing arguments. The lock now goes through two small helpers: fcntl.flock on POSIX (unchanged) and msvcrt.locking(LK_NBLCK, 1) on Windows, inside the same bounded retry loop. msvcrt.locking can lock a byte beyond EOF, so no pre-write is needed. Fixes #843. Remaining fcntl import sites are tracked in #845. Co-authored-by: dajiaohuang <dajiaohuang@users.noreply.github.com> |
||
|
|
8e4c877764 |
docs(changelog): backfill the #835 reviewer-calibration harness entry under [Unreleased] (#842)
check_changelog_covers_merges flagged #835 (merged 2026-09-07) as the one release-worthy commit since v3.21.2 without a CHANGELOG entry. Entry is derived from the PR body and the merged file set; it states that no calibration profile or measurement values shipped and that #653 / #828 stay open. Claude-Session: https://claude.ai/code/session_01AYAjWg2eBEz3UV7MZn7eFt Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
f832c89f60 |
docs: Gartenberg et al. (2026) fourth human-in-the-loop anchor, volume non-goal, cognitive-surrender note (#833) (#841)
- README.md / README.zh-TW.md motivation: fourth anchor paragraph for the Organization Science AI Task Force editorial "More versus better" (37(3):795-812). Scope stated: one journal, observational, aggregate, proprietary classifier. Cited as design rationale, not as evidence about ARS output. - POSITIONING.md "Rejected mechanisms": volume as an outcome. No batch manuscript generation, no fan-out of one run into several submissions, time-to-draft booked as a resource cost. - shared/collaboration_depth_rubric.md 1.0 -> 1.0.1: related-construct citation on Cognitive Vigilance (uncritical acceptance of AI output; "cognitive surrender" as the editorial cites Shaw & Nave 2026). Scoring, dimensions, and the descriptive-only reporting rule unchanged; a low score stays an observation, not a failure. - CHANGELOG [Unreleased] > Changed. The claim_strength_ladder.md item from #833 is byte-pinned by the revision-claim-drift suite (CLAIM_LADDER_SHA256 and the frozen v2 adjudication rubric); it is held for a separate decision. Claude-Session: https://claude.ai/code/session_01AYAjWg2eBEz3UV7MZn7eFt Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
27e6c9978d |
fix(socratic): F6 lists user directions unranked; round caps defer to the mentor agent (#834) (#840)
failure_paths.md § F6 preselected "[the most promising direction]", told the mentor to rank directions by "convergence potential", and prescribed "restrict discussion scope" — ranking and preselection of the user's own directions, which the #735 boundary forbids. F6 and socratic_mode_protocol.md also still said "round 15 → end" after #490 made the mentor agent's § Auto-End Conditions (Precise) the single authority. - F6: chronological, user-worded summary of expressed directions; the user chooses; the full-mode option names the visible exit marker; no scope restriction; round caps point at the agent file. - socratic_mode_protocol.md § Dialogue Management Rules: the 15-round line becomes a pointer to the agent authority. - test_socratic_rq_non_generation_contract.py: ranking/preselection vocabulary check on F6, own-round-count check on both reference files, pointer/heading parity with the agent file; mutation tests inject the pre-fix bytes and a stagnation-trigger control. - CHANGELOG [Unreleased]: contract-contradiction closure, no breadth claim. Claude-Session: https://claude.ai/code/session_01AYAjWg2eBEz3UV7MZn7eFt Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
75070eec84 |
feat(evals): #653/#828 add reviewer-calibration harness with isolated dispatch and audited scoring (#835)
* feat(evals): #653 reviewer-calibration suite scaffolding — corpus assembler, isolated dispatcher, deterministic scorer, pre-registered rubric/RUN_PLAN (corpus freeze pending PDF access) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H2iNYa6YYYaPUwD2Z2Jr5e * feat(evals): #653 freeze the ICLR 2026 calibration corpus manifest (12 papers) + shared PDF-text normalization Corpus freeze (PR-A of #653): `corpus/papers.json` (label-free, 6+6 ICLR 2026 papers by the pre-registered seed; pypdf 6.11.0; pool hashes unchanged from the 2026-08-07 selection) and `manifests/gold_labels.json` (public Decision note ids + strings). No page-cap exclusion fired; `verify` PASS. First real-PDF contact found a hashing defect: pypdf emits lone UTF-16 surrogates from math fonts (61 in one sampled manuscript) and strict UTF-8 encoding raised, so `extracted_text_sha256` was uncomputable. The normalization now lives in one shared module (`scripts/_calibration_pdf_text.py`: NFC + lone-surrogate -> U+FFFD), imported by both the assembler and the dispatcher so freeze/verify/dispatch hash identical bytes; the rule is recorded in the manifest's `extraction.text_normalization` and `verify` fails hard on rule drift (a rule, not a version). Two tests added (41 total). `scripts/fetch_calibration_corpus.py` is the authenticated OpenReview operator tool that produces the freeze input, so the "third-party reconstruction" claim in the README is backed by a runnable path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1 * refactor(evals): #653 simplify pass — shared hashing/fence/git-state, contract 1.1 docs /simplify findings applied (reuse, simplification, efficiency, altitude): - `_calibration_pdf_text.py` owns `sha256_hex` + `pdf_facts` (bytes hashed and parsed from one read via BytesIO; `extract_text=False` lets `verify` skip extraction when the pypdf version cannot be compared); surrogate replacement is one `re.sub` pass. Both the assembler and the dispatcher import it. - dispatcher reuses E4's closed data-fence grammar (`_delimited`), `_git_state` (declares unknown provenance dirty instead of raising), and the evidence path guard (`assert_plain_file`: rejects symlinked parent components, not just the leaf); one `_prepare` preamble for both stages; a text-hash mismatch now names its cause (installed vs manifest pypdf version). - assembler: exclusion rows stay dicts, `pool_list_mismatches` shared by freeze/verify, exclusion set built once. - scorer: `confusion`/`bootstrap_ci` take (predicted, gold) pairs (same RNG stream as before), `Counter` for the exact-mode vote, dead `_path` dropped. - RUN_PLAN/README: measurement contract 1.0 is closed to new rows (#664); the run publishes under 1.1 with its pre-registration record + write-once execution manifest (dispatcher/scorer support lands with the scored run). Re-freeze after the refactor reproduces papers[] and gold_labels byte-for-byte. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1 * fix(evals): #653 Iron Rule #7 at the two whole-file call boundaries + paper-id shape check Security review round 1 (first-party) found two below-threshold gaps and both are verified real: - The calibration dispatcher omitted E4's `DATA_BOUNDARY` sentence on the field-analyst call (the one E4 call that carries it, because `field_analyst_agent.md` states no untrusted-material rule of its own). Restored, and a fitted `REPORT_BOUNDARY` added on the synthesizer call, whose agent file is likewise dispatched whole with no such rule. Pinned by a transport-capture test that checks both sentences precede their fence. - Paper ids are spliced into file names (`<id>.pdf`, `cards/<id>/`) but `load_pool` accepted any non-empty string. Ids now must match `^[A-Za-z0-9_-]+$` (OpenReview's forum-id shape) in the assembler and the fetch tool; test pins the refusal. 43 tests pass. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1 * fix(evals): #653 codex round 1 — dispatch/verify invariant parity, card-path guard, scorer completeness Codex round 1 (gpt-6-astra xhigh) findings 2-7, 10, 11 and the cheap half of 9, each re-verified first-party before the change: - dispatcher: frozen cards go through the same plain-file guard as PDFs and agent files (a symlinked card1.md -> gold_labels.json was readable); the manifest's text_normalization rule and page_count are checked before dispatch, so dispatch and verify enforce the same manuscript invariants; transport-failure artifacts keep the partial stdout and stderr verbatim; every call attempt records RFC-3339 start/complete and prompt/output hashes into the panel record and cards freeze (the per-call evidence the heldout-measurement/1.1 execution manifest is built from). - verify: label must match decision_raw under the label transform; paper count and per-class label counts must equal the recorded quotas (synchronized paper+label removal no longer passes). - scorer: a second record for the same paper/replicate is a hard error, not a silent overwrite; a gold paper with no complete ensemble blocks the full tier; an A1 override needs its verbatim `raw` excerpt present in synthesis.md. Nine regression tests added (52 total). Real-corpus verify still PASS. Not addressed here (need a decision): finding 1 (camera-ready format leaks the accept label) and finding 8 (numeric seat scores vs categorical seat contract); finding 9's manifest/row builders land with the scored run. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1 * fix(evals): #653 drop the numeric score axis — protocol Phase 2 forbids AUC, seats are categorical Codex round 1 finding 8, verified against the source: the seat contract (eic/methodology/... agents) emits criterion-bound categorical judgements and states "Do not total, weight, average"; `calibration_mode_protocol.md` Phase 2 says "Do not report AUC: there is no continuous rubric score." The scorer nevertheless extracted a `Weighted Average` figure (a retired field) and RUN_PLAN promised AUC + score variance, so a conforming run would have published null numerics against a plan that promised them. The scorer now reports only what the protocol's full-tier table names: confusion matrix, balanced accuracy, FNR, FPR (bootstrap CIs), exact-label agreement (count/share/target-set size, with the binary-gold caveat), and replicate stability as categorical agreement (on side, on exact label). AUC is emitted as an explicit NOT REPORTED line. RUN_PLAN and the test fixtures follow. 52 tests pass. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1 * docs(evals): #653 mark the 2026-09-06 corpus SUPERSEDED (layout leaks the label, #828); RUN_PLAN model currency - README/RUN_PLAN: the frozen ICLR 2026 corpus is a harness-rehearsal corpus only — camera-ready replacement makes accepted PDFs visibly different from rejected submission PDFs (6/6 + 6/6; 30/30 in a fresh accepted-pool sample). No profile or measurement row may be published from it; the gold corpus becomes an ICLR 2027 submission-time capture. The "Why ICLR 2026" rationale is kept as pre-registered and annotated with the two facts that now cut against it (layout leak; Fable 5.1's 2026-06 cutoff covers the decisions). - RUN_PLAN + dispatcher default: subject `claude-fable-5` -> `claude-fable-5-1`, judge `gpt-5.6-sol` -> `gpt-6-astra` (provisional, #783 policy). Pre-dispatch edits, not amendments. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1 * feat(evals): #828 layout-tell guard at corpus freeze — refuse a corpus whose page-1 layout is not constant `assemble_calibration_corpus.py freeze` now reads page 1 of every cached PDF and evaluates four venue-template signals (published-as header, under-review header, "Anonymous authors", >=10 bare three-digit line numbers). Any signal that is not constant across the whole corpus refuses the freeze with the per-class counts; a uniform corpus records `layout_tell_check` in papers.json. `verify` recomputes the same check (skipped with a warning when a PDF is not cached; a manifest without the block warns). On the superseded 2026-09-06 ICLR 2026 corpus every signal is 6/0, so `verify` now FAILs on it by design. Shared `_open_reader` + `first_page_text` in the PDF helper. Six tests (signal detection, full and partial separation refused, uniform freeze + verify round-trip, missing-PDF skip, pre-check manifest warning). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh * feat(evals): #653/#828 rehearsal fixes + heldout-measurement/1.1 manifest and row builders Rehearsal 2026-09-06 (2 papers x 1 replicate, blocked at the first call by a rejected API key) exposed three dispatcher gaps, all fixed with tests: - credential preflight: zero-cost `GET /v1/models` before the first billed call; a definitive 401/403 refuses (key never echoed), network trouble is `inconclusive` and proceeds; outcome recorded in every record - credential rejection mid-run (`Failed to authenticate` / `API Error: 401` / `Not logged in`) is never retried (`CredentialRejected`); other transport failures keep the single retry - an aborted cards stage writes `runs/blocked-cards-<paper>.json` with its per-call rows instead of losing them; both stages share one record writer 1.1 contract substrate (RUN_PLAN "pre-registration record + execution manifest" item): - `dispatch_calibration_panel.py --stage manifest` folds the completed call rows of one attempt (frozen cards + panel records; `load_attempt` refuses mixed attempt identities) into a write-once, schema-validated `execution-manifest.json` - `build_calibration_measurement_row.py` composes the 1.1 row: plan and rubric hashed and compared against `frozen_commit` (drift refuses; dirty commit refuses), manifest re-derived from the records and compared field-for-field, judge rows required (no judges, no row), agreement recomputed by the checker's own `judge_divergence` (extracted from `check_heldout_measurement_report.py`, behaviour unchanged), validated by the checker before a write-once write - adjudication rubric gains `## Resolution direction` (flags_only, I13 lower-bound labelling); README tooling section; RUN_PLAN names the row builder; DATA_FLOWS names the dispatcher's preflight touchpoint; scorer docstring de-staled (no score axis); pytest manifest +1 No calibration number is recorded anywhere in the repository. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh * fix(evals): #653/#828 codex round 2 — bind every row input to its attempt, harden the guards 12 of 13 findings applied (gpt-6-astra xhigh, read-only exec): - P1 foreign metrics: scorer output is bound to the attempt (per_panel keys == the complete panel records here, attempt ids match, n_papers matches) - P1 raw drift: record admission re-hashes every completed call's raw output against output_sha256 (manifest stage and row builder alike); prompts are not retained (they embed the manuscript) - P1 preflight redirects: the probe uses a no-redirect opener (a 3xx is `inconclusive`) and skips a non-https ANTHROPIC_BASE_URL - P2 estimand: class-A adjudication is now pre-registered as bidirectional (every synthesis decision transcribed blind and compared with the grammar), so the row publishes a point_estimate instead of an I13 "lower bound" that only meant audit coverage - P2 pre-write parity with R5: manifest timestamps parsed and ordered before the write; declared claims checked against the local manifest - P2 strict JSON: inputs parsed with the checker's strict loader, outputs serialized with allow_nan=False and round-tripped - P2 judge failures: `--blocked-run` ledger entries merge into attempts.blocked_runs (I11) - P2 admission by content: suite/stage/status/provenance from the record body, never the filename; blocked records are identity-checked too - P2 cards re-run: a reused evidence dir refuses (write-once stage records) - P2 auth signature: anchored at the start of stdout/stderr and limited to exit-code failures; a timeout's partial prose is never a credential error - P2 layout signals: phrase tests run on whitespace-folded text - P2 partial PDF cache: verify checks every cached PDF (can refuse, cannot clear) instead of skipping the guard - P3 real `git show` test for sha256_at_commit on a temporary repository Partially applied: "distinguish unobservable signals from absence" (not built; the constancy rule is pre-registered as stricter by design). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh * fix(evals): #653/#828 shared transport — capture every assistant message, fence the subject's config Rehearsal take 2 (2026-09-06/07, 8 billed calls on the first paper) found two transport defects in `ClaudeCliTransport`, shared by the E4 and the calibration dispatchers: - text-mode `claude -p` prints only the LAST assistant message: the first paper's synthesis (long enough to be continued) came back starting mid-table, with the Editorial Decision Letter and its `### Decision:` line in the missing head. The transport now runs `--output-format stream-json --verbose` and concatenates the text blocks of every assistant message; an error result or an unreadable stream is a TransportFailure that keeps the raw bytes. - `--bare` does not fence the subject: a two-call probe on 2.1.260 showed the operator's whole global CLAUDE.md arriving as a system-reminder, plus `settings.json` `language` and the output style (the seats appended Traditional-Chinese "plain-language summary" sections). The subject now runs with an allowlisted environment (PATH/HOME/LANG/TMPDIR/TERM/USER/ SHELL + ANTHROPIC_*; no CLAUDE_* inherited from a parent session) and a per-transport empty `CLAUDE_CONFIG_DIR`; the same probe then reported no instruction beyond the SDK identity line and the date. E4 tests: one fake updated to emit stream-json; five new tests (message joining, error/junk results, unreadable-stream failure with bytes, environment allowlist, argv/env of a live call). Calibration docs and the panel record's `dispatch` field describe the new recipe (pre-dispatch change, no amendment). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh * fix(evals): #653/#828 codex round 3 on the shared transport — eviction signals, LF framing, network env, failure evidence Five P2 findings (gpt-6-astra xhigh, read-only exec), all applied: - refusal-fallback eviction: assistant `supersedes` and system `model_refusal_fallback.retracted_message_uuids` (wire fields verified in the installed CLI 2.1.260) drop retracted partials before concatenation - NDJSON split on LF only (`str.splitlines` also splits on U+0085 / U+2028 / U+2029 inside a JSON string); CRLF tolerated - environment allowlist keeps documented network/TLS inputs (proxies, NODE_EXTRA_CA_CERTS, SSL_CERT_*, CLAUDE_CODE_CLIENT_*); an apiKeyHelper that needs more is documented as unsupported behind the fence - transport failures carry assistant TEXT in `stdout` and the raw stream in `raw_stdout`; a framing-only stream is "no model response" (E4 no longer writes stream metadata as a partial response); both dispatchers preserve the raw stream as `*.transport-stream.jsonl` - a structured error result (`[TRANSPORT: result <subtype>]`, diagnostic in stdout) is classified by the calibration retry loop like the plain-text startup failure: a credential rejection is never retried E4 tests +6 (256), calibration +1. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh * feat(evals): #653/#828 keep the raw stream of successful calls as evidence `ClaudeCliTransport.last_raw_stdout` exposes the stream-json framing of the most recent successful call; the calibration dispatcher writes it next to the text as `<label>.transport-stream.jsonl`, so the next rehearsal shows how many assistant messages a deliverable spanned (the 2026-09-06 synthesis lost its head to exactly that). Probe 2026-09-07: a 12,000-line reply at effort low arrived as ONE text message after a thinking-only message, so the head loss is attributed to multiple text messages in one turn (likely interleaved thinking at xhigh), not to an output-length continuation; the parser covers both. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh * fix(evals): #653/#828 allow requiring a successful credential preflight * fix(calibration): bind audited decisions and preserve failed dispatch evidence * fix(transport): retain truncated UTF-8 output as byte evidence --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6b7ee6dcae |
fix: Astra request compat, no-delegation citation transport, hedge/quota prompt repairs, audit provenance (#823–#826) (#827)
* fix: Astra request compatibility, no-delegation citation transport, hedge/quota prompt repairs, audit provenance (#823 #824 #825 #826) #823 — OpenAI request builders (smoke entrypoint + documented example) drop `temperature`, which GPT-6 Astra rejects; the per-model effort vocabulary lives in scripts/cross_model_verification/openai_effort_guard.sh, sourced by both, and an unsupported explicit Astra value fails before curl. Hermetic fake-curl test runs both surfaces. #824 — the contained Codex citation transport rejects effort=ultra with REASONING_EFFORT_REQUIRES_DELEGATION before detection/auth/tempdir/launch on both entry paths (codex-cli 0.153.4 defines ultra as the multiAgentMode replacement). Model-independent by design. #825 — hedging can no longer rescue an unsupported claim (writer recovery tree, CER fallback row, temporal rule 5 in writer + both compiler mirrors, writer contract D2); universal prose quotas in the writer, compilers, writing_quality_check.md, academic-paper/SKILL.md, and contract D6 become diagnostics subordinate to author/venue requirements. Audit inventory corrected in place; held-out seed evals/heldout/unsupported_claim_recovery (NOT_RUN) registered. #826 — run_codex_audit.sh pins gpt-6-astra/xhigh and records both in a new sidecar `model` block; claim_audit_pipeline binds an unknown judge identity to a run-local cache key (no cross-run reuse) instead of defaulting to gpt-5.5-xhigh. Review: /simplify (4 angles), codex gpt-5.6-sol xhigh 2 rounds (r1: 1 P1 + 1 P2 + 2 P3 fixed; r2: 0 P1/P2), /security-review 0 findings; all 102 spec-consistency steps + pytest manifest replayed locally. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BNKiXpdHx1T5F5RbXT2Ueu * docs(claude): record the #824 ultra reversal in the v3.21.2 key-additions line The v3.21.2 bullet still said the contained Codex citation transport accepts ultra; #824 on this branch rejects it as a delegation request. Add the reversal so the live instruction surface matches the transport. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K7emV5r2aqZDJzAyYVuuDo --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
8fa3d651ad |
docs(release): v3.21.2 — model currency, checkpoint decision provenance, and CJK title-matching repairs [skip-closes-check] (#822)
Promotes the [Unreleased] block to v3.21.2 (2026-09-06) and aligns every version-bearing surface, following the v3.21.1 release-prep file set: CHANGELOG heading (empty [Unreleased] anchor kept), plugin/marketplace manifests, CITATION.cff, POSITIONING.md, MODE_REGISTRY.md, .claude/CLAUDE.md (table row, Key Additions, Version Info), academic-pipeline/SKILL.md plus its content-lock hash, docs/ARCHITECTURE.md current markers, the five README badges/headings/entries, and the spec-consistency lint pins with their fixtures. Claude-Session: https://claude.ai/code/session_011sWwwG3oCbtL4cGhRsr5US Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>v3.21.2 |
||
|
|
0861bc8538 |
chore(models): align docs and guardrails to Claude Fable 5.1 and GPT-6 Astra (#819) (#820)
* chore(models): align docs and guardrails to Claude Fable 5.1 and GPT-6 Astra Read both vendor system cards in full and applied the model-update pass: - Claude Fable 5.1 named as the current frontier model (PERFORMANCE en/zh-TW with a dated list-price re-derivation; cross-model primary-row example). - gpt-6-astra listed as a provisional cross-model verifier on both transports and recommended under the #783 lifecycle policy; gpt-5.6-sol keeps its validated status on the ChatGPT-subscription citation transport. Entry-gate smoke PASS on that transport (2026-09-05, codex-cli 0.153.4). SETUP en/zh-TW example sets, id-status allowlist, bakeoff baseline text, and .claude/CLAUDE.md move together. - Codex citation transport: `ultra` joins the closed reasoning-effort set as a named constant, with a test pinning turn/start forwarding and fail-closed rejection of unknown values. - New guardrail: checkpoint decision provenance (authority in the pipeline state machine, operational mirror in the orchestrator), indexed as risk R11; both content-lock hashes updated in this commit. - Provider-side monitoring / safety interventions named as a never-a-verdict case in the cross-model doc and the degradation registry row. - Model tiering records that the resolved tier is the declared model; risk register R1/R4/R5/R6 residual gaps updated. - Harness-retirement audit for the model change: audits/harness-retirement-2026-09-model-update.md (0 prompt retirements, 4 applied currency fixes, 2 deferred, 8 keep-as-debt annotations). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011sWwwG3oCbtL4cGhRsr5US * docs(changelog): align the model-update entries with the final text The [Unreleased] entries were written before the simplify pass moved the checkpoint-decision authority into the pipeline state machine, reused the existing transport-failure markers for provider-side interventions, and de-numbered the model-tiering note. Wording now matches the files. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011sWwwG3oCbtL4cGhRsr5US * test: scope the checkpoint-authority section out of the v3.6.7 orchestrator line budget The v3.6.7 Phase 6.6 budget test measures the orchestrator prompt minus every later independent extension, each with its own bounded cap. The new `## Checkpoint authority fidelity` section (13 lines) pushed the v3.6.7-attributed count to 652 against a 639 ceiling. Following the existing convention, the section gets its own measurement helper, an 18-line cap (5 lines of headroom), a dedicated test, and is subtracted from the historical budget. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011sWwwG3oCbtL4cGhRsr5US --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
9443623791 |
docs: de-stale RISK_REGISTER R10 residual gap — #769 rows shipped in the same release (#813) (#814)
R10 still claimed the guard-launcher degradations were 'not yet indexed in the degradation registry (#769)' although v3.21.1 itself shipped registry 1.3.0 with all five write_scope_guard_* rows. The stale clause is removed; the existing-controls line now points the guard's degrade posture at its registry rows; the residual gap keeps only the per-mechanism, per-channel loss description. check_risk_register and check_version_consistency pass locally. Closes #813 Claude-Session: https://claude.ai/code/session_011b4WKkZvjw7fCPDHQQhm6E Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
368576224b |
chore(audits): September harness-retirement audit — zero findings (#811) (#812)
Incremental audit over the 2026-08 baseline (
|
||
|
|
e8bf858be7 |
lint: skill-inventory parity across top-level dirs, skills/ symlinks, CLAUDE.md table, and marketplace.json (#809) (#810)
Closes #809. check_skill_inventory_parity.py: on-disk <name>/SKILL.md dirs are the authority; set-equality against skills/ symlinks, the CLAUDE.md Skills Overview table (exact unfenced H2, GFM header+separator, first-cell backticked names), and marketplace.json plugins[].skills; "N skills" count claims on plugin.json / marketplace.json / MODE_REGISTRY.md. check_spec_consistency.py derives its skill paths from disk at call time; table-row grammar single-sourced in _skill_lint. 60 mutation tests; wired into spec-consistency.yml and the pytest manifest. Dual-track pre-ship: /simplify (4 findings applied), /security-review (none), codex gpt-5.6-sol xhigh 7 rounds (12 findings fixed, round 7 clean). CHANGELOG also covers #805. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014U5nvjKy84twtB4VYsrex1 |
||
|
|
37bd060294 |
docs: fix duplicated word in MLA citation key rules (#805)
Docs-only: academic-paper/references/citation_format_switcher.md MLA key rules line read "No year in in-text"; now "No year in-text", matching the in-text format documented above it. No lint or hash lock pins this file. Contributed by @LeslieLi46. |
||
|
|
5debcd2efb |
fix: strip CJK outer wrapper marks only when they enclose one balanced unit (#800) (#804)
* fix: strip CJK outer wrapper marks only when they enclose one balanced unit (#800) The wrapper strip inherited from #431 was positional: it removed the first and last characters whenever they matched as a wrapper pair TYPE, without checking they belonged to the same bracket pair. 《红楼梦》与《金瓶梅》 — two titles joined in one string — normalized to 红楼梦》与《金瓶梅, leaving an orphaned 》 mid-key. Matching correctness was unaffected (both sides mangle identically; no exploitable asymmetry found in the #799 security pass), but the mangled key is a semantic anomaly for any future single-sided consumer. Fix adds _outer_pair_encloses: the outer marks are stripped only when the interior between them is itself balanced under all six wrapper pairs. 《围城》 still strips to 围城, nested balanced interiors still unwrap (《基于「ProEXC」的研究》 → 基于「ProEXC」的研究), while 《红楼梦》与《金瓶梅》, “研究”与“实践”, and “研究与“实践” keep their marks. Both consumers change together — the CJK client re-imports the shared function (#799), pinned behaviorally as well as by identity. The two new discriminating tests are mutation-verified to fail against the pre-fix module. Full suite: 9262 passed, 3 skipped. * fix: scope the CJK interior balance scan to the outer pair's own family (#804 review) Addresses the P1 and three advisories on PR #804. P1: `’` is the closer of `‘` AND the English apostrophe; `”` likewise appears unpaired in mixed typesetting. The family-blind interior scan read the lone `’` in `《Alzheimer’s病中ProEXC表达》` as an unbalanced quote and refused to strip a genuine `《…》` wrap. Verified against main: that pair went from exact=True / ratio 1.0 to exact=False / ratio 0.6818 — below the 0.70 floor, so the DOI-keyed ratio gate and the title-fallback exact gate failed together and a correct DOI could be reported as a mismatch. That is the failure class #798 repaired, so it is not covered by #800's conservative-direction blessing: that blessing is for titles whose outer marks are not one pair, and here they are. Fix scopes the scan to the outer pair's own family — a `《…》` wrap tracks only `《`/`》` and is blind to quote marks. This costs the check nothing it was buying: any mark that can orphan the OUTER pair is by definition of that pair's own family. The whole #800 behavior table survives (`《红楼梦》与《金瓶梅》` still trips on its stray `》`, `“研究”与“实践”` on its stray `”`, nested `《基于「ProEXC」的研究》` still unwraps, `《》` → ""), and the flagged interaction `《「研究』》` → `「研究』` now keeps its mismatched inner quotes as content, which is what distinguishes it from `《研究》`. Regressions added for both apostrophe shapes (`’s`, `’98`). The two lookup maps the family-blind scan needed are now dead and removed. Advisory 1: the depth-0 stray-closer branch is pinned. The first attempt did not discriminate — clamping absorbs one closer, so any title with equal opener/closer counts (every natural case, including `《红楼梦》与《金瓶梅》`) is refused by the trailing depth check anyway. Discriminating requires interior closers to outnumber openers by exactly the clamp count, which no natural title shape produces, so the test uses a documented synthetic asserted against the helper directly. All three branches are now mutation-verified: clamp-instead-of- refuse, drop-the-trailing-check, and family-blind (the P1 regression itself). Advisory 2: the module docstring's "behaviorally equivalent to the implementation this was promoted from" is qualified as historical — true of the #798/#799 promotion, false since #800 deliberately changed this behavior. Advisory 3: the changelog's non-CJK invariance claim is narrowed to what the code supports. The claim holds on the two `has_cjk`-gated paths; the client's `_cn_titles_match` is ungated, so a mark-carrying Han-free title can change verdict there (`《Hamlet》and《Macbeth》` vs its pre-mangled form — verified to flip). The empty-wrapper guarantee is likewise a `_cjk_titles_match` property: `exact_normalized_title("《》", "《》")` is still True through the base branch. Both narrowings are now pinned by tests so the prose cannot drift from the code. Full suite: 9270 passed, 4 skipped, 0 failed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9469fc4d07 |
fix: fail check_surface_form_parity with an environment error when pyyaml is missing (#801 follow-up) (#803)
With the manifest present but pyyaml unimportable, _load_manifest returned None and main() misdiagnosed it as "manifest ... empty / null / non-mapping", pointing the reader at the wrong file. The import failure is now a distinct _YamlUnavailableError; the lint exits 1 naming pyyaml and the requirements-dev.txt remedy. Regression test pins the message (45 tests, all green; lint itself still passes). Also de-enumerate the stale "(PyYAML + jsonschema ...)" dependency parenthetical in docs/SETUP.md and docs/SETUP.zh-TW.md (both language files together, per bilingual-parity discipline). Claude-Session: https://claude.ai/code/session_013R81d1YwGvJAznkPKk9gNw Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
30ad279cdf |
fix: declare markdown-it-py floor and make the autolink round-trip tail run visibly (#801) (#802)
* fix: declare markdown-it-py floor and make the autolink round-trip tail run visibly (#801) The no-link_open round-trip tail of test_gfm_bare_urls_emails_and_schemes_ cannot_autolink soft-imported markdown-it-py (undeclared in requirements- dev.txt) and silently returned when absent, so it had never run in CI, while ambient markdown-it-py 2.x failed it on clean main (2.2.0 + linkify-it-py 2.0.3, reported in #799). Verified dividing line: 2.2.0 fails, 3.0.0 and 4.0.0 pass with linkify-it-py held at 2.0.3. - Split the tail into test_escaped_markdown_yields_no_linkify_tokens_on_ round_trip, gated by pytest.importorskip minversions (markdown_it 3.0.0, linkify_it 2.0.3): ambient-old environments skip visibly. - Declare markdown-it-py>=3.0 + linkify-it-py>=2.0.3 in requirements-dev.txt with a reverse pointer at the consuming test, so CI exercises the round trip for the first time. - Move the identical soft-import tail in test_renderer_neutralizes_markdown_ active_inventory_path (newly activated in CI by the same declaration) to the same importorskip idiom; no floor needed (default CommonMark, no linkify) — verified passing under 2.2.0, 3.0.0, and 4.0.0. - Consolidate the triplicated hostile-row construction in test_evidence_rows.py into one _hostile_row helper. Renderer behavior and every renderer-side assertion are unchanged. Verification: both full files 414 passed under markdown-it-py 4.0.0; affected tests re-run under 2.2.0 (pass + visible skip) and 3.0.0 (pass). Closes #801 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013R81d1YwGvJAznkPKk9gNw * fix: flatten inline token children in the newly activated manifest markdown scan (#801) Cross-model review (codex, xhigh) on PR #802 flagged that the twin test's token scan iterated only top-level tokens, but markdown-it nests link_open / image / html_inline under inline tokens' children — so the assertion could only ever catch html_block. Verified empirically, then flattened children into the scan (same idiom as the evidence-rows round-trip test). Strengthened assertion passes under markdown-it-py 2.2.0, 3.0.0, and 4.0.0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013R81d1YwGvJAznkPKk9gNw --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e5718cbf58 |
fix: apply Chinese-aware title matching in the four index resolvers (#798) (#799)
* fix: apply Chinese-aware title matching in the four index resolvers (#798) `chinese_literature_client.py` already carried a Chinese-aware `normalize_cn_title` / `has_cjk`, but the four index resolvers (Semantic Scholar / OpenAlex / Crossref / arXiv) read the ASCII-centric helpers in `_text_similarity.py`, where `.lower()` folds case but never width (P U+FF30 never reaches P U+0050) and `string.punctuation` contains none of `。`, `《》`, or U+3000. A real Chinese paper served by an index in a different-but-legitimate typesetting therefore missed on two paths: the DOI-keyed cross-check, which gates on the fuzzy ratio alone and scored a fullwidth spelling of the identical title at 0.625 (under the 0.70 floor) reporting a correct DOI as DOI_MISMATCH; and the title-fallback search, which requires ratio AND exact equality and so fell to `unresolvable`. Both feed the `*_unmatched` contamination signals, so a genuine paper could render as CONTAMINATED-TRIANGULATION-UNMATCHED. Promotes `has_cjk` / `normalize_cn_title` into `_text_similarity.py` byte-identical (the CJK client now re-imports rather than keeping a private copy, per the #128 anti-drift goal), adds the Chinese-aware form to `exact_normalized_title` as an additive third branch, and folds it into `_similarity` through the existing `max`. Both gated on BOTH sides carrying a Han ideograph, so every non-CJK verdict is provably unchanged — pinned by an oracle test restating the pre-fix formula in full. 31 new tests, each verified to fail against the pre-fix module. Full suite: 9255 passed, 3 skipped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011qexY5ysaqaAyPp97byf4w * test: force the ratio-independence and non-destructiveness proofs (#798 review) Addresses the three requested changes on PR #799. 1. `_cn_titles_match` ratio-independence is now forced, not inferred, in the test that claims it. `test_legitimate_variants_match_despite_a_sub_threshold_ fuzzy_ratio` asserted the match on a pair the repaired `_similarity` scores 1.0, so a regression that ANDed the ratio back in as a necessary condition would still have passed. Its match assertions now run inside a `monkeypatch.context()` with `_similarity` replaced by a detonator, scoped so the ratio measurements above it still see the real function. A new `test_cn_titles_match_never_consults_the_fuzzy_ratio` adds the negative half under the same forced conditions, so the invariant cannot be satisfied by a helper that has stopped discriminating. The shared `_forbid_similarity` helper patches BOTH binding paths: the `_text_similarity` module attribute (a qualified call or lazy in-function import) and the client's own namespace (a module-level `from ... import _similarity`, already bound and blind to the first patch). Both styles were mutation-verified to trip it; before this change the named test passed the regression that the new one caught. 2. `test_ratio_never_lowered_off_the_cjk_path` asserted `>=` against the base ratio alone, so it passed a *raised* non-CJK score and never exercised the dotted-acronym branch. Renamed to `test_ratio_unchanged_off_the_cjk_path` and rewritten against a full `_pre_fix_similarity` oracle — the companion to the existing `_pre_fix_exact_normalized_title`, written out in full for the same anti-drift reason — asserting exact equality. Mutation-verified twice: one raising a base-branch score, one confined to the acronym branch; the old assertion caught neither. 3. "Byte-identical" corrected to "behaviorally equivalent" in the `normalize_cn_title` docstring and the CHANGELOG. The promotion hoists the wrapper/terminal-mark sets to module constants, precompiles the regex, and rewrites comments; behavioral equivalence is what the tests actually pin. CHANGELOG test count corrected 31 -> 32 and its oracle sentence updated to describe both oracles. Full suite: 9256 passed, 3 skipped (+1 test). The pre-existing `test_evidence_rows.py::test_gfm_bare_urls_emails_and_schemes_cannot_autolink` failure is unchanged and also fails on clean main. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011qexY5ysaqaAyPp97byf4w --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
127ff85e4b |
docs(release): v3.21.1 — bounded workflow substrates and transport hardening [skip-closes-check] (#797)
Promote the accumulated Unreleased changes, align all version-bearing surfaces and five localized README summaries, backfill merge coverage, and preserve the default-off/design-only/evidence-bound claim ceilings. Validated with 9225 passed, 3 skipped, 267 subtests; pre-tag version, changelog, spec, content-lock, distribution-claim, privacy, and staged secret checks all pass.v3.21.1 |
||
|
|
088d288a5e |
feat: add inquiry ledger, alternative-register design, and proving set (#796)
* feat: add opt-in inquiry branch ledger (#743) * docs: freeze alternative explanation register design (#744) * feat: add source-backed review criteria proving set (#575) |
||
|
|
385bc064e1 |
feat: add sealed bakeoff and workflow profile contracts (#795)
Implement the #789 sealed promotion-bakeoff lifecycle and the #742 research workflow profile contract. Consolidate the #794 Markdown link/anchor grammar, record #684 expert-stage readiness, and add the #575 closure-scope audit. |
||
|
|
7ef93e0cb5 |
docs: re-derive data_access_level for academic-paper and academic-paper-reviewer (#773) (#793)
* docs: re-derive data_access_level for academic-paper and academic-paper-reviewer under the dirtiest-input rule (#773) Applying the #756 derivation to the two carried-over pins the lint docstring flagged as un-derived: - academic-paper: redacted -> raw. Standalone modes ingest ungated user drafts and third-party reviewer comments, and literature_strategist's search-fills-gap flow ingests external-index search results inside the skill. The former value described only the post-Gate-2.5 pipeline path. - academic-paper-reviewer: verified_only -> raw. The standalone /ars-reviewer entry (Routing Step 1 routes "review my paper" directly) legitimately consumes an ungated pasted manuscript; the rule quantifies over ALL entry paths. Pipeline positioning unchanged. - deep-research: raw survives by a ceiling argument (no derivation can dirty the dirtiest value); recorded so no pin remains an un-derived carryover. EXPECTED_LEVELS provenance note rewritten per-pin; ARCHITECTURE §4 diagram + rules now separate the per-skill intake annotation from the per-stage output data level (§3 column, unchanged). Declarative only. Closes #773 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq * review: address codex findings on #773 — precise gate-sequencing claims, deep-research derivation - academic-paper's former 'redacted' is described as the orchestrated pipeline path (Stage-1 sanitized inputs), not "post-Gate-2.5" — Stage 2 precedes Gate 2.5. - academic-paper-reviewer's former 'verified_only' is stated as at best true for the initial Stage 3 dispatch; Stage 3' re-review consumes a freshly revised manuscript before Stage 4.5. - deep-research's raw is re-affirmed on its actual inputs (raw queries + unverified search results); the ceiling argument becomes supplementary rather than the derivation itself. Applied consistently across the lint docstring, ARCHITECTURE §4, and the CHANGELOG entry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
adc38300da |
feat: register write-scope guard launcher degradations in degradation_registry.json (#769) (#792)
* feat: register write-scope guard launcher degradations in the degradation registry (#769) The registry claims to index every graceful-degradation mechanism in the suite, but hooks/run_guard.sh's documented degraded states had no rows, and the #757 prose table in docs/CONTROL_AVAILABILITY.md stood up a second, unpinned authority for those facts. Four write_scope_guard_* rows added (no-python, no-git-bash, no-timeout-binary, subprocess-misbehaves), each with verbatim D3 authority anchors into hooks/run_guard.sh + the README Requirements bullet; pinned_by names scripts/test_run_guard_launcher.py where a CI-executable pin exists (the Windows-without-Git-Bash path never executes the launcher, so its row honestly carries no pin). Registry 1.2.0 -> 1.3.0; _EXPECTED_MECHANISMS updated in the same commit (D5 lock semantics). The CONTROL_AVAILABILITY degradations table now declares itself a convenience summary backpointing at the registry. Closes #769 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq * review: address codex findings on #769 — permission phrasing, no-timeout decision forwarding, launcher-internal failure coverage - Rows no longer claim "writes are never blocked" or relitigate what an 'allow' decision would do: the launcher emits no permissionDecision, so the session's normal permission rules still decide. - The no-timeout row's terminal_policy_effect states that the healthy watchdog fallback forwards the guard's real decision (including deny); only an overrun resolves to pass-through. - The misbehaves row now also covers the two remaining documented launcher-internal degradations (SELF_DIR self-resolution failure and the POSIX payload-length cap on multi-megabyte Writes), with verbatim anchors; CHANGELOG + CONTROL_AVAILABILITY backpointer updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq * review: round-2 codex findings on #769 — per-row quantifiers, actual validity-check shape, payload-edge honesty - terminal_policy_effect now speaks per failure path, not "every degraded path"; the no-Git-Bash row states the hook simply does not run. - The misbehaves row names the launcher's ACTUAL validity check (a JSON object carrying a top-level hookSpecificOutput key — deliberately shallow, not full hook-schema validation). - The multi-megabyte payload edge is recorded as a documented accepted, untested case with no pinned outcome — no deterministic claim. - The no_python row's authority anchor swaps to the no-permissionDecision pass-through line (the launcher's disputed 'allow' comment is pre-existing text this PR neither adds nor endorses). - CONTROL_AVAILABILITY prose quantifier fixed to match ("none of which ever blocks", with the no-timeout forwarding stated). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq * review: round-3 codex findings on #769 — payload edge split into its own no-pinned-outcome row - write_scope_guard_payload_capacity_edge becomes a dedicated row whose every field honestly declares "no pinned outcome" — the misbehaves row's pass-through claims are now unconditionally true for its own failure classes (registry 20 -> 21 rows, D5 lock updated). - CONTROL_AVAILABILITY prose reworded: degraded states never INTRODUCE a block; the no-timeout swap keeps the guard operating normally (real decisions, including deny, still apply). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
edb0265301 |
refactor: consolidate markdown-stripping helpers into scripts/_markdown_lint_util.py (#771) (#791)
* refactor: consolidate the duplicated markdown-stripping helpers into scripts/_markdown_lint_util.py (#771) The #757/#758 defrift lints shipped two diverging copies of the markdown non-rendering semantics; check_risk_register.py was already importing the siblings' private helpers as a stopgap. One shared module now carries the grammar; per #771 the #770 superset rules (inline code-span stripping + image exclusion in the link grammar) win, so CA-1..CA-3 inherit them too. Pure refactor, no invariant change. All three mutation suites (73 tests) pass unchanged against the shared module; the three lints pass on the real tree. Closes #771 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq * review: address codex findings on #771 — honest CHANGELOG framing, #759 attribution, three CA grammar tests - CHANGELOG no longer calls the change a pure refactor: the CA-side code-span/image grammar alignment is behavior-visible (strengthening direction) and is stated as such. - _markdown_lint_util docstring attributes check_risk_register to #759 (it shipped there; #760 is the governance change). - Three new CA mutation tests pin the inherited rules: broken image target / backticked pseudo-link do not fire CA-1; an image form of the README inbound link does not satisfy CA-3 (73 -> 76 tests). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq * simplify: complete the #771 consolidation per four-angle cleanup review - github_slug / heading_slugs move into _markdown_lint_util.py — the last cross-lint private import (check_risk_register -> check_control_ availability._heading_slugs) is gone; each lint now imports only the shared module plus genuinely data-owning dependencies. - New links_to() predicate absorbs the three copy-pasted CA-3/DF-3/RR-3 inbound-link loops; link_targets()/code_spans() named wrappers replace raw regex exports (regexes back to module-private); NON_RELATIVE_LINK_ PREFIXES replaces repeated literal tuples. - Fence state machine collapses to a single `fence` variable; helper strip passes back to module-private; extract_link_targets uses findall. - The grammar gains a direct, manifest-registered test suite (test__markdown_lint_util.py); the two CA tests that re-pinned grammar already pinned in the DF/RR suites are dropped, keeping the CA-3 image-exclusion mutation test. Consolidation history now lives in the CHANGELOG once instead of four prose sites. 87 tests green across the four suites; all three lints PASS on the tree. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq * docs: qualify the cross-lint import claim (codex P3) — markdown-helper imports only Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b6062c1401 |
feat: first Promotion Bakeoff run — gpt-5.6-sol validated for the codex subscription transport (#788)
* feat: first Promotion Bakeoff run — gpt-5.6-sol validated for the codex subscription transport (#787) Probe set: 30 refs (10 easy DOI-keyed journal articles; 10 hard: 3 arXiv, 2 DOI-less NeurIPS, 5 non-English; 10 fabrications), every real row resolver-confirmed same-day, every fabrication negative-checked. 180 same-day paired calls (30 x 3 repeats x 2 models), majority verdicts. Result: all five measures PASS with superiority — recall 1.00 vs 0.80, grounded completion 0.933 vs 0.900, p95 latency 26.5s vs 58.7s, zero guard misfires, false disagreement 0.00 = 0.00. Transport-qualified: gpt-5.6-sol stays provisional on the first-party API route (jq guards unexercised; allowlist unchanged). Report + probe-set sha256 under audits/; per-call index committed beside the probe set. Campaign side-product (transport): page-open webSearch items (action.type != "search") are skipped for binding instead of failing the stream (opened-page URLs still can never become bound sources), and DEVELOPER_INSTRUCTIONS requires an empty sources array for NOT_FOUND/NOT_SEARCHED. 52 transport tests green. Defective-tool run 1 archived unscored; three probe-row transcription errors were flagged MISMATCH by both models, independently re-verified, corrected, re-run. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j * fix: narrow the page-open exemption to the observed action.type == "other" shape (#788 codex P2) An empty action object, unknown action type, or non-dict action on a webSearch item is stream-fatal again; only the observed page-open shape is skipped. Mutation test sweeps four bad shapes (52 -> 53 tests). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j * fix: anchor the page-open exemption to the first-party closed WebSearchAction set Run-3 surfaced a third real shape ({"type": "openPage", "url": ...}) that the single-observation exemption rejected, tool-suppressing the baseline's measures (13 EVENT_STREAM_INVALID cells). The exempt set is now exactly the non-search members of the app-server protocol's closed WebSearchAction oneOf — {other, openPage, findInPage} plus the Responses-API spellings — verified against `codex app-server generate-json-schema` on 0.147.0. Unknown shapes stay stream-fatal (mutation sweep unchanged); 54 tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j * docs: score preregistered run 4 as the gate result; runs 1-3 recorded as exploratory Run 4 (frozen fixture @ |
||
|
|
5714f3a3eb |
fix: repair codex subscription transport against three codex-cli 0.147.0 drifts (#785) (#786)
1. Attestation stream: `codex login status` emits "Logged in using ChatGPT" on stderr in non-TTY invocation; detection read stdout only, so every detect failed AUTH_NOT_CHATGPT_SUBSCRIPTION. Accept the exact line on either stream (same idiom as the #684 harness); stderr-emitting fake-codex regression test added. 2. Structured-output schema: the provider now rejects "uniqueItems" (invalid_json_schema, HTTP 400). Dropped from the provider-sent MODEL_OUTPUT_SCHEMA; the local validator already refuses duplicated source URLs fail-closed. 3. Disable list: --disable code_mode_host silently removes the standalone web-search tool on this build (search executes through the code-mode host; isolated by live bisection), failing every call closed as MODEL_RETURNED_NOT_SEARCHED. The host is no longer disabled; code_mode stays disabled and the forbidden-event scan still rejects any item type outside the {userMessage, reasoning, agentMessage, webSearch} allowlist. Live cross_model_smoke_test_codex.sh: PASS for gpt-5.5 and gpt-5.6-sol (2026-08-19). 51 transport tests green; #630 guard green. Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
075390a5c3 |
docs: cross-model recommendation surfaces follow generation currency (#784)
* docs: cross-model recommendation surfaces follow generation currency (#783) Recommendation decoupled from validation status: gpt-5.6-sol (current OpenAI flagship) becomes the lead OpenAI example while staying provisional — a dated lifecycle note records the flip carries no measurement claim; the Promotion Bakeoff remains the only route to validated. gpt-5.5 / gpt-5.5-pro demoted to validated previous-generation rows (measured bakeoff baseline unchanged). Gemini 3.1 Pro stays recommended (first-party check 2026-08-19: still Google's most capable Pro model). SETUP en/zh-TW quick-setup blocks updated in lockstep (parity lint green); the #630 guard's recommendation witness re-pinned to the new policy sentence with its mutation test updated in the same commit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j * fix: close codex round-1 P2s — evidence-ceiling wording + policy-body witness P2-1: "measured bakeoff baseline" overstated the evidence (no bakeoff run has ever been recorded); gpt-5.5 is the designated baseline, validated = allowlist status only. Reworded on all four surfaces (canonical doc, SETUP en/zh-TW, CHANGELOG). P2-2: the #630 recommendation-policy witness pinned only the heading; the guard now pins the two load-bearing body clauses (no-measurement-claim, bakeoff-only route to validated) with mutation tests for each (29 -> 31). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j * fix: close codex round-2 P2 — superiority claim requires an observed measure The rewritten outcome bullet's "or operational benefit" branch let a measured-superiority claim rest on an unmeasured benefit; superiority now requires observed superiority on one of the five measures, and operational benefits are scoped to recommendation policy. Header no longer says "two distinct promotions" for what is now one promotion plus a claim rule. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j * docs: close codex round-3 P2 — name the subscription-transport exception The citation-only codex subscription blocks (canonical + SETUP en/zh-TW) keep gpt-5.5 deliberately; a comment now names this as a transport-specific exception to the generation-currency recommendation and points at the codex smoke test before swapping ids. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3f14c8e16f |
Add OrcaRouter to THIRD_PARTY.md community directory (#782)
Adds a listing row under 'Listed projects' for OrcaRouter, an OpenAI- and Anthropic-compatible gateway that can serve as the cross-model verification provider via ARS_OPENAI_COMPAT_BASE_URL + ARS_CROSS_MODEL with namespaced model IDs. Credits ARS and describes the integration faithfully, per the 'How to get listed' bar. Closes #781. Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
2b639c12ee |
docs(release): v3.21.0 — ISO/IEC 42001-spirit track [skip-closes-check] (#779)
* docs(release): v3.21.0 — ISO/IEC 42001-spirit track [skip-closes-check] - CHANGELOG: promote [Unreleased] -> 3.21.0 (2026-08-18); backfill the six uncovered merges (#762 assessment, #757/#768, #758/#770, #755/#774, #754/PR #763 ref, #764 hotfix); fresh empty [Unreleased] anchor - version bump across all pinned surfaces: 5 READMEs (badge, pipeline section, new v3.21.0 what's-new block), MODE_REGISTRY, .claude/CLAUDE.md (suite version, Last Updated, v3.21.0 Key Additions), CITATION.cff (version + date-released), POSITIONING citation prose, plugin.json, marketplace.json, academic-pipeline/SKILL.md (+#528 content-lock rehash), ARCHITECTURE (7 current-component markers + Last Updated), check_spec_consistency pins + test fixtures - READMEs x5: live pipeline-guarantee sentence reworded from 'cannot be skipped' to the #753 matrix-licensed form (MANDATORY, no unrecorded bypass, override reasoning recorded); historical release-notes sections untouched Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC * docs(release): fix stale-tense sentence in #754 changelog entry (codex R1) [skip-closes-check] Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>v3.21.0 |
||
|
|
4b8427f6d5 |
docs: solo-maintainer governance statement + SECURITY triage procedure (#760) (#778)
* docs: solo-maintainer governance statement + SECURITY triage procedure (#760) GOVERNANCE.md: decision authority, honest cross-model scope (error-detection control, not organizational independence), release authority, EOL posture, and the operating-principles section (three distilled principles with informative ISO/IEC 42001 anchors, the #753-#760 coverage map, Annex C not-applicable assessments). SECURITY.md: 7-day promise becomes acknowledgement-only hard promise plus a written severity-tiered best-effort triage procedure. NOTICE.md: governance pointer. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC * docs: apply dual-agent review to #760 (order-safe claims, honest ceilings, sustainable promises) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC * docs: close codex R1 on #760 (per-operator cross-model attribution, fairness signals disclosed, EOL-scoped promises) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC * docs: narrow fairness scoring absolute to persons/real-world allocation (codex R2) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c5f0ac69af |
feat: lightweight risk register + RR-1..3 mirroring lint (#759) (#777)
* feat: lightweight risk register + RR-1..3 mirroring lint (#759) docs/RISK_REGISTER.md links ten standing risks to existing controls, matrix-mirrored evidence statuses, and residual gaps with tracking issues. scripts/check_risk_register.py pins pointer integrity (RR-1), verbatim status mirroring against stage_capability_matrix.json with a malformed-citation guard (RR-2), and README discoverability (RR-3); 12 mutation tests, wired into spec-consistency CI and the pytest manifest. Helpers imported from check_data_flows / check_stage_capability_matrix rather than copied (#771). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC * refactor: apply /simplify review to #759 lint (reuse sibling helpers, close side-doors) - import _LINK_RE/_CODE_SPAN_RE/_strip_non_rendering from check_data_flows (capture group added there, harmless for its .sub use), _load/_status from check_stage_capability_matrix, github_slug/_heading_slugs from check_control_availability - no third copies (#771) - RR-1: repo-containment + anchor-fragment validation (CA-1 parity) - RR-2: asserted-status ceiling (matrix is sole authority for MEASURED/MIXED), inventory lock on shipped matrix-row citations (D5/M13 house pattern), per-segment malformed-citation reporting - RR-3: resolved-path + code-span-stripped inbound link (DF-3 parity) - missing-doc single fatal error (sibling contract); ERROR: prefix in main - tests: 12 -> 18 incl. real-tree pass; _MIRRORED_FILES derived from imported constants Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC * fix: anchor-on-non-markdown-target guard (CA-1 parity) + test; count fixes Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC * fix: close codex R1 findings (lock scoped to evidence bullets, space-path opt-out removed, R9 pin wording) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC * docs: align R10 residual-gap wording with the availability matrix (codex R2) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
17bf063456 |
docs: CI workflow enforcement-class table + inventory lint (#755) (#774)
* docs: CI workflow enforcement-class table + WC-1/WC-2 lint (#755) docs/ARCHITECTURE.md gains §7.1: all 14 workflows classified by trigger / what it checks / enforcement class (blocking / advisory / administrative / post-push detection) / bypass token, with the honest count line (8 blocking on at least one event class, 2 advisory, 1 administrative, 3 post-push detection) and the explicit statement that tag workflows detect after the push — their stop-power is the maintainer acting on the failure. Per-workflow facts verified against the workflow files (eval-harness ack token + PR-only gating; changelog gate release/** head scope; pytest path filters; the three tag triggers). CONTRIBUTING release-checklist prose now points at the classification instead of implying uniform CI enforcement. Lint (same-PR drift-point discipline): check_workflow_classification.py — WC-1 inventory sync both directions (a new, renamed, or removed workflow fails CI until the table matches; duplicates refused), WC-2 class cells begin with the closed four-term vocabulary. Class semantics stay review-owned (degradation-registry posture). 9 mutation tests; wired into spec-consistency.yml + the pytest manifest. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * refactor: apply /simplify + codex R1 — table accuracy + lint hardening (#755) Review round (3 cleanup agents + codex gpt-5.6-sol xhigh R1), findings deduped and applied: Table accuracy (codex 3 P2 + 1 P3, cleanup F2/F3): every trigger cell now states its actual branch/path/tag filters (repository-hygiene and command-invariants had birth-drifted cells; several rows omitted targeting-main scopes); freshness-check reclassified honestly (Advisory for staleness, but malformed protocol metadata is a hard failure); bypass cells say "justification requested, not machine-validated" (both workflows accept the bare token); command-invariants "what it checks" gains its other two enforced checks; bypass column normalized to "none"; the legend absorbs the tag-workflows sentence and the duplicated qualifier prose is trimmed. Lint hardening: section extraction switches to the shared _skill_lint.heading_section (exact full-line heading incl. the #755 anchor, fence-aware — 15 fewer bespoke lines); rows parse once with escaped-pipe-aware cell splitting; the inventory glob covers *.yaml; WC-2 matches vocabulary terms as whole words (Blockingg fails); the arity guard moves under WC-1 with a test; new WC-3 recomputes the bolded count line from the Class column (the honesty sentence can no longer self-invalidate when a workflow is added); new WC-4 pins every [bypass-token] in a Bypass cell to verbatim presence in its workflow file. 14 mutation tests. Surfaces: docs/CONTROL_AVAILABILITY.md corrects its "on every change" claim and links §7.1; the ARCHITECTURE "How to read" §7 bullet indexes the CI sub-view. Skipped with reason: read_or_exit2 exit-2 convention (sibling lints in this fleet use the exit-1 missing-doc violation shape; consistency wins). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R2 findings — .yaml fixture parity + comment-blind WC-4 (#755) - The mutation fixture copies *.yaml alongside *.yml, so a future .yaml workflow with a valid row passes the fixture as it passes the real lint. - WC-4 strips full-comment lines before the token search: a renamed executable token surviving only in a YAML comment no longer satisfies the pin (token in a non-comment echo/log string recorded as an accepted edge). Mutation test added (15 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R3 finding — tag pushes reach three more workflows (#755) GitHub Actions matches tag pushes on unfiltered or paths-only push: triggers (paths filters are not evaluated for tags), so spec-consistency, command-invariants, and freshness-check also run on every v* tag push — where their failures are post-push detection like the tag-only workflows. Trigger cells amended and a subtlety note added above the table; "three tag workflows" narrowed to "three tag-only workflows". Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R4 finding — malformed token spellings fail loudly (#755) Any bracketed span in a Bypass cell must be a well-formed [lowercase-hyphen] token: a typo like [skip_cooldown] now yields a WC-4 violation instead of silently falling out of the token grammar. Mutation test added (16 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R5 finding — whitespace token typos caught (#755) The any-bracket span matcher now accepts any non-] content, so [skip cooldown] (space typo) reaches the well-formedness check and fails loudly. Mutation test added (17 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R6 finding — bogus rows fail instead of dropping out (#755) Every pipe row in the section that is not the header or the separator must open with a backticked workflow filename; a malformed row now yields a WC-1 violation instead of silently leaving the inventory and the WC-3 count. Mutation test added (18 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * test: mirror docs/ARCHITECTURE.md into the CA fixture (#755) The new CONTROL_AVAILABILITY link to ARCHITECTURE §7.1 made the #768 test fixture (which mirrors only the files the doc links) miss its target, failing CA-1 in the fixture tree while the real tree passes — caught by CI, not locally, because the local sweep re-ran the lint but not its sibling test file. ARCHITECTURE.md joins the mirrored list. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9ccf4a9c9f |
fix: academic-pipeline data_access_level verified_only → raw (#756) (#772)
* fix: academic-pipeline data_access_level verified_only -> raw (#756) The orchestrator legitimately consumes raw input — Stage 1 accepts raw user requests, mid-entry accepts raw existing papers — and the governing dirtiest-input rule (ground_truth_isolation_pattern.md) requires the annotation to reflect that. The integrity gates run INSIDE the pipeline, downstream of its intake, so verified_only was internally inconsistent on the suite's most prominent consumer (option (a) of the issue: honest minimal relabel; no trust-domain split). Surfaces aligned in the same commit: - academic-pipeline/SKILL.md frontmatter (its #528 content-lock sha256 recomputed in check_pipeline_boundary_semantics.py, same commit per the lock discipline). - docs/ARCHITECTURE.md §4: pipeline node moves to the raw class, the User -> pipeline intake edge is drawn, and the rules block states the dirtiest-input rationale with a pointer to the pins. - check_data_access_level.py grows an EXPECTED_LEVELS per-skill pin layer (acceptance criterion 2): a silent flip back to verified_only, an unregistered new skill, or a stale pin now fails CI; vocabulary check unchanged. 7 mutation tests + manifest entry (152). Not touched: CHANGELOG history (records what v3.x declared at the time); shared/agents/compliance_agent.md (agent-level declaration, runs at the gates); academic-paper-reviewer verified_only (possible same-class question for standalone /ars-reviewer raw-paper input — out of #756 scope, reported separately). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: restore the pre-existing CLI-level test layer I overwrote (#756) The previous commit replaced scripts/test_check_data_access_level.py wholesale, dropping six original unittest cases (CLI subprocess via --path, including the three malformed-frontmatter stdout-reporting contracts) and breaking the run_skill_linter --path interface by removing argparse from the lint. Both restored: main takes --path again, the original unittest class is back (its valid-root case now builds the four registered skills, since the pin layer correctly rejects an unregistered synthetic skill), and the #756 pin-layer mutation tests ride alongside. 11 tests green; manifest entry verified through the CI runner. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * refactor: apply /simplify + codex R1 — single-pass lint, honest pins, aligned surfaces (#756) Review round (3 cleanup agents + codex gpt-5.6-sol xhigh R1), findings deduped and applied: - check_data_access_level.py rewritten as a SINGLE pass: one violation per problem (the vocabulary layer had zero unique failure coverage and double-reported every failure, including a twice-printed YAML traceback); non-mapping metadata (e.g. "metadata: active") is now a reported violation instead of an AttributeError crash (codex P2); LEGAL_VALUES survives as a pin-vocabulary assertion; main() stays local (run_lint no longer fits once check_metadata_field drops out) and run_lint's stale "both check scripts" docstring is corrected. - Pin provenance honesty: the docstring now says only the academic-pipeline pin is #756-derived; the other three freeze pre-existing declarations against silent drift. Follow-up derivation for reviewer/paper opened as #773. - ARCHITECTURE §4: rule restatement dropped (the pattern doc owns the rule), the two competing one-line lint descriptions merged into one, and the §2 legend disambiguates §3's per-stage "Data level" column from the skill-level declaration (the four VERIFIED_ONLY stage cells are a different, per-stage claim — left as-is). - CHANGELOG [Unreleased] gains the #756 Fixed entry (the pre-tag covers-merges gate is fail-closed). - handoff_schemas.md data_access_level block now names the pin layer. - write_skill fixture helper migrated to tests/test_helpers.py (migrate-at-next-edit convention); both skill-lint test files import it; CLI scenarios updated to registered skill names (the pin layer correctly pre-empts unregistered synthetic skills); new one-violation-per-problem and non-mapping-metadata regression tests (18 green across both files). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: restore the #753 CHANGELOG bullet heading (codex R2) (#756) The #756 entry insertion had consumed the #753 bullet opening and absorbed its body into the new bullet; the #753 heading is restored as its own bullet. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: legend example had the gate boundary reversed (codex R3) (#756) Gates consume unverified drafts and PRODUCE verified artifacts; the Data-level column is documented as a postcondition on stage outputs, not material the gate "operates on". Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
cdd48d916d |
docs: DATA_FLOWS.md — single map of network touchpoints + local stores (#758) (#770)
* docs: single data-flow map + DF-1..DF-3 coverage lint (#758) Add docs/DATA_FLOWS.md — one row per network touchpoint (trigger, payload class, recipient, credentials, off switch) and one row per local store (path, content, TTL, deletion), with an explicit scope statement (the Claude session itself is platform-governed; nothing publishes autonomously). Covers the four gate resolvers, the standalone Chinese-literature resolver (NOT in the gate), the consent-bound claim-standing discovery adapters, both cross-model transports (API and citation-only Codex subscription), the SessionStart update check, the manual smoke tests, and the v3.9.4 timeline bootstrap — the last one surfaced by the new lint itself on first run (it was absent from the #758 issue enumeration). Inbound links from README, SECURITY.md (in-scope exfiltration anchor), and THIRD_PARTY.md (core-suite vs third-party contrast). Lint (same-PR drift-point discipline): scripts/check_data_flows.py — DF-1 every non-test scripts/*.py importing a network module (AST scan, so no-call guards naming urllib.request in strings do not count) must be named on the map; DF-2 same for curl-invoking shell scripts; DF-3 README/SECURITY/THIRD_PARTY keep a rendered resolving inbound link (fences + HTML comments stripped with the semantics converged in the PR #768 review; consolidation into a shared helper is follow-up). 12 mutation tests; wired into spec-consistency.yml + pytest manifest. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * refactor: apply /simplify pass (4-agent, deduped) (#758) Doc: the four gate resolvers collapse into a 4-column sub-table under one shared trigger/payload/off-switch lead (the wide table kept only heterogeneous touchpoints); the exhaustiveness sentence is bounded to what DF-1/DF-2 actually detect (direct imports + curl; spawned-CLI and session-tooling paths held by review); "Nothing here publishes" now inherits POSITIONING.md and its not-a-runtime-guarantee qualifier; the subscription-free note is trimmed to its rationale; Related gains the SETUP bullet as the tunables authority. Coverage: docs/SETUP.md becomes the fourth DF-3-pinned inbound surface (pointer added in the cache section); the four translated READMEs mirror the README pointer; docs/DATA_FLOWS.md registers into check_spec_consistency.py relative-link validation. Lint: DF-1 module vocabulary rebuilt as the network subset of the no-call envelope FORBIDDEN_IMPORTS (deviations documented: dotted urllib.request/http.client instead of bare urllib/http; ssl excluded); scan is now recursive into scripts/ subpackages. Tunable constants in verification_cache.py gain update-both comments. Tests: the three hollow assert-baseline tests become real mutations (name-based test exemption, uncomment-curl, from-urllib idiom); recursive-scan and SETUP-surface tests added (17 total). Skipped with reason: endpoint-hostname lint (near-zero event rate, composed-URL false-fire risk); row-id shrink constant (review-owned per degradation-registry precedent); markdown-helper consolidation with check_control_availability.py (whichever PR merges second extracts the shared module — recorded in both PR bodies). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R1 findings — 8 P2 + 4 P3 (#758) Doc accuracy (6): Chinese-literature row rewritten (callable client, no CLI; PubMed path sends the required NCBI contact email + bibliographic search coordinates); codex-transport payload names citation_context (can contain unpublished manuscript text); update check documented as one curl transfer per 24 h with redirects and the ARS_UPDATE_CHECK_REMOTE_URL override; retraction-status SQLite cache added to local stores (caller-supplied path, 30-day stale threshold, no auto-expiry); discovery adapters credentials corrected (fixed User-Agent, resolver env keys not consumed); resolver payload narrowed to identifiers + title query strings; update-check state content corrected (state label + two version strings). Lint mis-pass/mis-fire (4): DF-1/DF-2 coverage now requires the full repo-relative path (basename-substring collision closed); DF-2 is recursive over scripts/ and hooks/, recognizes path-qualified curl, and masks quoted spans before the comment strip; DF-3 strips inline code spans before link extraction (a backticked link does not render). Five mutation tests added (22 total). The code-span rule is a divergence from check_control_availability.py to be carried over at the declared helper consolidation. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R2 findings — 4 P2 + 1 P3 (#758) - DF-2 scans command-substitution bodies BEFORE quote masking, so resp="$(curl ...)" — a real network call inside double quotes — fires (mutation test added; suite now genuinely 22, correcting the prior commit message which said 22 when 21 were collected). - Map gains the Codex audit wrapper row (scripts/run_codex_audit.sh: human/CI/hook-invoked only, sends deliverable + supporting file contents through the local Codex CLI login). - Update check re-bounded: at most one SUCCESSFUL check per 24 h; a failed attempt writes no state and may retry next session. - Cache TTL wording corrected: expiry is a cache miss, not deletion; expired rows persist until invalidated or the file is deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: full comment lines execute nothing — DF-2 substitution scan (#758) The R2 command-substitution scan ran before any comment handling, so a full comment line containing $(curl ...) false-fired — surfaced by the codex R3 pass (timed out mid-review, but its transcript had already demonstrated the false fire). Comment-only lines are now skipped before the substitution scan; a $(curl) inside a trailing inline comment remains a documented accepted edge. Mutation test added (23 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R3 findings — command-position curl + image links + retention wording (#758) - DF-2 rebuilt around COMMAND POSITION: curl counts only as the first non-assignment token of a segment (pipes/separators/substitution openers), so `command -v curl` preflights and `echo curl` no longer false-fire; VAR=x curl still fires; wrapper-prefixed invocations (sudo/timeout) are documented accepted edges. - DF-3 link grammar excludes image syntax —  renders no anchor and cannot keep the acceptance criterion green. - Cache retention wording includes the overwrite path: expired rows persist until overwritten by re-verification, invalidated, or the file is deleted. - Three mutation tests added (26 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R4 finding — curl behind shell control words (#758) The command-position head-token scan now skips shell control words (if/elif/while/until/then/else/do/!/time/exec) before naming the head, so `if curl …; then` and `while ! curl …; do` fire while `if true; then` stays quiet. Two mutation tests (28 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R5 finding — option tokens after control words (#758) `time -p curl …` / option-bearing exec forms: the head scan now skips `-`-prefixed option tokens alongside assignments and control words, so the option cannot shadow the command head. Mutation test added (29 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R6 finding — harness spawned-CLI paths scoped out (#758) Three maintainer-only measurement scripts reach the network through locally authenticated CLIs (dispatch_e4_panel via claude -p, run_review_criteria_constructive_value via Codex, check_ranking_lift via gh api). They are not user-facing feature paths, so instead of diluting the touchpoint tables they are now an explicit named scope exclusion — the exhaustiveness claim no longer silently spans them. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R7 finding — boundary count wording (#758) "Two boundaries" became three after the R6 harness exclusion; the count is removed rather than maintained. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
43a02bf7e2 |
docs: per-channel control-availability matrix (#757) (#768)
* docs: per-channel control-availability matrix (#757) Add docs/CONTROL_AVAILABILITY.md — one row per enforcement mechanism, one column per install channel (plugin / skills copy / repo clone / Cowork / claude.ai Project / Claude Science / Pi), with honest active / conditional / absent cells, per-channel notes citing the existing scattered sources (README Requirements, SETUP methods, pi/README.md, hooks/run_guard.sh), and the guard's environment degradation table. Linked from README (Requirements + SETUP pointer) and SETUP (Installation methods intro). Evidence re-verified against the working tree: the channel set has grown past the six named in the issue (SETUP now also documents Cowork and the claude.ai 4a/4b split), so the matrix covers all seven documented channels. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * refactor: apply /simplify pass + add CA-1..CA-3 defrift lint (#757) Simplify round (4-agent review, findings deduped): - Drop the 'How to read an integrity claim' section (it had already drifted from the matrix) and the all-identical Upstream row; both replaced by one legend sentence and one paragraph. - Move channel-scoped caveats (Cowork / claude.ai / Claude Science / Pi) from per-cell footnotes into a 'Channel-wide limitation' column of the channel table; notes drop from 11 to 7. - De-drift row labels: no inline allowlist contents (canonical list is pinned by check_tools_allowlist.py), no exhaustive feature list, no hard-coded Claude Code minimum version (lives in SETUP Method 0). - README: single slimmed pointer (second link and both enumerations removed); pointer mirrored to the four translated READMEs and docs/SETUP.zh-TW.md. - Degradations table scoped to actual guard degradations (the slash-form version row was misfiled); registry backpointer added; guard-launcher registry registration split to #769. Lint (per the new-claim-surface-needs-lint-in-same-PR discipline): - scripts/check_control_availability.py — CA-1 links/anchors resolve, CA-2 every SETUP '### Method' heading reachable from the channel table, CA-3 README + SETUP inbound links pinned. Cell semantics stay owned by code review (degradation-registry posture). - 9 mutation tests; wired into spec-consistency.yml + pytest manifest (150 entries). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R1 findings — 4 P2 accuracy corrections (#757) - SessionStart announce/update-reminder row: Conditional, not Active (bash launcher on Windows needs Git Bash; reminder needs curl) — new note 8. - Cross-model note 6 no longer claims credentials+curl universally; the citation-only Codex subscription transport is named as the alternative transport behind the same consent boundary. - Pi channel limitation reworded: the wrapper supplies no orchestration but uses an installed Pi capability when available. - 'Enforcement mechanisms' claim language aligned to 'controls' in the purpose statement and all five README pointers (consistent with note 7's trust-based posture). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R2 findings — lint mis-pass cases + note-8 wording (#757) - CA-1 link grammar accepts optional quoted titles so a titled dead link cannot silently skip the check. - CA-2 counts only fragments on links whose resolved destination IS docs/SETUP.md — a same-slug anchor into a copied file no longer satisfies method coverage. - CA-3 checks resolved link destinations, not a filename substring — a label that keeps the filename while the target moves now fails. - Note 8: singular SessionStart hook (hooks.json defines one; the announce script runs the update check internally). - 3 new mutation tests pinning each mis-pass case (12 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R3 finding — commented-out markdown counts for nothing (#757) Strip HTML comments before extracting links and headings in all three invariants: a commented-out inbound link no longer satisfies CA-3, a commented-out SETUP method heading no longer demands CA-2 coverage, and a commented-out dead link no longer fires CA-1. Two mutation tests pin both directions (14 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R4 finding — GFM type-2 HTML-block semantics (#757) A line beginning with <!-- opens a raw-HTML block through the --> line (including trailing text on the closing line) or to EOF if unclosed; nothing on those lines renders. The comment stripper now models that line-level behavior before the inline-span strip, so a link after --> on a comment line cannot satisfy CA-3 and a dead link after an unclosed comment cannot fire CA-1. Two mutation tests pin both (16 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * test: fix R4 mutation scenario — line-start vs inline comment (#757) The previous commit's CA-3 HTML-block test inserted the comment mid-line (inside the blockquote), where GFM renders the link normally and the lint correctly stays quiet — the test scenario was wrong, not the lint. Replaced with a whole-line mutation that actually begins with <!--, and added the inline-comment symmetry case (link still renders → CA-3 satisfied). 17 tests green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R5 finding — block-quoted HTML-block lines (#757) The type-2 HTML-block rule applies to block-quote content: the stripper now looks through leading '> ' markers before the line-start test, so '> <!-- note --> [link]' cannot satisfy CA-3. Deeper CommonMark laminations are declared out of scope in the docstring (the surfaces do not use them; a full parser is out of proportion for a maintainer-slip guard). 18 tests green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R6 finding — repo-containment on CA-1 targets (#757) A relative link that resolves outside the repository root now fails CA-1 even when the host path exists — an over-deep ../.. slip must not be masked by an existing host file. Mutation test added (19 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R7 finding — fenced code excluded from extraction (#757) Fenced code regions render literally, and README/SETUP use fences today, so they are in-scope: a link inside a fence no longer satisfies CA-3, and a sample "### Method" heading inside a SETUP fence no longer demands CA-2 coverage. Fence stripping runs before the comment pass so a comment opener inside a fence stays literal. Two mutation tests (21 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc * fix: close codex R8 finding — CommonMark fence-length closing rule (#757) The fence stripper now tracks the opening run character and length: a closer must be a same-character run at least that long with only trailing whitespace, so a four-backtick fence demonstrating an inner triple-backtick block is no longer closed early. Mutation test added (22 total). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e9759dc4f4 |
Align distribution-surface claims with evidence ceilings (#753) (#766)
* fix(claims): align distribution-surface claims with evidence ceilings (#753) - plugin.json / marketplace.json: drop 'Production-grade' / '39-agent ensemble' for matrix-licensed wording ('contract-audited', '39 prompt roles (3 plugin-exposed agents; the rest run inline by default)') - academic-pipeline/SKILL.md: no-bypass prose rewritten to the actual mechanism (mandatory checkpoints; overrides require recorded user reasoning); #528 content-lock hash updated in the same commit - shared/cross_model_verification.md: 31%->5-10% relabeled as an unvalidated working hypothesis - shared/ground_truth_isolation_pattern.md: gold-labels rule rewritten to the intended boundary (no unconditional loading into operational agent context) - version-consistency invariant 8: binds 'N prompt roles' spelling too, checks every count token (finditer) - new scripts/check_distribution_surface_claims.py (D1-D5, 20 mutation tests, CI-wired): fail-closed manifest load, shared claim vocabulary imported from check_stage_capability_matrix, percentage refusal, mandatory bindable count token, plugin-exposed count bound to MIRRORS Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki * fix(claims): codex R1 — integrity-family must-PASS sweep + lint case/boundary fixes (#753) - integrity 'must PASS with zero issues' absolutes now name the recorded 3-round FAIL-loop exit (integrity_review_protocol, reinforcement_content, team_collaboration_protocol, integrity_verification_agent, SKILL.md flow row); 'recorded with reasoning' weakened to 'recorded user decision' (rationale escalates per compliance override ladder) - D3 percent check lowercases input (matrix caller parity) - D5 plugin-exposed regex case-insensitive - AGENT_CLAIM_RE gains trailing boundaries (39-agentic / singular 'prompt role' no longer count as bound); 4 new mutation tests (20 -> 24) - SKILL.md #528 content-lock hash rebumped Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki * fix(claims): codex R2 — Stage 2.5 routing parity, passport-state honesty, gold-set scope, strict JSON (#753) - Stage 2.5 flow row + both state-machine checkpoint triggers name the recorded FAIL-loop exit (SKILL.md + pipeline_state_machine.md, both content-lock hashes rebumped) - team protocol handoff checklist: FAIL-loop continuation keeps passport verification_status UNVERIFIED; VERIFIED only on zero-issue PASS - ground-truth gold exception scoped to synthetic/public-safe content; live-reviewer calibration sets stay runtime-supplied - D1 rejects non-standard JSON constants (NaN/Infinity) via parse_constant; 2 new tests (24 -> 26) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki * fix(claims): codex R3 — prerequisite checker + handoff materials accept the recorded FAIL-loop route (#753) - state_tracker_agent prerequisite table: Stage 3 / Stage 5 entry rows accept a recorded Integrity Check FAIL Loop resolution (previously the documented continuation route was unreachable at the checker) - SKILL.md handoff lines 2.5->3 and 4.5->5 no longer mislabel a FAIL-loop continuation draft as verified; team protocol Materials/Approval rows aligned the same way - SKILL.md + state_tracker_agent content-lock hashes rebumped Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki * fix(claims): codex R4 — orchestrator transfer rows + advisory dispatch accept the recorded FAIL-loop route (#753) - orchestrator 2.5->3 and 4.5->5 transfer rows no longer require a 'Verified'-labeled draft on a recorded FAIL-loop continuation - #660/#672 advisory dispatch anchors to the Stage 4.5 terminal resolution (PASS, or recorded FAIL-loop continuation) instead of exact PASS only - orchestrator content-lock hash rebumped Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
704b46d247 |
fix: citation-surface version drift + version-consistency invariant 12 (#763)
* fix: citation-surface version drift + version-consistency invariant 12 (#754) CITATION.cff and POSITIONING.md citation prose sat at 3.14.0 while the suite moved to 3.20.1 (Zenodo v3.20.1 deposit exists — pure metadata drift). Bump both to 3.20.1 and add lint invariant 12 so the drift class fails CI: CITATION.cff version: and every (Version X.Y.Z) token in POSITIONING.md must equal the suite version. Absent file = skip (invariant-8 posture); present-but-versionless CITATION.cff = error. 6 new mutation tests (53 -> 59), full suite green, real-repo lint green. Closes #754 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki * refactor: harden invariant 12 per 4-angle cleanup review - CITATION.cff parsed as YAML (regex scrape misread quoted versions as drift); absence now errors like README.md so deletion cannot silently disable the invariant; version goes through the broad-capture + strict-semver idiom (non-canonical reported as such, not as drift) - date-released gated against the CHANGELOG latest-entry date with the invariant-10 +/-7-day window (the second half of the #754 drift) - POSITIONING token uses broad capture + strict validation so v-prefixed or truncated edits error instead of being filtered (pre-#169 lesson) - regex moved to the numbered constants block; helpers moved to the test fixture header; surfaces wired into _write_aligned_fixture so every pass-case exercises invariant 12; class renamed TestCitationSurfaces - 11 targeted tests (53 -> 64), green under pytest and unittest runners Refs #754 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki * fix: harden invariant 12 date handling per codex review (2 P2) - catch ValueError from PyYAML's timestamp constructor (impossible dates like 2026-02-30 raise ValueError, not YAMLError) so the lint reports instead of crashing - reject datetime values (unquoted timestamps parse as a date SUBCLASS that would TypeError against the date baseline) as not-strict- YYYY-MM-DD errors - 2 regression tests (64 -> 66), both verified red before the fix Refs #754 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki * fix: require date-released in invariant 12 (codex round-2 P2) An absent or null date-released silently skipped the freshness check, so deleting the field disabled the invariant's date half. Absence now errors (same posture as file absence). 1 regression test (66 -> 67), verified red before the fix. Refs #754 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki * fix: close invariant-12 silent-pass paths per codex round 3 (2 of 3 P2) - POSITIONING clause captured whole then stripped + strictly validated, so a malformed edit ('(Version 3.4.0 )') errors instead of dropping out of the capture and passing a stale citation - missing-version regression test isolated from the missing-date error (fixture keeps a valid date-released; asserts the specific diagnostic) - duplicate-YAML-key finding adjudicated NOT-FIX under the threat model: last-wins matches every CFF consumer (GitHub cite widget / Zenodo / cffconvert), so a duplicate-key file renders the identical citation everywhere - untidiness, not drift; documented as a known limitation in the checker docstring - 1 new test (68 total), suite green under both runners Refs #754 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki * fix: catch empty POSITIONING payload + refresh CHANGELOG counts (codex round 4) - capture class widened to [^)]* so '(Version )' reaches strict validation and errors instead of falling out of the capture - CHANGELOG test counts refreshed (53 -> 69) and review provenance updated to the 5-round threat-model-bounded trajectory - 1 regression test, verified red before the fix Refs #754 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
07a2afd94c |
docs(audits): ISO 42001-spirit gap assessment (2026-08-17) (#762)
Dual-track audit (in-session structural review + independent cross-model GPT-5.6 xhigh full-repo audit), all findings re-verified first-party. Scope decision: adopt the standard's spirit (transparency / verifiability / feasibility), explicitly not pursuing certification. Files the verified findings register behind meta-issue #761 (#753-#760) and records the explicitly-not-adopted list. Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f6ffc70312 |
fix(docs): co-locate passport_as_reset_boundary reference in #743 design doc (#764)
The #743 inquiry-branch-ledger design doc mentions ARS_PASSPORT_RESET without referencing passport_as_reset_boundary, so check_passport_reset_contract.py fails on main and blocks every open PR's spec-consistency gate. Add the protocol reference at the mention site per the v3.6.3 co-location contract. Refs #743 [skip-closes-check] CI hotfix restoring a green main; #743 remains open as the feature's tracking issue and must not be auto-closed by a one-line doc-reference fix. Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
946c2494af |
docs(design): #743 inquiry branch ledger design freeze (#752)
* docs(design): #743 inquiry branch ledger design freeze — event-sourced states, adoption receipts, reopen invalidation Freezes inquiry-branch-ledger/1.0: append-only hash-chained event log with deterministic replay, the closed five-status state machine (AI-surfaced facets enter parked and never active), the adoption receipt that keeps AI provenance immutable, author-only reopen with visible stale-not-rewritten invalidation over recorded downstream refs, ARS_INQUIRY_LEDGER opt-in with zero simple-path prompts, additive passport-pointer storage, and the #745 registration gate before any alpha ships. No implementation or evaluation is authorized. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753 * docs(design): #743 review round 1 — facet exit rules, origin-bound receipts, event-sourced staleness, budget as replay invariant, frozen payloads and storage Closes the 7 P1 + 5 P2 codex findings: an unadopted AI facet can only be adopted or rejected (rejected facets terminal), so reopen cannot bypass the adoption receipt; receipts bind source_event_id and the attestation boundary is stated (recorded attestation, not authentication — /ars-mark-read precedent); §3.1 freezes every payload shape with replacement semantics and merge dedup ordering; replay order and canonical form specified, with the passport pointer digest as the trusted head that closes the truncation hole; staleness becomes event-sourced (artifact_marked_stale / reconfirmed / superseded) with first-degree scope stated honestly; the live budget is a replay invariant so the ledger can never record an over-budget state; profile binding is a projection over an immutable initial binding; reopen conditions get stable ids and signals bind exactly one; §7 freezes the passport aggregate shape, atomic writes, closed absence semantics, and the user-workspace/never-committed data boundary; §8 records transport limits and the same-PR CI-gated registration ordering. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753 * docs(design): #743 review round 2 — post-state budget on every event, cause-bound stale resolution, supersession link maintenance, explicit ledger digest, archived status, condition-id identity Closes the round-2 findings: the budget invariant now constrains every event's post-state incl. profile_rebound (lower-budget rebinds require prior dispositions); artifact staleness projects a SET of outstanding causes and clearing events bind resolves_stale_event_id to exactly one cause; artifact_superseded deterministically replaces the retired ref in every listing branch's downstream_refs (first-degree link maintenance); the passport-head digest is defined exactly as SHA-256 over JCS of the ledger document with no placeholder; branch_archived/archived join the closed lifecycle (terminal, facet-exit-lawful); condition_ids are per-branch unique, non-rebindable, and retired permanently on removal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753 --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
07d4833f50 |
feat: #745 stage capability/evidence matrix with enforceable claim ceilings (#751)
* feat: #745 stage capability/evidence matrix with enforceable claim ceilings stage-capability-matrix/1.0 data source + M1-M10 lint (frozen task-family vocabulary shared with the #742 profile contract, non-collapsible evidence statuses, in-repo eval refs, verbatim claim anchors, stale-evidence notes, effectiveness-language discipline on unmeasured rows, byte-pinned generated view), 38 mutation tests, CI + pytest-manifest wiring, contracts README section. Seeded with 13 rows over all nine task families: 4 measured/mixed (incl. the seeded-defect panel's currently-failing severity gate recorded as MIXED) and 9 designed/not-run whose ceilings say exactly that. Refs #745 (matrix + lint slice; the matched stage-substitution evaluation program and README claim-language migration remain open). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753 * fix: #745 review round 1 — report bindings, falsifiable conformance pins, anchor hardening, single-source view path Applies the 4-agent simplify review: M11 measurement-report binding (date-equal to the bound report, accepts the pre-#654 legacy measured_at field, sibling supersession requires a staleness_note) closes the hand-typed-provenance drift the matrix exists to prevent; M12 conformance_pinned_by makes CI_GATED/TESTED falsifiable (D4-style path/function pins, all 13 rows populated); M7 anchors gain containment-checked memoized reads and three rows now bind real README capability sentences; M4 refuses measurement provenance on unmeasured rows in both directions; generated_view leaves the JSON (DEFAULT_VIEW is the single source and --render validates before writing); invalid stale_after_days skips M8 instead of fabricating a default; _flat() replaces table-cell escaping; the Minnesota-colliding 'sota' stem is dropped. 52 mutation tests. Skipped by judgment: task_families stays in the JSON (self-describing contract + lock, degradation-registry precedent); shared anchor-verifier extraction deferred until a third consumer exists; per-row staleness half-lives and a warn tier deferred. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753 * fix: #745 review round 2 — containment on eval_ref, report-pattern binding, two-tier claim stems, inventory locks, honest revision row Closes the codex round-1 findings: eval_ref must be a repo-contained suite directory; measurement_report must match measurement-*.json and unreadable sibling reports fail supersession instead of skipping it; claim language splits never-licensed stems (guarantee/proven/state-of-the-art, refused on every row) from measured-licensed stems, and percentages are refused outside MEASURED/MIXED; shipped row ids and anchor minimums are D5-style locked; future measured_at refused; bool stale_after_days refused; non-object behavioral_evidence reports instead of crashing; anchors render into the byte-pinned view. The revision.claim_drift_guard row is narrowed to what the 2026-08-07 record actually measured — the condensed guard-block prompt, 6 pressure items + 2 controls — with a ceiling that does not transfer to the shipped pipeline wiring. 64 mutation tests. Deferred by judgment (documented open half of #745): reverse inventory of README/CHANGELOG claim sentences. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753 * fix: #745 review round 3 — percent rule covers mechanism text, non-object reports refused, exact report-name match, robust anchor lock Closes codex round-2: the percentage discipline no longer exempts the mechanism field; a JSON-list bound or sibling measurement report yields an M11 violation instead of an AttributeError; _REPORT_NAME_RE uses fullmatch; the M13 anchor-minimum lock counts only well-formed lists after M7 has diagnosed the malformed value. 69 mutation tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753 * test: #745 pin the task-family vocabulary to the merged #742 design doc §2 table The frozen TASK_FAMILIES constant and the profile contract's stage/task-family table can no longer drift apart silently — the cross-consumer test parses the doc's table ids and requires exact equality including order. 70 mutation tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753 * docs: align CHANGELOG test count (70) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753 --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |