771 Commits

Author SHA1 Message Date
Edward Cheng-I Wu d5accd6b1f docs(changelog): record the #856 es-ES trigger phrases under [Unreleased] (#866)
Extends the #855 entry with the merged companion trigger change (a366e39):
body Español lines plus description subsets, the CONTENT_LOCKS re-pin, the
code-point description counts, the eval and routing-smoke evidence, and the
follow-ups #864 and #865.


Claude-Session: https://claude.ai/code/session_01CckFaPj7hPxWjCqn1dhbXt

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 11:57:37 +08:00
Dídac Rios a366e39e2e Add conservative es-ES trigger keywords to the four skills (#856)
* add conservative es-ES trigger keywords to the four skills

* address review: re-pin pipeline content lock, add verificar citas, fix reseñas typo

* address review items 4-6: trim deep-research es subset under 1,024, drop standalone investigación, add es-ES routing fixtures
2026-09-14 11:21:50 +08:00
Edward Cheng-I Wu 91fc74d37e docs(contributing): provisional single-owner locale applications (#862) (#863)
Add a provisional-application route under the "Two named owners" locale-pack
condition: a single-owner application is recorded in a dedicated issue, is not
a supported pack, and gets a 14-day backup window opening with the first minor
release after both the locale mechanism and the primary owner's recorded
acceptance. An unfilled window lapses the application; a supported pack that
loses either owner leaves the supported list. Point the interim
activation-layer sentence at #862 (Phase 1) alongside #850.

Record the #861 locale-pack policy and this amendment together under
[Unreleased] Changed so the pre-tag changelog-covers-merges gate sees both.


Claude-Session: https://claude.ai/code/session_01HwTR5WFp4E4gNU8RF2CjNy

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 12:56:36 +08:00
Edward Cheng-I Wu 7cf67a657d docs(contributing): locale packs are community-maintained; translation fast-merge excludes operative text (#861)
States the ownership model for non-default output locales ahead of the
#850 mechanism: two named owners, recorded currency with a 14-day window
per minor release, visible staleness in CI that never delays a core
release, configuration-and-presentation scope only, and #509 trigger
discipline including the Agent Skills 1,024-character description cap.
Narrows the translation fast-merge category so operative instructions
follow the review tier of the file they touch. Adds README.es-ES.md to
the README sync list.


Claude-Session: https://claude.ai/code/session_01EEbmXsHyQ3dzYoep74zSmH

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 11:09:02 +08:00
Edward Cheng-I Wu 4be3d11ca9 docs(changelog): record the #855 es-ES README under [Unreleased] (#860)
Claude-Session: https://claude.ai/code/session_01EEbmXsHyQ3dzYoep74zSmH

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 10:50:26 +08:00
Dídac Rios 665b3ca0a0 add es-ES README with language nav links and CI registration (#855) 2026-09-13 10:07:42 +08:00
Edward Cheng-I Wu 90f2176cd5 evals: add claude plugin eval suite for the academic-paper citation-check flow (#859)
* evals: add `claude plugin eval` suite for the academic-paper citation-check flow

Eight cases (six fire, two negative) under plugin-evals-citation-check/,
graded as a with/without-plugin ablation. Fire cases carry a synthetic
source pack so the four author-defined citation failures (no source on
file, wrong authors, hedged finding cited as established, retracted or
concern-flagged paper cited as live) are detectable offline. Styles: APA 7
(en / zh-TW mixed / es), IEEE, Vancouver (style unnamed), Chicago NB.
Cases pin model: sonnet; run with --judge-model opus.

Calibration (two pilots, 2026-09-13) and caveats are in the suite README.
The with-plugin arm cannot load the mode prompt in the eval sandbox
because the command stub references plugin files by relative path
(#857), and plain-language prompts fired the skill in 3 of 6 cases (#858;
the Spanish case is one data point for #850). No uplift figure is claimed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EEbmXsHyQ3dzYoep74zSmH

* evals(citation-check): apply cross-model review findings and re-calibrate

Twelve of thirteen review findings applied: replace the disputable
four-author "et al." planting in 04 with a year mismatch; make the clean
citations in 02 and 06 supported by their abstracts; replace real Taiwan
journal names in 02 with fictional ones; tie the 08 presence regexes to an
"unused" statement; require metadata preservation and reject audit content
on the 07 conversion negative; add no-overreach to 05; turn the 02 language
check into an llm grader; exempt unchanged entries in a complete corrected
list from no-false-positive and drop its DOI claim; drop the 04
style-name regex (both arms fixed the year without naming the style).
The one rejected finding (08 skill check "display-only") was wrong: it
carries arm: both and is scored; README says so.

Pilot 3 on the revised suite: $4.65, max 142 s / 7 turns / $0.45 per run;
04 and 08 re-run clean after the last two grader fixes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EEbmXsHyQ3dzYoep74zSmH

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-13 10:01:19 +08:00
Edward Cheng-I Wu 88725b8a55 fix(academic-paper): advertise revision-coach rebuttal triggers in the SKILL.md description (#851) (#853)
The revision-coach phrases "I got reviewer comments", "revision roadmap",
"should we push back", "conference rebuttal", "grant panel response" lived
only in the SKILL.md body, which the model reads after deciding to load
the skill. Add them (plus zh-TW/ko equivalents) to the frontmatter
description (699 chars, under the 1,024 Claude Code allowance) and add
the three missing English phrases to the body Trigger Keywords line.

Verification: plugin-evals/03-iclr-rebuttal-en skill-fired 0/2 -> 7/7.
The case's two llm rubrics are rewritten in enumerate-then-quote style;
the earlier claim-list phrasing drew 3-vote FAILs from the runner judge
on outputs a reasoning judge passed. README caveats updated.

Closes #851


Claude-Session: https://claude.ai/code/session_013fbc5qpXkAac1o4HLMGinE

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 23:55:20 +08:00
Edward Cheng-I Wu cfbd2c6f63 evals: add claude plugin eval suite for the academic-paper revision-coach flow (#852)
Seven cases (five fire, two should-not-fire), twenty graders, run as a
with/without-plugin ablation. Primary quality axis is "no unauthorised
rewriting" (no manuscript prose drafted, nothing changed that no reviewer
asked for, no results or changes asserted that have not happened).
Calibrated against five pilots on 2026-09-12; README records the run
command, ceilings, pilot cost and caveats. plugin-evals/results/ is
gitignored. The ICLR case surfaced the trigger gap filed as #851 and is
its acceptance check.


Claude-Session: https://claude.ai/code/session_013fbc5qpXkAac1o4HLMGinE

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 22:54:35 +08:00
Edward Cheng-I Wu f1a57bbcab fix: shared file-lock helper with msvcrt backend for the remaining fcntl sites (#845) (#847)
* fix: shared file-lock helper with msvcrt backend for the six fcntl sites (#845)

scripts/file_lock.py owns the backend choice (fcntl.flock on POSIX,
msvcrt.locking on byte 0 on Windows) and routes adjudication_activity,
inquiry_branch_ledger, review_criteria_binding, and ars_mark_read through
acquire()/release(). POSIX lock sequences are unchanged. Per-site Windows
decisions: adjudication reads degrade to exclusive with a 5 s bounded wait;
the review-criteria manifest lock is capped at 30 s on Windows only; the
inquiry ledger alpha keeps refusing non-POSIX hosts. Two finally blocks that
released an unacquired lock now release only what they acquired. SETUP docs
state the best-effort Windows posture; no Windows CI job is added.

Refs #845, #843, #844.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2

* fix(file_lock): interrupted attempts honour the deadline; pin adjudication wait policy (#845)

Cross-model review round 1 (gpt-6-astra, xhigh): a persistent
InterruptedError could retry past the bound; the Windows-shape test did
not exercise adjudication's reader-waits / writer-does-not-wait policy;
the adjudication contention message now names LockTimeout instead of
BlockingIOError, recorded in the CHANGELOG rather than masked.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2

* refactor(file_lock): held() context manager, single BACKEND source, one fake msvcrt (#845)

/simplify pass (four cleanup reviewers): the release-only-if-acquired
invariant moves into file_lock.held() and review_criteria_binding /
inquiry_branch_ledger use it; runtime branches key off BACKEND and
SHARED_LOCKS_SUPPORTED is dropped; EINTR joins the retryable errno set and
the unreachable EDEADLK entry goes; backend calls are deduplicated; all four
consumers try the sibling import first so one module instance is shared;
the Windows fake lives once in tests/fake_msvcrt.py; test scaffolding is
folded into a lock_pair fixture and a parametrized wait test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2

* fix(file_lock): keep lock acquisition and the guarded body in separate try blocks (#845)

Cross-model review round 3 (gpt-6-astra, xhigh): wrapping the body in the
same handler that translates LockTimeout meant a contended inner lock inside
the body was reported as the outer manifest/passport lock failing. Both
consumers now acquire in their own try block and release only after a
successful acquire; held() is dropped from the helper. The subprocess test
pins that a LockTimeout raised inside the binding body surfaces as itself.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2

* test(file_lock): let the body LockTimeout leave _locked() so the attribution check bites (#845)

Cross-model review round 4: the inner LockTimeout was caught inside the
binding body, so the erroneous outer translation would still have passed.
Verified by mutation: restoring the outer translation fails this test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2

* ci(673): whitelist scripts/test_file_lock.py as a non-consumer importer of the activity runtime (#845)

The shared file-lock test imports adjudication_activity in a subprocess to
exercise its lock backend under a fake msvcrt; it never reads or writes an
activity store. The exact-owner whitelist is the lint's route for that.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 14:47:26 +08:00
Edward Cheng-I Wu c7af8b9017 docs(changelog): record the #844 Windows msvcrt lock backend for /ars-mark-read under [Unreleased] (#846)
Claude-Session: https://claude.ai/code/session_0131cZMWBPPeEFiqgEPFZ3X2

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-09 17:55:20 +09:00
Wu Shuwen 32f754aae5 fix: make /ars-mark-read work on Windows via msvcrt lock backend (#843) (#844)
scripts/ars_mark_read.py imported POSIX-only fcntl at module load, so the
documented /ars-mark-read CLI failed on Windows before parsing arguments.
The lock now goes through two small helpers: fcntl.flock on POSIX (unchanged)
and msvcrt.locking(LK_NBLCK, 1) on Windows, inside the same bounded retry
loop. msvcrt.locking can lock a byte beyond EOF, so no pre-write is needed.

Fixes #843. Remaining fcntl import sites are tracked in #845.

Co-authored-by: dajiaohuang <dajiaohuang@users.noreply.github.com>
2026-09-09 17:21:44 +09:00
Edward Cheng-I Wu 8e4c877764 docs(changelog): backfill the #835 reviewer-calibration harness entry under [Unreleased] (#842)
check_changelog_covers_merges flagged #835 (merged 2026-09-07) as the one
release-worthy commit since v3.21.2 without a CHANGELOG entry. Entry is
derived from the PR body and the merged file set; it states that no
calibration profile or measurement values shipped and that #653 / #828
stay open.


Claude-Session: https://claude.ai/code/session_01AYAjWg2eBEz3UV7MZn7eFt

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-08 16:50:40 +09:00
Edward Cheng-I Wu f832c89f60 docs: Gartenberg et al. (2026) fourth human-in-the-loop anchor, volume non-goal, cognitive-surrender note (#833) (#841)
- README.md / README.zh-TW.md motivation: fourth anchor paragraph for the
  Organization Science AI Task Force editorial "More versus better"
  (37(3):795-812). Scope stated: one journal, observational, aggregate,
  proprietary classifier. Cited as design rationale, not as evidence about
  ARS output.
- POSITIONING.md "Rejected mechanisms": volume as an outcome. No batch
  manuscript generation, no fan-out of one run into several submissions,
  time-to-draft booked as a resource cost.
- shared/collaboration_depth_rubric.md 1.0 -> 1.0.1: related-construct
  citation on Cognitive Vigilance (uncritical acceptance of AI output;
  "cognitive surrender" as the editorial cites Shaw & Nave 2026). Scoring,
  dimensions, and the descriptive-only reporting rule unchanged; a low
  score stays an observation, not a failure.
- CHANGELOG [Unreleased] > Changed.

The claim_strength_ladder.md item from #833 is byte-pinned by the
revision-claim-drift suite (CLAIM_LADDER_SHA256 and the frozen v2
adjudication rubric); it is held for a separate decision.


Claude-Session: https://claude.ai/code/session_01AYAjWg2eBEz3UV7MZn7eFt

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-08 15:50:58 +09:00
Edward Cheng-I Wu 27e6c9978d fix(socratic): F6 lists user directions unranked; round caps defer to the mentor agent (#834) (#840)
failure_paths.md § F6 preselected "[the most promising direction]", told the
mentor to rank directions by "convergence potential", and prescribed
"restrict discussion scope" — ranking and preselection of the user's own
directions, which the #735 boundary forbids. F6 and socratic_mode_protocol.md
also still said "round 15 → end" after #490 made the mentor agent's
§ Auto-End Conditions (Precise) the single authority.

- F6: chronological, user-worded summary of expressed directions; the user
  chooses; the full-mode option names the visible exit marker; no scope
  restriction; round caps point at the agent file.
- socratic_mode_protocol.md § Dialogue Management Rules: the 15-round line
  becomes a pointer to the agent authority.
- test_socratic_rq_non_generation_contract.py: ranking/preselection
  vocabulary check on F6, own-round-count check on both reference files,
  pointer/heading parity with the agent file; mutation tests inject the
  pre-fix bytes and a stagnation-trigger control.
- CHANGELOG [Unreleased]: contract-contradiction closure, no breadth claim.


Claude-Session: https://claude.ai/code/session_01AYAjWg2eBEz3UV7MZn7eFt

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-08 15:50:52 +09:00
Edward Cheng-I Wu 75070eec84 feat(evals): #653/#828 add reviewer-calibration harness with isolated dispatch and audited scoring (#835)
* feat(evals): #653 reviewer-calibration suite scaffolding — corpus assembler, isolated dispatcher, deterministic scorer, pre-registered rubric/RUN_PLAN (corpus freeze pending PDF access)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H2iNYa6YYYaPUwD2Z2Jr5e

* feat(evals): #653 freeze the ICLR 2026 calibration corpus manifest (12 papers) + shared PDF-text normalization

Corpus freeze (PR-A of #653): `corpus/papers.json` (label-free, 6+6 ICLR 2026
papers by the pre-registered seed; pypdf 6.11.0; pool hashes unchanged from
the 2026-08-07 selection) and `manifests/gold_labels.json` (public Decision
note ids + strings). No page-cap exclusion fired; `verify` PASS.

First real-PDF contact found a hashing defect: pypdf emits lone UTF-16
surrogates from math fonts (61 in one sampled manuscript) and strict UTF-8
encoding raised, so `extracted_text_sha256` was uncomputable. The
normalization now lives in one shared module (`scripts/_calibration_pdf_text.py`:
NFC + lone-surrogate -> U+FFFD), imported by both the assembler and the
dispatcher so freeze/verify/dispatch hash identical bytes; the rule is
recorded in the manifest's `extraction.text_normalization` and `verify`
fails hard on rule drift (a rule, not a version). Two tests added (41 total).

`scripts/fetch_calibration_corpus.py` is the authenticated OpenReview
operator tool that produces the freeze input, so the "third-party
reconstruction" claim in the README is backed by a runnable path.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1

* refactor(evals): #653 simplify pass — shared hashing/fence/git-state, contract 1.1 docs

/simplify findings applied (reuse, simplification, efficiency, altitude):
- `_calibration_pdf_text.py` owns `sha256_hex` + `pdf_facts` (bytes hashed and
  parsed from one read via BytesIO; `extract_text=False` lets `verify` skip
  extraction when the pypdf version cannot be compared); surrogate replacement
  is one `re.sub` pass. Both the assembler and the dispatcher import it.
- dispatcher reuses E4's closed data-fence grammar (`_delimited`), `_git_state`
  (declares unknown provenance dirty instead of raising), and the evidence
  path guard (`assert_plain_file`: rejects symlinked parent components, not
  just the leaf); one `_prepare` preamble for both stages; a text-hash
  mismatch now names its cause (installed vs manifest pypdf version).
- assembler: exclusion rows stay dicts, `pool_list_mismatches` shared by
  freeze/verify, exclusion set built once.
- scorer: `confusion`/`bootstrap_ci` take (predicted, gold) pairs (same RNG
  stream as before), `Counter` for the exact-mode vote, dead `_path` dropped.
- RUN_PLAN/README: measurement contract 1.0 is closed to new rows (#664);
  the run publishes under 1.1 with its pre-registration record + write-once
  execution manifest (dispatcher/scorer support lands with the scored run).

Re-freeze after the refactor reproduces papers[] and gold_labels byte-for-byte.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1

* fix(evals): #653 Iron Rule #7 at the two whole-file call boundaries + paper-id shape check

Security review round 1 (first-party) found two below-threshold gaps and
both are verified real:
- The calibration dispatcher omitted E4's `DATA_BOUNDARY` sentence on the
  field-analyst call (the one E4 call that carries it, because
  `field_analyst_agent.md` states no untrusted-material rule of its own).
  Restored, and a fitted `REPORT_BOUNDARY` added on the synthesizer call,
  whose agent file is likewise dispatched whole with no such rule. Pinned by
  a transport-capture test that checks both sentences precede their fence.
- Paper ids are spliced into file names (`<id>.pdf`, `cards/<id>/`) but
  `load_pool` accepted any non-empty string. Ids now must match
  `^[A-Za-z0-9_-]+$` (OpenReview's forum-id shape) in the assembler and the
  fetch tool; test pins the refusal.

43 tests pass.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1

* fix(evals): #653 codex round 1 — dispatch/verify invariant parity, card-path guard, scorer completeness

Codex round 1 (gpt-6-astra xhigh) findings 2-7, 10, 11 and the cheap half of 9,
each re-verified first-party before the change:
- dispatcher: frozen cards go through the same plain-file guard as PDFs and
  agent files (a symlinked card1.md -> gold_labels.json was readable); the
  manifest's text_normalization rule and page_count are checked before
  dispatch, so dispatch and verify enforce the same manuscript invariants;
  transport-failure artifacts keep the partial stdout and stderr verbatim;
  every call attempt records RFC-3339 start/complete and prompt/output
  hashes into the panel record and cards freeze (the per-call evidence the
  heldout-measurement/1.1 execution manifest is built from).
- verify: label must match decision_raw under the label transform; paper
  count and per-class label counts must equal the recorded quotas
  (synchronized paper+label removal no longer passes).
- scorer: a second record for the same paper/replicate is a hard error, not
  a silent overwrite; a gold paper with no complete ensemble blocks the full
  tier; an A1 override needs its verbatim `raw` excerpt present in
  synthesis.md.
Nine regression tests added (52 total). Real-corpus verify still PASS.

Not addressed here (need a decision): finding 1 (camera-ready format leaks
the accept label) and finding 8 (numeric seat scores vs categorical seat
contract); finding 9's manifest/row builders land with the scored run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1

* fix(evals): #653 drop the numeric score axis — protocol Phase 2 forbids AUC, seats are categorical

Codex round 1 finding 8, verified against the source: the seat contract
(eic/methodology/... agents) emits criterion-bound categorical judgements and
states "Do not total, weight, average"; `calibration_mode_protocol.md`
Phase 2 says "Do not report AUC: there is no continuous rubric score." The
scorer nevertheless extracted a `Weighted Average` figure (a retired field)
and RUN_PLAN promised AUC + score variance, so a conforming run would have
published null numerics against a plan that promised them.

The scorer now reports only what the protocol's full-tier table names:
confusion matrix, balanced accuracy, FNR, FPR (bootstrap CIs), exact-label
agreement (count/share/target-set size, with the binary-gold caveat), and
replicate stability as categorical agreement (on side, on exact label).
AUC is emitted as an explicit NOT REPORTED line. RUN_PLAN and the test
fixtures follow. 52 tests pass.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1

* docs(evals): #653 mark the 2026-09-06 corpus SUPERSEDED (layout leaks the label, #828); RUN_PLAN model currency

- README/RUN_PLAN: the frozen ICLR 2026 corpus is a harness-rehearsal corpus
  only — camera-ready replacement makes accepted PDFs visibly different from
  rejected submission PDFs (6/6 + 6/6; 30/30 in a fresh accepted-pool sample).
  No profile or measurement row may be published from it; the gold corpus
  becomes an ICLR 2027 submission-time capture. The "Why ICLR 2026" rationale
  is kept as pre-registered and annotated with the two facts that now cut
  against it (layout leak; Fable 5.1's 2026-06 cutoff covers the decisions).
- RUN_PLAN + dispatcher default: subject `claude-fable-5` -> `claude-fable-5-1`,
  judge `gpt-5.6-sol` -> `gpt-6-astra` (provisional, #783 policy). Pre-dispatch
  edits, not amendments.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EehvYnmG8Xym5tXhTLfVm1

* feat(evals): #828 layout-tell guard at corpus freeze — refuse a corpus whose page-1 layout is not constant

`assemble_calibration_corpus.py freeze` now reads page 1 of every cached PDF
and evaluates four venue-template signals (published-as header, under-review
header, "Anonymous authors", >=10 bare three-digit line numbers). Any signal
that is not constant across the whole corpus refuses the freeze with the
per-class counts; a uniform corpus records `layout_tell_check` in
papers.json. `verify` recomputes the same check (skipped with a warning when
a PDF is not cached; a manifest without the block warns). On the superseded
2026-09-06 ICLR 2026 corpus every signal is 6/0, so `verify` now FAILs on it
by design. Shared `_open_reader` + `first_page_text` in the PDF helper.

Six tests (signal detection, full and partial separation refused, uniform
freeze + verify round-trip, missing-PDF skip, pre-check manifest warning).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh

* feat(evals): #653/#828 rehearsal fixes + heldout-measurement/1.1 manifest and row builders

Rehearsal 2026-09-06 (2 papers x 1 replicate, blocked at the first call by a
rejected API key) exposed three dispatcher gaps, all fixed with tests:

- credential preflight: zero-cost `GET /v1/models` before the first billed
  call; a definitive 401/403 refuses (key never echoed), network trouble is
  `inconclusive` and proceeds; outcome recorded in every record
- credential rejection mid-run (`Failed to authenticate` / `API Error: 401`
  / `Not logged in`) is never retried (`CredentialRejected`); other
  transport failures keep the single retry
- an aborted cards stage writes `runs/blocked-cards-<paper>.json` with its
  per-call rows instead of losing them; both stages share one record writer

1.1 contract substrate (RUN_PLAN "pre-registration record + execution
manifest" item):

- `dispatch_calibration_panel.py --stage manifest` folds the completed call
  rows of one attempt (frozen cards + panel records; `load_attempt` refuses
  mixed attempt identities) into a write-once, schema-validated
  `execution-manifest.json`
- `build_calibration_measurement_row.py` composes the 1.1 row: plan and
  rubric hashed and compared against `frozen_commit` (drift refuses; dirty
  commit refuses), manifest re-derived from the records and compared
  field-for-field, judge rows required (no judges, no row), agreement
  recomputed by the checker's own `judge_divergence` (extracted from
  `check_heldout_measurement_report.py`, behaviour unchanged), validated by
  the checker before a write-once write
- adjudication rubric gains `## Resolution direction` (flags_only, I13
  lower-bound labelling); README tooling section; RUN_PLAN names the row
  builder; DATA_FLOWS names the dispatcher's preflight touchpoint; scorer
  docstring de-staled (no score axis); pytest manifest +1

No calibration number is recorded anywhere in the repository.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh

* fix(evals): #653/#828 codex round 2 — bind every row input to its attempt, harden the guards

12 of 13 findings applied (gpt-6-astra xhigh, read-only exec):

- P1 foreign metrics: scorer output is bound to the attempt (per_panel keys ==
  the complete panel records here, attempt ids match, n_papers matches)
- P1 raw drift: record admission re-hashes every completed call's raw output
  against output_sha256 (manifest stage and row builder alike); prompts are
  not retained (they embed the manuscript)
- P1 preflight redirects: the probe uses a no-redirect opener (a 3xx is
  `inconclusive`) and skips a non-https ANTHROPIC_BASE_URL
- P2 estimand: class-A adjudication is now pre-registered as bidirectional
  (every synthesis decision transcribed blind and compared with the grammar),
  so the row publishes a point_estimate instead of an I13 "lower bound" that
  only meant audit coverage
- P2 pre-write parity with R5: manifest timestamps parsed and ordered before
  the write; declared claims checked against the local manifest
- P2 strict JSON: inputs parsed with the checker's strict loader, outputs
  serialized with allow_nan=False and round-tripped
- P2 judge failures: `--blocked-run` ledger entries merge into
  attempts.blocked_runs (I11)
- P2 admission by content: suite/stage/status/provenance from the record
  body, never the filename; blocked records are identity-checked too
- P2 cards re-run: a reused evidence dir refuses (write-once stage records)
- P2 auth signature: anchored at the start of stdout/stderr and limited to
  exit-code failures; a timeout's partial prose is never a credential error
- P2 layout signals: phrase tests run on whitespace-folded text
- P2 partial PDF cache: verify checks every cached PDF (can refuse, cannot
  clear) instead of skipping the guard
- P3 real `git show` test for sha256_at_commit on a temporary repository

Partially applied: "distinguish unobservable signals from absence" (not
built; the constancy rule is pre-registered as stricter by design).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh

* fix(evals): #653/#828 shared transport — capture every assistant message, fence the subject's config

Rehearsal take 2 (2026-09-06/07, 8 billed calls on the first paper) found
two transport defects in `ClaudeCliTransport`, shared by the E4 and the
calibration dispatchers:

- text-mode `claude -p` prints only the LAST assistant message: the
  first paper's synthesis (long enough to be continued) came back starting
  mid-table, with the Editorial Decision Letter and its `### Decision:`
  line in the missing head. The transport now runs `--output-format
  stream-json --verbose` and concatenates the text blocks of every
  assistant message; an error result or an unreadable stream is a
  TransportFailure that keeps the raw bytes.
- `--bare` does not fence the subject: a two-call probe on 2.1.260 showed
  the operator's whole global CLAUDE.md arriving as a system-reminder,
  plus `settings.json` `language` and the output style (the seats appended
  Traditional-Chinese "plain-language summary" sections). The subject now
  runs with an allowlisted environment (PATH/HOME/LANG/TMPDIR/TERM/USER/
  SHELL + ANTHROPIC_*; no CLAUDE_* inherited from a parent session) and a
  per-transport empty `CLAUDE_CONFIG_DIR`; the same probe then reported no
  instruction beyond the SDK identity line and the date.

E4 tests: one fake updated to emit stream-json; five new tests (message
joining, error/junk results, unreadable-stream failure with bytes,
environment allowlist, argv/env of a live call). Calibration docs and the
panel record's `dispatch` field describe the new recipe (pre-dispatch
change, no amendment).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh

* fix(evals): #653/#828 codex round 3 on the shared transport — eviction signals, LF framing, network env, failure evidence

Five P2 findings (gpt-6-astra xhigh, read-only exec), all applied:

- refusal-fallback eviction: assistant `supersedes` and system
  `model_refusal_fallback.retracted_message_uuids` (wire fields verified in
  the installed CLI 2.1.260) drop retracted partials before concatenation
- NDJSON split on LF only (`str.splitlines` also splits on U+0085 /
  U+2028 / U+2029 inside a JSON string); CRLF tolerated
- environment allowlist keeps documented network/TLS inputs (proxies,
  NODE_EXTRA_CA_CERTS, SSL_CERT_*, CLAUDE_CODE_CLIENT_*); an apiKeyHelper
  that needs more is documented as unsupported behind the fence
- transport failures carry assistant TEXT in `stdout` and the raw stream
  in `raw_stdout`; a framing-only stream is "no model response" (E4 no
  longer writes stream metadata as a partial response); both dispatchers
  preserve the raw stream as `*.transport-stream.jsonl`
- a structured error result (`[TRANSPORT: result <subtype>]`, diagnostic
  in stdout) is classified by the calibration retry loop like the
  plain-text startup failure: a credential rejection is never retried

E4 tests +6 (256), calibration +1.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh

* feat(evals): #653/#828 keep the raw stream of successful calls as evidence

`ClaudeCliTransport.last_raw_stdout` exposes the stream-json framing of the
most recent successful call; the calibration dispatcher writes it next to
the text as `<label>.transport-stream.jsonl`, so the next rehearsal shows
how many assistant messages a deliverable spanned (the 2026-09-06 synthesis
lost its head to exactly that). Probe 2026-09-07: a 12,000-line reply at
effort low arrived as ONE text message after a thinking-only message, so
the head loss is attributed to multiple text messages in one turn (likely
interleaved thinking at xhigh), not to an output-length continuation; the
parser covers both.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014bJSoFkGeRaWTMJodk4VGh

* fix(evals): #653/#828 allow requiring a successful credential preflight

* fix(calibration): bind audited decisions and preserve failed dispatch evidence

* fix(transport): retain truncated UTF-8 output as byte evidence

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-08 08:38:58 +09:00
Edward Cheng-I Wu 6b7ee6dcae fix: Astra request compat, no-delegation citation transport, hedge/quota prompt repairs, audit provenance (#823–#826) (#827)
* fix: Astra request compatibility, no-delegation citation transport, hedge/quota prompt repairs, audit provenance (#823 #824 #825 #826)

#823 — OpenAI request builders (smoke entrypoint + documented example) drop
`temperature`, which GPT-6 Astra rejects; the per-model effort vocabulary
lives in scripts/cross_model_verification/openai_effort_guard.sh, sourced by
both, and an unsupported explicit Astra value fails before curl. Hermetic
fake-curl test runs both surfaces.

#824 — the contained Codex citation transport rejects effort=ultra with
REASONING_EFFORT_REQUIRES_DELEGATION before detection/auth/tempdir/launch on
both entry paths (codex-cli 0.153.4 defines ultra as the multiAgentMode
replacement). Model-independent by design.

#825 — hedging can no longer rescue an unsupported claim (writer recovery
tree, CER fallback row, temporal rule 5 in writer + both compiler mirrors,
writer contract D2); universal prose quotas in the writer, compilers,
writing_quality_check.md, academic-paper/SKILL.md, and contract D6 become
diagnostics subordinate to author/venue requirements. Audit inventory
corrected in place; held-out seed evals/heldout/unsupported_claim_recovery
(NOT_RUN) registered.

#826 — run_codex_audit.sh pins gpt-6-astra/xhigh and records both in a new
sidecar `model` block; claim_audit_pipeline binds an unknown judge identity
to a run-local cache key (no cross-run reuse) instead of defaulting to
gpt-5.5-xhigh.

Review: /simplify (4 angles), codex gpt-5.6-sol xhigh 2 rounds (r1: 1 P1 +
1 P2 + 2 P3 fixed; r2: 0 P1/P2), /security-review 0 findings; all 102
spec-consistency steps + pytest manifest replayed locally.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BNKiXpdHx1T5F5RbXT2Ueu

* docs(claude): record the #824 ultra reversal in the v3.21.2 key-additions line

The v3.21.2 bullet still said the contained Codex citation transport accepts
ultra; #824 on this branch rejects it as a delegation request. Add the
reversal so the live instruction surface matches the transport.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K7emV5r2aqZDJzAyYVuuDo

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 08:08:30 +09:00
Edward Cheng-I Wu 8fa3d651ad docs(release): v3.21.2 — model currency, checkpoint decision provenance, and CJK title-matching repairs [skip-closes-check] (#822)
Promotes the [Unreleased] block to v3.21.2 (2026-09-06) and aligns every
version-bearing surface, following the v3.21.1 release-prep file set:
CHANGELOG heading (empty [Unreleased] anchor kept), plugin/marketplace
manifests, CITATION.cff, POSITIONING.md, MODE_REGISTRY.md, .claude/CLAUDE.md
(table row, Key Additions, Version Info), academic-pipeline/SKILL.md plus
its content-lock hash, docs/ARCHITECTURE.md current markers, the five
README badges/headings/entries, and the spec-consistency lint pins with
their fixtures.


Claude-Session: https://claude.ai/code/session_011sWwwG3oCbtL4cGhRsr5US

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
v3.21.2
2026-09-06 01:20:39 +09:00
Edward Cheng-I Wu 0861bc8538 chore(models): align docs and guardrails to Claude Fable 5.1 and GPT-6 Astra (#819) (#820)
* chore(models): align docs and guardrails to Claude Fable 5.1 and GPT-6 Astra

Read both vendor system cards in full and applied the model-update pass:

- Claude Fable 5.1 named as the current frontier model (PERFORMANCE en/zh-TW
  with a dated list-price re-derivation; cross-model primary-row example).
- gpt-6-astra listed as a provisional cross-model verifier on both transports
  and recommended under the #783 lifecycle policy; gpt-5.6-sol keeps its
  validated status on the ChatGPT-subscription citation transport. Entry-gate
  smoke PASS on that transport (2026-09-05, codex-cli 0.153.4). SETUP en/zh-TW
  example sets, id-status allowlist, bakeoff baseline text, and .claude/CLAUDE.md
  move together.
- Codex citation transport: `ultra` joins the closed reasoning-effort set as a
  named constant, with a test pinning turn/start forwarding and fail-closed
  rejection of unknown values.
- New guardrail: checkpoint decision provenance (authority in the pipeline
  state machine, operational mirror in the orchestrator), indexed as risk R11;
  both content-lock hashes updated in this commit.
- Provider-side monitoring / safety interventions named as a never-a-verdict
  case in the cross-model doc and the degradation registry row.
- Model tiering records that the resolved tier is the declared model; risk
  register R1/R4/R5/R6 residual gaps updated.
- Harness-retirement audit for the model change:
  audits/harness-retirement-2026-09-model-update.md (0 prompt retirements,
  4 applied currency fixes, 2 deferred, 8 keep-as-debt annotations).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011sWwwG3oCbtL4cGhRsr5US

* docs(changelog): align the model-update entries with the final text

The [Unreleased] entries were written before the simplify pass moved the
checkpoint-decision authority into the pipeline state machine, reused the
existing transport-failure markers for provider-side interventions, and
de-numbered the model-tiering note. Wording now matches the files.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011sWwwG3oCbtL4cGhRsr5US

* test: scope the checkpoint-authority section out of the v3.6.7 orchestrator line budget

The v3.6.7 Phase 6.6 budget test measures the orchestrator prompt minus every
later independent extension, each with its own bounded cap. The new
`## Checkpoint authority fidelity` section (13 lines) pushed the v3.6.7-attributed
count to 652 against a 639 ceiling. Following the existing convention, the
section gets its own measurement helper, an 18-line cap (5 lines of headroom),
a dedicated test, and is subtracted from the historical budget.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011sWwwG3oCbtL4cGhRsr5US

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 00:46:14 +09:00
Edward Cheng-I Wu 9443623791 docs: de-stale RISK_REGISTER R10 residual gap — #769 rows shipped in the same release (#813) (#814)
R10 still claimed the guard-launcher degradations were 'not yet indexed
in the degradation registry (#769)' although v3.21.1 itself shipped
registry 1.3.0 with all five write_scope_guard_* rows. The stale clause
is removed; the existing-controls line now points the guard's degrade
posture at its registry rows; the residual gap keeps only the
per-mechanism, per-channel loss description. check_risk_register and
check_version_consistency pass locally.

Closes #813


Claude-Session: https://claude.ai/code/session_011b4WKkZvjw7fCPDHQQhm6E

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-02 01:57:38 +08:00
Edward Cheng-I Wu 368576224b chore(audits): September harness-retirement audit — zero findings (#811) (#812)
Incremental audit over the 2026-08 baseline (1bd287f): full-diff review of
the 15 Bucket A agent bodies changed since then, mechanical pattern scan
across all 23, and an 87-path referenced-file existence check. 0 P0/P1/P2 —
every change traces to a documented shipping contract feature, and several
of the month's PRs already retired debt-like patterns (numeric scoring
rubric, confidence weighting, priority/effort ranking, universal source
quotas) through their own review. August keep-list unchanged, with two
documented additions.


Claude-Session: https://claude.ai/code/session_011b4WKkZvjw7fCPDHQQhm6E

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 20:48:01 +08:00
Edward Cheng-I Wu e8bf858be7 lint: skill-inventory parity across top-level dirs, skills/ symlinks, CLAUDE.md table, and marketplace.json (#809) (#810)
Closes #809.

check_skill_inventory_parity.py: on-disk <name>/SKILL.md dirs are the
authority; set-equality against skills/ symlinks, the CLAUDE.md Skills
Overview table (exact unfenced H2, GFM header+separator, first-cell
backticked names), and marketplace.json plugins[].skills; "N skills"
count claims on plugin.json / marketplace.json / MODE_REGISTRY.md.
check_spec_consistency.py derives its skill paths from disk at call
time; table-row grammar single-sourced in _skill_lint. 60 mutation
tests; wired into spec-consistency.yml and the pytest manifest.

Dual-track pre-ship: /simplify (4 findings applied), /security-review
(none), codex gpt-5.6-sol xhigh 7 rounds (12 findings fixed, round 7
clean). CHANGELOG also covers #805.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014U5nvjKy84twtB4VYsrex1
2026-08-31 02:02:28 +08:00
LeslieLi46 37bd060294 docs: fix duplicated word in MLA citation key rules (#805)
Docs-only: academic-paper/references/citation_format_switcher.md MLA key rules line read "No year in in-text"; now "No year in-text", matching the in-text format documented above it. No lint or hash lock pins this file. Contributed by @LeslieLi46.
2026-08-30 23:36:43 +08:00
Akshath Rajkumar 5debcd2efb fix: strip CJK outer wrapper marks only when they enclose one balanced unit (#800) (#804)
* fix: strip CJK outer wrapper marks only when they enclose one balanced unit (#800)

The wrapper strip inherited from #431 was positional: it removed the first
and last characters whenever they matched as a wrapper pair TYPE, without
checking they belonged to the same bracket pair. 《红楼梦》与《金瓶梅》 —
two titles joined in one string — normalized to 红楼梦》与《金瓶梅, leaving
an orphaned 》 mid-key.

Matching correctness was unaffected (both sides mangle identically; no
exploitable asymmetry found in the #799 security pass), but the mangled key
is a semantic anomaly for any future single-sided consumer.

Fix adds _outer_pair_encloses: the outer marks are stripped only when the
interior between them is itself balanced under all six wrapper pairs.
《围城》 still strips to 围城, nested balanced interiors still unwrap
(《基于「ProEXC」的研究》 → 基于「ProEXC」的研究), while 《红楼梦》与《金瓶梅》,
“研究”与“实践”, and “研究与“实践” keep their marks.

Both consumers change together — the CJK client re-imports the shared
function (#799), pinned behaviorally as well as by identity. The two new
discriminating tests are mutation-verified to fail against the pre-fix
module. Full suite: 9262 passed, 3 skipped.

* fix: scope the CJK interior balance scan to the outer pair's own family (#804 review)

Addresses the P1 and three advisories on PR #804.

P1: `’` is the closer of `‘` AND the English apostrophe; `”` likewise appears
unpaired in mixed typesetting. The family-blind interior scan read the lone `’`
in `《Alzheimer’s病中ProEXC表达》` as an unbalanced quote and refused to
strip a genuine `《…》` wrap. Verified against main: that pair went from
exact=True / ratio 1.0 to exact=False / ratio 0.6818 — below the 0.70 floor, so
the DOI-keyed ratio gate and the title-fallback exact gate failed together and a
correct DOI could be reported as a mismatch. That is the failure class #798
repaired, so it is not covered by #800's conservative-direction blessing: that
blessing is for titles whose outer marks are not one pair, and here they are.

Fix scopes the scan to the outer pair's own family — a `《…》` wrap tracks only
`《`/`》` and is blind to quote marks. This costs the check nothing it was
buying: any mark that can orphan the OUTER pair is by definition of that pair's
own family. The whole #800 behavior table survives (`《红楼梦》与《金瓶梅》`
still trips on its stray `》`, `“研究”与“实践”` on its stray `”`, nested
`《基于「ProEXC」的研究》` still unwraps, `《》` → ""), and the flagged
interaction `《「研究』》` → `「研究』` now keeps its mismatched inner quotes as
content, which is what distinguishes it from `《研究》`. Regressions added for
both apostrophe shapes (`’s`, `’98`).

The two lookup maps the family-blind scan needed are now dead and removed.

Advisory 1: the depth-0 stray-closer branch is pinned. The first attempt did
not discriminate — clamping absorbs one closer, so any title with equal
opener/closer counts (every natural case, including `《红楼梦》与《金瓶梅》`)
is refused by the trailing depth check anyway. Discriminating requires interior
closers to outnumber openers by exactly the clamp count, which no natural title
shape produces, so the test uses a documented synthetic asserted against the
helper directly. All three branches are now mutation-verified: clamp-instead-of-
refuse, drop-the-trailing-check, and family-blind (the P1 regression itself).

Advisory 2: the module docstring's "behaviorally equivalent to the
implementation this was promoted from" is qualified as historical — true of the
#798/#799 promotion, false since #800 deliberately changed this behavior.

Advisory 3: the changelog's non-CJK invariance claim is narrowed to what the
code supports. The claim holds on the two `has_cjk`-gated paths; the client's
`_cn_titles_match` is ungated, so a mark-carrying Han-free title can change
verdict there (`《Hamlet》and《Macbeth》` vs its pre-mangled form — verified to
flip). The empty-wrapper guarantee is likewise a `_cjk_titles_match` property:
`exact_normalized_title("《》", "《》")` is still True through the base branch.
Both narrowings are now pinned by tests so the prose cannot drift from the code.

Full suite: 9270 passed, 4 skipped, 0 failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 17:55:24 +08:00
Edward Cheng-I Wu 9469fc4d07 fix: fail check_surface_form_parity with an environment error when pyyaml is missing (#801 follow-up) (#803)
With the manifest present but pyyaml unimportable, _load_manifest
returned None and main() misdiagnosed it as "manifest ... empty / null /
non-mapping", pointing the reader at the wrong file. The import failure
is now a distinct _YamlUnavailableError; the lint exits 1 naming pyyaml
and the requirements-dev.txt remedy. Regression test pins the message
(45 tests, all green; lint itself still passes).

Also de-enumerate the stale "(PyYAML + jsonschema ...)" dependency
parenthetical in docs/SETUP.md and docs/SETUP.zh-TW.md (both language
files together, per bilingual-parity discipline).


Claude-Session: https://claude.ai/code/session_013R81d1YwGvJAznkPKk9gNw

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 11:32:24 +08:00
Edward Cheng-I Wu 30ad279cdf fix: declare markdown-it-py floor and make the autolink round-trip tail run visibly (#801) (#802)
* fix: declare markdown-it-py floor and make the autolink round-trip tail run visibly (#801)

The no-link_open round-trip tail of test_gfm_bare_urls_emails_and_schemes_
cannot_autolink soft-imported markdown-it-py (undeclared in requirements-
dev.txt) and silently returned when absent, so it had never run in CI, while
ambient markdown-it-py 2.x failed it on clean main (2.2.0 + linkify-it-py
2.0.3, reported in #799). Verified dividing line: 2.2.0 fails, 3.0.0 and
4.0.0 pass with linkify-it-py held at 2.0.3.

- Split the tail into test_escaped_markdown_yields_no_linkify_tokens_on_
  round_trip, gated by pytest.importorskip minversions (markdown_it 3.0.0,
  linkify_it 2.0.3): ambient-old environments skip visibly.
- Declare markdown-it-py>=3.0 + linkify-it-py>=2.0.3 in requirements-dev.txt
  with a reverse pointer at the consuming test, so CI exercises the round
  trip for the first time.
- Move the identical soft-import tail in test_renderer_neutralizes_markdown_
  active_inventory_path (newly activated in CI by the same declaration) to
  the same importorskip idiom; no floor needed (default CommonMark, no
  linkify) — verified passing under 2.2.0, 3.0.0, and 4.0.0.
- Consolidate the triplicated hostile-row construction in
  test_evidence_rows.py into one _hostile_row helper.

Renderer behavior and every renderer-side assertion are unchanged.
Verification: both full files 414 passed under markdown-it-py 4.0.0;
affected tests re-run under 2.2.0 (pass + visible skip) and 3.0.0 (pass).

Closes #801

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013R81d1YwGvJAznkPKk9gNw

* fix: flatten inline token children in the newly activated manifest markdown scan (#801)

Cross-model review (codex, xhigh) on PR #802 flagged that the twin test's
token scan iterated only top-level tokens, but markdown-it nests link_open /
image / html_inline under inline tokens' children — so the assertion could
only ever catch html_block. Verified empirically, then flattened children
into the scan (same idiom as the evidence-rows round-trip test).
Strengthened assertion passes under markdown-it-py 2.2.0, 3.0.0, and 4.0.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013R81d1YwGvJAznkPKk9gNw

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 10:35:29 +08:00
Akshath Rajkumar e5718cbf58 fix: apply Chinese-aware title matching in the four index resolvers (#798) (#799)
* fix: apply Chinese-aware title matching in the four index resolvers (#798)

`chinese_literature_client.py` already carried a Chinese-aware
`normalize_cn_title` / `has_cjk`, but the four index resolvers (Semantic
Scholar / OpenAlex / Crossref / arXiv) read the ASCII-centric helpers in
`_text_similarity.py`, where `.lower()` folds case but never width (P
U+FF30 never reaches P U+0050) and `string.punctuation` contains none of
`。`, `《》`, or U+3000.

A real Chinese paper served by an index in a different-but-legitimate
typesetting therefore missed on two paths: the DOI-keyed cross-check,
which gates on the fuzzy ratio alone and scored a fullwidth spelling of
the identical title at 0.625 (under the 0.70 floor) reporting a correct
DOI as DOI_MISMATCH; and the title-fallback search, which requires ratio
AND exact equality and so fell to `unresolvable`. Both feed the
`*_unmatched` contamination signals, so a genuine paper could render as
CONTAMINATED-TRIANGULATION-UNMATCHED.

Promotes `has_cjk` / `normalize_cn_title` into `_text_similarity.py`
byte-identical (the CJK client now re-imports rather than keeping a
private copy, per the #128 anti-drift goal), adds the Chinese-aware form
to `exact_normalized_title` as an additive third branch, and folds it
into `_similarity` through the existing `max`. Both gated on BOTH sides
carrying a Han ideograph, so every non-CJK verdict is provably unchanged
— pinned by an oracle test restating the pre-fix formula in full.

31 new tests, each verified to fail against the pre-fix module.
Full suite: 9255 passed, 3 skipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qexY5ysaqaAyPp97byf4w

* test: force the ratio-independence and non-destructiveness proofs (#798 review)

Addresses the three requested changes on PR #799.

1. `_cn_titles_match` ratio-independence is now forced, not inferred, in the
   test that claims it. `test_legitimate_variants_match_despite_a_sub_threshold_
   fuzzy_ratio` asserted the match on a pair the repaired `_similarity` scores
   1.0, so a regression that ANDed the ratio back in as a necessary condition
   would still have passed. Its match assertions now run inside a
   `monkeypatch.context()` with `_similarity` replaced by a detonator, scoped so
   the ratio measurements above it still see the real function. A new
   `test_cn_titles_match_never_consults_the_fuzzy_ratio` adds the negative half
   under the same forced conditions, so the invariant cannot be satisfied by a
   helper that has stopped discriminating.

   The shared `_forbid_similarity` helper patches BOTH binding paths: the
   `_text_similarity` module attribute (a qualified call or lazy in-function
   import) and the client's own namespace (a module-level `from ... import
   _similarity`, already bound and blind to the first patch). Both styles were
   mutation-verified to trip it; before this change the named test passed the
   regression that the new one caught.

2. `test_ratio_never_lowered_off_the_cjk_path` asserted `>=` against the base
   ratio alone, so it passed a *raised* non-CJK score and never exercised the
   dotted-acronym branch. Renamed to `test_ratio_unchanged_off_the_cjk_path`
   and rewritten against a full `_pre_fix_similarity` oracle — the companion to
   the existing `_pre_fix_exact_normalized_title`, written out in full for the
   same anti-drift reason — asserting exact equality. Mutation-verified twice:
   one raising a base-branch score, one confined to the acronym branch; the old
   assertion caught neither.

3. "Byte-identical" corrected to "behaviorally equivalent" in the
   `normalize_cn_title` docstring and the CHANGELOG. The promotion hoists the
   wrapper/terminal-mark sets to module constants, precompiles the regex, and
   rewrites comments; behavioral equivalence is what the tests actually pin.
   CHANGELOG test count corrected 31 -> 32 and its oracle sentence updated to
   describe both oracles.

Full suite: 9256 passed, 3 skipped (+1 test). The pre-existing
`test_evidence_rows.py::test_gfm_bare_urls_emails_and_schemes_cannot_autolink`
failure is unchanged and also fails on clean main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011qexY5ysaqaAyPp97byf4w

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 18:35:16 +08:00
Edward Cheng-I Wu 127ff85e4b docs(release): v3.21.1 — bounded workflow substrates and transport hardening [skip-closes-check] (#797)
Promote the accumulated Unreleased changes, align all version-bearing surfaces and five localized README summaries, backfill merge coverage, and preserve the default-off/design-only/evidence-bound claim ceilings.

Validated with 9225 passed, 3 skipped, 267 subtests; pre-tag version, changelog, spec, content-lock, distribution-claim, privacy, and staged secret checks all pass.
v3.21.1
2026-08-24 17:36:07 +08:00
Edward Cheng-I Wu 088d288a5e feat: add inquiry ledger, alternative-register design, and proving set (#796)
* feat: add opt-in inquiry branch ledger (#743)

* docs: freeze alternative explanation register design (#744)

* feat: add source-backed review criteria proving set (#575)
2026-08-24 15:49:18 +08:00
Edward Cheng-I Wu 385bc064e1 feat: add sealed bakeoff and workflow profile contracts (#795)
Implement the #789 sealed promotion-bakeoff lifecycle and the #742 research workflow profile contract. Consolidate the #794 Markdown link/anchor grammar, record #684 expert-stage readiness, and add the #575 closure-scope audit.
2026-08-24 12:08:01 +08:00
Edward Cheng-I Wu 7ef93e0cb5 docs: re-derive data_access_level for academic-paper and academic-paper-reviewer (#773) (#793)
* docs: re-derive data_access_level for academic-paper and academic-paper-reviewer under the dirtiest-input rule (#773)

Applying the #756 derivation to the two carried-over pins the lint
docstring flagged as un-derived:

- academic-paper: redacted -> raw. Standalone modes ingest ungated user
  drafts and third-party reviewer comments, and literature_strategist's
  search-fills-gap flow ingests external-index search results inside the
  skill. The former value described only the post-Gate-2.5 pipeline path.
- academic-paper-reviewer: verified_only -> raw. The standalone
  /ars-reviewer entry (Routing Step 1 routes "review my paper" directly)
  legitimately consumes an ungated pasted manuscript; the rule
  quantifies over ALL entry paths. Pipeline positioning unchanged.
- deep-research: raw survives by a ceiling argument (no derivation can
  dirty the dirtiest value); recorded so no pin remains an un-derived
  carryover.

EXPECTED_LEVELS provenance note rewritten per-pin; ARCHITECTURE §4
diagram + rules now separate the per-skill intake annotation from the
per-stage output data level (§3 column, unchanged). Declarative only.

Closes #773

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq

* review: address codex findings on #773 — precise gate-sequencing claims, deep-research derivation

- academic-paper's former 'redacted' is described as the orchestrated
  pipeline path (Stage-1 sanitized inputs), not "post-Gate-2.5" — Stage 2
  precedes Gate 2.5.
- academic-paper-reviewer's former 'verified_only' is stated as at best
  true for the initial Stage 3 dispatch; Stage 3' re-review consumes a
  freshly revised manuscript before Stage 4.5.
- deep-research's raw is re-affirmed on its actual inputs (raw queries +
  unverified search results); the ceiling argument becomes supplementary
  rather than the derivation itself.

Applied consistently across the lint docstring, ARCHITECTURE §4, and the
CHANGELOG entry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 13:26:07 +08:00
Edward Cheng-I Wu adc38300da feat: register write-scope guard launcher degradations in degradation_registry.json (#769) (#792)
* feat: register write-scope guard launcher degradations in the degradation registry (#769)

The registry claims to index every graceful-degradation mechanism in the
suite, but hooks/run_guard.sh's documented degraded states had no rows,
and the #757 prose table in docs/CONTROL_AVAILABILITY.md stood up a
second, unpinned authority for those facts.

Four write_scope_guard_* rows added (no-python, no-git-bash,
no-timeout-binary, subprocess-misbehaves), each with verbatim D3
authority anchors into hooks/run_guard.sh + the README Requirements
bullet; pinned_by names scripts/test_run_guard_launcher.py where a
CI-executable pin exists (the Windows-without-Git-Bash path never
executes the launcher, so its row honestly carries no pin). Registry
1.2.0 -> 1.3.0; _EXPECTED_MECHANISMS updated in the same commit (D5
lock semantics). The CONTROL_AVAILABILITY degradations table now
declares itself a convenience summary backpointing at the registry.

Closes #769

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq

* review: address codex findings on #769 — permission phrasing, no-timeout decision forwarding, launcher-internal failure coverage

- Rows no longer claim "writes are never blocked" or relitigate what an
  'allow' decision would do: the launcher emits no permissionDecision,
  so the session's normal permission rules still decide.
- The no-timeout row's terminal_policy_effect states that the healthy
  watchdog fallback forwards the guard's real decision (including deny);
  only an overrun resolves to pass-through.
- The misbehaves row now also covers the two remaining documented
  launcher-internal degradations (SELF_DIR self-resolution failure and
  the POSIX payload-length cap on multi-megabyte Writes), with verbatim
  anchors; CHANGELOG + CONTROL_AVAILABILITY backpointer updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq

* review: round-2 codex findings on #769 — per-row quantifiers, actual validity-check shape, payload-edge honesty

- terminal_policy_effect now speaks per failure path, not "every degraded
  path"; the no-Git-Bash row states the hook simply does not run.
- The misbehaves row names the launcher's ACTUAL validity check (a JSON
  object carrying a top-level hookSpecificOutput key — deliberately
  shallow, not full hook-schema validation).
- The multi-megabyte payload edge is recorded as a documented accepted,
  untested case with no pinned outcome — no deterministic claim.
- The no_python row's authority anchor swaps to the no-permissionDecision
  pass-through line (the launcher's disputed 'allow' comment is
  pre-existing text this PR neither adds nor endorses).
- CONTROL_AVAILABILITY prose quantifier fixed to match ("none of which
  ever blocks", with the no-timeout forwarding stated).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq

* review: round-3 codex findings on #769 — payload edge split into its own no-pinned-outcome row

- write_scope_guard_payload_capacity_edge becomes a dedicated row whose
  every field honestly declares "no pinned outcome" — the misbehaves
  row's pass-through claims are now unconditionally true for its own
  failure classes (registry 20 -> 21 rows, D5 lock updated).
- CONTROL_AVAILABILITY prose reworded: degraded states never INTRODUCE a
  block; the no-timeout swap keeps the guard operating normally (real
  decisions, including deny, still apply).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 12:54:36 +08:00
Edward Cheng-I Wu edb0265301 refactor: consolidate markdown-stripping helpers into scripts/_markdown_lint_util.py (#771) (#791)
* refactor: consolidate the duplicated markdown-stripping helpers into scripts/_markdown_lint_util.py (#771)

The #757/#758 defrift lints shipped two diverging copies of the markdown
non-rendering semantics; check_risk_register.py was already importing the
siblings' private helpers as a stopgap. One shared module now carries the
grammar; per #771 the #770 superset rules (inline code-span stripping +
image exclusion in the link grammar) win, so CA-1..CA-3 inherit them too.

Pure refactor, no invariant change. All three mutation suites (73 tests)
pass unchanged against the shared module; the three lints pass on the
real tree.

Closes #771

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq

* review: address codex findings on #771 — honest CHANGELOG framing, #759 attribution, three CA grammar tests

- CHANGELOG no longer calls the change a pure refactor: the CA-side
  code-span/image grammar alignment is behavior-visible (strengthening
  direction) and is stated as such.
- _markdown_lint_util docstring attributes check_risk_register to #759
  (it shipped there; #760 is the governance change).
- Three new CA mutation tests pin the inherited rules: broken image
  target / backticked pseudo-link do not fire CA-1; an image form of the
  README inbound link does not satisfy CA-3 (73 -> 76 tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq

* simplify: complete the #771 consolidation per four-angle cleanup review

- github_slug / heading_slugs move into _markdown_lint_util.py — the last
  cross-lint private import (check_risk_register -> check_control_
  availability._heading_slugs) is gone; each lint now imports only the
  shared module plus genuinely data-owning dependencies.
- New links_to() predicate absorbs the three copy-pasted CA-3/DF-3/RR-3
  inbound-link loops; link_targets()/code_spans() named wrappers replace
  raw regex exports (regexes back to module-private); NON_RELATIVE_LINK_
  PREFIXES replaces repeated literal tuples.
- Fence state machine collapses to a single `fence` variable; helper
  strip passes back to module-private; extract_link_targets uses findall.
- The grammar gains a direct, manifest-registered test suite
  (test__markdown_lint_util.py); the two CA tests that re-pinned grammar
  already pinned in the DF/RR suites are dropped, keeping the CA-3
  image-exclusion mutation test. Consolidation history now lives in the
  CHANGELOG once instead of four prose sites.

87 tests green across the four suites; all three lints PASS on the tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq

* docs: qualify the cross-lint import claim (codex P3) — markdown-helper imports only

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015NZwcSFBwiJBZEtsSTcCxq

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 12:53:31 +08:00
Edward Cheng-I Wu b6062c1401 feat: first Promotion Bakeoff run — gpt-5.6-sol validated for the codex subscription transport (#788)
* feat: first Promotion Bakeoff run — gpt-5.6-sol validated for the codex subscription transport (#787)

Probe set: 30 refs (10 easy DOI-keyed journal articles; 10 hard: 3 arXiv,
2 DOI-less NeurIPS, 5 non-English; 10 fabrications), every real row
resolver-confirmed same-day, every fabrication negative-checked. 180
same-day paired calls (30 x 3 repeats x 2 models), majority verdicts.

Result: all five measures PASS with superiority — recall 1.00 vs 0.80,
grounded completion 0.933 vs 0.900, p95 latency 26.5s vs 58.7s, zero
guard misfires, false disagreement 0.00 = 0.00. Transport-qualified:
gpt-5.6-sol stays provisional on the first-party API route (jq guards
unexercised; allowlist unchanged). Report + probe-set sha256 under
audits/; per-call index committed beside the probe set.

Campaign side-product (transport): page-open webSearch items
(action.type != "search") are skipped for binding instead of failing the
stream (opened-page URLs still can never become bound sources), and
DEVELOPER_INSTRUCTIONS requires an empty sources array for
NOT_FOUND/NOT_SEARCHED. 52 transport tests green. Defective-tool run 1
archived unscored; three probe-row transcription errors were flagged
MISMATCH by both models, independently re-verified, corrected, re-run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: narrow the page-open exemption to the observed action.type == "other" shape (#788 codex P2)

An empty action object, unknown action type, or non-dict action on a
webSearch item is stream-fatal again; only the observed page-open shape
is skipped. Mutation test sweeps four bad shapes (52 -> 53 tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: anchor the page-open exemption to the first-party closed WebSearchAction set

Run-3 surfaced a third real shape ({"type": "openPage", "url": ...}) that
the single-observation exemption rejected, tool-suppressing the baseline's
measures (13 EVENT_STREAM_INVALID cells). The exempt set is now exactly
the non-search members of the app-server protocol's closed WebSearchAction
oneOf — {other, openPage, findInPage} plus the Responses-API spellings —
verified against `codex app-server generate-json-schema` on 0.147.0.
Unknown shapes stay stream-fatal (mutation sweep unchanged); 54 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* docs: score preregistered run 4 as the gate result; runs 1-3 recorded as exploratory

Run 4 (frozen fixture @ 3fc6ddb, frozen parser @ c9c865d, both pushed
pre-run): all five measures PASS, zero misfires on BOTH models, recall
1.00 vs 0.80, grounded completion 0.933 vs 0.867, p95 27.5s vs 51.1s.
Report rewritten with the preregistration statement and the full
exploratory-round accounting; call index replaced with run-4 data;
claim surfaces and CHANGELOG updated to run-4 numbers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* docs(code): pin the bare-discriminator decision against the first-party schema (#788 codex round-2 P2 rejected with evidence)

The round-2 finding claimed openPage/findInPage require url/pattern; the
protocol schema (generate-json-schema, 0.147.0) marks every non-search
variant required:["type"] with url/pattern nullable optionals. Demanding
optional fields is the exact false-fatality class that invalidated
bakeoff runs 1 and 3. Decision recorded in the comment and pinned by
bare-discriminator test rows (54 tests, +2 param rows).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix+docs: close codex round-3 findings — ordering, replayability, exposure analysis

P2 (ordering): webSearch action-shape validation now runs BEFORE the
MODEL_RETURNED_NOT_SEARCHED early return, so a model NOT_SEARCHED verdict
can never mask response-shape drift; mutation test added (55 tests).
P2 (replayability): the 180 full receipt rows, the offline scorer
(verified to reproduce the gate byte-for-byte from committed artifacts
alone), and the parameterized fleet runner are committed beside the
probe set; the report states the replayability boundary plainly (raw
event streams are digest-only by transport design).
P1 (answer-key exposure): empirical scan across all 540 retained
receipts finds zero repo-referencing bound queries/sources; report gains
an exposure-analysis section with scope caveats and corroboration; the
structural fix (sealed hash-commit preregistration, fresh fabrication
pool per run) is filed as #789 for future bakeoffs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-4 P2s — fleet gate, timeout margin, fresh-probe rule, nearest-rank p95

Scorer refuses truncated/duplicated/partial fleets (exactly one row per
(ref_id, repeat) across 30x3) before computing any measure; p95 moves to
the nearest-rank order statistic (51.13/27.46 -> 51.20/28.09, matching
the review's own recomputation; gate unchanged) and the method is named
on every surface. Runner outer timeout raised above the transport's
inner 300s deadline so its finally-block cleanup always fires first.
The canonical recorded-run note and report outcome now require a FRESH
probe set for the API-route run per #789 (this set's labels are public),
resolving the self-contradiction with the exposure analysis.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-5 P2s — scorer consumes runner output + binds receipt identity

score_run.py now scores either the committed run-4 JSONLs (default) or a
fleet-runner output directory (argv[1]), so a reproduced fleet can never
silently re-report the old result; every scored row must pass identity
binding (outer model/ref/repeat, receipt.model, receipt.request_id, and
a request_digest recomputed from the probe set via the transport's
canonical form), refusing mis-associated or edited fleets. Verified:
committed data reproduces the gate unchanged, results-dir mode scores
the live run-4 cells, and a cross-model receipt swap is refused.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-6 P2s — dual-fleet hard-zero + unambiguous shape code

Measure 4 now requires zero guard misfires in BOTH fleets (a baseline
suppressed by tool misfires cannot anchor a fair comparison — the run-3
lesson, now enforced by the scorer). The transport emits
EVENT_STREAM_INVALID for a non-null non-list search `results` value
instead of silently skipping into NO_BOUND_SEARCH_RESULTS
(wrong_search_shape fixture expectation updated in lockstep), and the
scorer's shape family is trimmed to exactly the emitted shape codes with
the behavior-family classification documented. Provably no effect on the
scored run: run-4 contains zero rows in any affected code family (only
SOURCE_NOT_IN_SEARCH_RESULTS 12/6, a behavior code) and both fleets
already sit at zero misfires; all gate numbers unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-7 P2s — results-shape check before verdict return + same-day fleet enforcement

The results-shape validation joins the pre-verdict scan loop so a model
NOT_SEARCHED verdict can never mask dict-shaped search results (mutation
test added: wrong_search_shape + NOT_SEARCHED -> EVENT_STREAM_INVALID;
56 tests). The runner refuses to resume over cells from an earlier date,
and the scorer refuses mixed-date fleets across both models (run-4 is
single-date; gate numbers unchanged).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-8 P2s — complete pre-verdict search validation + runner failure semantics

The pre-verdict scan now performs the COMPLETE search-item strict
validation (query type/length/control chars + results shape), covering
legacy action-less items — the round-7 placement validated only
action-typed items, which also made the round-7 mutation test fail (a
red test my verification pipeline masked via tail; committed here only
with PYTEST_EXIT=0 verified directly). Runner: a fleet with any failed
call now exits nonzero instead of printing ALL DONE, and an outer-
timeout kill sweeps the adapter's orphaned temp dirs (the detached
app-server exits on stdin EOF; the ephemeral auth copy is what the
verifier's skipped finally-block would have removed). 56 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-9 P2s — NOT_SEARCHED source contract, retry-not-skip, quiescent sweep, dated rows

Transport: NOT_SEARCHED with a populated sources array fails closed as
FINAL_OUTPUT_INVALID before the early return (mutation test; 57 tests).
Runner: a same-day cell that recorded a failure is discarded and retried
on resume instead of silently counting as complete; the orphan-tempdir
sweep runs only after the executor drains so it can never delete a live
worker's ephemeral CODEX_HOME. Scorer: every row must carry a real ISO
date — an undated fleet cannot satisfy the same-day gate on empty
strings. Committed run-4 data re-verified green under all new gates.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-10 findings — single-path validation, scorer gate exit, probe-hash pin

P1 root treatment: the entire search-processing pipeline (cap, strict
per-item validation, reference-bound filter, URL binding incl. the
result-entry object-shape check) now runs BEFORE any verdict branch, so
every shape-fatal path fires identically regardless of the model's
answer — the verdict-masking bug class (rounds 3/7/8/9/10) is closed by
construction, not by another patch. Emptiness outcomes stay verdict-
conditional (an honest NOT_SEARCHED with no bound search remains model
behavior). Mutation test: bound search with a non-object result entry +
NOT_SEARCHED verdict -> EVENT_STREAM_INVALID (58 tests).
Scorer: refuses a probe set whose whole-file sha256 differs from the
frozen hash (labels now inside the scoring identity), and exits nonzero
when any gate fails. The P1's rerun demand is accepted: a run-5 fleet
under this frozen parser follows as the scored gate run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* docs: run 5 under the frozen final instrument is the gate result (#788 codex round-10 P1 accepted)

Fleet rerun 2026-08-20 under parser+scorer db6ed67 (pushed pre-run):
all five measures PASS with zero misfires on both fleets — recall 0.90
vs 0.80, grounded completion 0.911 vs 0.889, p95 26.1s vs 47.6s. Run 4
reclassified as a prior-instrument exploratory round; committed
receipts/index/scorer default swapped to run-5 data (committed scorer
replays the gate from repo artifacts alone, exit 0); all claim surfaces
carry run-5 numbers and the cross-fleet consistency note (candidate led
measures 1/2/5 in every full paired fleet).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-11 P1s — no failed-trial erasure + full-fleet entry validation (instrument fixpoint)

Runner: a recorded failed trial is never deleted on resume; it is
carried into the failure count and forces a nonzero exit, so the only
path past a failure is rerunning the ENTIRE fleet fresh — selective
retry-until-green is structurally impossible. (Provably no scored fleet
was affected: runs 4 and 5 each completed in a single invocation with
zero failures and no retry/carried lines in their logs.)
Transport: the strict pre-verdict loop now validates every consumed
field of EVERY search item — id, query, results-list shape, and each
entry's object shape, bound or unbound — reaching the instrument
fixpoint: no field the pipeline reads is unvalidated, so no future
verdict-masking variant of this class exists. Mutation tests for
unbound-malformed-entries and id shapes (60 tests). A run-6 fleet under
this frozen instrument follows as the gate run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* docs: run 6 under the fixpoint instrument is the gate result

Fleet rerun 2026-08-20 under adf18f9 (parser fixpoint + no-erasure
runner + full-gate scorer, all pushed pre-run): 5/5 PASS, zero misfires
both fleets — recall 1.00 vs 0.70, grounded completion 0.911 vs 0.867,
p95 28.8s vs 43.3s. Run 5 reclassified prior-instrument; artifacts and
scorer default swapped to run-6; leak scan clean across all 900 retained
receipts (runs 2-6); candidate led measures 1/2/5 in all four full
paired fleets.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-12 P2s — pinned effort, fleet-private temp root, split disclosure

Runner pins ARS_CROSS_MODEL_REASONING_EFFORT to the provider default
(explicitly unset per call, recorded per row) and routes all transport
temp dirs into a fleet-private mkdtemp root so the timeout sweep can
never touch another invocation's dirs. The audit now discloses the one
run-6 1-1-1 split (baseline fab-05: MISMATCH/NOT_SEARCHED/NOT_FOUND ->
INDETERMINATE, conservative miss) and names the actual baseline misses
(fab-01, fab-08 majority NOT_SEARCHED; fab-05 split) — verified against
the committed receipts, correcting a stale carried-over sentence. The
effort variable was verified unset for every fleet (shell env + profile
carry no export); gate numbers unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-13 P2s — receipt-contract validation, effort-marker gates, always-sweep

Scorer validates every non-null receipt against the closed contract
(required keys, verdict/searched types, positive-verdict grounding with
fully-bound sources, empty sources on NOT_FOUND/NOT_SEARCHED, queries
present when searched) before any metric trusts it, and results-dir
scoring requires the pinned-effort marker on every row (committed
gate-run rows predate the marker; the audit attests their configuration).
Runner refuses carried cells without the marker and sweeps the
fleet-private temp root in a finally-block on every outcome — a
signal/OOM-killed verifier no longer leaves its ephemeral auth copy.
Committed gate scoring still exits 0 unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-14 P2s — per-verdict receipt invariants + runner rejects malformed parsed receipts

Scorer: grounded verdicts (VERIFIED/MISMATCH/NOT_FOUND) require
searched=true and a null reason_code; NOT_SEARCHED requires
searched=false and a reason from the transport's closed emitted set —
a fabricated NOT_FOUND-without-search or NOT_SEARCHED-with-search row
can no longer contribute to recall or completion. Runner: a verifier
exiting 0 with parsed-but-malformed output records RECEIPT_INVALID,
counts as a failure, and forces nonzero exit. Committed run-6 gate
scoring re-verified: exit 0, numbers unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-15 P2s — uniform item-field validation + shared receipt contract (axis terminal)

Transport: every webSearch item — page-opens included — now has its
action payload validated against the closed WebSearchAction variant
types (url/pattern/query string-or-null, queries string-array), plus
uniform id and results/entry shape checks; a recognized discriminator
with a wrong-typed payload is stream-fatal (mutation sweep; 60 tests).
Tooling: the full closed receipt contract (exact key set, transport/
auth_mode/containment, digest formats, per-verdict cross-field
invariants) moves into a shared receipt_contract.py imported by BOTH
run_fleet.py and score_run.py — one implementation, applied to fresh
cells, resumed cells, and every scored row, so the two consumers cannot
diverge. This terminates the validation axis: every field of every
webSearch item and every key of every receipt is now checked; committed
run-6 gate scoring re-verified exit 0 with numbers unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix+docs: close codex round-16 — strict source bindings; instrument-freeze boundary pinned (P1 declined with recorded rationale)

receipt_contract.py enforces the full canonical source-binding shape
(closed 4-key object, non-trivial https URL, non-empty item id,
result_index 0-127 with bool exclusion) — committed run-6 data passes
unchanged. The round-16 rerun demand is DECLINED under a pinned
maintainer boundary, recorded in the report and the canonical
recorded-run note: runs 4/5 were discarded because consumed-data gaps
could alter scored outcomes; post-run-6 hardening validates only
surfaces outside every consumed path and cannot change any verdict,
binding, latency, or measure of a past fleet — such hardening applies
from the next fleet. The disagreement is recorded, not hidden.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-17 P2s — complete schema mirror + latency-sample validation

receipt_contract.py is now a COMPLETE stdlib mirror of the canonical
receipt schema: identifier/event-id/https-url patterns and length
bounds, array caps (queries<=32, sources<=16), closed entry objects,
auth_mode const, NOT_SEARCHED => empty queries+sources with a mandatory
reason, unknown reason codes refused globally. Scorer refuses boolean,
negative, non-numeric, or absurd wall_seconds before the percentile
gate. Committed run-6 gate scoring re-verified: exit 0, numbers
unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-18 P2s — hashability, boolean identity, SIGTERM cleanup

Validator: array/object verdict or reason_code becomes a contract
failure instead of an uncaught TypeError (which would have escaped the
runner's SystemExit handling and re-opened the no-reroll gap);
containment flags are checked by identity (`is True`) so integer 1
cannot satisfy the schema's boolean constants. Runner: SIGTERM/SIGINT
raise SystemExit so the finally-block sweep of the fleet-private auth
copies also runs on cancellation. Mutation-verified (3/3 caught);
committed run-6 scoring exit 0 unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-19 P2s — discriminator hashability, integer repeats, real ISO timestamps

Transport: the WebSearchAction discriminator is type-checked before set
membership in both _is_page_open and the uniform loop — an array/object
type fails closed as EVENT_STREAM_INVALID instead of crashing the
verifier past shape accounting. Scorer: repeat must be an exact int in
1..3 (1.0 satisfied the completeness Counter while minting ...-r1.0),
and ts must parse as a full ISO timestamp with offset instead of a
digit-shaped prefix. Committed run-6 scoring exit 0 unchanged; 60
transport tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-20 — source-query binding, midnight guard; null-action P2 declined with schema evidence

receipt_contract.py rejects sources whose search_item_id has no retained
entry in search_queries (unretained evidence never counts as grounding).
run_fleet.py fails visibly BEFORE reporting success when a fresh fleet's
cells span two calendar dates. The explicit-null-action P2 is declined
on first-party evidence: ThreadItem types action as
anyOf[WebSearchAction, null] (generate-json-schema, 0.147.0), so null is
protocol-legal and follows the legacy path where the item still faces
the complete strict validation — fatal-izing it is the run-1/run-3
false-fatality class; decision pinned in the code comment. Run-6
scoring exit 0 unchanged; 60 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-21 — cancellation stops queued quota burn, worker failures persist; open-variant P2 declined with schema evidence

P1: SIGTERM/SIGINT set a STOP event making every queued job a no-op
(marked [CANCELLED], counted as failure), so shutdown waits only for
the at-most-3 in-flight calls instead of burning the rest of a paid
180-call fleet. P2: a worker exception after the paid call persists a
failed cell with the job identity, so a resume can never treat the
consumed trial as missing and re-roll it. The closed-variant P2 is
declined on first-party evidence: no WebSearchAction variant sets
additionalProperties, so extra fields are schema-legal and rejecting
them would make any future informational field fleet-fatal; decision
pinned in the code comment, known fields stay type-checked.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-22 P1 — stray STOP=None placeholder no longer nullifies the cancellation event

The round-21 placeholder assignment landed AFTER the Event creation in
module order, resetting STOP to None and disabling queued-call
cancellation exactly as the review read it. The placeholder is removed;
the Event created before worker start is the one the signal handler
sets. Static check pins that no STOP=None assignment remains.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-23 P2 — refuse contradictory receipt+error rows

A row carrying both a valid receipt and a truthy error is structurally
impossible from the runner and is refused as corrupted/external instead
of being scored as grounded evidence; error rows with a null receipt
stay counted as misfires. Committed run-6 scoring exit 0 unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-24 P2s — output-contract precedence, portable probe digest, runner preflight

Transport: NOT_FOUND carrying sources is FINAL_OUTPUT_INVALID even when
the stream also lacks a bound search — output-contract violations now
outrank emptiness outcomes so the shape event cannot be misfiled as a
behavior code. Probe digest verification moves into the shared module
with CRLF->LF normalization (a Windows autocrlf checkout is not probe
drift) and the runner runs the same preflight BEFORE any paid call, so
180 subscription calls can never be spent on a fixture the scorer will
refuse. Run-6 scoring exit 0 unchanged; 60 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-25 — scorer-equivalent resume preflight + exclusive fleet lock

validate_row (shared receipt_contract) now carries the COMPLETE row
validation — contradiction, outer identity, exact-integer repeat, full
ISO timestamp, sane latency, receipt identity binding to the probe row,
closed receipt contract — and is the single implementation used by both
the scorer and the runner's resume preflight, so a misnamed or copied
cell fails before any further quota is spent. The runner takes an
exclusive flock on the output dir, refusing a second concurrent
invocation that would duplicate paid calls and race cell writes.
Committed run-6 scoring exit 0 unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-26 — URL-binding drift hits measure 4; cross-platform fleet lock

Transport: a bound search whose non-empty result entries yield no
extractable URL is EVENT_STREAM_INVALID (provider moved/renamed the URL
key = response-shape drift), no longer the behavioral
NO_BOUND_SEARCH_RESULTS; the pre-existing pin of the old classification
is updated in lockstep and a canonical_url regression test added (62
tests). Runner: the fleet lock falls back to msvcrt.locking on Windows,
keeping the documented reproduction path viable.

Note: run-6's receipts contain zero NO_BOUND_SEARCH_RESULTS /
SOURCE_NOT-with-empty-binding rows of the reclassified kind (reason
distribution: only SOURCE_NOT_IN_SEARCH_RESULTS with non-empty
bindings and clean rows), so the gate numbers are provably unaffected;
the change also falls under the pinned instrument-freeze boundary.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-27 P2s — key-drift/value split, full resume validation, Windows invoke, UTF-8 I/O

Transport: EVENT_STREAM_INVALID for empty bindings now requires that NO
recognized URL key exists across the bound entries (true key drift); a
recognized key with an unusable value stays behavioral — both sides
test-pinned (62 tests). Runner: every resumed cell, failed ones
included, faces validate_row + the effort check before further quota is
spent; the transport is invoked via sys.executable (Windows honors no
shebang); all subprocess/artifact text I/O pinned to strict UTF-8 in
runner and scorer. Run-6 scoring exit 0 unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-28 P2 — URL-key drift determined pre-verdict

The key-drift determination (bound entries present, no recognized URL
key anywhere) moves before the NOT_SEARCHED early return, so a model
NOT_SEARCHED answer can no longer mask renamed-URL-key response drift;
the post-verdict emptiness branch keeps only behavioral outcomes.
Masking regression test added (63 tests); run-6 scoring exit 0
unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-29 P1 — fleet runner gated to POSIX

The #630 transport's process-group containment (start_new_session +
os.killpg in _stop_process) is POSIX-only, so a native-Windows fleet
would consume paid calls while every cell fails during cleanup — the
rounds-26/27 surface accommodations implied support the deeper stack
never had. The runner now refuses non-POSIX up front with a WSL
pointer; the dead msvcrt lock branch is removed (the scorer, which is
genuinely offline and portable, keeps its CRLF-tolerant digest and
UTF-8 reads). Run-6 scoring exit 0 unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix+docs: close codex round-30 — receipt-level invariance proof replaces live-validation claim; session isolation; stray-file preflight

P1 resolved by PROOF rather than a third rerun: run 6's 180 receipts
carry exactly two reason states (null; SOURCE_NOT_IN_SEARCH_RESULTS
12/8) with zero receipt-less, error, NO_BOUND, NO_REFERENCE,
FINAL_OUTPUT_INVALID, MODEL_RETURNED_NOT_SEARCHED, or
EVENT_STREAM_INVALID rows — each post-run-6 transport change either
touches unconsumed surfaces or only relabels cells in code families
that provably never occurred, so no run-6 cell can differ under the
shipped parser. The stale "run 6 live-validates the shipped parser" and
"instrument FIXPOINT / no masking path remains" sentences are replaced
with the precise provable statements; the freeze policy now REQUIRES
this proof standard (no proof on a consumed path = rerun, as runs 4/5
were). P2s: verifier subprocesses start in their own session so an
interactive Ctrl-C cannot turn in-flight calls into resume-poisoning
EXIT failures; the runner refuses unexpected result files before
spending quota.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-31 P2s — orphaned .tmp cells refused, fresh cells identity-bound

The preflight refuses orphaned atomic-write temp files (a crash between
write and rename must not silently re-roll a completed paid trial), and
fresh cells face the same validate_row identity binding as resumed
cells and the scorer before being persisted as success — a receipt with
the wrong model/request_id/digest becomes a recorded RECEIPT_INVALID
failure. Run-6 scoring exit 0 unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* feat: counterbalanced interleaved scheduling for the bakeoff fleet (#788 round-32 P1 accepted)

The two models' calls for each (reference, repeat) cell are adjacent in
the queue with deterministic parity-alternating pair order, so model
identity is decorrelated from execution time — provider load or
web-search drift during the fleet can no longer masquerade as a model
effect. A counterbalanced run-7 follows as the scored gate run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* docs: counterbalanced run 7 is the gate result; run 6 superseded for order confound

Run 7 under frozen instrument 69cd04a (interleaved parity-alternating
pair scheduling): all five measures PASS — recall 0.90 vs 0.80, p95
25.0s vs 49.6s (median 14.8 vs 17.5), grounded completion tied at
0.900, zero misfires both fleets, two 1-1-1 splits disclosed and scored
as conservative misses. Honesty note carried on every claim surface:
the sequential fleets' grounded-completion edge did NOT survive
counterbalancing and is not claimed; recall and latency led in all five
paired fleets. Artifacts and scorer default swapped to run-7 (committed
scorer replays the gate, exit 0); leak scan clean across 1,080 retained
receipts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-33 P1 — resume refuses half-complete counterbalanced pairs

Every (reference, repeat) pair must be wholly present or wholly missing
on resume: a one-sided pair would run the counterpart far from its
partner and silently reintroduce the model-vs-time confound. Run 7 is
unaffected (single uninterrupted invocation); scoring exit 0 unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* docs: close codex round-34 P2 — recorded-run note limited to the measured superiority (2 and 5, tie on 1)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 07:13:23 +08:00
Edward Cheng-I Wu 5714f3a3eb fix: repair codex subscription transport against three codex-cli 0.147.0 drifts (#785) (#786)
1. Attestation stream: `codex login status` emits "Logged in using ChatGPT"
   on stderr in non-TTY invocation; detection read stdout only, so every
   detect failed AUTH_NOT_CHATGPT_SUBSCRIPTION. Accept the exact line on
   either stream (same idiom as the #684 harness); stderr-emitting
   fake-codex regression test added.
2. Structured-output schema: the provider now rejects "uniqueItems"
   (invalid_json_schema, HTTP 400). Dropped from the provider-sent
   MODEL_OUTPUT_SCHEMA; the local validator already refuses duplicated
   source URLs fail-closed.
3. Disable list: --disable code_mode_host silently removes the standalone
   web-search tool on this build (search executes through the code-mode
   host; isolated by live bisection), failing every call closed as
   MODEL_RETURNED_NOT_SEARCHED. The host is no longer disabled; code_mode
   stays disabled and the forbidden-event scan still rejects any item type
   outside the {userMessage, reasoning, agentMessage, webSearch} allowlist.

Live cross_model_smoke_test_codex.sh: PASS for gpt-5.5 and gpt-5.6-sol
(2026-08-19). 51 transport tests green; #630 guard green.


Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 20:12:35 +08:00
Edward Cheng-I Wu 075390a5c3 docs: cross-model recommendation surfaces follow generation currency (#784)
* docs: cross-model recommendation surfaces follow generation currency (#783)

Recommendation decoupled from validation status: gpt-5.6-sol (current
OpenAI flagship) becomes the lead OpenAI example while staying
provisional — a dated lifecycle note records the flip carries no
measurement claim; the Promotion Bakeoff remains the only route to
validated. gpt-5.5 / gpt-5.5-pro demoted to validated previous-generation
rows (measured bakeoff baseline unchanged). Gemini 3.1 Pro stays
recommended (first-party check 2026-08-19: still Google's most capable
Pro model). SETUP en/zh-TW quick-setup blocks updated in lockstep
(parity lint green); the #630 guard's recommendation witness re-pinned
to the new policy sentence with its mutation test updated in the same
commit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-1 P2s — evidence-ceiling wording + policy-body witness

P2-1: "measured bakeoff baseline" overstated the evidence (no bakeoff run
has ever been recorded); gpt-5.5 is the designated baseline, validated =
allowlist status only. Reworded on all four surfaces (canonical doc,
SETUP en/zh-TW, CHANGELOG).
P2-2: the #630 recommendation-policy witness pinned only the heading; the
guard now pins the two load-bearing body clauses (no-measurement-claim,
bakeoff-only route to validated) with mutation tests for each (29 -> 31).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* fix: close codex round-2 P2 — superiority claim requires an observed measure

The rewritten outcome bullet's "or operational benefit" branch let a
measured-superiority claim rest on an unmeasured benefit; superiority now
requires observed superiority on one of the five measures, and operational
benefits are scoped to recommendation policy. Header no longer says "two
distinct promotions" for what is now one promotion plus a claim rule.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

* docs: close codex round-3 P2 — name the subscription-transport exception

The citation-only codex subscription blocks (canonical + SETUP en/zh-TW)
keep gpt-5.5 deliberately; a comment now names this as a transport-specific
exception to the generation-currency recommendation and points at the codex
smoke test before swapping ids.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mo3QXHj2yzwNQ3VQKVEM2j

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 18:00:30 +08:00
Marc-oss-hub 3f14c8e16f Add OrcaRouter to THIRD_PARTY.md community directory (#782)
Adds a listing row under 'Listed projects' for OrcaRouter, an OpenAI- and
Anthropic-compatible gateway that can serve as the cross-model verification
provider via ARS_OPENAI_COMPAT_BASE_URL + ARS_CROSS_MODEL with namespaced
model IDs. Credits ARS and describes the integration faithfully, per the
'How to get listed' bar. Closes #781.

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-19 17:21:08 +08:00
Edward Cheng-I Wu 2b639c12ee docs(release): v3.21.0 — ISO/IEC 42001-spirit track [skip-closes-check] (#779)
* docs(release): v3.21.0 — ISO/IEC 42001-spirit track [skip-closes-check]

- CHANGELOG: promote [Unreleased] -> 3.21.0 (2026-08-18); backfill the six
  uncovered merges (#762 assessment, #757/#768, #758/#770, #755/#774,
  #754/PR #763 ref, #764 hotfix); fresh empty [Unreleased] anchor
- version bump across all pinned surfaces: 5 READMEs (badge, pipeline
  section, new v3.21.0 what's-new block), MODE_REGISTRY, .claude/CLAUDE.md
  (suite version, Last Updated, v3.21.0 Key Additions), CITATION.cff
  (version + date-released), POSITIONING citation prose, plugin.json,
  marketplace.json, academic-pipeline/SKILL.md (+#528 content-lock rehash),
  ARCHITECTURE (7 current-component markers + Last Updated),
  check_spec_consistency pins + test fixtures
- READMEs x5: live pipeline-guarantee sentence reworded from 'cannot be
  skipped' to the #753 matrix-licensed form (MANDATORY, no unrecorded
  bypass, override reasoning recorded); historical release-notes sections
  untouched

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC

* docs(release): fix stale-tense sentence in #754 changelog entry (codex R1) [skip-closes-check]

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
v3.21.0
2026-08-18 13:45:36 +08:00
Edward Cheng-I Wu 4b8427f6d5 docs: solo-maintainer governance statement + SECURITY triage procedure (#760) (#778)
* docs: solo-maintainer governance statement + SECURITY triage procedure (#760)

GOVERNANCE.md: decision authority, honest cross-model scope
(error-detection control, not organizational independence), release
authority, EOL posture, and the operating-principles section (three
distilled principles with informative ISO/IEC 42001 anchors, the
#753-#760 coverage map, Annex C not-applicable assessments).
SECURITY.md: 7-day promise becomes acknowledgement-only hard promise
plus a written severity-tiered best-effort triage procedure.
NOTICE.md: governance pointer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC

* docs: apply dual-agent review to #760 (order-safe claims, honest ceilings, sustainable promises)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC

* docs: close codex R1 on #760 (per-operator cross-model attribution, fairness signals disclosed, EOL-scoped promises)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC

* docs: narrow fairness scoring absolute to persons/real-world allocation (codex R2)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 10:57:38 +08:00
Edward Cheng-I Wu c5f0ac69af feat: lightweight risk register + RR-1..3 mirroring lint (#759) (#777)
* feat: lightweight risk register + RR-1..3 mirroring lint (#759)

docs/RISK_REGISTER.md links ten standing risks to existing controls,
matrix-mirrored evidence statuses, and residual gaps with tracking
issues. scripts/check_risk_register.py pins pointer integrity (RR-1),
verbatim status mirroring against stage_capability_matrix.json with a
malformed-citation guard (RR-2), and README discoverability (RR-3);
12 mutation tests, wired into spec-consistency CI and the pytest
manifest. Helpers imported from check_data_flows / check_stage_capability_matrix rather than copied (#771).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC

* refactor: apply /simplify review to #759 lint (reuse sibling helpers, close side-doors)

- import _LINK_RE/_CODE_SPAN_RE/_strip_non_rendering from check_data_flows
  (capture group added there, harmless for its .sub use), _load/_status
  from check_stage_capability_matrix, github_slug/_heading_slugs from
  check_control_availability - no third copies (#771)
- RR-1: repo-containment + anchor-fragment validation (CA-1 parity)
- RR-2: asserted-status ceiling (matrix is sole authority for
  MEASURED/MIXED), inventory lock on shipped matrix-row citations
  (D5/M13 house pattern), per-segment malformed-citation reporting
- RR-3: resolved-path + code-span-stripped inbound link (DF-3 parity)
- missing-doc single fatal error (sibling contract); ERROR: prefix in main
- tests: 12 -> 18 incl. real-tree pass; _MIRRORED_FILES derived from
  imported constants

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC

* fix: anchor-on-non-markdown-target guard (CA-1 parity) + test; count fixes

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC

* fix: close codex R1 findings (lock scoped to evidence bullets, space-path opt-out removed, R9 pin wording)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC

* docs: align R10 residual-gap wording with the availability matrix (codex R2)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AFJbYZyrJmSPFMJKQZhVHC

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 10:57:18 +08:00
Edward Cheng-I Wu 17bf063456 docs: CI workflow enforcement-class table + inventory lint (#755) (#774)
* docs: CI workflow enforcement-class table + WC-1/WC-2 lint (#755)

docs/ARCHITECTURE.md gains §7.1: all 14 workflows classified by
trigger / what it checks / enforcement class (blocking / advisory /
administrative / post-push detection) / bypass token, with the honest
count line (8 blocking on at least one event class, 2 advisory, 1
administrative, 3 post-push detection) and the explicit statement that
tag workflows detect after the push — their stop-power is the
maintainer acting on the failure. Per-workflow facts verified against
the workflow files (eval-harness ack token + PR-only gating; changelog
gate release/** head scope; pytest path filters; the three tag
triggers).

CONTRIBUTING release-checklist prose now points at the classification
instead of implying uniform CI enforcement.

Lint (same-PR drift-point discipline): check_workflow_classification.py
— WC-1 inventory sync both directions (a new, renamed, or removed
workflow fails CI until the table matches; duplicates refused), WC-2
class cells begin with the closed four-term vocabulary. Class
semantics stay review-owned (degradation-registry posture). 9 mutation
tests; wired into spec-consistency.yml + the pytest manifest.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* refactor: apply /simplify + codex R1 — table accuracy + lint hardening (#755)

Review round (3 cleanup agents + codex gpt-5.6-sol xhigh R1), findings
deduped and applied:

Table accuracy (codex 3 P2 + 1 P3, cleanup F2/F3): every trigger cell
now states its actual branch/path/tag filters (repository-hygiene and
command-invariants had birth-drifted cells; several rows omitted
targeting-main scopes); freshness-check reclassified honestly
(Advisory for staleness, but malformed protocol metadata is a hard
failure); bypass cells say "justification requested, not
machine-validated" (both workflows accept the bare token);
command-invariants "what it checks" gains its other two enforced
checks; bypass column normalized to "none"; the legend absorbs the
tag-workflows sentence and the duplicated qualifier prose is trimmed.

Lint hardening: section extraction switches to the shared
_skill_lint.heading_section (exact full-line heading incl. the #755
anchor, fence-aware — 15 fewer bespoke lines); rows parse once with
escaped-pipe-aware cell splitting; the inventory glob covers *.yaml;
WC-2 matches vocabulary terms as whole words (Blockingg fails); the
arity guard moves under WC-1 with a test; new WC-3 recomputes the
bolded count line from the Class column (the honesty sentence can no
longer self-invalidate when a workflow is added); new WC-4 pins every
[bypass-token] in a Bypass cell to verbatim presence in its workflow
file. 14 mutation tests.

Surfaces: docs/CONTROL_AVAILABILITY.md corrects its "on every change"
claim and links §7.1; the ARCHITECTURE "How to read" §7 bullet indexes
the CI sub-view.

Skipped with reason: read_or_exit2 exit-2 convention (sibling lints in
this fleet use the exit-1 missing-doc violation shape; consistency
wins).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R2 findings — .yaml fixture parity + comment-blind WC-4 (#755)

- The mutation fixture copies *.yaml alongside *.yml, so a future
  .yaml workflow with a valid row passes the fixture as it passes the
  real lint.
- WC-4 strips full-comment lines before the token search: a renamed
  executable token surviving only in a YAML comment no longer
  satisfies the pin (token in a non-comment echo/log string recorded
  as an accepted edge). Mutation test added (15 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R3 finding — tag pushes reach three more workflows (#755)

GitHub Actions matches tag pushes on unfiltered or paths-only push:
triggers (paths filters are not evaluated for tags), so
spec-consistency, command-invariants, and freshness-check also run on
every v* tag push — where their failures are post-push detection like
the tag-only workflows. Trigger cells amended and a subtlety note
added above the table; "three tag workflows" narrowed to "three
tag-only workflows".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R4 finding — malformed token spellings fail loudly (#755)

Any bracketed span in a Bypass cell must be a well-formed
[lowercase-hyphen] token: a typo like [skip_cooldown] now yields a
WC-4 violation instead of silently falling out of the token grammar.
Mutation test added (16 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R5 finding — whitespace token typos caught (#755)

The any-bracket span matcher now accepts any non-] content, so
[skip cooldown] (space typo) reaches the well-formedness check and
fails loudly. Mutation test added (17 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R6 finding — bogus rows fail instead of dropping out (#755)

Every pipe row in the section that is not the header or the separator
must open with a backticked workflow filename; a malformed row now
yields a WC-1 violation instead of silently leaving the inventory and
the WC-3 count. Mutation test added (18 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* test: mirror docs/ARCHITECTURE.md into the CA fixture (#755)

The new CONTROL_AVAILABILITY link to ARCHITECTURE §7.1 made the #768
test fixture (which mirrors only the files the doc links) miss its
target, failing CA-1 in the fixture tree while the real tree passes —
caught by CI, not locally, because the local sweep re-ran the lint but
not its sibling test file. ARCHITECTURE.md joins the mirrored list.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 09:03:53 +08:00
Edward Cheng-I Wu 9ccf4a9c9f fix: academic-pipeline data_access_level verified_only → raw (#756) (#772)
* fix: academic-pipeline data_access_level verified_only -> raw (#756)

The orchestrator legitimately consumes raw input — Stage 1 accepts raw
user requests, mid-entry accepts raw existing papers — and the
governing dirtiest-input rule (ground_truth_isolation_pattern.md)
requires the annotation to reflect that. The integrity gates run
INSIDE the pipeline, downstream of its intake, so verified_only was
internally inconsistent on the suite's most prominent consumer
(option (a) of the issue: honest minimal relabel; no trust-domain
split).

Surfaces aligned in the same commit:
- academic-pipeline/SKILL.md frontmatter (its #528 content-lock sha256
  recomputed in check_pipeline_boundary_semantics.py, same commit per
  the lock discipline).
- docs/ARCHITECTURE.md §4: pipeline node moves to the raw class, the
  User -> pipeline intake edge is drawn, and the rules block states the
  dirtiest-input rationale with a pointer to the pins.
- check_data_access_level.py grows an EXPECTED_LEVELS per-skill pin
  layer (acceptance criterion 2): a silent flip back to verified_only,
  an unregistered new skill, or a stale pin now fails CI; vocabulary
  check unchanged. 7 mutation tests + manifest entry (152).

Not touched: CHANGELOG history (records what v3.x declared at the
time); shared/agents/compliance_agent.md (agent-level declaration,
runs at the gates); academic-paper-reviewer verified_only (possible
same-class question for standalone /ars-reviewer raw-paper input —
out of #756 scope, reported separately).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: restore the pre-existing CLI-level test layer I overwrote (#756)

The previous commit replaced scripts/test_check_data_access_level.py
wholesale, dropping six original unittest cases (CLI subprocess via
--path, including the three malformed-frontmatter stdout-reporting
contracts) and breaking the run_skill_linter --path interface by
removing argparse from the lint. Both restored: main takes --path
again, the original unittest class is back (its valid-root case now
builds the four registered skills, since the pin layer correctly
rejects an unregistered synthetic skill), and the #756 pin-layer
mutation tests ride alongside. 11 tests green; manifest entry verified
through the CI runner.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* refactor: apply /simplify + codex R1 — single-pass lint, honest pins, aligned surfaces (#756)

Review round (3 cleanup agents + codex gpt-5.6-sol xhigh R1), findings
deduped and applied:

- check_data_access_level.py rewritten as a SINGLE pass: one violation
  per problem (the vocabulary layer had zero unique failure coverage
  and double-reported every failure, including a twice-printed YAML
  traceback); non-mapping metadata (e.g. "metadata: active") is now a
  reported violation instead of an AttributeError crash (codex P2);
  LEGAL_VALUES survives as a pin-vocabulary assertion; main() stays
  local (run_lint no longer fits once check_metadata_field drops out)
  and run_lint's stale "both check scripts" docstring is corrected.
- Pin provenance honesty: the docstring now says only the
  academic-pipeline pin is #756-derived; the other three freeze
  pre-existing declarations against silent drift. Follow-up derivation
  for reviewer/paper opened as #773.
- ARCHITECTURE §4: rule restatement dropped (the pattern doc owns the
  rule), the two competing one-line lint descriptions merged into one,
  and the §2 legend disambiguates §3's per-stage "Data level" column
  from the skill-level declaration (the four VERIFIED_ONLY stage cells
  are a different, per-stage claim — left as-is).
- CHANGELOG [Unreleased] gains the #756 Fixed entry (the pre-tag
  covers-merges gate is fail-closed).
- handoff_schemas.md data_access_level block now names the pin layer.
- write_skill fixture helper migrated to tests/test_helpers.py
  (migrate-at-next-edit convention); both skill-lint test files import
  it; CLI scenarios updated to registered skill names (the pin layer
  correctly pre-empts unregistered synthetic skills); new
  one-violation-per-problem and non-mapping-metadata regression tests
  (18 green across both files).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: restore the #753 CHANGELOG bullet heading (codex R2) (#756)

The #756 entry insertion had consumed the #753 bullet opening and
absorbed its body into the new bullet; the #753 heading is restored as
its own bullet.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: legend example had the gate boundary reversed (codex R3) (#756)

Gates consume unverified drafts and PRODUCE verified artifacts; the
Data-level column is documented as a postcondition on stage outputs,
not material the gate "operates on".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 09:02:58 +08:00
Edward Cheng-I Wu cdd48d916d docs: DATA_FLOWS.md — single map of network touchpoints + local stores (#758) (#770)
* docs: single data-flow map + DF-1..DF-3 coverage lint (#758)

Add docs/DATA_FLOWS.md — one row per network touchpoint (trigger,
payload class, recipient, credentials, off switch) and one row per
local store (path, content, TTL, deletion), with an explicit scope
statement (the Claude session itself is platform-governed; nothing
publishes autonomously). Covers the four gate resolvers, the
standalone Chinese-literature resolver (NOT in the gate), the
consent-bound claim-standing discovery adapters, both cross-model
transports (API and citation-only Codex subscription), the SessionStart
update check, the manual smoke tests, and the v3.9.4 timeline
bootstrap — the last one surfaced by the new lint itself on first run
(it was absent from the #758 issue enumeration).

Inbound links from README, SECURITY.md (in-scope exfiltration anchor),
and THIRD_PARTY.md (core-suite vs third-party contrast).

Lint (same-PR drift-point discipline): scripts/check_data_flows.py —
DF-1 every non-test scripts/*.py importing a network module (AST scan,
so no-call guards naming urllib.request in strings do not count) must
be named on the map; DF-2 same for curl-invoking shell scripts; DF-3
README/SECURITY/THIRD_PARTY keep a rendered resolving inbound link
(fences + HTML comments stripped with the semantics converged in the
PR #768 review; consolidation into a shared helper is follow-up).
12 mutation tests; wired into spec-consistency.yml + pytest manifest.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* refactor: apply /simplify pass (4-agent, deduped) (#758)

Doc: the four gate resolvers collapse into a 4-column sub-table under
one shared trigger/payload/off-switch lead (the wide table kept only
heterogeneous touchpoints); the exhaustiveness sentence is bounded to
what DF-1/DF-2 actually detect (direct imports + curl; spawned-CLI and
session-tooling paths held by review); "Nothing here publishes" now
inherits POSITIONING.md and its not-a-runtime-guarantee qualifier; the
subscription-free note is trimmed to its rationale; Related gains the
SETUP bullet as the tunables authority.

Coverage: docs/SETUP.md becomes the fourth DF-3-pinned inbound surface
(pointer added in the cache section); the four translated READMEs
mirror the README pointer; docs/DATA_FLOWS.md registers into
check_spec_consistency.py relative-link validation.

Lint: DF-1 module vocabulary rebuilt as the network subset of the
no-call envelope FORBIDDEN_IMPORTS (deviations documented: dotted
urllib.request/http.client instead of bare urllib/http; ssl excluded);
scan is now recursive into scripts/ subpackages. Tunable constants in
verification_cache.py gain update-both comments.

Tests: the three hollow assert-baseline tests become real mutations
(name-based test exemption, uncomment-curl, from-urllib idiom);
recursive-scan and SETUP-surface tests added (17 total).

Skipped with reason: endpoint-hostname lint (near-zero event rate,
composed-URL false-fire risk); row-id shrink constant (review-owned per
degradation-registry precedent); markdown-helper consolidation with
check_control_availability.py (whichever PR merges second extracts the
shared module — recorded in both PR bodies).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R1 findings — 8 P2 + 4 P3 (#758)

Doc accuracy (6): Chinese-literature row rewritten (callable client, no
CLI; PubMed path sends the required NCBI contact email + bibliographic
search coordinates); codex-transport payload names citation_context
(can contain unpublished manuscript text); update check documented as
one curl transfer per 24 h with redirects and the
ARS_UPDATE_CHECK_REMOTE_URL override; retraction-status SQLite cache
added to local stores (caller-supplied path, 30-day stale threshold,
no auto-expiry); discovery adapters credentials corrected (fixed
User-Agent, resolver env keys not consumed); resolver payload narrowed
to identifiers + title query strings; update-check state content
corrected (state label + two version strings).

Lint mis-pass/mis-fire (4): DF-1/DF-2 coverage now requires the full
repo-relative path (basename-substring collision closed); DF-2 is
recursive over scripts/ and hooks/, recognizes path-qualified curl,
and masks quoted spans before the comment strip; DF-3 strips inline
code spans before link extraction (a backticked link does not render).
Five mutation tests added (22 total). The code-span rule is a
divergence from check_control_availability.py to be carried over at
the declared helper consolidation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R2 findings — 4 P2 + 1 P3 (#758)

- DF-2 scans command-substitution bodies BEFORE quote masking, so
  resp="$(curl ...)" — a real network call inside double quotes — fires
  (mutation test added; suite now genuinely 22, correcting the prior
  commit message which said 22 when 21 were collected).
- Map gains the Codex audit wrapper row (scripts/run_codex_audit.sh:
  human/CI/hook-invoked only, sends deliverable + supporting file
  contents through the local Codex CLI login).
- Update check re-bounded: at most one SUCCESSFUL check per 24 h; a
  failed attempt writes no state and may retry next session.
- Cache TTL wording corrected: expiry is a cache miss, not deletion;
  expired rows persist until invalidated or the file is deleted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: full comment lines execute nothing — DF-2 substitution scan (#758)

The R2 command-substitution scan ran before any comment handling, so a
full comment line containing $(curl ...) false-fired — surfaced by the
codex R3 pass (timed out mid-review, but its transcript had already
demonstrated the false fire). Comment-only lines are now skipped before
the substitution scan; a $(curl) inside a trailing inline comment
remains a documented accepted edge. Mutation test added (23 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R3 findings — command-position curl + image links + retention wording (#758)

- DF-2 rebuilt around COMMAND POSITION: curl counts only as the first
  non-assignment token of a segment (pipes/separators/substitution
  openers), so `command -v curl` preflights and `echo curl` no longer
  false-fire; VAR=x curl still fires; wrapper-prefixed invocations
  (sudo/timeout) are documented accepted edges.
- DF-3 link grammar excludes image syntax — ![map](...) renders no
  anchor and cannot keep the acceptance criterion green.
- Cache retention wording includes the overwrite path: expired rows
  persist until overwritten by re-verification, invalidated, or the
  file is deleted.
- Three mutation tests added (26 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R4 finding — curl behind shell control words (#758)

The command-position head-token scan now skips shell control words
(if/elif/while/until/then/else/do/!/time/exec) before naming the head,
so `if curl …; then` and `while ! curl …; do` fire while `if true;
then` stays quiet. Two mutation tests (28 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R5 finding — option tokens after control words (#758)

`time -p curl …` / option-bearing exec forms: the head scan now skips
`-`-prefixed option tokens alongside assignments and control words, so
the option cannot shadow the command head. Mutation test added (29
total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R6 finding — harness spawned-CLI paths scoped out (#758)

Three maintainer-only measurement scripts reach the network through
locally authenticated CLIs (dispatch_e4_panel via claude -p,
run_review_criteria_constructive_value via Codex, check_ranking_lift
via gh api). They are not user-facing feature paths, so instead of
diluting the touchpoint tables they are now an explicit named scope
exclusion — the exhaustiveness claim no longer silently spans them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R7 finding — boundary count wording (#758)

"Two boundaries" became three after the R6 harness exclusion; the count
is removed rather than maintained.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 00:28:24 +08:00
Edward Cheng-I Wu 43a02bf7e2 docs: per-channel control-availability matrix (#757) (#768)
* docs: per-channel control-availability matrix (#757)

Add docs/CONTROL_AVAILABILITY.md — one row per enforcement mechanism,
one column per install channel (plugin / skills copy / repo clone /
Cowork / claude.ai Project / Claude Science / Pi), with honest
active / conditional / absent cells, per-channel notes citing the
existing scattered sources (README Requirements, SETUP methods,
pi/README.md, hooks/run_guard.sh), and the guard's environment
degradation table. Linked from README (Requirements + SETUP pointer)
and SETUP (Installation methods intro).

Evidence re-verified against the working tree: the channel set has
grown past the six named in the issue (SETUP now also documents
Cowork and the claude.ai 4a/4b split), so the matrix covers all
seven documented channels.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* refactor: apply /simplify pass + add CA-1..CA-3 defrift lint (#757)

Simplify round (4-agent review, findings deduped):
- Drop the 'How to read an integrity claim' section (it had already
  drifted from the matrix) and the all-identical Upstream row; both
  replaced by one legend sentence and one paragraph.
- Move channel-scoped caveats (Cowork / claude.ai / Claude Science / Pi)
  from per-cell footnotes into a 'Channel-wide limitation' column of the
  channel table; notes drop from 11 to 7.
- De-drift row labels: no inline allowlist contents (canonical list is
  pinned by check_tools_allowlist.py), no exhaustive feature list, no
  hard-coded Claude Code minimum version (lives in SETUP Method 0).
- README: single slimmed pointer (second link and both enumerations
  removed); pointer mirrored to the four translated READMEs and
  docs/SETUP.zh-TW.md.
- Degradations table scoped to actual guard degradations (the slash-form
  version row was misfiled); registry backpointer added; guard-launcher
  registry registration split to #769.

Lint (per the new-claim-surface-needs-lint-in-same-PR discipline):
- scripts/check_control_availability.py — CA-1 links/anchors resolve,
  CA-2 every SETUP '### Method' heading reachable from the channel
  table, CA-3 README + SETUP inbound links pinned. Cell semantics stay
  owned by code review (degradation-registry posture).
- 9 mutation tests; wired into spec-consistency.yml + pytest manifest
  (150 entries).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R1 findings — 4 P2 accuracy corrections (#757)

- SessionStart announce/update-reminder row: Conditional, not Active
  (bash launcher on Windows needs Git Bash; reminder needs curl) — new
  note 8.
- Cross-model note 6 no longer claims credentials+curl universally; the
  citation-only Codex subscription transport is named as the alternative
  transport behind the same consent boundary.
- Pi channel limitation reworded: the wrapper supplies no orchestration
  but uses an installed Pi capability when available.
- 'Enforcement mechanisms' claim language aligned to 'controls' in the
  purpose statement and all five README pointers (consistent with note
  7's trust-based posture).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R2 findings — lint mis-pass cases + note-8 wording (#757)

- CA-1 link grammar accepts optional quoted titles so a titled dead
  link cannot silently skip the check.
- CA-2 counts only fragments on links whose resolved destination IS
  docs/SETUP.md — a same-slug anchor into a copied file no longer
  satisfies method coverage.
- CA-3 checks resolved link destinations, not a filename substring — a
  label that keeps the filename while the target moves now fails.
- Note 8: singular SessionStart hook (hooks.json defines one; the
  announce script runs the update check internally).
- 3 new mutation tests pinning each mis-pass case (12 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R3 finding — commented-out markdown counts for nothing (#757)

Strip HTML comments before extracting links and headings in all three
invariants: a commented-out inbound link no longer satisfies CA-3, a
commented-out SETUP method heading no longer demands CA-2 coverage, and
a commented-out dead link no longer fires CA-1. Two mutation tests pin
both directions (14 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R4 finding — GFM type-2 HTML-block semantics (#757)

A line beginning with <!-- opens a raw-HTML block through the --> line
(including trailing text on the closing line) or to EOF if unclosed;
nothing on those lines renders. The comment stripper now models that
line-level behavior before the inline-span strip, so a link after -->
on a comment line cannot satisfy CA-3 and a dead link after an unclosed
comment cannot fire CA-1. Two mutation tests pin both (16 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* test: fix R4 mutation scenario — line-start vs inline comment (#757)

The previous commit's CA-3 HTML-block test inserted the comment mid-line
(inside the blockquote), where GFM renders the link normally and the
lint correctly stays quiet — the test scenario was wrong, not the lint.
Replaced with a whole-line mutation that actually begins with <!--, and
added the inline-comment symmetry case (link still renders → CA-3
satisfied). 17 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R5 finding — block-quoted HTML-block lines (#757)

The type-2 HTML-block rule applies to block-quote content: the stripper
now looks through leading '> ' markers before the line-start test, so
'> <!-- note --> [link]' cannot satisfy CA-3. Deeper CommonMark
laminations are declared out of scope in the docstring (the surfaces do
not use them; a full parser is out of proportion for a maintainer-slip
guard). 18 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R6 finding — repo-containment on CA-1 targets (#757)

A relative link that resolves outside the repository root now fails
CA-1 even when the host path exists — an over-deep ../.. slip must not
be masked by an existing host file. Mutation test added (19 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R7 finding — fenced code excluded from extraction (#757)

Fenced code regions render literally, and README/SETUP use fences
today, so they are in-scope: a link inside a fence no longer satisfies
CA-3, and a sample "### Method" heading inside a SETUP fence no longer
demands CA-2 coverage. Fence stripping runs before the comment pass so
a comment opener inside a fence stays literal. Two mutation tests (21
total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

* fix: close codex R8 finding — CommonMark fence-length closing rule (#757)

The fence stripper now tracks the opening run character and length: a
closer must be a same-character run at least that long with only
trailing whitespace, so a four-backtick fence demonstrating an inner
triple-backtick block is no longer closed early. Mutation test added
(22 total).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EosnA4RdUYgbF2KmZ1DTmc

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 23:57:07 +08:00
Edward Cheng-I Wu e9759dc4f4 Align distribution-surface claims with evidence ceilings (#753) (#766)
* fix(claims): align distribution-surface claims with evidence ceilings (#753)

- plugin.json / marketplace.json: drop 'Production-grade' / '39-agent
  ensemble' for matrix-licensed wording ('contract-audited', '39 prompt
  roles (3 plugin-exposed agents; the rest run inline by default)')
- academic-pipeline/SKILL.md: no-bypass prose rewritten to the actual
  mechanism (mandatory checkpoints; overrides require recorded user
  reasoning); #528 content-lock hash updated in the same commit
- shared/cross_model_verification.md: 31%->5-10% relabeled as an
  unvalidated working hypothesis
- shared/ground_truth_isolation_pattern.md: gold-labels rule rewritten to
  the intended boundary (no unconditional loading into operational agent
  context)
- version-consistency invariant 8: binds 'N prompt roles' spelling too,
  checks every count token (finditer)
- new scripts/check_distribution_surface_claims.py (D1-D5, 20 mutation
  tests, CI-wired): fail-closed manifest load, shared claim vocabulary
  imported from check_stage_capability_matrix, percentage refusal,
  mandatory bindable count token, plugin-exposed count bound to MIRRORS

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

* fix(claims): codex R1 — integrity-family must-PASS sweep + lint case/boundary fixes (#753)

- integrity 'must PASS with zero issues' absolutes now name the recorded
  3-round FAIL-loop exit (integrity_review_protocol, reinforcement_content,
  team_collaboration_protocol, integrity_verification_agent, SKILL.md flow
  row); 'recorded with reasoning' weakened to 'recorded user decision'
  (rationale escalates per compliance override ladder)
- D3 percent check lowercases input (matrix caller parity)
- D5 plugin-exposed regex case-insensitive
- AGENT_CLAIM_RE gains trailing boundaries (39-agentic / singular 'prompt
  role' no longer count as bound); 4 new mutation tests (20 -> 24)
- SKILL.md #528 content-lock hash rebumped

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

* fix(claims): codex R2 — Stage 2.5 routing parity, passport-state honesty, gold-set scope, strict JSON (#753)

- Stage 2.5 flow row + both state-machine checkpoint triggers name the
  recorded FAIL-loop exit (SKILL.md + pipeline_state_machine.md, both
  content-lock hashes rebumped)
- team protocol handoff checklist: FAIL-loop continuation keeps passport
  verification_status UNVERIFIED; VERIFIED only on zero-issue PASS
- ground-truth gold exception scoped to synthetic/public-safe content;
  live-reviewer calibration sets stay runtime-supplied
- D1 rejects non-standard JSON constants (NaN/Infinity) via parse_constant;
  2 new tests (24 -> 26)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

* fix(claims): codex R3 — prerequisite checker + handoff materials accept the recorded FAIL-loop route (#753)

- state_tracker_agent prerequisite table: Stage 3 / Stage 5 entry rows
  accept a recorded Integrity Check FAIL Loop resolution (previously the
  documented continuation route was unreachable at the checker)
- SKILL.md handoff lines 2.5->3 and 4.5->5 no longer mislabel a FAIL-loop
  continuation draft as verified; team protocol Materials/Approval rows
  aligned the same way
- SKILL.md + state_tracker_agent content-lock hashes rebumped

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

* fix(claims): codex R4 — orchestrator transfer rows + advisory dispatch accept the recorded FAIL-loop route (#753)

- orchestrator 2.5->3 and 4.5->5 transfer rows no longer require a
  'Verified'-labeled draft on a recorded FAIL-loop continuation
- #660/#672 advisory dispatch anchors to the Stage 4.5 terminal resolution
  (PASS, or recorded FAIL-loop continuation) instead of exact PASS only
- orchestrator content-lock hash rebumped

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 20:02:59 +08:00
Edward Cheng-I Wu 704b46d247 fix: citation-surface version drift + version-consistency invariant 12 (#763)
* fix: citation-surface version drift + version-consistency invariant 12 (#754)

CITATION.cff and POSITIONING.md citation prose sat at 3.14.0 while the
suite moved to 3.20.1 (Zenodo v3.20.1 deposit exists — pure metadata
drift). Bump both to 3.20.1 and add lint invariant 12 so the drift class
fails CI: CITATION.cff version: and every (Version X.Y.Z) token in
POSITIONING.md must equal the suite version. Absent file = skip
(invariant-8 posture); present-but-versionless CITATION.cff = error.
6 new mutation tests (53 -> 59), full suite green, real-repo lint green.

Closes #754

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

* refactor: harden invariant 12 per 4-angle cleanup review

- CITATION.cff parsed as YAML (regex scrape misread quoted versions as
  drift); absence now errors like README.md so deletion cannot silently
  disable the invariant; version goes through the broad-capture +
  strict-semver idiom (non-canonical reported as such, not as drift)
- date-released gated against the CHANGELOG latest-entry date with the
  invariant-10 +/-7-day window (the second half of the #754 drift)
- POSITIONING token uses broad capture + strict validation so v-prefixed
  or truncated edits error instead of being filtered (pre-#169 lesson)
- regex moved to the numbered constants block; helpers moved to the test
  fixture header; surfaces wired into _write_aligned_fixture so every
  pass-case exercises invariant 12; class renamed TestCitationSurfaces
- 11 targeted tests (53 -> 64), green under pytest and unittest runners

Refs #754

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

* fix: harden invariant 12 date handling per codex review (2 P2)

- catch ValueError from PyYAML's timestamp constructor (impossible
  dates like 2026-02-30 raise ValueError, not YAMLError) so the lint
  reports instead of crashing
- reject datetime values (unquoted timestamps parse as a date SUBCLASS
  that would TypeError against the date baseline) as not-strict-
  YYYY-MM-DD errors
- 2 regression tests (64 -> 66), both verified red before the fix

Refs #754

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

* fix: require date-released in invariant 12 (codex round-2 P2)

An absent or null date-released silently skipped the freshness check,
so deleting the field disabled the invariant's date half. Absence now
errors (same posture as file absence). 1 regression test (66 -> 67),
verified red before the fix.

Refs #754

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

* fix: close invariant-12 silent-pass paths per codex round 3 (2 of 3 P2)

- POSITIONING clause captured whole then stripped + strictly validated,
  so a malformed edit ('(Version 3.4.0 )') errors instead of dropping
  out of the capture and passing a stale citation
- missing-version regression test isolated from the missing-date error
  (fixture keeps a valid date-released; asserts the specific diagnostic)
- duplicate-YAML-key finding adjudicated NOT-FIX under the threat model:
  last-wins matches every CFF consumer (GitHub cite widget / Zenodo /
  cffconvert), so a duplicate-key file renders the identical citation
  everywhere - untidiness, not drift; documented as a known limitation
  in the checker docstring
- 1 new test (68 total), suite green under both runners

Refs #754

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

* fix: catch empty POSITIONING payload + refresh CHANGELOG counts (codex round 4)

- capture class widened to [^)]* so '(Version )' reaches strict
  validation and errors instead of falling out of the capture
- CHANGELOG test counts refreshed (53 -> 69) and review provenance
  updated to the 5-round threat-model-bounded trajectory
- 1 regression test, verified red before the fix

Refs #754

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 15:57:19 +08:00
Edward Cheng-I Wu 07a2afd94c docs(audits): ISO 42001-spirit gap assessment (2026-08-17) (#762)
Dual-track audit (in-session structural review + independent cross-model
GPT-5.6 xhigh full-repo audit), all findings re-verified first-party.
Scope decision: adopt the standard's spirit (transparency / verifiability /
feasibility), explicitly not pursuing certification. Files the verified
findings register behind meta-issue #761 (#753-#760) and records the
explicitly-not-adopted list.


Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 15:57:16 +08:00
Edward Cheng-I Wu f6ffc70312 fix(docs): co-locate passport_as_reset_boundary reference in #743 design doc (#764)
The #743 inquiry-branch-ledger design doc mentions ARS_PASSPORT_RESET
without referencing passport_as_reset_boundary, so
check_passport_reset_contract.py fails on main and blocks every open
PR's spec-consistency gate. Add the protocol reference at the mention
site per the v3.6.3 co-location contract.

Refs #743

[skip-closes-check] CI hotfix restoring a green main; #743 remains open
as the feature's tracking issue and must not be auto-closed by a one-line
doc-reference fix.


Claude-Session: https://claude.ai/code/session_011n3WG8Z3Us8fX8UhXS51Ki

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 15:56:40 +08:00
Edward Cheng-I Wu 946c2494af docs(design): #743 inquiry branch ledger design freeze (#752)
* docs(design): #743 inquiry branch ledger design freeze — event-sourced states, adoption receipts, reopen invalidation

Freezes inquiry-branch-ledger/1.0: append-only hash-chained event log with
deterministic replay, the closed five-status state machine (AI-surfaced facets
enter parked and never active), the adoption receipt that keeps AI provenance
immutable, author-only reopen with visible stale-not-rewritten invalidation
over recorded downstream refs, ARS_INQUIRY_LEDGER opt-in with zero simple-path
prompts, additive passport-pointer storage, and the #745 registration gate
before any alpha ships. No implementation or evaluation is authorized.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753

* docs(design): #743 review round 1 — facet exit rules, origin-bound receipts, event-sourced staleness, budget as replay invariant, frozen payloads and storage

Closes the 7 P1 + 5 P2 codex findings: an unadopted AI facet can only be
adopted or rejected (rejected facets terminal), so reopen cannot bypass the
adoption receipt; receipts bind source_event_id and the attestation boundary
is stated (recorded attestation, not authentication — /ars-mark-read
precedent); §3.1 freezes every payload shape with replacement semantics and
merge dedup ordering; replay order and canonical form specified, with the
passport pointer digest as the trusted head that closes the truncation hole;
staleness becomes event-sourced (artifact_marked_stale / reconfirmed /
superseded) with first-degree scope stated honestly; the live budget is a
replay invariant so the ledger can never record an over-budget state;
profile binding is a projection over an immutable initial binding;
reopen conditions get stable ids and signals bind exactly one; §7 freezes
the passport aggregate shape, atomic writes, closed absence semantics, and
the user-workspace/never-committed data boundary; §8 records transport
limits and the same-PR CI-gated registration ordering.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753

* docs(design): #743 review round 2 — post-state budget on every event, cause-bound stale resolution, supersession link maintenance, explicit ledger digest, archived status, condition-id identity

Closes the round-2 findings: the budget invariant now constrains every
event's post-state incl. profile_rebound (lower-budget rebinds require prior
dispositions); artifact staleness projects a SET of outstanding causes and
clearing events bind resolves_stale_event_id to exactly one cause;
artifact_superseded deterministically replaces the retired ref in every
listing branch's downstream_refs (first-degree link maintenance); the
passport-head digest is defined exactly as SHA-256 over JCS of the ledger
document with no placeholder; branch_archived/archived join the closed
lifecycle (terminal, facet-exit-lawful); condition_ids are per-branch
unique, non-rebindable, and retired permanently on removal.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 12:12:10 +08:00
Edward Cheng-I Wu 07d4833f50 feat: #745 stage capability/evidence matrix with enforceable claim ceilings (#751)
* feat: #745 stage capability/evidence matrix with enforceable claim ceilings

stage-capability-matrix/1.0 data source + M1-M10 lint (frozen task-family
vocabulary shared with the #742 profile contract, non-collapsible evidence
statuses, in-repo eval refs, verbatim claim anchors, stale-evidence notes,
effectiveness-language discipline on unmeasured rows, byte-pinned generated
view), 38 mutation tests, CI + pytest-manifest wiring, contracts README
section. Seeded with 13 rows over all nine task families: 4 measured/mixed
(incl. the seeded-defect panel's currently-failing severity gate recorded as
MIXED) and 9 designed/not-run whose ceilings say exactly that.

Refs #745 (matrix + lint slice; the matched stage-substitution evaluation
program and README claim-language migration remain open).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753

* fix: #745 review round 1 — report bindings, falsifiable conformance pins, anchor hardening, single-source view path

Applies the 4-agent simplify review: M11 measurement-report binding (date-equal
to the bound report, accepts the pre-#654 legacy measured_at field, sibling
supersession requires a staleness_note) closes the hand-typed-provenance drift
the matrix exists to prevent; M12 conformance_pinned_by makes CI_GATED/TESTED
falsifiable (D4-style path/function pins, all 13 rows populated); M7 anchors
gain containment-checked memoized reads and three rows now bind real README
capability sentences; M4 refuses measurement provenance on unmeasured rows in
both directions; generated_view leaves the JSON (DEFAULT_VIEW is the single
source and --render validates before writing); invalid stale_after_days skips
M8 instead of fabricating a default; _flat() replaces table-cell escaping;
the Minnesota-colliding 'sota' stem is dropped. 52 mutation tests.

Skipped by judgment: task_families stays in the JSON (self-describing contract
+ lock, degradation-registry precedent); shared anchor-verifier extraction
deferred until a third consumer exists; per-row staleness half-lives and a
warn tier deferred.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753

* fix: #745 review round 2 — containment on eval_ref, report-pattern binding, two-tier claim stems, inventory locks, honest revision row

Closes the codex round-1 findings: eval_ref must be a repo-contained suite
directory; measurement_report must match measurement-*.json and unreadable
sibling reports fail supersession instead of skipping it; claim language
splits never-licensed stems (guarantee/proven/state-of-the-art, refused on
every row) from measured-licensed stems, and percentages are refused outside
MEASURED/MIXED; shipped row ids and anchor minimums are D5-style locked;
future measured_at refused; bool stale_after_days refused; non-object
behavioral_evidence reports instead of crashing; anchors render into the
byte-pinned view. The revision.claim_drift_guard row is narrowed to what the
2026-08-07 record actually measured — the condensed guard-block prompt, 6
pressure items + 2 controls — with a ceiling that does not transfer to the
shipped pipeline wiring. 64 mutation tests.

Deferred by judgment (documented open half of #745): reverse inventory of
README/CHANGELOG claim sentences.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753

* fix: #745 review round 3 — percent rule covers mechanism text, non-object reports refused, exact report-name match, robust anchor lock

Closes codex round-2: the percentage discipline no longer exempts the
mechanism field; a JSON-list bound or sibling measurement report yields an
M11 violation instead of an AttributeError; _REPORT_NAME_RE uses fullmatch;
the M13 anchor-minimum lock counts only well-formed lists after M7 has
diagnosed the malformed value. 69 mutation tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753

* test: #745 pin the task-family vocabulary to the merged #742 design doc §2 table

The frozen TASK_FAMILIES constant and the profile contract's stage/task-family
table can no longer drift apart silently — the cross-consumer test parses the
doc's table ids and requires exact equality including order. 70 mutation tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753

* docs: align CHANGELOG test count (70)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BPU2Hg2WufTjMVyBoje753

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 12:05:22 +08:00