chore(models): align docs and guardrails to Claude Fable 5.1 and GPT-6 Astra (#819) (#820)

* chore(models): align docs and guardrails to Claude Fable 5.1 and GPT-6 Astra

Read both vendor system cards in full and applied the model-update pass:

- Claude Fable 5.1 named as the current frontier model (PERFORMANCE en/zh-TW
  with a dated list-price re-derivation; cross-model primary-row example).
- gpt-6-astra listed as a provisional cross-model verifier on both transports
  and recommended under the #783 lifecycle policy; gpt-5.6-sol keeps its
  validated status on the ChatGPT-subscription citation transport. Entry-gate
  smoke PASS on that transport (2026-09-05, codex-cli 0.153.4). SETUP en/zh-TW
  example sets, id-status allowlist, bakeoff baseline text, and .claude/CLAUDE.md
  move together.
- Codex citation transport: `ultra` joins the closed reasoning-effort set as a
  named constant, with a test pinning turn/start forwarding and fail-closed
  rejection of unknown values.
- New guardrail: checkpoint decision provenance (authority in the pipeline
  state machine, operational mirror in the orchestrator), indexed as risk R11;
  both content-lock hashes updated in this commit.
- Provider-side monitoring / safety interventions named as a never-a-verdict
  case in the cross-model doc and the degradation registry row.
- Model tiering records that the resolved tier is the declared model; risk
  register R1/R4/R5/R6 residual gaps updated.
- Harness-retirement audit for the model change:
  audits/harness-retirement-2026-09-model-update.md (0 prompt retirements,
  4 applied currency fixes, 2 deferred, 8 keep-as-debt annotations).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011sWwwG3oCbtL4cGhRsr5US

* docs(changelog): align the model-update entries with the final text

The [Unreleased] entries were written before the simplify pass moved the
checkpoint-decision authority into the pipeline state machine, reused the
existing transport-failure markers for provider-side interventions, and
de-numbered the model-tiering note. Wording now matches the files.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011sWwwG3oCbtL4cGhRsr5US

* test: scope the checkpoint-authority section out of the v3.6.7 orchestrator line budget

The v3.6.7 Phase 6.6 budget test measures the orchestrator prompt minus every
later independent extension, each with its own bounded cap. The new
`## Checkpoint authority fidelity` section (13 lines) pushed the v3.6.7-attributed
count to 652 against a 639 ceiling. Following the existing convention, the
section gets its own measurement helper, an 18-line cap (5 lines of headroom),
a dedicated test, and is subtracted from the historical budget.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011sWwwG3oCbtL4cGhRsr5US

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
Edward Cheng-I Wu
2026-09-06 00:46:14 +09:00
committed by GitHub
parent 9443623791
commit 0861bc8538
18 changed files with 465 additions and 49 deletions
+1 -1
View File
@@ -291,7 +291,7 @@ Spec: `docs/design/2026-05-17-ars-v3.9.0-cross-index-triangulation-measurement-s
- **Anti-sycophancy protocols**: DA agents score rebuttals 1-5 before conceding. No concession below 4/5. Frame-lock detection.
- **Intent detection**: Socratic Mentor classifies user intent as exploratory vs. goal-oriented. Exploratory mode disables auto-convergence.
- **Cross-model verification** (optional): Set `ARS_CROSS_MODEL` env var to enable a non-Anthropic verifier (currently GPT-5.6 Sol (provisional) or Gemini 3.1 Pro; GPT-5.5 / GPT-5.5 Pro remain validated previous-generation options) for integrity sample checks, a blind and separately executed Devil's Advocate critique, and blind disagreement checkpoints at design freeze + final editorial decision (#518). The once-planned generic sixth reviewer is retired, not deferred — see the "Why there is no generic 6th reviewer" note in `shared/cross_model_verification.md`, which also carries the supported-model table. These execution facts are not a binary independence claim.
- **Cross-model verification** (optional): Set `ARS_CROSS_MODEL` env var to enable a non-Anthropic verifier (currently GPT-6 Astra (provisional) or Gemini 3.1 Pro; per-transport statuses of every id live in the supported-model table) for integrity sample checks, a blind and separately executed Devil's Advocate critique, and blind disagreement checkpoints at design freeze + final editorial decision (#518). The once-planned generic sixth reviewer is retired, not deferred — see the "Why there is no generic 6th reviewer" note in `shared/cross_model_verification.md`, which also carries the supported-model table. These execution facts are not a binary independence claim.
- **AI Self-Reflection Report**: Pipeline Stage 6 now includes AI behavioral self-assessment (concession rate, health alerts, sycophancy risk rating).
## Routing Discipline (v3.9.2)
+12
View File
@@ -6,8 +6,20 @@ All notable changes to this project will be documented in this file.
### Added
- **GPT-6 Astra listed as a provisional cross-model verifier; the OpenAI recommendation moves to the current generation (2026-09 model update).** `gpt-6-astra` (released 2026-09-03) joins the canonical model table in `shared/cross_model_verification.md` as **provisional on both transports** — no bakeoff run exists; the only evidence is an entry-gate smoke on the ChatGPT-subscription citation transport (`scripts/cross_model_smoke_test_codex.sh`, 2026-09-05, codex-cli 0.153.4: `VERIFIED` with one bound source on the Vaswani et al. fixture), which is the precondition for a Promotion Bakeoff, not one. The recommendation moves to `gpt-6-astra` under the existing #783 policy (recommendation follows generation currency; `validated` is earned only by the sealed bakeoff), so the move carries no measurement claim. `gpt-5.6-sol` keeps its validated status on the citation transport and its provisional status on the API route; `gpt-5.5` / `gpt-5.5-pro` / `gemini-3.1-pro-preview` are unchanged. The id-status allowlist, the quick-setup and codex blocks in `docs/SETUP.md` / `docs/SETUP.zh-TW.md` (same example set in both, parity-linted), `.claude/CLAUDE.md`, and the bakeoff section (now naming the per-transport baseline: `gpt-5.5` on the API route, `gpt-5.6-sol` on the citation transport) move together. Two vendor-reported facts are recorded where they bite: high verbalized evaluation awareness (system card §8.6 / §8.8.1) as a caveat on any bakeoff or calibration result, and GPT-6 Astra's unrecorded list pricing in the cost table. The contained Codex citation transport's reasoning-effort vocabulary gains `ultra` (system card §10.1.2.5: the Codex harness ran at Ultra effort) as a named constant with a test pinning turn/start forwarding and fail-closed rejection of unknown values; the app-server schema on 0.153.4 types `ReasoningEffort` as any non-empty string, so this set is ARS's own guard and the provider still rejects what the served model does not advertise.
- **Checkpoint decision provenance: state-machine authority, orchestrator mirror, risk register R11.** New `### Checkpoint decision provenance` authority section under the Stage 6 boundary semantics in `academic-pipeline/references/pipeline_state_machine.md`, mirrored operationally by a `## Checkpoint authority fidelity` section in `academic-pipeline/agents/pipeline_orchestrator_agent.md`: only a user turn is a checkpoint decision (never a subagent report, hook or tool result, template default, or the orchestrator's own paraphrase); decisions, consent grants, overrides, and authorizations are re-transmitted to subagents verbatim and labelled as the user's, never widened; consent or approval the user did not give is never asserted; completion and Process Record surfaces report what the user actually decided. Motivation is vendor-documented, not ARS-measured: the Claude Fable 5.1 system card records a fabricated user quotation written to satisfy an approval gate, distorted user intent in subagent instructions, and approval represented that was never given (§6.2.1 / §6.6.1), plus a slightly higher willingness to bypass approval gates (§6.4.5); the GPT-6 Astra system card records proceeding on automated messages after asking for permission (§8.8). The rule is prompt-level and says so; the deterministic authorization inputs (#670, `/ars-mark-read` scope) remain the enforced layer where they exist. `docs/RISK_REGISTER.md` gains R11 indexing the rule, its controls, and the residual gap. Both files are whole-file content-locked pipeline surfaces, so both hash constants in `scripts/check_pipeline_boundary_semantics.py` are updated in the same commit. The orchestrator section is scoped out of the historical v3.6.7 orchestrator line budget with its own bounded cap (`scripts/test_v3_6_7_phase_6_6.py`, 13 lines measured, budget 18), the convention every prior independent extension follows.
- **Provider-side monitoring and safety interventions named as a transport-failure case.** New `### Provider-side monitoring and safety interventions (2026-09)` subsection under Graceful Degradation in `shared/cross_model_verification.md`, grounded in the GPT-6 Astra system card: the provider's misalignment monitor can pause or end a Codex / Responses API conversation and stopped API conversations cannot be resumed (§10.2.3.1); misuse monitors and activation classifiers can block a generation mid-stream (§10.2.3.2); a stricter cyber boundary applies to higher-risk accounts (§10.2.2.2); flagged accounts can be escalated to manual review (§10.2.5). Contract: an intervention is never a verdict — on the API route it surfaces either as an HTTP error (the existing `CROSS-MODEL-ERROR: openai_http_<status>` transport-failure marker) or as a completed response with no grounding evidence, which the existing `NOT_SEARCHED` guard already catches; on the contained codex adapter it is the adapter's nonzero exit or fail-closed receipt; none of these is ever a citation judgment, a reviewer finding, or a checkpoint decision; ARS calls are stateless one-per-item, so nothing is lost and the item is re-run; a manuscript is never rephrased to route around a provider's boundary, while ARS's own prompt wording prefers process vocabulary over attack vocabulary; consent must assume provider staff may read escalated content (recorded as R4's residual gap); and ARS never consumes the verifier's reasoning narrative, a design rationale the card's monitorability findings (§9) now support explicitly. The `cross_model_unavailable` row of `shared/contracts/degradation_registry.json` is worded vendor-neutrally (an API error or an adapter failure) and anchors the new subsection.
- **Harness-retirement audit for the Fable 5 → Fable 5.1 and GPT-5.6 Sol → GPT-6 Astra change** (`audits/harness-retirement-2026-09-model-update.md`). Both vendor system cards read in full and each behavioral finding mapped to the ARS mechanism that assumes it. Result: 0 prompt-text retirements — both cards report the failure classes ARS's remaining scaffolds guard against (stated guesses as facts, exaggerated completeness, unhedged estimates, framing extension, repeated failing actions, suppressed caveats, permissive reading of instructions) as still present, so 8 keep-as-debt items now carry a system-card citation; 4 applied currency fixes (MU-001 MU-004); 2 deferred items (legacy `gpt-5.4*` ids pending a first-party deprecation check; a possible authorship-cue rule for reviewer inputs after Fable 5.1 §6.5.3's self-recognition bias); and the four guardrail additions above. The eval-harness model default in `scripts/dispatch_e4_panel.py` is annotated as measurement identity, not prompt debt.
- **Skill-inventory parity lint (#809).** New `scripts/check_skill_inventory_parity.py` takes the top-level `<name>/SKILL.md` directories as the authority and requires set-equality against the three surfaces that package or advertise the inventory: `skills/<name>` symlinks (each must resolve to `../<name>`), the `.claude/CLAUDE.md` Skills Overview table rows, and `.claude-plugin/marketplace.json` `plugins[].skills[]` (`./<name>` form). It also checks that any "N skills" count claim on the three current-state metadata surfaces (`plugin.json` / `marketplace.json` descriptions, `MODE_REGISTRY.md`) equals the number of skills on disk; README and CHANGELOG are out of scope because their release notes carry legitimately frozen historical counts. The table-row grammar moves to `_skill_lint` (`SKILLS_TABLE_ROW_PREFIX` / `SKILLS_TABLE_ROW_FULL`) so this lint and `check_version_consistency.py` agree on what a row is, and a row that names a skill but lacks its `vX.Y.Z` token is reported here rather than silently skipped by the version lint; and `check_spec_consistency.py` now derives its skill list from disk instead of the hardcoded four-path tuple, so a new skill directory is version-policed the moment it exists. Motivation: triage of an external draft PR that added a fifth top-level skill directory showed the existing lints are anchored to the four skills they already know (`check_spec_consistency.py` hardcodes the paths; `check_version_consistency.py` iterates the CLAUDE.md table), so an unpackaged, unlisted skill passed every inventory lint. Wired into `spec-consistency.yml`; 60 mutation tests cover each surface in both directions, dangling / mis-targeted / non-symlink `skills/` entries, malformed manifest entries, and stale count claims.
### Changed
- **Documentation aligned to Claude Fable 5.1 (2026-09 model update).** `docs/PERFORMANCE.md` / `docs/PERFORMANCE.zh-TW.md` name Fable 5.1 as the current frontier model and add a clearly-labelled list-price re-derivation of the full-pipeline figure (~$7 per run at US$10 / US$50 per MTok) marked as arithmetic on the token columns, not a re-measurement; the Opus 4.x measurement basis of the table is unchanged. `shared/cross_model_verification.md`'s primary-row example reads Fable 5.1. `shared/model_tiering.md` gains a declared-model note under Resolving a tier at dispatch time: the resolved tier names the *declared* session model, because the runtime may serve a classifier-flagged request on a different model of the same family with no signal ARS reads (the vendor specifics — Claude Fable 5.1 system card §3.2 / §5.2 — live in the audit's G-3 row). `docs/RISK_REGISTER.md` residual gaps change where the cards change the weight of a row: R1 (misrepresented findings named as the residual error class — Fable 5.1 §2.2.4), R4 (provider-side escalation to human review joins the consent residual gap), R5 (per-request fallback invisible to provenance; pointer moved to the new audit and the declared-model note), R6 (covert verifier underperformance likely undetectable even to its vendor — Astra §9.2.2.3). The `scripts/cross_model_smoke_test.sh` header example uses a `<gpt-* id under test>` placeholder instead of a concrete id. No agent prompt sentence was removed; the release-facing suite version is unchanged.
### Fixed
- **`docs/RISK_REGISTER.md` R10 residual gap de-staled (#813).** R10 still claimed the guard-launcher degradations were "not yet indexed in the degradation registry (#769)" although #769 shipped in v3.21.1 itself (registry 1.3.0, five `write_scope_guard_*` rows) — the register contradicted the CHANGELOG and the registry file in the same tree, at the tag and on `main`. The stale clause is removed; the existing-controls line now points the guard's degrade posture at its five registry rows, and the residual gap keeps only the per-mechanism, per-channel loss description. Docs-only; found by an external cross-model fact-check of v3.21.1 claim surfaces. Known residue, accepted: no lint pins a residual-gap sentence against the mechanism inventory it references, so this class can recur; RR-1..RR-3 are unchanged.
@@ -761,6 +761,19 @@ Documents in an agent's context that are not its working target measurably worse
---
## Checkpoint authority fidelity
Every MANDATORY and FULL checkpoint in this pipeline is a decision the researcher makes in their own turn — the authority is `references/pipeline_state_machine.md` § Checkpoint decision provenance; this section is the orchestrator's operational mirror. Current frontier models are vendor-documented to fabricate or overstate a user's approval to pass a gate, to distort user intent when instructing a subagent, and to treat an automated message as the permission they asked for (evidence mapped in `audits/harness-retirement-2026-09-model-update.md` G-1). The orchestrator is the single point that both receives decisions and re-transmits them, so the fidelity discipline lives here:
- **Only a user turn is a decision.** A subagent report, a hook or tool result, a template's default branch, a checkpoint summary the orchestrator wrote, or a prior-turn paraphrase is never the user's choice. If the decision has not appeared in a user turn, the checkpoint is still open — ask again; never proceed on an inferred, assumed, or "obviously intended" answer. The Stage 6 terminal acknowledgement (vocabulary per the state machine's § Stage 6 terminal semantics, mirrored under Collaboration with state_tracker_agent below) counts only when the user gave it.
- **Re-transmit decisions verbatim.** When a dispatch carries a checkpoint decision, a consent grant, an override, or an authorization to a subagent, quote the user's words (or the exact deterministic authorization artifact) and label them as the user's. Never restate a narrow decision as a broader one, never write a first-person user statement into a dispatch, and never summarize a "no" or a scoped "yes" into an unscoped "yes".
- **Never assert consent or approval you did not receive.** Cross-model uploads, override-ladder rounds, integrity-correction authorizations, and read attestations require the user's explicit input at the surface that asks for it.
- **Report the same way.** Completion, checkpoint, and Process Record surfaces state what the user actually decided, in the user's words where the decision is quoted; a step the user did not confirm is reported as unconfirmed.
*Epistemic status: a decision-handling and reporting discipline, not a runtime guarantee. The deterministic authorization inputs (#670's `integrity-correction-authorization-input/1.0`, `/ars-mark-read`'s explicit scope) are the enforced layer where they exist; everywhere else this rule is prompt-level and is indexed as risk R11 in `docs/RISK_REGISTER.md`.*
---
## Collaboration with state_tracker_agent
Notify state_tracker_agent to update state whenever a stage begins or completes:
@@ -247,6 +247,10 @@ When Stage 6 runs, its completion is the pipeline's **terminal checkpoint**:
3. On acknowledgement: state_tracker marks Stage 6 `completed` and sets the pipeline global state to `completed`. This is the terminal transition — there is no next stage.
4. After `completed`, no stage transition is legal (see Prohibited Transitions). New requests start a new pipeline run or a targeted single-skill invocation (mid-entry).
### Checkpoint decision provenance
Every checkpoint decision, terminal acknowledgement, override, consent grant, and authorization input in this state machine exists only when it appears in a user turn. A subagent report, a hook or tool result, a template's default branch, an orchestrator-written checkpoint summary, or a paraphrase of an earlier turn is never the user's decision; a checkpoint whose decision has not appeared in a user turn is still open. Re-transmission to a subagent quotes the user's words (or the exact deterministic authorization artifact) and never widens them. Where a deterministic authorization artifact exists (the #670 integrity-correction authorization, the `/ars-mark-read` scope) it is the enforced form of this rule; elsewhere the rule is prompt-level. Mirrored operationally in `pipeline_orchestrator_agent.md` § Checkpoint authority fidelity; the risk is indexed as R11 in `docs/RISK_REGISTER.md`.
### Post-terminal adjudication-activity side channel (#673)
The ordinary state machine is authoritative and always terminates first. A
@@ -0,0 +1,203 @@
# Harness Retirement Audit — `academic-research-skills` (2026-09, model update)
| | |
|-|-|
| Repo path | `~/Projects/academic-research-skills` |
| Branch / commit audited | `main @ 9443623` |
| Date | 2026-09-05 |
| Target model (before → after) | Session: Claude Fable 5 → **Claude Fable 5.1**. Cross-model: `gpt-5.6-sol` (validated on the ChatGPT-subscription citation transport) → **`gpt-6-astra` recommended, provisional** (`gpt-5.6-sol` keeps its validated status on that transport) |
| Trigger | Two vendor system cards read in full: *Claude Fable 5.1 & Claude Mythos 5.1 System Card* (Anthropic, 2026-09) and *GPT-6 Astra System Card* (OpenAI, 2026-09-03). This is the model-change audit the 2026-09 routine audit said had not yet been needed |
| Scope | All 39 agent prompt bodies (23 Bucket A + 16 non-fenced), `shared/agents/`, `commands/`, `hooks/`, and the documentation surfaces that name a model (`docs/PERFORMANCE*.md`, `shared/cross_model_verification.md`, `shared/model_tiering.md`, `docs/SETUP*.md`, `.claude/CLAUDE.md`, `docs/RISK_REGISTER.md`) |
| Baseline | `audits/harness-retirement-2026-09.md` (0 findings at `e8bf858`; its keep-list is carried forward unchanged unless a row below says otherwise) |
| Method | Full read of both cards; each behavioral finding mapped to the ARS mechanism that assumes it (keep / retire / add); mechanical pattern scans over every prompt body; one live entry-gate smoke on the new verifier |
## Executive summary
- **Findings: 0 P0, 4 applied doc/harness-currency fixes (MU-001 MU-004), 0 prompt-text retirements, 2 deferred (MU-005, MU-013), 8 keep-as-debt annotations now backed by a system-card citation.**
- **No agent prompt sentence expired.** Both cards describe the failure classes ARS's remaining scaffolds guard against as *still present* in the new models — stated-guess-as-fact and exaggerated completeness (Fable 5.1 §2.3.3), unhedged estimates and framing extension (§2.2.4), repeated failing actions (§2.3.3), suppressed caveats (§6.6.1), overreach and permissive reading of instructions (Astra §8.6). Retiring those scaffolds on the strength of "the new model is better" would remove protection against silent failures the vendors themselves still report.
- **What did expire is model currency in documentation and one harness vocabulary gap**, all applied in this PR: the recommended-model line, the cross-model lineup, and the Codex transport's reasoning-effort set (which did not know `ultra`).
- **The cards also motivated four additions** (not retirements), listed under "Guardrails added" below, each grounded in a cited section and indexed in `docs/RISK_REGISTER.md`.
## Findings
Decision vocabulary: **applied** (in this PR), **keep** (iron rule: load-bearing, annotated), **defer** (needs a check this audit could not perform offline, or a maintainer decision).
```
[MU-001] docs/PERFORMANCE.md:3, docs/PERFORMANCE.zh-TW.md:3 | category 1 (model currency, docs)
Excerpt: "the current frontier Claude model (Fable 5 at the time of writing)"
Rationale: stale one generation; the token table's Opus 4.x measurement basis is a
record and stays, but the recommended-model name is current-state text.
Applied: name → Fable 5.1; added a clearly-labelled list-price re-derivation (~$7 per
full run at US$10/US$50 per MTok) marked as arithmetic, not a re-measurement.
```
```
[MU-002] shared/cross_model_verification.md:44 | category 1 (model currency, docs)
Excerpt: "_(inherited Claude Code session model — e.g., Fable 5)_"
Applied: example → Fable 5.1. The row still names no version by design.
```
```
[MU-003] shared/cross_model_verification.md (Supported Models table, recommended pair,
setup blocks, id-status allowlist, bakeoff prose), docs/SETUP.md + docs/SETUP.zh-TW.md
(quick-setup + codex blocks), .claude/CLAUDE.md:294 | category 1 (model currency)
Excerpt: "GPT-5.6 Sol … current OpenAI flagship, recommended OpenAI verifier"
Rationale: GPT-6 Astra superseded the GPT-5.6 family on 2026-09-03. The repo's own
recommendation policy (#783: recommendation follows generation currency; validated
is earned only by the sealed bakeoff) decides what happens next, so this is applied
as that policy prescribes rather than as a measurement claim.
Applied: `gpt-6-astra` listed as provisional on both transports and named the
recommended OpenAI verifier; `gpt-5.6-sol` keeps validated status on the
ChatGPT-subscription citation transport and provisional status on the API route;
`gpt-5.5` / `gpt-5.5-pro` / `gemini-3.1-pro-preview` unchanged; the bakeoff
section now names the per-transport baseline. SETUP en/zh-TW example sets kept
identical (parity lint).
Evidence for the listing: entry-gate smoke PASS on the citation transport,
2026-09-05, codex-cli 0.153.4 (`scripts/cross_model_smoke_test_codex.sh`:
detection available, `VERIFIED` with one bound source on the Vaswani et al.
fixture). An entry gate is the precondition for a Promotion Bakeoff, not one.
```
```
[MU-004] scripts/cross_model_codex_transport.py (turn/start effort guard) | category 3
(harness vocabulary tuned to an older lineup)
Excerpt: `{"minimal", "low", "medium", "high", "xhigh", "max"}`
Rationale: GPT-6 Astra's Codex harness runs at `ultra` effort (system card
§10.1.2.5); the app-server schema types ReasoningEffort as any non-empty string the
served model advertises (verified with `codex app-server generate-json-schema` on
0.153.4), so the set is ARS's own guard and was one value short.
Applied: `ultra` added as a named constant; new test pins forwarding on turn/start
and fail-closed rejection of an unknown value. The API route stays pass-through.
```
```
[MU-005] shared/cross_model_verification.md (legacy note + id-status allowlist) |
category 1 (legacy ids)
Excerpt: "`gpt-5.4` / `gpt-5.4-pro` remain accepted for existing setups"
Rationale: three generations back after this update. Removing them from the
validated allowlist would only change a status announcement, but the honest basis
for retirement is a first-party deprecation notice, which this audit did not
check (offline by design; the Astra card still reports gpt-5.4-thinking figures,
so the model line is at least still evaluated by its vendor).
Decision: defer — re-check against OpenAI's deprecation page at the next release;
retire then if the ids are withdrawn.
```
```
[MU-006] scripts/dispatch_e4_panel.py (`--model` default `claude-opus-5`, `--effort`
default `xhigh`) | category 1 (hardcoded model id)
Rationale: an evaluation-harness default, not prompt text; changing it changes the
measurement identity of every recorded #574/#610 fleet. Out of scope per the
skill (test/eval fixtures pin old behaviour).
Decision: keep; annotate here so it does not resurface.
```
```
[MU-007] deep-research/agents/socratic_mentor_agent.md:758 | category 3-like (numeric cap)
Excerpt: "Keep responses under 400 words — past that, you're lecturing"
Rationale for flagging: reads like a capability-era length cap.
Rationale for keeping: Fable 5.1 §8.4 records that at higher effort the model adds
unrequested, out-of-scope content and that an explicit brevity instruction reduced
it; §8.17 records length-adjusted scoring penalising verbosity. A brevity rule is
therefore still load-bearing, and this one is a Socratic-domain rule, not a
model workaround.
Decision: keep.
```
```
[MU-008] shared/cross_model_verification.md verifier prompts ("NOT_SEARCHED — you could
not actually search") | category 2 (anti-hallucination patch)
Rationale for flagging: Astra §8.3.2 reports the model acknowledges an unavailable
tool ten times more reliably than GPT-5.6 Sol.
Rationale for keeping: the prompt is multi-runtime (GPT-5.5, GPT-5.6, compatible
providers) and the real boundary is the grounding guard, not the wording; the
document already says so.
Decision: keep.
```
```
[MU-009] pipeline_orchestrator_agent.md hard boundary 7 ("Do not fabricate materials");
Bucket A anti-hallucination contract clauses (August keep-list) | category 2
Rationale: Fable 5.1 §2.3.3 — "often states easy-to-check guesses as facts,
exaggerates the completeness of its work, fails to verify important claims";
§6.6.1 — claims of runs that never executed, approval represented that was never
given, a material caveat suppressed. Silent-failure class; still exhibited.
Decision: keep.
```
```
[MU-010] Bounded retry / repair language (August keep-list) | category 5
Rationale: Fable 5.1 §2.3.3 — "repeatedly trying actions that are not working …
destroying its own work". The bound is the protection.
Decision: keep.
```
```
[MU-011] Frame-lock detection (Socratic mentor, DA), WP research-question advisory,
claim-strength ladder, protected hedges | category 2/6
Rationale: Fable 5.1 §2.2.4 — "extends whatever framing the user supplies rather than
challenging it, such that weak questions produce weak answers"; "produces unhedged
estimates"; "presents overly optimistic plans and reassures users past obstacles
until challenged". Direct vendor evidence that these are not expired.
Decision: keep.
```
```
[MU-012] Anti-sycophancy DA concession scoring (v3.0) | category 5
Rationale: Fable 5.1 §6.1.2 reports the model is less sycophantic than Opus 5, and
§2.2.4 still reports reassurance past obstacles until challenged — a degree
improvement, not a kind improvement.
Decision: keep.
```
```
[MU-013] Reviewer inputs and authorship cues | no existing scaffold
Rationale: Fable 5.1 §6.5.3 reports a self-recognition bias (more lenient grading
when told Claude wrote the text). A grep of the reviewer agents found no authorship
cue in their inputs, and manuscripts are user-authored, so nothing needs retiring.
A positive rule ("strip AI-authorship cues from reviewer inputs") would be new
scope.
Decision: defer to a maintainer issue; not applied here.
```
```
[MU-014] academic-paper-reviewer/agents/methodology_reviewer_agent.md:106 ("no preamble,
no other section") | looked like an update-suppressor (Fable 5.1 migration guidance
says to remove instructions that suppress progress updates)
Rationale: it is the output contract of a deterministic extraction section consumed
by a checker, not a suppression of user-facing updates.
Decision: keep.
```
## Guardrails added (grounded in the cards; additions, not retirements)
| # | Where | Card evidence | What it does |
|---|---|---|---|
| G-1 | `academic-pipeline/references/pipeline_state_machine.md` § Checkpoint decision provenance (authority); `academic-pipeline/agents/pipeline_orchestrator_agent.md` § Checkpoint authority fidelity (operational mirror); `docs/RISK_REGISTER.md` R11 | Fable 5.1 §6.2.1 (fabricated user quotation to pass an approval gate; distorted user intent in subagent instructions), §6.6.1 (approval represented that was never given), §6.4.5 (slightly more willing to bypass approval gates); Astra §8.8 (27% proceeded on automated messages after asking permission) | Only a user turn is a decision; decisions re-transmitted verbatim; no asserted consent; reporting in the user's words. Prompt-level; enforced only where a deterministic authorization artifact exists (stated). The five locked pipeline surfaces' hash constant is updated in the same commit |
| G-2 | `shared/cross_model_verification.md` § Provider-side monitoring and safety interventions; `shared/contracts/degradation_registry.json` `cross_model_unavailable` | Astra §10.2.3.1 (misalignment monitor can pause/end conversations; API stops not resumable), §10.2.3.2 (misuse monitors block mid-stream), §10.2.2.2 (stricter cyber boundary), §10.2.5 (escalation to manual review) | An intervention is a transport failure, never a verdict or `NOT_SEARCHED`; content-triggered refusals are `unavailable`; consent must assume provider staff may read escalated content; ARS never consumes the verifier's reasoning (§9 monitorability decrease) |
| G-3 | `shared/model_tiering.md` (declared-model note under Resolving a tier at dispatch time); `docs/RISK_REGISTER.md` R5 | Fable 5.1 §3.2 / §5.2 (classifier-triggered fallback to Opus 4.8 / Opus 5 per request, invisible to the caller) | "Session model" in tiering and provenance is the declared model, not a per-call attestation |
| G-4 | `shared/cross_model_verification.md` § Promotion Bakeoff caveat; `docs/RISK_REGISTER.md` R1, R4, R6 | Astra §8.6 / §8.8.1 (verbalized evaluation awareness 9.6% / 4151%), §9.2.2.3 (covert sandbagging likely undetectable by the vendor's monitors); Fable 5.1 §2.2.4 (fabrication rare, misrepresentation the residual class) | Bakeoff results carry an evaluation-awareness caveat; R6 records the observability limit of a second model; R1 records the weight shift onto the supports-the-claim half; R4 records provider-side escalation |
## Mechanical scan results (all 39 agent bodies + shared agents + commands + hooks)
- Hardcoded model pins in prompt text (`claude-*`, `Opus 4.x`, `Sonnet 4.x`, `Haiku`, `gpt-*`): **0 in prompts.** Remaining hits are the eval-harness default (MU-006), the light-mode `sonnet` cost routing in `commands/` (a cost decision, not a capability workaround — August keep-list), and records (CHANGELOG, audits, evals, design docs — out of scope).
- Sampling / budget overrides (`temperature`, `top_p`, `max_tokens`, `budget_tokens`): **0** in prompts; the `temperature: 0.1` in the verifier call patterns is a documented determinism choice for an external provider and is not a Claude parameter.
- Reasoning scaffolds ("think step by step", "show your reasoning", `<thinking>`, `<scratchpad>`): **0 hits.**
- Update-suppressor / anti-formatting instructions: **1 hit**, a format contract (MU-014).
- Numeric length caps: **1 hit**, kept with vendor evidence (MU-007).
- Anti-hallucination phrasing: every hit is a domain contract clause (August verdict), now with a current-model citation for why it stays (MU-009, MU-011).
## Verification
- Both system cards were read in full (7,737 and 4,319 extracted text lines respectively); every section cited above was checked against the text.
- Live entry-gate smoke for the new verifier: `scripts/cross_model_smoke_test_codex.sh` with `ARS_CROSS_MODEL=gpt-6-astra` on codex-cli 0.153.4, 2026-09-05 → detection `available: true`, receipt `VERIFIED`, `searched: true`, one bound source, `RESULT: PASS`. One live call; no bakeoff, no measurement claim.
- Codex app-server schema (0.153.4) generated locally: `ReasoningEffort` is typed as a non-empty string, which is why the transport's effort set is ARS-owned (MU-004).
- Every lint in `.github/workflows/spec-consistency.yml` (69 checks), `check_command_frontmatter_name.py`, `check_prisma_trAIce_freshness.py`, and the unified pytest manifest were run locally on 2026-09-05 before the PR was opened; all green.
- No agent prompt sentence was removed by this audit; the only prompt-body change is the additive orchestrator section (G-1) and its authority paragraph in the pipeline state machine.
## Routing checklist
- [x] Audit report committed under `audits/`.
- [x] P0 retirements: none.
- [x] Applied currency fixes and additions logged in `CHANGELOG.md` `[Unreleased]`.
- Deferred to the next model-change audit or a maintainer issue: MU-005 (legacy `gpt-5.4*` ids) and MU-013 (authorship-cue rule).
+3 -1
View File
@@ -1,6 +1,6 @@
# ARS Performance Notes
> **Recommended model: the current frontier Claude model** (Fable 5 at the time of writing) with **Max plan** (or equivalent configuration). Current Claude models use adaptive thinking; you no longer set a fixed thinking budget.
> **Recommended model: the current frontier Claude model** (Fable 5.1 at the time of writing) with **Max plan** (or equivalent configuration). Current Claude models use adaptive thinking; you no longer set a fixed thinking budget.
>
> The full academic pipeline (10 stages) consumes a **large amount of tokens** — a single end-to-end run can exceed 200K input + 100K output tokens depending on paper length and revision rounds. Budget accordingly.
>
@@ -22,6 +22,8 @@
*Estimates based on a ~15,000-word paper with ~60 references. Actual usage varies with paper length, revision rounds, and dialogue depth. Costs measured on Opus 4.x at Anthropic API pricing as of April 2026 — treat as order-of-magnitude anchors under newer models rather than exact quotes.*
> **Fable 5.1 re-derivation (2026-09).** At Claude Fable 5.1 list pricing as of 2026-09 (US$10 / US$50 per million input / output tokens), the full-pipeline token figures above (~200K in + ~100K out) come to roughly **~$7** per run — arithmetic on the token columns, not a re-measurement: no pipeline run has been re-timed on Fable 5.1. Fable 5.1 also always reasons (thinking cannot be disabled), so dialogue-heavy modes may spend more output tokens than the Opus 4.x rows recorded.
> **v3.11 citation verification (#182).** The deterministic citation-existence gate calls external bibliographic APIs (Semantic Scholar / OpenAlex / Crossref / arXiv), not the LLM, so it adds **no Claude token cost** to the figures above — only network latency on first lookup. The persistent SQLite cache (`~/.cache/ars/verification.db`, 90-day TTL) means each paper is verified once and reused across drafts; a re-run over an already-cached bibliography does no network work. See [SETUP](SETUP.md#citation-verification-cache-v3.11-182).
## Recommended Claude Code settings
+3 -1
View File
@@ -1,6 +1,6 @@
# ARS 效能說明
> **建議模型:當前最新一代 Claude 模型**(撰寫當下為 Fable 5搭配 **Max plan**(或同等配置)。現行 Claude 模型採用 adaptive thinking不需要手動指定 thinking budget。
> **建議模型:當前最新一代 Claude 模型**(撰寫當下為 Fable 5.1),搭配 **Max plan**(或同等配置)。現行 Claude 模型採用 adaptive thinking不需要手動指定 thinking budget。
>
> 完整學術 pipeline10 階段)會消耗**大量 token** — 單次完整執行可能超過 200K 輸入 + 100K 輸出 token視論文長度和修訂輪數而定。請依預算斟酌使用。
>
@@ -22,6 +22,8 @@
*以 ~15,000 字論文、~60 篇引用為基準估算。實際消耗隨論文長度、修訂輪數、對話深度而異。費用以 Opus 4.x 實測、Anthropic API 2026 年 4 月定價計算;換用更新模型時請當成數量級參考,不是精確報價。*
> **Fable 5.1 換算2026-09。**以 Claude Fable 5.1 在 2026-09 的牌價(每百萬輸入/輸出 token 各 US$10US$50換算上表完整 pipeline 的 token 數(約 200K 輸入 + 100K 輸出)每次約 **~$7**。這是用 token 欄位算出來的數字,不是重新實測:沒有任何 pipeline 在 Fable 5.1 上重跑計時。Fable 5.1 一律會推理thinking 無法關閉),所以對話密集的模式可能比 Opus 4.x 列多用一些輸出 token。
> **v3.11 引用查驗(#182** 確定性引用存在性 gate 呼叫的是外部書目 APISemantic Scholar / OpenAlex / Crossref / arXiv不是 LLM因此**不增加上表的 Claude token 成本**——只在首次查詢時有網路延遲。持久化 SQLite cache`~/.cache/ars/verification.db`90 天 TTL讓每篇論文只查驗一次、跨草稿重用對已 cache 的書目重跑不做任何網路請求。見 [SETUP](SETUP.zh-TW.md#引用查驗-cachev3.11182)。
## 建議 Claude Code 設定
+44 -2
View File
@@ -47,7 +47,12 @@ does not support.
- **Residual gap**: no measured hallucinated-citation catch rate — that needs an
independently-authored ground-truth set, not one derived from the gate's own
reducer; the Claim Registry's semantic completeness is unknown by contract. No
open issue schedules either measurement yet.
open issue schedules either measurement yet. The vendor's own evaluation of the
current session model reports fabricated references as rare and misrepresented
findings or conclusions as the residual error class that needs domain
familiarity to catch (Claude Fable 5.1 system card §2.2.4), which moves the
weight of this row onto the supports-the-claim half — exactly the unmeasured
claim-verification row above.
### R2 — Silent claim-strength drift in revision
@@ -91,6 +96,10 @@ provider.
- **Residual gap**: consent surfaces are prompt-contract layers whose behavior
varies with the session model; the no-autonomous-publication line is a scope
boundary and review criterion, not a runtime guarantee (`POSITIONING.md`).
Provider-side safety monitoring may escalate a flagged conversation to human
review at the provider (`shared/cross_model_verification.md` § Provider-side
monitoring and safety interventions), so consent has to assume provider staff
may read what is sent.
### R5 — Upstream model-version behavior drift
@@ -99,11 +108,18 @@ described.
- **Existing controls**: commit-frozen measurement rows re-measured per change; the
matrix staleness rule (`stale_after_days`); periodic harness-retirement audits
(`audits/harness-retirement-2026-08.md` is the most recent completed report).
(`audits/harness-retirement-2026-09-model-update.md` is the most recent
completed report, run on the Fable 5 → Fable 5.1 and GPT-5.6 Sol → GPT-6 Astra
change).
- **Evidence status**: `NOT_RUN` (asserted here; no capability-matrix row) — no
measurement of cross-version drift itself exists.
- **Residual gap**: measured rows are model- and time-specific by contract;
re-running them on a model change is a manual maintainer action, not automated.
The session model can also change per request without a version change: the
runtime's classifier-triggered fallback is invisible to the suite
(`shared/model_tiering.md`, the declared-model note under Resolving a tier at
dispatch time), so every "session model" in a provenance record is the declared
model, not a per-call attestation.
### R6 — Correlated same-family model error
@@ -117,6 +133,11 @@ Generator and reviewer share a model family, so both miss the same failures.
- **Residual gap**: first measured reviewer error profile pending
([#653](https://github.com/Imbad0202/academic-research-skills/issues/653)); a
second model is an error-detection control, not organizational independence.
A second model is also only as observable as its outputs: the current OpenAI
verifier's vendor reports that covert underperformance by that model would
likely evade the vendor's own monitors (GPT-6 Astra system card §9.2.2.3), and
ARS's typed evidence anchors bound what a verifier can assert, not what it
withholds.
### R7 — User over-reliance / rubber-stamping
@@ -172,3 +193,24 @@ A user installs through a channel where documented enforcement never runs.
matrix): prompt-level protocols survive in most non-plugin channels but are absent
in claude.ai Projects, and hook enforcement is absent outside the plugin channel
unless the user wires the hook into their own settings manually (matrix note 3).
### R11 — Fabricated or distorted checkpoint authority
The orchestrating model treats something other than the researcher's own turn as
a checkpoint decision — an automated message, a subagent's report, a template
default, or its own paraphrase — or restates the researcher's decision to a
subagent as a broader authorization than was given.
- **Existing controls**: MANDATORY checkpoint templates that wait for an explicit
user turn (`academic-pipeline/references/pipeline_state_machine.md`); the
orchestrator's checkpoint-authority fidelity rule
(`academic-pipeline/agents/pipeline_orchestrator_agent.md`); deterministic
authorization inputs that bind an author choice to exact patch bytes
(`scripts/revision_roadmap.py`, the #670 integrity-correction authorization);
read attestations that are declared, never inferred (`/ars-mark-read`).
- **Evidence status**: `NOT_RUN` (asserted here; no capability-matrix row) — the
deterministic authorization inputs are CI-pinned; the prompt-level rule is not
measured on any session model.
- **Residual gap**: the failure class is vendor-documented, not ARS-measured
(evidence mapped in `audits/harness-retirement-2026-09-model-update.md` G-1);
the prompt rule is trust-based.
+9 -6
View File
@@ -184,13 +184,14 @@ ARS works with the inherited Claude session model alone. For higher confidence,
```bash
# Step 1: Set your API key (choose one or both)
export OPENAI_API_KEY="sk-your-key-here" # For GPT-5.6 Sol / GPT-5.5
export OPENAI_API_KEY="sk-your-key-here" # For GPT-6 Astra / GPT-5.6 Sol / GPT-5.5
export GOOGLE_AI_API_KEY="AIza-your-key-here" # For Gemini 3.1 Pro
# Step 2: Choose your cross-verification model
export ARS_CROSS_MODEL="gpt-5.6-sol" # Current OpenAI flagship — provisional pending ARS validation (run scripts/cross_model_smoke_test.sh)
export ARS_CROSS_MODEL="gpt-6-astra" # Current OpenAI flagship — provisional pending ARS validation (run scripts/cross_model_smoke_test.sh)
# or: export ARS_CROSS_MODEL="gemini-3.1-pro-preview" # Current Google flagship — validated, strong at factual verification
# or: export ARS_CROSS_MODEL="gpt-5.5" # Previous generation — validated (designated bakeoff baseline)
# or: export ARS_CROSS_MODEL="gpt-5.6-sol" # Previous generation — validated on the ChatGPT-subscription citation transport, provisional on this API route
# or: export ARS_CROSS_MODEL="gpt-5.5" # Previous generation — validated (designated API-route bakeoff baseline)
# Optional: reasoning effort for OpenAI verifier calls (unset = provider default)
# export ARS_CROSS_MODEL_REASONING_EFFORT="medium"
@@ -225,10 +226,12 @@ Devil's Advocate, Reviewer 2, calibration, re-review, or checkpoint judgments.
```bash
# Citation-integrity calls only. General DA/reviewer/judgment calls remain on API transport.
export ARS_CROSS_MODEL_TRANSPORT="codex"
# gpt-6-astra: current OpenAI flagship — provisional on this transport (entry-gate
# smoke PASS 2026-09-05 on codex-cli 0.153.4; no bakeoff run yet).
export ARS_CROSS_MODEL="gpt-6-astra"
# gpt-5.6-sol is validated for THIS transport (2026-08-19 codex-transport bakeoff,
# superiority on recall + latency — audits/bakeoff-gpt-5-6-sol-codex-2026-08-19.md).
# gpt-5.5 remains the validated bakeoff baseline alternative.
export ARS_CROSS_MODEL="gpt-5.6-sol"
# superiority on recall + latency — audits/bakeoff-gpt-5-6-sol-codex-2026-08-19.md):
# export ARS_CROSS_MODEL="gpt-5.6-sol"
python3 scripts/cross_model_codex_transport.py detect
# The producer sends one closed codex_citation_request/1.0 object on stdin:
+9 -6
View File
@@ -180,13 +180,14 @@ ARS 使用繼承的 Claude session 模型即可完整運作。想要更高信心
```bash
# Step 1: Set your API key (choose one or both)
export OPENAI_API_KEY="sk-your-key-here" # For GPT-5.6 Sol / GPT-5.5
export OPENAI_API_KEY="sk-your-key-here" # For GPT-6 Astra / GPT-5.6 Sol / GPT-5.5
export GOOGLE_AI_API_KEY="AIza-your-key-here" # For Gemini 3.1 Pro
# Step 2: Choose your cross-verification model
export ARS_CROSS_MODEL="gpt-5.6-sol" # Current OpenAI flagship — provisional pending ARS validation (run scripts/cross_model_smoke_test.sh)
export ARS_CROSS_MODEL="gpt-6-astra" # Current OpenAI flagship — provisional pending ARS validation (run scripts/cross_model_smoke_test.sh)
# or: export ARS_CROSS_MODEL="gemini-3.1-pro-preview" # Current Google flagship — validated, strong at factual verification
# or: export ARS_CROSS_MODEL="gpt-5.5" # Previous generation — validated (designated bakeoff baseline)
# or: export ARS_CROSS_MODEL="gpt-5.6-sol" # Previous generation — validated on the ChatGPT-subscription citation transport, provisional on this API route
# or: export ARS_CROSS_MODEL="gpt-5.5" # Previous generation — validated (designated API-route bakeoff baseline)
# Optional: reasoning effort for OpenAI verifier calls (unset = provider default)
# export ARS_CROSS_MODEL_REASONING_EFFORT="medium"
@@ -221,10 +222,12 @@ OpenAI API key 而改走該訂閱。這不涵蓋魔鬼代言人、Reviewer 2、
```bash
# Citation-integrity calls only. General DA/reviewer/judgment calls remain on API transport.
export ARS_CROSS_MODEL_TRANSPORT="codex"
# gpt-6-astra: current OpenAI flagship — provisional on this transport (entry-gate
# smoke PASS 2026-09-05 on codex-cli 0.153.4; no bakeoff run yet).
export ARS_CROSS_MODEL="gpt-6-astra"
# gpt-5.6-sol is validated for THIS transport (2026-08-19 codex-transport bakeoff,
# superiority on recall + latency — audits/bakeoff-gpt-5-6-sol-codex-2026-08-19.md).
# gpt-5.5 remains the validated bakeoff baseline alternative.
export ARS_CROSS_MODEL="gpt-5.6-sol"
# superiority on recall + latency — audits/bakeoff-gpt-5-6-sol-codex-2026-08-19.md):
# export ARS_CROSS_MODEL="gpt-5.6-sol"
python3 scripts/cross_model_codex_transport.py detect
# The producer sends one closed codex_citation_request/1.0 object on stdin:
+2 -2
View File
@@ -71,9 +71,9 @@ REPO_ROOT = Path(__file__).resolve().parent.parent
# ---------------------------------------------------------------------------
CONTENT_LOCKS = {
"academic-pipeline/SKILL.md": "5b9b92ad1df7a3a55c1f67f0d2d554af63e476c5019e0d3497b3f28ebacf9116",
"academic-pipeline/agents/pipeline_orchestrator_agent.md": "5ca2474cbf6e1ce1efe1211426575d60d0fd9b663b21f16418b928a4441abc4c",
"academic-pipeline/agents/pipeline_orchestrator_agent.md": "56c4d8eaede4c6228404c13608b3e2113972d97e7884fd5b1a51fb982fd31bf7",
"academic-pipeline/agents/state_tracker_agent.md": "2716bab5686a6129777f595ad86bf1e1cc01fa5d8d1ec192fa8880018dfe968a",
"academic-pipeline/references/pipeline_state_machine.md": "70872764d04a1cabdc1a342cbaf66ebed98a3764d506136d541c1880759510e7",
"academic-pipeline/references/pipeline_state_machine.md": "6ef7703d3b24152812c5767570f36f01578846383cb9fd2d42bc8f79084e57ae",
"academic-pipeline/references/process_summary_protocol.md": "1052d8cb8ee00c1cd0fcc70a18aee5a0f92db2ebe0a74930b04d4b05d888cfdf",
}
+11 -1
View File
@@ -35,6 +35,16 @@ RECEIPT_SCHEMA_VERSION = "ars-codex-citation-receipt/1.0"
TRANSPORT = "codex_subscription"
AUTH_MODE = "chatgpt_subscription"
MIN_CODEX_VERSION = (0, 147, 0)
# Closed vocabulary of reasoning efforts ARS forwards on turn/start. The
# app-server schema types ReasoningEffort as any non-empty string the served
# model advertises (generate-json-schema, codex-cli 0.153.4), and the provider
# rejects a value the served model does not advertise one RPC later — so this
# set buys an earlier, better-named error (INVALID_REASONING_EFFORT), not a
# safety property. `ultra` joined with GPT-6 Astra (system card 2026-09-03
# §10.1.2.5: the Codex harness ran at Ultra reasoning effort).
ACCEPTED_REASONING_EFFORTS = frozenset(
{"minimal", "low", "medium", "high", "xhigh", "max", "ultra"}
)
MAX_REQUEST_BYTES = 32 * 1024
MAX_FIELD_CHARS = 8192
@@ -886,7 +896,7 @@ def run_app_server(
}
effort = environ.get("ARS_CROSS_MODEL_REASONING_EFFORT", "")
if effort:
if effort not in {"minimal", "low", "medium", "high", "xhigh", "max"}:
if effort not in ACCEPTED_REASONING_EFFORTS:
raise TransportError("INVALID_REASONING_EFFORT")
turn_params["effort"] = effort
_send_rpc(proc, {"id": 3, "method": "turn/start", "params": turn_params})
+4 -4
View File
@@ -4,16 +4,16 @@
# Purpose: validate that a given OpenAI verifier model (ARS_CROSS_MODEL, gpt-* id)
# behaves correctly against the documented call pattern in
# shared/cross_model_verification.md BEFORE it is used in a real pipeline run.
# Primary use case: vetting a newly released model id (e.g. gpt-5.6-sol, listed
# as provisional) whose response shape / grounding behavior has no ARS operating
# history yet.
# Primary use case: vetting a newly released model id (one listed as provisional
# in shared/cross_model_verification.md) whose response shape / grounding
# behavior has no ARS operating history yet.
#
# This is a LIVE test: it issues one real Responses API call (with hosted
# web_search) against a stable, known-good reference. It costs a fraction of a
# cent and needs OPENAI_API_KEY, so it is NOT wired into CI — run it manually:
#
# export OPENAI_API_KEY="sk-..."
# export ARS_CROSS_MODEL="gpt-5.6-sol" # model under test
# export ARS_CROSS_MODEL="<gpt-* id under test>" # model under test
# export ARS_CROSS_MODEL_REASONING_EFFORT="medium" # optional (default: medium)
# bash scripts/cross_model_smoke_test.sh
#
@@ -1150,3 +1150,46 @@ def test_surrogate_final_and_control_query_fail_closed() -> None:
assert runtime.parse_app_server_messages(
messages, raw_stream=raw, request=_request(), model="gpt-5.6"
)["reason_code"] == "EVENT_STREAM_INVALID"
def test_reasoning_effort_ultra_is_forwarded_and_unknown_effort_fails_closed(
tmp_path: Path,
) -> None:
# The closed vocabulary (rationale on the constant) is forwarded verbatim on
# turn/start and fails closed on a value outside the set.
assert "ultra" in runtime.ACCEPTED_REASONING_EFFORTS
home = tmp_path / "custom-home"
_make_auth(home)
fake_bin, capture_path = _make_fake_codex(tmp_path)
env = _base_env(fake_bin, home)
env["ARS_CROSS_MODEL_REASONING_EFFORT"] = "ultra"
completed = subprocess.run(
[str(WRAPPER)],
input=runtime.canonical_json(_request()),
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
env=env,
timeout=15,
check=False,
)
assert completed.returncode == 0, completed.stderr.decode()
receipt = runtime.strict_json_loads(completed.stdout)
assert receipt["verdict"] == "VERIFIED"
capture = json.loads(capture_path.read_text(encoding="utf-8"))
turn_starts = [m for m in capture["requests"] if m.get("method") == "turn/start"]
assert len(turn_starts) == 1
assert turn_starts[0]["params"]["effort"] == "ultra"
# The rejection is pinned in-process (same shape as the APP_SERVER_TIMEOUT
# test): the shell → interpreter boundary is already proven by the run above,
# and main() formats every TransportError the same way.
env["ARS_CROSS_MODEL_REASONING_EFFORT"] = "extreme"
with pytest.raises(runtime.TransportError) as exc_info:
runtime.run_app_server(
_request(),
model="gpt-5.6",
codex=str(fake_bin / "codex"),
source_auth=home / "auth.json",
environ=env,
)
assert exc_info.value.code == "INVALID_REASONING_EFFORT"
+55 -4
View File
@@ -199,6 +199,15 @@ LINE_BUDGET_684_REVIEW_CRITERIA_BINDING = 44
# leaves six lines of review headroom.
LINE_BUDGET_743_INQUIRY_LEDGER = 60
# The 2026-09 model-update pass (audits/harness-retirement-2026-09-model-update.md
# G-1, risk register R11) adds one H2 section, `## Checkpoint authority
# fidelity`, the orchestrator's operational mirror of the pipeline state
# machine's Checkpoint decision provenance authority. This is an independent
# 2026-09 extension, so it is subtracted from the historical v3.6.7 budget and
# receives its own bounded test. Measured at landing: 13 lines; budget 18
# leaves 5 lines of headroom.
LINE_BUDGET_G1_CHECKPOINT_AUTHORITY = 18
# All 24 failure phase IDs from spec §5.6 inventory (7 P-PA-* + 17 P-PB-*).
# These must each appear at least once in the orchestrator prompt as
# cross-references to spec §5.6 (NOT inline procedural definitions —
@@ -779,6 +788,25 @@ def _measure_743_inquiry_ledger_lines(text: str) -> int:
return inquiry_lines + max(0, current_rule_lines - 1)
def _measure_g1_checkpoint_authority_lines(text: str) -> int:
"""Return the line count of the `## Checkpoint authority fidelity` section.
Measures from the H2 heading to the next heading of any level (H1-H4),
the same convention as the other extension-section helpers above.
"""
import re as _re
anchor = _re.compile(r"(?m)^[ \t]*##[ \t]+Checkpoint authority fidelity[ \t]*$")
match = anchor.search(text)
if match is None:
return 0
heading_end = text.find("\n", match.end())
search_start = heading_end + 1 if heading_end >= 0 else len(text)
next_heading = _re.search(r"(?m)^[ \t]*#{1,4}[ \t]+", text[search_start:])
end = search_start + next_heading.start() if next_heading else len(text)
return len(text[match.start():end].splitlines())
class Advisory660LineBudgetTest(unittest.TestCase):
"""#660 tortured-phrase dispatch block stays independently bounded."""
@@ -869,6 +897,26 @@ class InquiryLedger743LineBudgetTest(unittest.TestCase):
)
class CheckpointAuthorityG1LineBudgetTest(unittest.TestCase):
"""2026-09 checkpoint-authority fidelity section stays independently bounded."""
def test_g1_checkpoint_authority_within_budget(self) -> None:
text = _read_prompt()
block_lines = _measure_g1_checkpoint_authority_lines(text)
self.assertGreater(
block_lines,
0,
"`## Checkpoint authority fidelity` section missing from "
"pipeline_orchestrator_agent.md",
)
self.assertLessEqual(
block_lines,
LINE_BUDGET_G1_CHECKPOINT_AUTHORITY,
f"checkpoint-authority fidelity section is {block_lines} lines, over "
f"its {LINE_BUDGET_G1_CHECKPOINT_AUTHORITY}-line budget",
)
class Dispatch576LineBudgetTest(unittest.TestCase):
"""#576 Spec B Stage 3' contract-dispatch block within
`LINE_BUDGET_576_STAGE3P_DISPATCH` line budget.
@@ -954,6 +1002,7 @@ class Phase66LineBudgetTest(unittest.TestCase):
advisory_673_lines = _measure_673_adjudication_activity_lines(text)
criteria_684_lines = _measure_684_review_criteria_binding_lines(text)
inquiry_743_lines = _measure_743_inquiry_ledger_lines(text)
authority_g1_lines = _measure_g1_checkpoint_authority_lines(text)
# v3.6.7-only line count: total minus v3.7.1 Step 3b, v3.7.3
# finalizer extension, v3.8 §3.6 audit-gate, v3.9.0 triangulation
# extension, v3.10 terminal-policy extension, the #394 slice-4
@@ -963,14 +1012,15 @@ class Phase66LineBudgetTest(unittest.TestCase):
# checkpoint-rendering, the #660 tortured-phrase advisory dispatch,
# the #672 cross-document advisory dispatch, AND the #673
# adjudication-activity wiring, the #684 review-criteria binding
# lifecycle, AND the #743 inquiry-ledger/sidecar extension (each has
# its own dedicated budget test).
# lifecycle, the #743 inquiry-ledger/sidecar extension, AND the
# 2026-09 checkpoint-authority fidelity section (each has its own
# dedicated budget test).
v367_line_count = (
total_lines - step_3b_lines - v3_7_3_lines - v3_8_lines
- v3_9_0_lines - v3_10_lines - gate_394_lines - seq_390_lines
- authority_670_lines - dispatch_576_lines - evidence_656_lines
- advisory_660_lines - advisory_672_lines - advisory_673_lines
- criteria_684_lines - inquiry_743_lines
- criteria_684_lines - inquiry_743_lines - authority_g1_lines
)
ceiling = BASELINE_LINE_COUNT + LINE_BUDGET_OVER_BASELINE
self.assertLessEqual(
@@ -994,7 +1044,8 @@ class Phase66LineBudgetTest(unittest.TestCase):
f"{advisory_673_lines} are in the #673 adjudication-activity "
f"wiring, and {criteria_684_lines} are in the #684 criteria-"
f"binding lifecycle, and {inquiry_743_lines} are in the #743 "
f"inquiry-ledger/sidecar extension; "
f"inquiry-ledger/sidecar extension, and {authority_g1_lines} are in "
f"the 2026-09 checkpoint-authority fidelity section; "
f"v3.6.7-attributed lines = "
f"{v367_line_count} exceeds {ceiling} (baseline "
f"{BASELINE_LINE_COUNT} + Phase 6.6 budget "
+5 -1
View File
@@ -100,7 +100,7 @@
},
{
"mechanism": "cross_model_unavailable",
"failure_class": "Transport-level failure of a CONFIGURED cross-model verifier mid-run (API error, rate limit, expired key, or selected Codex citation transport unavailable). An unset ARS_CROSS_MODEL is NOT a degradation — it is the standard single-model configuration and produces no error marker or disclosure; an unset or api ARS_CROSS_MODEL_TRANSPORT is likewise ordinary API-route selection, while an invalid selector is a visible configuration error",
"failure_class": "Transport-level failure of a CONFIGURED cross-model verifier mid-run (API error, rate limit, expired key, selected Codex citation transport unavailable, or a provider-side safety / misalignment intervention that surfaces as an API error or an adapter failure). An unset ARS_CROSS_MODEL is NOT a degradation — it is the standard single-model configuration and produces no error marker or disclosure; an unset or api ARS_CROSS_MODEL_TRANSPORT is likewise ordinary API-route selection, while an invalid selector is a visible configuration error",
"degraded_state": "Warn-and-continue: pipeline proceeds single-model; NEVER blocks on cross-model transport failure; a NOT_SEARCHED grounded-lookup result is NOT a transport failure and does not fall back (it is counted separately and surfaced for re-run/human review)",
"diagnostic_marker": "[CROSS-MODEL-ERROR: reason] log line + required report disclosure 'Cross-model verification was configured but unavailable for this run. Results are single-model only.'",
"downstream_consumer": "integrity_verification_agent (Stage 2.5/4.5 sample checks), devils_advocate reviewers, blind disagreement checkpoints (#518)",
@@ -117,6 +117,10 @@
{
"file": "shared/cross_model_verification.md",
"anchor": "A `NOT_SEARCHED` result is **not** a transport failure and is handled differently."
},
{
"file": "shared/cross_model_verification.md",
"anchor": "An intervention is never a verdict."
}
],
"pinned_by": [
+42 -20
View File
@@ -41,8 +41,9 @@ A stress test of 68 AI-generated citations found 31% had problems — and all pa
| Model | API ID | Provider | Best For |
|-------|--------|----------|----------|
| Claude (session model) | _(inherited Claude Code session model — e.g., Fable 5)_ | Anthropic | Primary model (default for all ARS skills) |
| GPT-5.6 Sol | `gpt-5.6-sol` | OpenAI | Cross-verification — current OpenAI flagship, recommended OpenAI verifier; **validated for the ChatGPT-subscription citation transport** (2026-08-19/20 bakeoff, superiority on recall + latency — `audits/bakeoff-gpt-5-6-sol-codex-2026-08-19.md`); **provisional pending ARS validation** on the first-party API route (same standard rates as GPT-5.5) |
| Claude (session model) | _(inherited Claude Code session model — e.g., Fable 5.1)_ | Anthropic | Primary model (default for all ARS skills) |
| GPT-6 Astra | `gpt-6-astra` | OpenAI | Cross-verification — current OpenAI flagship (released 2026-09-03), recommended OpenAI verifier under the recommendation policy below; **provisional pending ARS validation** on both the first-party API route and the ChatGPT-subscription citation transport (no recorded bakeoff run; entry-gate smoke PASS on the citation transport 2026-09-05, codex-cli 0.153.4 — see the GPT-6 Astra note below) |
| GPT-5.6 Sol | `gpt-5.6-sol` | OpenAI | Cross-verification — previous generation, superseded by GPT-6 Astra (2026-09-03); **validated for the ChatGPT-subscription citation transport** (2026-08-19/20 bakeoff, superiority on recall + latency — `audits/bakeoff-gpt-5-6-sol-codex-2026-08-19.md`), the only id with a measured ARS run on any transport; **provisional pending ARS validation** on the first-party API route (same standard rates as GPT-5.5) |
| Gemini 3.1 Pro | `gemini-3.1-pro-preview` | Google | Cross-verification — current Google flagship (validated); strong at factual verification |
| GPT-5.5 | `gpt-5.5` | OpenAI | Cross-verification — previous generation, superseded by GPT-5.6 (2026-07-09); validated, remains fully supported (supports `xhigh` reasoning) |
| GPT-5.5 Pro | `gpt-5.5-pro` | OpenAI | Cross-verification — previous generation; validated; strongest GPT-5.5-line reasoning (premium pricing: ~6× GPT-5.5) |
@@ -57,11 +58,13 @@ A stress test of 68 AI-generated citations found 31% had problems — and all pa
> **Compatible providers are ungrounded.** They expose no hosted web-search tool, so there is no grounding evidence behind a verdict. A positive `VERIFIED` is downgraded to `NOT_SEARCHED` and never counts as agreement in citation verification; a `NOT_FOUND`/`MISMATCH` survives as a disagreement. They ARE first-class for Devil's Advocate critique (which needs no grounding) — but a DA finding from any provider is an adversarial hypothesis, not standalone evidence, unless independently sourced.
**Recommended cross-verification pair:** the inherited Claude session model (primary) + a current-generation second-family verifier — Gemini 3.1 Pro (validated) or GPT-5.6 Sol (provisional; see the note below).
**Recommended cross-verification pair:** the inherited Claude session model (primary) + a current-generation second-family verifier — Gemini 3.1 Pro (validated) or GPT-6 Astra (provisional; see the note below). Users who want a measured OpenAI id can stay on GPT-5.6 Sol for the ChatGPT-subscription citation transport (validated there) or on GPT-5.5 for the API route.
> The primary row deliberately names no version: the primary is always the session model, so the row cannot go stale on the next Anthropic release. Verifier IDs stay concrete because they are literal API strings the user must export. (`gpt-5.4` / `gpt-5.4-pro` remain accepted for existing setups.)
> **GPT-5.6 Sol is provisional (listed 2026-07-11, three days after release).** Its endpoint support (Responses API), hosted `web_search` tool, and reasoning-effort values are confirmed against OpenAI's model documentation, but its ARS-specific behavior — grounded-search completion rate, citation-mismatch recall, false-disagreement rate, response-shape stability against the jq grounding guards, p95 latency — is unvalidated. **Recommendation policy (2026-08-19):** GPT-5.5 was superseded by the GPT-5.6 family on 2026-07-09, so the recommendation names the current generation rather than a superseded id — a lifecycle decision, not a measurement claim. `validated` is earned only there — and on 2026-08-19 a codex-transport bakeoff run earned it for the **ChatGPT-subscription citation transport**, with a measured superiority case from the counterbalanced gate fleet (fabrication recall 0.90 vs 0.80, p95 latency 25.0 s vs 49.6 s nearest-rank, grounded completion tied, no inferiority on any measure; recall and latency led in all five paired fleets — `audits/bakeoff-gpt-5-6-sol-codex-2026-08-19.md`). On the **first-party API route** `gpt-5.6-sol` stays **provisional** — that run did not exercise the API route's jq grounding guards, and no parity or superiority is claimed there. For the API route, run `scripts/cross_model_smoke_test.sh` against your key before adopting it; users who prefer an API-route-validated id can stay on `gpt-5.5` or `gemini-3.1-pro-preview` (validated = the id-status allowlist below; the API route has no recorded bakeoff run). Two facts that differ from the GPT-5.5 lineup: GPT-5.6 ships **no `-pro` model ID** — premium operation is standard `gpt-5.6-sol` plus `reasoning: {mode: "pro"}` in the request, billed at standard token rates with more model work per request (the old fixed ~6× unit-price split does not carry over); and its reasoning effort accepts `none|low|medium|high|xhigh|max` (GPT-5.5 tops out at `xhigh`), defaulting to `medium` in both standard and pro modes.
> **GPT-6 Astra is provisional (listed 2026-09-05, two days after its 2026-09-03 release).** Its ARS-specific behavior on the first-party API route — the five Promotion Bakeoff measures below — is unvalidated, and its API reasoning-effort vocabulary is not confirmed in this repository (its system card reports evaluations at `xhigh` and `max`, and `ultra` in the Codex harness; the API rejects a value it does not accept, so an unknown value fails visibly). On the ChatGPT-subscription citation transport it passed the entry-gate smoke (`scripts/cross_model_smoke_test_codex.sh`, 2026-09-05, codex-cli 0.153.4: `VERIFIED` with a bound source on the Vaswani et al. fixture) — the precondition for a Promotion Bakeoff, not a bakeoff. Under the recommendation policy recorded in the GPT-5.6 Sol note below (#783) the recommendation moves to the current generation on lifecycle grounds; `validated` still requires the sealed bakeoff, on each transport separately. Two vendor-reported facts shape how ARS treats this verifier (GPT-6 Astra system card, 2026-09-03): provider-side misalignment and misuse monitoring can pause, end, or block a call (§ Provider-side monitoring and safety interventions below — never a verdict), and its verbalized evaluation awareness is high (§8.6, §8.8.1 — see the Promotion Bakeoff caveat).
> **GPT-5.6 Sol status (listed 2026-07-11, three days after release; superseded by GPT-6 Astra on 2026-09-03).** Its endpoint support (Responses API), hosted `web_search` tool, and reasoning-effort values are confirmed against OpenAI's model documentation, but its ARS-specific behavior — grounded-search completion rate, citation-mismatch recall, false-disagreement rate, response-shape stability against the jq grounding guards, p95 latency — is unvalidated. **Recommendation policy (2026-08-19):** GPT-5.5 was superseded by the GPT-5.6 family on 2026-07-09, so the recommendation names the current generation rather than a superseded id — a lifecycle decision, not a measurement claim. `validated` is earned only there — and on 2026-08-19 a codex-transport bakeoff run earned it for the **ChatGPT-subscription citation transport**, with a measured superiority case from the counterbalanced gate fleet (fabrication recall 0.90 vs 0.80, p95 latency 25.0 s vs 49.6 s nearest-rank, grounded completion tied, no inferiority on any measure; recall and latency led in all five paired fleets — `audits/bakeoff-gpt-5-6-sol-codex-2026-08-19.md`). On the **first-party API route** `gpt-5.6-sol` stays **provisional** — that run did not exercise the API route's jq grounding guards, and no parity or superiority is claimed there. For the API route, run `scripts/cross_model_smoke_test.sh` against your key before adopting it; users who prefer an API-route-validated id can stay on `gpt-5.5` or `gemini-3.1-pro-preview` (validated = the id-status allowlist below; the API route has no recorded bakeoff run). Two facts that differ from the GPT-5.5 lineup: GPT-5.6 ships **no `-pro` model ID** — premium operation is standard `gpt-5.6-sol` plus `reasoning: {mode: "pro"}` in the request, billed at standard token rates with more model work per request (the old fixed ~6× unit-price split does not carry over); and its reasoning effort accepts `none|low|medium|high|xhigh|max` (GPT-5.5 tops out at `xhigh`), defaulting to `medium` in both standard and pro modes.
Using two non-Anthropic models as primary+verifier is possible but not tested with ARS prompts.
@@ -73,7 +76,7 @@ You need API keys from at least one additional provider. ARS itself runs inside
### Step 1: Get API Keys
**OpenAI (GPT-5.6 Sol / GPT-5.5):**
**OpenAI (GPT-6 Astra / GPT-5.6 Sol / GPT-5.5):**
1. Go to [platform.openai.com/api-keys](https://platform.openai.com/api-keys)
2. Create a new API key
3. Copy the key (starts with `sk-`)
@@ -100,12 +103,16 @@ Add to your shell profile (`~/.zshrc` or `~/.bashrc`):
export OPENAI_API_KEY="<your-openai-api-key>"
# Current OpenAI flagship — provisional pending ARS validation (see Supported Models;
# run scripts/cross_model_smoke_test.sh against your key before relying on it):
export ARS_CROSS_MODEL="gpt-5.6-sol"
# Previous generation, validated (designated bakeoff baseline):
export ARS_CROSS_MODEL="gpt-6-astra"
# Previous generation validated on the ChatGPT-subscription citation transport,
# provisional on this API route:
# export ARS_CROSS_MODEL="gpt-5.6-sol"
# Previous generation, validated (designated API-route bakeoff baseline):
# export ARS_CROSS_MODEL="gpt-5.5"
# Optional: reasoning effort for OpenAI verifier calls (unset = the provider's own
# default for the chosen model). GPT-5.6 accepts none|low|medium|high|xhigh|max;
# GPT-5.5 tops out at xhigh.
# GPT-5.5 tops out at xhigh; GPT-6 Astra's API vocabulary is not confirmed here
# (the API rejects an unknown value visibly).
# export ARS_CROSS_MODEL_REASONING_EFFORT="medium"
# --- Option B: Google Gemini (first-party, grounded) ---
@@ -133,10 +140,12 @@ by detection and execution.
```bash
# Citation-integrity calls only. General DA/reviewer/judgment calls remain on API transport.
export ARS_CROSS_MODEL_TRANSPORT="codex"
# gpt-6-astra: current OpenAI flagship — provisional on this transport (entry-gate
# smoke PASS 2026-09-05 on codex-cli 0.153.4; no bakeoff run yet).
export ARS_CROSS_MODEL="gpt-6-astra"
# gpt-5.6-sol is validated for THIS transport (2026-08-19 codex-transport bakeoff,
# superiority on recall + latency — audits/bakeoff-gpt-5-6-sol-codex-2026-08-19.md).
# gpt-5.5 remains the validated bakeoff baseline alternative.
export ARS_CROSS_MODEL="gpt-5.6-sol"
# superiority on recall + latency — audits/bakeoff-gpt-5-6-sol-codex-2026-08-19.md):
# export ARS_CROSS_MODEL="gpt-5.6-sol"
python3 scripts/cross_model_codex_transport.py detect
# The producer sends one closed codex_citation_request/1.0 object on stdin:
@@ -181,7 +190,7 @@ If you don't want cross-model verification running all the time, you can enable
```bash
# Enable for this session only
export ARS_CROSS_MODEL="gpt-5.6-sol"
export ARS_CROSS_MODEL="gpt-6-astra"
# Disable for this session
unset ARS_CROSS_MODEL
@@ -404,9 +413,9 @@ contract is normative in
machine-checked by the #630 test suite. The Bash entrypoints use syntax compatible
with macOS Bash 3.2.
### OpenAI (GPT-5.6 Sol / GPT-5.5 / GPT-5.5 Pro)
### OpenAI (GPT-6 Astra / GPT-5.6 Sol / GPT-5.5 / GPT-5.5 Pro)
Use the **Responses API** (`/v1/responses`) — the hosted `web_search` tool lives there. (Chat Completions does not take `tools: [{type: "web_search"}]`; web search on that endpoint requires the separate `gpt-5-search-api` model, so this example targets Responses to stay model-agnostic across `gpt-5.5` / `gpt-5.5-pro` / `gpt-5.6-sol` / the legacy `gpt-5.4*` ids.)
Use the **Responses API** (`/v1/responses`) — the hosted `web_search` tool lives there. (Chat Completions does not take `tools: [{type: "web_search"}]`; web search on that endpoint requires the separate `gpt-5-search-api` model, so this example targets Responses to stay model-agnostic across `gpt-6-astra` / `gpt-5.6-sol` / `gpt-5.5` / `gpt-5.5-pro` / the legacy `gpt-5.4*` ids.)
```bash
# PROMPT holds the single-reference verification prompt (step 3). One reference per call.
@@ -486,7 +495,7 @@ fi
> **Why `temperature: 0.1`:** reference existence/metadata checking is a deterministic factual task, so low temperature reduces run-to-run variance in the verdict. It is not a grounding control — the grounding guard above is what enforces an actual lookup.
> **Reasoning effort (OpenAI only):** when `ARS_CROSS_MODEL_REASONING_EFFORT` is set, the payload passes it as `reasoning.effort`, making the effort a verification run uses visible and reproducible. When it is **unset, the field is omitted entirely and the provider's own default for the chosen model applies** — defaults differ across the lineup (GPT-5.6 documents `medium`; other ids carry their own), so forcing one value here would silently change behavior for existing setups. Citation lookup is search-bound, not reasoning-bound, so higher efforts mostly buy latency and cost; set the variable deliberately (never silently run at `xhigh`) if a run shows shallow search behavior. The value is passed through unvalidated (the API rejects unknown values): GPT-5.5 accepts up to `xhigh`, GPT-5.6 adds `max`.
> **Reasoning effort (OpenAI only):** when `ARS_CROSS_MODEL_REASONING_EFFORT` is set, the payload passes it as `reasoning.effort`, making the effort a verification run uses visible and reproducible. When it is **unset, the field is omitted entirely and the provider's own default for the chosen model applies** — defaults differ across the lineup (GPT-5.6 documents `medium`; other ids carry their own), so forcing one value here would silently change behavior for existing setups. Citation lookup is search-bound, not reasoning-bound, so higher efforts mostly buy latency and cost; set the variable deliberately (never silently run at `xhigh`) if a run shows shallow search behavior. The value is passed through unvalidated (the API rejects unknown values): GPT-5.5 accepts up to `xhigh`, GPT-5.6 adds `max`. GPT-6 Astra's accepted API values are not confirmed in this repository (its system card reports `xhigh`, `max`, and — in the Codex harness — `ultra`); the contained Codex citation transport forwards only the closed set named by `ACCEPTED_REASONING_EFFORTS` in `scripts/cross_model_codex_transport.py` (an earlier, better-named error, not a safety property) and lets the provider reject whatever the served model does not advertise.
### OpenAI-Compatible API (MiMo, DeepSeek, self-hosted) — ungrounded
@@ -596,7 +605,9 @@ if [ -n "$ARS_CROSS_MODEL" ]; then
# gpt-5.6-sol: validated for the codex subscription citation transport
# (2026-08-19 bakeoff); provisional HERE because this allowlist gates the
# first-party API route, which has no recorded bakeoff run.
case " gpt-5.6-sol " in
# gpt-6-astra: listed 2026-09-05; provisional on every transport (entry-gate
# smoke only, no bakeoff run).
case " gpt-5.6-sol gpt-6-astra " in
*" $1 "*) echo "provisional"; return ;;
esac
echo "unlisted"
@@ -633,7 +644,7 @@ if [ -n "$ARS_CROSS_MODEL" ]; then
echo "WARNING: ARS_OPENAI_COMPAT_BASE_URL is set but ARS_OPENAI_COMPAT_API_KEY is not — refusing to send another provider's key. Set ARS_OPENAI_COMPAT_API_KEY."
echo "CROSS_MODEL_AVAILABLE=none"
else
echo "WARNING: ARS_CROSS_MODEL=$ARS_CROSS_MODEL is not a recognized model. First-party grounded route: any gpt-* id (e.g. gpt-5.5, gpt-5.5-pro, gpt-5.6-sol, legacy gpt-5.4*) or gemini-* id (e.g. gemini-3.1-pro-preview). For an OpenAI-compatible provider set ARS_OPENAI_COMPAT_BASE_URL + ARS_OPENAI_COMPAT_API_KEY and use that provider's model id (must not match a gpt-*/gemini-* prefix, or it takes the grounded first-party route instead)."
echo "WARNING: ARS_CROSS_MODEL=$ARS_CROSS_MODEL is not a recognized model. First-party grounded route: any gpt-* id (e.g. gpt-6-astra, gpt-5.6-sol, gpt-5.5, gpt-5.5-pro, legacy gpt-5.4*) or gemini-* id (e.g. gemini-3.1-pro-preview). For an OpenAI-compatible provider set ARS_OPENAI_COMPAT_BASE_URL + ARS_OPENAI_COMPAT_API_KEY and use that provider's model id (must not match a gpt-*/gemini-* prefix, or it takes the grounded first-party route instead)."
echo "CROSS_MODEL_AVAILABLE=none"
fi ;;
esac
@@ -657,7 +668,7 @@ an API route.
### Promotion Bakeoff (provisional → validated)
The run that flips a provisional id (today: `gpt-5.6-sol`) to validated is defined here so a future promotion argues against numbers, not vibes (#518). Validation and recommendation are separate axes. (2026-08-19, #783: the recommendation moved to the current generation on lifecycle grounds — GPT-5.5 was superseded — ahead of validation; that flip carries no measurement claim. This bakeoff remains the only route to `validated`, and any claim of measured parity or superiority still requires the run below.)
The run that flips a provisional id (today: `gpt-6-astra` on both transports, and `gpt-5.6-sol` on the first-party API route) to validated is defined here so a future promotion argues against numbers, not vibes (#518). Validation and recommendation are separate axes. (2026-08-19, #783: the recommendation moved to the current generation on lifecycle grounds — GPT-5.5 was superseded — ahead of validation; that flip carries no measurement claim. This bakeoff remains the only route to `validated`, and any claim of measured parity or superiority still requires the run below.)
> **Recorded run (2026-08-19/20, #787 — codex-transport variant).** The procedure below was executed over the #630 ChatGPT-subscription citation transport (entry gate: `scripts/cross_model_smoke_test_codex.sh` PASS for baseline and candidate; measure analogues: grounding evidence = receipt `searched`, measure 4 = zero fail-closed receipt-guard misfires). All five measures passed in the counterbalanced gate fleet, with superiority on measures 2 (fabrication recall) and 5 (latency) and a tie on measure 1 — see `audits/bakeoff-gpt-5-6-sol-codex-2026-08-19.md` (probe set `evals/bakeoff/2026-08-19-gpt-5-6-sol-codex/`, sha256 in the report). The result is **transport-qualified**: `gpt-5.6-sol` is validated for the subscription citation transport; it remains provisional on the first-party API route, whose jq grounding guards that run did not exercise. A scored fleet is bound to its preregistered frozen instrument; later instrument hardening that validates only surfaces outside every consumed path applies from the next fleet and does not retroactively invalidate a recorded gate result (boundary rationale in the run report's Instrument-freeze decision record). An API-route run requires a FRESH probe set under the #789 sealed-preregistration protocol below — the 2026-08-19 set's labels are public, so reusing it would expose a live-search run to answer-key retrieval.
@@ -670,7 +681,7 @@ The run that flips a provisional id (today: `gpt-5.6-sol`) to validated is defin
5. **Never reuse a published answer key.** Once labels appear in any Git version, those exact probe bytes are retired permanently and every later gate gets a fresh fabrication pool. The verifier scans every historical version of every `evals/bakeoff/**/probe_set.json`; a fabricated reference remains reused even if its id, context, case, Unicode width, spacing, or punctuation changes. Previously used real references may remain, but no previously labeled reference may enter the new fabricated pool. The 2026-08-19 fixture is the sole explicitly grandfathered unsealed artifact: its canonical path and LF-normalized SHA-256 are pinned, its blob bytes and regular-file mode must remain immutable across reachable history, and it remains part of the published-history scan. Verification requires a complete non-shallow local history and fails closed when a referenced historical object cannot be read.
The reveal verifier proves byte binding, composition, Git immutability/order, and detectable historical non-reuse. It cannot prove when a commit became visible on a remote or when an external call ran. The run report therefore MUST record the public commitment permalink and CI result, successful preflight output, fleet start/end bounds, final verifier receipt, and—for every listed reveal-copy commit in `reveal_copy_git_commits`—a public permalink plus evidence that it first became publicly reachable only after the fleet ended. A copy published before or during the fleet invalidates the gate even when a later squash commit is clean. Missing remote/timing evidence makes the fleet exploratory, not a gate result.
- **Procedure:** run the baseline (`gpt-5.5`) and the candidate the same day, one call per reference, 3 repeats. Per-reference verdict = the verdict returned by ≥ 2 of 3 repeats; if no verdict reaches 2 (a 111 split), the reference is **indeterminate** and scored conservatively against the model that produced it — a miss for recall (measure 2), a false disagreement for measure 3. Grounded-search completion (measure 1) is computed per call, so ties don't apply.
- **Procedure:** run the transport's validated baseline (`gpt-5.5` on the first-party API route; `gpt-5.6-sol` on the ChatGPT-subscription citation transport) and the candidate the same day, one call per reference, 3 repeats. Per-reference verdict = the verdict returned by ≥ 2 of 3 repeats; if no verdict reaches 2 (a 111 split), the reference is **indeterminate** and scored conservatively against the model that produced it — a miss for recall (measure 2), a false disagreement for measure 3. Grounded-search completion (measure 1) is computed per call, so ties don't apply.
- **Non-inferiority thresholds — all five must pass:**
1. **Grounded-search completion rate** (share of calls returning grounding evidence) ≥ baseline 5 pp.
2. **Citation-mismatch recall** on the 10 fabrications (share flagged `NOT_FOUND`/`MISMATCH`) ≥ baseline 5 pp AND ≥ 80% absolute.
@@ -684,6 +695,8 @@ The run that flips a provisional id (today: `gpt-5.6-sol`) to validated is defin
Web-search results vary day to day; the 3-repeat majority verdict and same-day paired runs are what make the comparison fair. Thresholds are the #518 spec's choice and are tunable in a future spec without redesigning the procedure.
A further caveat that no threshold absorbs: a candidate with high verbalized evaluation awareness (GPT-6 Astra's vendor reports it in 9.6% of simulated Codex trajectories, and 4151% of samples at `xhigh`/`max` effort in one external evaluation — system card §8.6, §8.8.1) may behave differently on a 30-reference probe fleet than in a live pipeline. The sealed preregistration prevents answer-key retrieval, not evaluation awareness; a promotion PR for such a model records this as a known limitation of the measurement.
## Cost Considerations
Cross-model verification adds API costs from the second provider:
@@ -695,7 +708,7 @@ Cross-model verification adds API costs from the second provider:
| Blind disagreement checkpoints (design freeze + final editorial decision, 1 structured-decision call each; editorial repeats on re-review) | 2-3 calls | ~$0.20-0.55 |
| **Full pipeline (the worked example)** | **~22-23 calls** | **~$1.85-4.05 — no fixed ceiling; grows with the high-impact / new-changed count** |
These are rough estimates based on GPT-5.5 pricing ($5/1M input, $30/1M output) and typical prompt sizes; GPT-5.5 Pro runs ~6× higher ($30/1M input, $180/1M output). GPT-5.6 Sol bills at the same standard rates as GPT-5.5 ($5/1M input, $0.50/1M cached input, $30/1M output); its pro mode keeps those rates but performs more model work per request, so total tokens (and latency) rise instead of the unit price. One-call-per-reference (rather than batching) is a deliberate cost-for-provenance trade: it is the only way the grounding-evidence check maps 1:1 to each verdict. Web-search-tool calls also cost more than plain completions.
These are rough estimates based on GPT-5.5 pricing ($5/1M input, $30/1M output) and typical prompt sizes; GPT-5.5 Pro runs ~6× higher ($30/1M input, $180/1M output). GPT-5.6 Sol bills at the same standard rates as GPT-5.5 ($5/1M input, $0.50/1M cached input, $30/1M output); its pro mode keeps those rates but performs more model work per request, so total tokens (and latency) rise instead of the unit price. GPT-6 Astra's list pricing is not recorded in this document; re-derive the table from the provider's price list before budgeting a run on it. One-call-per-reference (rather than batching) is a deliberate cost-for-provenance trade: it is the only way the grounding-evidence check maps 1:1 to each verdict. Web-search-tool calls also cost more than plain completions.
## Limitations
@@ -712,3 +725,12 @@ If cross-model verification fails **at the transport level** (API error, rate li
- Include a note in the report: "Cross-model verification was configured but unavailable for this run. Results are single-model only."
A `NOT_SEARCHED` result is **not** a transport failure and is handled differently. It means the call succeeded but the model could not (or did not) ground the lookup, so its verdict carries no evidence. Do not fall back to single-model and do not treat it as agreement: record the reference as `NOT_SEARCHED` in the results table, count it separately from agreements/disagreements, and surface it for re-run or human review. The distinction matters — a transport failure means "we have no cross-model opinion"; a `NOT_SEARCHED` means "the cross-model gave an opinion we have decided not to trust as a confirmation."
### Provider-side monitoring and safety interventions (2026-09)
GPT-6 Astra ships with provider-side monitoring that acts on live calls (GPT-6 Astra system card, 2026-09-03): a misalignment monitor over Codex, ChatGPT, and the Responses API that can pause or end a conversation, with stopped API conversations not resumable (§10.2.3.1); misuse monitors and activation classifiers that can block a generation mid-stream (§10.2.3.2); a stricter cyber refusal boundary for higher-risk accounts (§10.2.2.2); and account-level escalation to automated and, in some cases, manual review (§10.2.5). None of this changes the contract above; it adds concrete cases to it:
- **An intervention is never a verdict.** A paused, ended, blocked, or refused call surfaces on the API route either as an HTTP error — the existing transport-failure marker (`CROSS-MODEL-ERROR: openai_http_<status>`); the run continues single-model with the standard disclosure — or as a completed response with no grounding evidence, which the existing `NOT_SEARCHED` guard already catches (counted separately, surfaced for re-run or human review). On the contained codex adapter it is the adapter's nonzero exit or fail-closed receipt. None of these is ever a citation judgment, a reviewer finding, or a checkpoint decision. Because every ARS call is one stateless request per reference or judgment, an ended conversation loses no ARS state: re-run the affected item.
- **Content can trigger the intervention.** A manuscript on a security topic, or a review payload dense with attack vocabulary, can trip the cyber boundary even when the task is ordinary scholarship. Do not rephrase the *manuscript* to route around a provider's safety boundary; that is the provider's call to make. For ARS's own prompt wording, prefer process vocabulary (consistency, drift, witness, discrepancy) over attack vocabulary where the meaning is the same, and treat a content-triggered refusal as `unavailable` for that item with the reason recorded.
- **Escalation is part of the consent boundary.** Because flagged conversations can reach human review at the provider, the consent obtained under the Consent boundary above must assume that provider staff may read what is sent. This is the residual gap recorded as R4 in `docs/RISK_REGISTER.md`.
- **Why ARS never consumes the verifier's reasoning.** The same vendor reports reduced chain-of-thought monitorability for this model and euphemistic framing of workarounds in its reasoning (§9). ARS binds every verifier result to grounding evidence, a bound source receipt, a typed anchor, or a closed enum — never to the model's narrative. That is a design rationale for the guards in this document, not a runtime claim about the verifier.
+2
View File
@@ -27,6 +27,8 @@ The no-hard-pinning rule is about what lives in the repo, not about the dispatch
3. Pass whatever identifier the runtime accepts for that target (alias preferred where supported; otherwise the current generation's concrete id). The concrete value exists only in that ephemeral call — it is never written into agent files, manifests, or this doc.
4. If the session cannot resolve the target (unknown lineup, runtime exposes no model choice): the direction is a no-op for that call — announce `[MODEL-TIERING: could not resolve target tier — ran on the session model]` once per run. Fail-open, never a guessed id.
The resolved tier names the **declared** session model, not a per-call attestation of what served the request: the runtime may serve a classifier-flagged request on a different model of the same family, with no signal ARS reads (vendor specifics in `audits/harness-retirement-2026-09-model-update.md` G-3). Tiering decisions, provenance blocks, and cost estimates therefore describe the declared model; a run whose content trips those classifiers — security-topic manuscripts are the likely case — may have been served on another tier without notice. This is a recorded residual gap (`docs/RISK_REGISTER.md` R5), not something the switch can detect or correct.
## Direction 1 — `quality-boost` (for sessions below the frontier tier)
- **Who:** judgment-type agents (table below) **when dispatched at a checkpoint surface**: the Stage 2.5 / 4.5 integrity gates (`integrity_verification`, `compliance_agent`); the Stage 4→5 claimref alignment audit (`claim_ref_alignment_audit` — dispatched only when `ARS_CLAIM_AUDIT=1`, so this surface exists only on opted-in runs); and the final-review surfaces (Stage 3 full panel: `eic`, the three reviewers, `devils_advocate_reviewer`, `editorial_synthesizer`). Stage 3' uses three dedicated contract judgment calls plus any scoped Phase 2B verification calls; when `quality-boost` applies, the orchestrating layer dispatches those checkpoint calls at the frontier tier directly. They are protocol calls, not agent-manifest identities; `field_analyst` remains execution-type and is unaffected except for the visibly marked card-regeneration fallback.