feat: add hermetic tortured-phrase screening (#699)

Refs #660. Measurement and issue closure follow in the separately preregistered post-main mechanical conformance PR.
This commit is contained in:
Edward Cheng-I Wu
2026-08-10 12:28:42 +08:00
committed by GitHub
parent f3cfdb4936
commit 86bf0e5c2c
42 changed files with 15560 additions and 26 deletions
+9
View File
@@ -409,6 +409,15 @@ jobs:
# the unified pytest manifest above.
run: python3 scripts/check_bibliographic_integrity_signals.py
- name: Check tortured-phrase screening integration (#660)
# Locks the closed snapshot/manifest/advisory schemas, v1.2 carrier,
# exact-byte synthetic fixture, hermetic runtime, one-page renderer,
# UNMEASURED claim ceiling, documentation, and CI bindings. Behavioral
# and mutation tests run through the unified pytest manifest above.
env:
PYTHONPATH: .
run: python3 scripts/check_tortured_phrase_screening_integration.py
- name: Run v3.9.2 Phase Boundary coverage lint (#133)
# Enforces 22 Bucket A agents have ## Phase Boundary (v3.9.2)
# block, 16 Bucket B/C/D agents DON'T, and each Bucket A block
+12
View File
@@ -48,3 +48,15 @@ paths = [
description = "Frozen E4 evidence-bundle model prose misread as generic-api-key (venue descriptors, not credentials)"
targetRules = ["generic-api-key"]
paths = ['''evals/heldout/reviewer_seeded_defects/runs/raw/.*\.md''']
# Class 3 (rule-targeted): #660's repository-owned synthetic conformance
# fixtures use citation_key values such as fixture_missing_2026. These are
# deterministic literature-record identifiers, not credentials. All other
# secret rules remain active over the same fixture bytes.
[[allowlists]]
description = "Synthetic tortured-phrase fixture citation keys misread as generic-api-key"
targetRules = ["generic-api-key"]
paths = [
'''scripts/fixtures/tortured_phrase_screening/.*''',
'''scripts/fixtures/bibliographic_integrity_signals/tortured_phrase_v1_2_.*\.json''',
]
+2
View File
@@ -6,6 +6,8 @@ All notable changes to this project will be documented in this file.
### Added
- **Hermetic tortured-phrase screening for own drafts and cited metadata (#660).** Three closed contracts and a hermetic runtime accept only an explicit user-supplied or clearly synthetic-fixture snapshot with a detached manifest bound to the exact raw snapshot SHA-256 and an explicit rights declaration. The repository ships no native PPS content, PPS importer, network fetcher, or redistributed phrase list, and the checker uses no live model, external API, human or model judge, ambient clock, source-file timestamp, or Git/network time. Own-draft output is a bounded `HEURISTIC-ADVISORY` / `UNMEASURED` artifact rendered immediately before formatting; cited-metadata enrichment is non-in-place and emits one current v1.2 row for each title and abstract surface, including explicit `ABSTRACT_MISSING` and `ABSTRACT_EMPTY` states. Read-only consumers compose all rows in the one existing `Bibliographic Integrity Advisories` section; the advisory mints no marker, changes no gate or finalizer policy, proposes no rewrite, and makes no clean-draft, AI/author/papermill/origin, contextual-validity, publisher-acceptance, or matcher-accuracy claim. Shipped fixtures are synthetic conformance inputs only, not an accuracy evaluation.
- **Non-ranking, explicit-author revision authority (#670).** Schema 7 now has a closed immutable `revision-roadmap/1.0` core that keeps transported reviewer severity, editorial `obligation_class`, typed cost surface, bounded consequence, deterministic source order, and exact proposed block/operation targets independent. Explicit session-author choices live in a separately raw-hash-bound `author-adjudication/1.0` sidecar, with complete triage, decline reasons, exact target subsets, exact registered-claim replacement bytes/rungs, and exact declined-overlap collateral authority; presentation views never alter `R<n>` transport references or gate arithmetic. Claim surfaces are exact UTF-8 spans whose protected text must equal the referenced ClaimIntent `claim_text`, preventing a valid claim id from being paired with unrelated prose. Integrity issue lists and gate failures are proposal-only: a distinct explicit-author input must approve both exact targets/operations and the complete patch SHA-256 before the deterministic builder can emit a base/list/round-bound integrity authorization sidecar. Current patch 1.1 and apply-report 1.3 replay the appropriate disjoint authority branch before any write, reject patch 1.0 on the current CLI, protect registered claim surfaces exactly, and disclose the remaining unregistered-semantic-drift E6 boundary. A closed Revision-Evidence Bundle carries an exact integrity-PASS start through continuous review-write, all-declined no-op, and explicitly authorized integrity-correction rounds to the final draft, with contained read-once artifacts, deterministic patch replay, and byte-exact post-draft/report verification. The #576 current family moves in lockstep to contract 1.1, `obligation_class`/renamed residual and rate fields, exact Schema 11 author-field copies, and five hard-required artifacts: original manuscript, revised manuscript, roadmap, author adjudication, and Revision-Evidence Bundle. Missing-original, `first_link_not_run`, and old-report patch-binding degradations remain archived 1.0 behavior; current/legacy schema IDs and runtimes are isolated. All acceptance is hermetic—no live model, external API, judge, or scored evaluation.
- **Replay-bound authority-profile content-coverage advisories (#681).** A separately versioned `evidence-row/1.1` surface and closed `content-coverage-advisory/1.0` carrier now bind bounded passages from explicit session-held strings to exact #666 authority requirements and `structured_expectations[]` pointers after full #667 packet-artifact replay. The `LLM-ADVISORY` layer preserves every deterministic packet status, readiness value, caller-supplied authorization field, institutional-acceptance boundary, pointer, and digest; structural gaps, external dependencies, waiver/exception boundaries, applicability-false exclusions, and missing session content cannot be laundered into semantic missing-element findings. Every profiled expectation is explicitly checked or explicitly `not_checked`, quotes retain #656 strict decoding, UTF-8 span/hash, 25-word/1,000-character, rights, inert-rendering, and human-read boundaries, and the finalizer opens only named inputs without directory scans, retrieval, network, subprocess, cache, or model calls. The surface is honestly `UNMEASURED`—no scored held-out row or efficacy claim is fabricated.
+4
View File
@@ -414,6 +414,10 @@ Existing `CONTAMINATED-*` tokens remain legacy-carrier outputs during dual write
Render `finding: unresolved` or status `not_checked`, `unknown`, or `degraded` as **NOT CLEAN — UNRESOLVED**; never call those states passed, absent, or clean.
`terminal_policy.eligible: false` cannot trigger refusal or `HIGH-BLOCK`. A v1.1 eligible retraction row is still only an input to the finalizer; this formatter never evaluates it. Migration/deprecation authority is `shared/bibliographic_integrity_signals.md`.
For a v1.2 `tortured_phrase_match` row, transcribe `HEURISTIC-ADVISORY`, `UNMEASURED`, its exact `cited_title` or `cited_abstract` surface, counts, snapshot source/version/as-of/raw SHA-256, detached-manifest SHA-256, and any reason code. The category label is the neutral **Phrase-list screening advisory**; use **phrase-list match requiring review** only when `finding: detected`. A checked zero means only that no match was observed on that checked surface; a missing abstract stays `not_checked` / `unresolved` with `ABSTRACT_MISSING`, and a whitespace-only abstract uses `ABSTRACT_EMPTY`. These rows remain `display.marker_token: null` and `terminal_policy.eligible: false`: they never refuse conversion, alter a gate, mint or rewrite a citation marker, suggest replacement prose, or support an AI-, author-, papermill-, misconduct-, origin-, contextual-validity-, false-positive/false-negative-, precision-, recall-, coverage-, publisher-acceptance-, or clean-draft claim.
For an own-draft `tortured-phrase-advisory/1.0` artifact, accept only the already validated artifact produced over the exact conversion input and render its single bounded page of at most 25 match-detail rows. Do not re-run matching, infer context, alter counts, traverse an unbounded detail view, or edit the manuscript. Its `HEURISTIC-ADVISORY` / `UNMEASURED` result is advisory only and never a formatter refusal condition.
## Cite-Time Terminal Policy Gate (v3.10) — STAMP-ONLY freshness + rule 11
Per spec §3 PR-B item 10 (R1 P0-C + R2-P0). The finalizer is the SOLE policy evaluator; the formatter is a **dumb stamp-checking gate** — it MUST NOT re-evaluate `strict_articles_only` DOI/venue/provenance logic (that would duplicate the finalizer and invite drift, Invariant 13). It only (1) recomputes the passport's current `terminal_policies` slug and compares stamps, and (2) refuses on `severity=HIGH-BLOCK` tokens.
+6
View File
@@ -347,6 +347,12 @@ Stage 2.5 (pre-review) and Stage 4.5 (post-revision) verification. 5-phase proto
- [v3.4.0] `compliance_agent` runs mode-aware PRISMA-trAIce + RAISE compliance check; tier-based block semantics. See `shared/compliance_checkpoint_protocol.md`.
### Tortured-phrase advisory (#660)
After the exact Stage 4.5 pass and immediately before Stage 5 formatting, the orchestrator runs the deterministic #660 checker over the exact accepted working draft using an explicit user-supplied or synthetic-fixture snapshot and detached manifest bound to the raw snapshot SHA-256; omitted supply produces an explicit `not_checked` artifact. The path ships no native PPS content/importer/fetcher or redistributed phrase list and uses no live model, external API, human or model judge, or ambient clock; timestamps are explicit inputs. Its own-draft result is `HEURISTIC-ADVISORY` / `UNMEASURED`, never changes the Stage 4.5 PASS or Stage 5 gate, never rewrites prose, and must be re-run only after a revision has re-entered the existing integrity/screen sequence.
For the literature corpus, a non-in-place producer emits one current v1.2 advisory row per `cited_title` and `cited_abstract`; a missing abstract remains explicitly `not_checked` / `unresolved` with `ABSTRACT_MISSING`. Downstream consumers are read-only and compose every row into the one existing `Bibliographic Integrity Advisories` section. The advisory mints no marker, triggers no terminal policy, gate, finalizer promotion, ranking, citation rewrite, or replacement text, and supports no clean-draft, origin, papermill, contextual-validity, publisher-acceptance, or matcher-accuracy claim.
---
## Two-Stage Review Protocol
@@ -737,6 +737,61 @@ Mid-Entry Material Passport Check:
---
## Tortured-Phrase Advisory Dispatch (#660)
After Stage 4.5 passes, and immediately before Stage 5 converts the exact
accepted working draft, dispatch `scripts/tortured_phrase_screening.py` on that
draft. The input draft and output advisory paths must differ. Supply only an
explicitly named local canonical snapshot and detached manifest; the manifest
must bind the exact raw snapshot bytes by SHA-256 and declare
`user_supplied` or `synthetic_fixture`. If the pair is intentionally absent,
retain the runtime's explicit `not_checked` result rather than skipping the
record or calling the draft clean. A partial, invalid, mismatched, unsupported,
or nonzero-unsupported-rule pair is degraded/unresolved and never becomes a
zero-match result.
A scan may atomically write a schema-valid degraded advisory and then exit 1.
Preserve and validate that exact artifact, surface its reason, and never delete it,
skip its handoff, or reinterpret the nonzero status as a new Stage 4.5 or terminal
gate.
Pass `checked_at` and `recorded_at` as explicit RFC 3339 inputs. This path must
not read the system clock, file times, Git time, timezone, or network time. It
has no native PPS parser/importer, URL fetch path, or redistributed PPS list
content and invokes no model, external API, human/model judge, contextual
classifier, or expensive evaluation. Do not attempt to manufacture or repair a
snapshot from remembered phrases or fetched content.
Validate the complete `tortured-phrase-advisory/1.0` before handoff to the
formatter. It is always `layer: HEURISTIC-ADVISORY` and
`evaluation_status: UNMEASURED`. The fixed positive meaning is **phrase-list
match requiring review**; a zero match means only no match was observed on the
exact checked bytes and is not a clean certification. It never establishes
AI/author origin, paper-mill production, misconduct, contextual validity,
precision/recall, false-positive/false-negative rate, list coverage, or
publisher acceptance.
This advisory does not alter Stage 4.5 PASS, the mandatory Stage-5 checkpoint,
or any terminal-policy state. It never edits the draft or suggests replacement
text. Surface the validated report and let the user preserve, revise, or
proceed. A revision changes the input bytes, invalidates the current advisory,
and must return through the existing integrity/final-screen sequence; the
checker itself never rewrites. The formatter only renders the already validated
artifact and does not rerun matching or change counts.
Carry schema-valid `bibliographic-integrity-signal/1.2` cited-source rows
forward unchanged. They are independently bound to `cited_title` and
`cited_abstract`; a missing abstract stays `not_checked` / `unresolved` with
`ABSTRACT_MISSING`. These rows compose lexically in the one existing
`Bibliographic Integrity Advisories` section. They have
`display.marker_token: null` and `terminal_policy.eligible: false`; the
Cite-Time Provenance Finalizer must not promote them, and neither finalizer nor
formatter may create a marker, gate, terminal token, rewrite, or replacement
from them. Corpus enrichment is producer-owned, returns a new passport copy,
and never authorizes this read-only orchestrator to mutate a passport in place.
---
## Cite-Time Provenance Finalizer (v3.7.1)
When `academic-pipeline` mode is active, the orchestrator runs the **Cite-Time Provenance Finalizer** at every Stage 4 → Stage 5 transition (and on every revision loop pass back through Stage 4) to resolve the two-layer citation markers emitted by `synthesis_agent`, `draft_writer_agent`, and `report_compiler_agent` per Step 3a.
@@ -44,6 +44,34 @@ Execution steps:
5. ⚠️ **IRON RULE**: Must PASS with zero issues to proceed to Stage 5
```
## Tortured-Phrase Advisory Boundary (#660)
Tortured-phrase screening is not a sixth integrity phase and never changes a
Stage 2.5 or Stage 4.5 PASS/FAIL decision. After the exact final draft passes
Stage 4.5, the orchestrator builds the separate own-draft
`tortured-phrase-advisory/1.0` immediately before formatting. It remains
`HEURISTIC-ADVISORY` / `UNMEASURED`; a match requires review but establishes no
AI/author origin, paper-mill production, misconduct, cleanliness, contextual
false-positive/false-negative status, accuracy, or publisher acceptance. A
zero-match result states only that no list match was observed on the checked
bytes and is not a clean certification.
The checker consumes only the explicitly named local draft and, when supplied,
canonical snapshot and detached-manifest paths. The manifest's
`snapshot_sha256` binds the exact raw snapshot bytes and declares
`user_supplied` or `synthetic_fixture` supply; omitted supply produces an
explicit `not_checked` artifact. The #660 path
has no native PPS import/fetch or redistributed PPS content and invokes no
model, external API, human/model judge, ambient clock, file time, or network
time; timestamps are explicit inputs. It never edits the draft. A user-chosen
revision changes the checked bytes and must re-enter the existing integrity
and screening sequence rather than being auto-rewritten by the advisory.
Cited-source v1.2 rows remain separate per title and abstract. A missing
abstract is explicit `not_checked` / `unresolved` (`ABSTRACT_MISSING`). They
render only in the single `Bibliographic Integrity Advisories` section and
never mint a marker, trigger a gate, or supply replacement text.
## Score Trajectory Tracking (v3.3)
Reference: `academic-pipeline/references/score_trajectory_protocol.md`
@@ -106,6 +106,39 @@ Consumer agents never modify, backfill, or derive new content into `literature_c
Consumer agents do NOT re-validate schema, do NOT parse JSON Schema at runtime, and do NOT dereference `source_pointer` URIs. The v3.6.4 input-port lint validates adapter output, but a passport may reach a Phase 1 agent through other paths (hand-edits, `resume_from_passport`, assembled passports). When a consumer cannot parse `literature_corpus[]`, emit `[CORPUS PARSE FAILURE: <cause>]` in the Search Strategy Report and fall back to external-DB-only flow. Do not abort Phase 1, do not attempt schema repair, do not invent contents.
## Tortured-phrase advisory rows (#660)
`bibliographic-integrity-signal/1.2` tortured-phrase rows are read-only
advisories, not pre-screening criteria. A consumer may preserve and surface the
row, but must not use it to include, exclude, rank, downgrade, relabel, or
rewrite a corpus entry. In particular, a phrase-list match is not evidence of
AI/author origin, paper-mill production, misconduct, source quality, or
contextual invalidity. A checked zero-match is not a clean certificate and is
not evidence that screening coverage is complete.
The two cited surfaces remain independent: `cited_title` and
`cited_abstract` each have one current row when the check is invoked. A missing
abstract remains visible as `not_checked` / `unresolved` with
`ABSTRACT_MISSING`; a consumer must not copy the title result into the abstract
slot or silently drop the row. Manual entries have no exemption from this
local metadata check.
Only the dedicated producer may call the non-in-place enricher. It consumes an
explicitly named local canonical snapshot plus detached manifest whose SHA-256
binds the exact snapshot bytes; accepted supply is `user_supplied` or
`synthetic_fixture`. Phase 1 consumers do not build, fetch, import, repair, or
redistribute a PPS list, do not dereference `source_pointer`, and do not run a
model, external API, human/model judge, or ambient clock for this advisory.
They never modify `literature_corpus[]`; a producer returns a separate passport
copy and takes explicit timestamps.
Every v1.2 row stays `HEURISTIC-INDICATOR` with
`HEURISTIC-ADVISORY` / `UNMEASURED` context. It composes only in the existing
single `Bibliographic Integrity Advisories` section. It creates no ref-marker
token, terminal policy, gate, replacement text, or automatic rewrite. The
separate own-draft advisory follows the same claim boundary and is not a corpus
screening input.
## Zero-hit and provenance reporting (F3 / F4)
Two reproducibility surfaces sit inside the PRE-SCREENED block. Every consumer agent must emit each one when the corresponding trigger fires; both are non-blocking and independent of which Step 2 case dispatches next.
@@ -406,6 +406,16 @@ Use `RetractionStatusCache`'s separate `retraction_status_cache_v1` namespace.
A row over 30 days old requires live revalidation before it can be strict
eligible. Browser fallback must not be used to evade API limits.
## Tortured-Phrase Advisory Production (#660)
This #660 path is a deterministic local metadata enricher, not a search or source-retrieval path. Given an exact local `literature_corpus[]` passport plus an explicit snapshot/manifest pair, invoke `scripts/tortured_phrase_screening.py enrich-passport` with distinct input and output paths. The command writes a separate passport, preserves every other signal and legacy row, and leaves consumers read-only; never update a passport in place. Consumers validate and render the resulting rows but do not re-run the matcher.
Scan only each entry's literal local `title` and optional `abstract` strings. Do not dereference `source_pointer`. Emit exactly one current v1.2 `tortured_phrase_match` row for `cited_title` and one for `cited_abstract`; an absent abstract is an explicit `not_checked` / `unresolved` row with `ABSTRACT_MISSING`, while a present whitespace-only abstract uses `ABSTRACT_EMPTY`. A manual corpus entry receives no exemption. Missing, invalid, or unavailable snapshot state stays explicit and cannot become a checked zero.
The only supported supplies are a user-supplied or clearly synthetic fixture snapshot plus its detached manifest. Verify the manifest-declared `snapshot_sha256` against the exact raw snapshot bytes, retain the detached-manifest hash, and retain its rights declaration. This repository supplies no native PPS content, PPS importer, network fetcher, or redistributed phrase list. This #660 path uses no live model, external API, human or model judge, ambient clock, source file timestamp, or Git/network time; the caller must provide the required RFC 3339 timestamps.
Every row is `HEURISTIC-ADVISORY` and `UNMEASURED`. Describe a positive only as a **phrase-list match requiring review**; a checked zero means only that no match was observed on the named metadata surface. Never infer AI, author, papermill, misconduct, or other origin; contextual validity; false-positive/false-negative rate, accuracy, precision, recall, or coverage; cleanliness; or publisher acceptance. Do not propose replacement text, change citation selection or ranking, mint a marker, evaluate terminal policy, alter an integrity gate, or rewrite metadata. The formatter composes all rows into the one existing `Bibliographic Integrity Advisories` section.
## APA 7.0 Quick Reference
Reference: `references/apa7_style_guide.md`
@@ -0,0 +1,713 @@
# #660 — Tortured-phrase screening for own drafts and cited sources
> **Status:** DESIGN-FROZEN / UNMEASURED
> **Issue:** #660
> **Dependency:** #678 is closed and its canonical bibliographic-integrity carrier
> is the only corpus-side carrier.
> **Execution boundary:** this design authorizes no live model, external API,
> network fetch, human or model judge, native PPS import, redistributed phrase-list
> content, or expensive evaluation.
## 1. Decision and claim boundary
ARS will add a local, deterministic phrase-pattern matcher and expose its output on
two advisory surfaces:
1. an own-draft report produced immediately before final formatting; and
2. one canonical bibliographic-integrity signal row for each citation and each
metadata surface (`title` and `abstract`), including explicit not-checked rows.
The matcher is **mechanically deterministic**: the same manuscript or metadata
bytes, canonical-AST snapshot bytes, detached manifest bytes, explicit timestamps,
and runtime version produce byte-identical machine output. That execution property
does not raise the epistemic strength of the result. A phrase-list rule is a risk
heuristic, not an authority about origin, misconduct, paper-mill production, source
quality, or contextual legitimacy. Corpus rows therefore use:
```text
epistemic_class = heuristic_advisory
epistemic_label = HEURISTIC-INDICATOR
```
The own-draft artifact likewise carries:
```text
layer = HEURISTIC-ADVISORY
evaluation_status = UNMEASURED
```
The fixed user-facing positive wording is **“phrase-list match requiring review.”**
A checked zero-match result says only **“no phrase-list match observed on the
checked surface; absence is not a clean certification.”** No output may call a work
contaminated, generated, fraudulent, tortured, paper-mill-produced, or clean.
The carrier's always-present category label is the neutral **“Phrase-list screening
advisory”**; positive wording appears only on a `detected` row, never on a zero,
missing, or degraded row.
V1 performs no contextual false-positive judgment. It does not suggest replacement
text and does not edit a manuscript or citation. A future human or model judgment
would require a separately versioned, explicitly judgment-labelled artifact; it
must never be folded into the deterministic matcher transcript.
## 2. Authority and compatibility
### 2.1 One corpus carrier
`literature_corpus[].bibliographic_integrity_signals[]`, governed by
`shared/contracts/passport/bibliographic_integrity_signal.schema.json`, remains the
single corpus-side authority. #660 does not introduce another corpus aggregate,
reuse the closed boolean `contamination_signals` object, or add an advisory token to
the reference-marker grammar.
Implementation adds `bibliographic-integrity-signal/1.2` to that schema. Version
1.2 is a feature-specific profile within the existing carrier, not a new carrier.
It requires `signal_type: tortured_phrase_match`, the heuristic class and label
above, a closed `tortured_phrase_context` block, advisory-only terminal policy, and
`display.marker_token: null`. Shipped canonical v1.0 and v1.1 fixtures and current
producer outputs retain byte and validation identity. The new signal-type invariant
intentionally rejects previously underconstrained tortured-phrase mutations that
claimed `deterministic_fact` or terminal-policy eligibility; no canonical producer
emitted those shapes. The v1.1 retraction profile and its terminal policy are
unchanged.
The existing #678 tortured-phrase fixture is a scaffold, not evidence that #660 is
implemented. New #660 producers emit only v1.2 rows. No migration may reinterpret
old heuristic rows as a completed title or abstract check.
### 2.2 Separate citation-by-surface rows
Each corpus entry holds exactly one current v1.2 row for each Cartesian key
`(citation_key, surface)`. The id is hash-bound:
```text
bis:<schema-safe-citation-key>:tpm_title_<20-hex-binding-prefix>
bis:<schema-safe-citation-key>:tpm_abstract_<20-hex-binding-prefix>
```
The existing corpus schema already restricts `citation_key` to a signal-id-safe
alphabet. Inputs outside that alphabet fail corpus validation; the producer does not
invent a fallback identity. The binding prefix is derived deterministically from the
citation key, surface, exact snapshot SHA-256 (or explicit unavailable state), and
exact surface-content SHA-256 (or null for an absent or empty surface).
`signal_type` remains `tortured_phrase_match`; the surface-specific id makes the rows
independently addressable and prevents one aggregate status from hiding partial
coverage.
The non-in-place enricher supersedes exactly one prior v1.2 row for the same
citation×surface when any binding changes. It preserves every legacy v1.0 row and
every other signal type, emits no history row, and rejects two or more pre-existing
v1.2 current rows for one citation×surface as ambiguous. An unchanged replay yields
the same id and byte-identical row.
The corpus schema requires a title, so every schema-valid entry receives a title
row whenever the check is invoked. An abstract is available only when the field is
present and contains at least one non-whitespace Unicode code point. The exact
surface rules are:
| Surface state | `check_status` | `finding` | Meaning |
|---|---|---|---|
| valid snapshot and one or more retained instances | `checked` | `detected` | phrase-list match requiring review |
| valid snapshot and zero retained instances | `checked` | `not_detected` | no match on this checked surface; not a clean certificate |
| abstract field absent | `not_checked` | `unresolved` | `ABSTRACT_MISSING` |
| abstract present but empty/whitespace | `not_checked` | `unresolved` | `ABSTRACT_EMPTY` |
| snapshot not supplied | `not_checked` | `unresolved` | `SNAPSHOT_NOT_PROVIDED` |
| manifest/snapshot/AST/parser integrity failure | `degraded` | `unresolved` | exact closed failure code |
Manual corpus entries have no exemption. Their local title is checked exactly like
any other title, and their abstract follows the same present/absent rule. No DOI,
resolver, acquired PDF, or network request is required.
### 2.3 Formatter and finalizer ownership
The bibliography producer appends schema-valid rows; it does not evaluate policy.
The formatter transcribes them into the existing, single `Bibliographic Integrity
Advisories` section in lexical `signal_id` order. The Cite-Time Provenance Finalizer
does not promote a v1.2 row, and the formatter does not create or refuse a marker
because of it. Retraction, citation-existence, and legacy-contamination terminal,
finalizer, and marker-policy behavior remains unchanged. The additive advisory-table
columns and stronger inert escaping intentionally change rendered table bytes, not
those earlier policy semantics.
The generic #678 advisory table remains complete: the max-25 detail rule in §9 does
not authorize dropping canonical signal rows from that table.
## 3. Snapshot and detached-manifest contract
### 3.1 No native PPS ingestion or redistribution
ARS V1 accepts only an already prepared canonical-AST JSON snapshot from a local
path. It contains no parser, converter, downloader, URL option, API client, or
native importer for PPS syntax. The repository carries no PPS fingerprint content.
Hermetic fixtures use invented, license-clear synthetic phrases and may not be
described as PPS coverage.
The user is responsible for lawfully preparing any real snapshot outside ARS. V1
supports exactly two supply modes:
- `user_supplied`: a local snapshot that ARS neither vendors nor
republishes; and
- `synthetic_fixture`: repository-owned invented test content.
Adding vendored third-party content, a fetch-at-run path, or a native source-format
importer requires a new design version plus an in-repository license or written
permission record. Public visibility alone is not redistribution authority.
### 3.2 Exact-byte binding
Every snapshot travels with a separate strict-JSON manifest. The manifest schema is
`tortured-phrase-snapshot-manifest/1.0` and is closed recursively. Its exact field
shape is:
```text
schema_version
snapshot_id
source {
name
version
as_of
locator
}
supply_mode
snapshot_schema_version
snapshot_sha256
grammar_profile = ars-tortured-phrase-canonical-ast/1.0
normalizer_profile = ars-nfkc-casefold-token/1.0
preprocessor {
name
version
native_grammar
reduction_notes
}
unsupported_rule_count
rule_count
rights {
basis
redistribution_status
reference
user_declaration
}
```
`snapshot_sha256` is SHA-256 over the exact raw snapshot bytes, before decoding,
newline conversion, JSON parsing, key sorting, or serialization. The matcher validates
this binding before it reads any rule content. It then parses both files as strict
UTF-8 without a BOM,
rejecting duplicate decoded keys, non-finite numbers, invalid Unicode, unknown
fields, and folded near-miss version markers. JSON reserialization is never a
substitute for the raw-byte hash.
`source.locator` is provenance text and is never dereferenced. `source.as_of` is an
explicit session-held ISO date (`YYYY-MM-DD`), never a file time or clock read.
`snapshot_schema_version`, `grammar_profile`, `normalizer_profile`, `snapshot_id`,
and `rule_count` must equal the corresponding values in the exact snapshot.
`preprocessor.native_grammar` is a nullable bounded grammar-profile string and
`reduction_notes` is an array of bounded strings; together they disclose what
external source grammar, if any, was reduced into the canonical AST. They do not
activate an ARS importer. `unsupported_rule_count` is top-level and must be zero for
a checked run.
`rights.basis`, `rights.redistribution_status`, `rights.reference`, and
`rights.user_declaration` travel together. The repository fixture uses exactly
`basis: synthetic_fixture`, `redistribution_status: permitted`, `reference: null`,
and `user_declaration: null`; its `preprocessor.native_grammar` is
`synthetic-canonical-ast/1.0`. The closed `basis` values are
`synthetic_fixture`, `user_declared_authorized`, `written_permission`, and
`unresolved`; redistribution is `permitted`, `not_permitted`, or `unresolved`.
`user_declared_authorized` requires a non-empty user declaration;
`written_permission` requires a non-empty reference; and `unresolved` requires
unresolved redistribution status. A session-only user snapshot may be checked while
declaring no redistribution authority, and is never copied into the repository. #660
itself vendors no third-party content. The manifest records a claim and its reference;
it does not manufacture legal authority.
The report records both the raw snapshot hash and the raw detached-manifest hash.
A manifest is not proof that the supplied snapshot is a complete or authoritative
copy of any external list. It establishes only which exact bytes ARS checked.
All rules in a snapshot must validate. An unknown operator, invalid node, duplicate
rule id, rule-count mismatch, unsupported source reduction, or manifest mismatch
rejects the entire snapshot before matching. ARS never drops individual rules and
then reports the remainder as a checked list.
## 4. Canonical pattern AST
### 4.1 Closed rule shape
Each snapshot contains a lexically unique ASCII `rule_id`, one positive expression,
and a rule-level `exclude_if` array. The canonical snapshot schema is the single
authority for the exact JSON member names and numeric bounds. The runtime must consume
that public representation directly; a private alternate AST shape is forbidden.
A schema/runtime integration guard and fixtures lock their agreement before any
snapshot hash is frozen.
The only expression operators in `ars-tortured-phrase-canonical-ast/1.0` are
`literal`, `all`, `any`, and `near`. There is no native `not` node. Source-level negation may be
represented only by the rule-level, segment-scoped `exclude_if` reduction below.
Any source rule that cannot be represented without changing its semantics must be
reported through nonzero `unsupported_rule_count`; such a snapshot cannot authorize a
checked run. ARS itself makes no completeness claim.
Runtime work is closed and bounded before matching: at most 512 rules, 12 AST levels,
64 AST nodes per rule, eight normalized tokens per literal, 512 witnesses per node,
100,000 witness combinations per node, 4,096 parser intervals, 4,096 output segments,
`MAX_PARSE_WORK_UNITS = 100_000` shared across every parser opener, closer,
context candidate, and excluded candidate in the complete document,
100,000 rule-by-segment evaluations, a shared 5,000,000-unit literal/composition/
exclusion work budget for the complete input artifact (including all corpus entries),
500,000 input tokens, a pre-normalization maximum of 4,096 raw code points per token,
and 4,096 persisted match rows. The raw-token ceiling is enforced while collecting
source code points, before NFKC/casefold can allocate an expanded normalized value.
Corpus admission additionally
freezes `MAX_CORPUS_ENTRIES = 512` and
`MAX_CORPUS_EXISTING_SIGNALS = 8192`, where the latter is the aggregate number of
pre-existing `bibliographic_integrity_signals[]` rows across the complete input
corpus. A schema-valid input at exactly either limit is not rejected by that
cardinality guard; `N+1` raises the resource-limit failure before the passport is
copied, any surface is matched, or any output write is attempted. The direct enricher
never mutates its input document. The CLI creates no output on either admission
failure and leaves a pre-existing output byte-identical.
Decoded JSON/YAML structure is also bounded before schema traversal or copying:
`MAX_STRUCTURE_DEPTH = 64` and `MAX_STRUCTURE_NODES = 200_000`. The strict JSON
loader and the direct passport enricher both apply the same iterative guard. Parser
recursion, depth `N+1`, node `N+1`, or a shared/recursive YAML alias is a command-level
resource/structure failure: it emits no traceback, creates no new output, and leaves
an existing output byte-identical.
Snapshot, manifest, draft, advisory output, and passport input/output bytes are
independently capped and every successful output is readable by the corresponding
validator under the same cap. Every parser helper enforces the interval ceiling while
collecting intervals, not after an unbounded temporary list is built. Crossing an
artifact-level matcher ceiling discards partial matches and produces the closed
degraded/unresolved resource-limit state; it can never retain the remaining subset
and report a checked zero-match result. Crossing a command-level corpus admission or
serialized-output ceiling instead aborts before atomic replacement, so no partial
passport is published.
When a loaded-snapshot scan nevertheless emits an artifact-level `degraded` result
(including parse, empty-input, or matcher-resource degradation), the CLI first writes
the complete replayable degraded artifact atomically and then exits 1. Corpus
enrichment follows the same rule if any newly projected current surface is degraded.
An intentional `not_checked` result, including a missing abstract or explicitly
omitted snapshot, is not relabelled as a command failure and may exit 0.
The schema-authoritative node serialization and semantics are:
- `literal` is `{op, value}`. `value` is a non-empty bounded literal; after §5
normalization it must contain at least one token and matches one contiguous equal
token sequence.
- `all` is `{op, terms, max_span_tokens}` with 28 recursive `terms`. Every term
must have a witness in the same parsed segment. The minimal covering half-open
token span of a combination must be no wider than `max_span_tokens`. Term order is
not an ordering claim.
- `any` is `{op, alternatives}` with 28 recursive `alternatives`. Its witness set
is the union of the alternative witness sets.
- `near` is `{op, left, right, max_gap_tokens, ordered}`. Both witnesses must occur
in the same segment. Gap is the number of normalized tokens strictly between the
half-open spans; overlap or contact has gap zero. A candidate is retained only when
the gap is at most `max_gap_tokens`; when `ordered` is true, the complete left
witness must precede the right witness. The combined witness is the minimal
covering window.
All AST objects are closed. Empty `all`/`any`, a literal that normalizes to zero
tokens, an invalid token window, excessive recursion or witness expansion caught by
the documented safety ceilings, or an unknown field rejects the whole snapshot.
### 4.2 Segment-scoped exclusion
`exclude_if`, when present, contains 18 rule-level items of exact shape
`{expression, within_tokens}`. `expression` is a positive recursive AST; an exclusion
cannot contain another exclusion. A candidate positive witness is suppressed only
when an exclusion witness is in the exact same parsed segment and its token gap from
the candidate is at most `within_tokens`. An exclusion in another paragraph, quote,
reference entry, abstract, title, or file has no effect.
This is the only V1 negation semantics. It is intentionally narrower than an
unbounded document-level NOT. Snapshot preparation must reject, not approximate, a
source expression whose negation scope cannot be represented this way.
### 4.3 Witnesses, overlaps, and counts
Every witness carries the input artifact id, surface, segment id and class,
half-open normalized-token span, half-open raw Unicode-code-point and UTF-8 byte
spans, rule id, and SHA-256 of the exact matched raw bytes. Raw spans always point
into the original unmodified input; normalized text is never written back. Any
human-readable evidence excerpt is bounded to at most 25 whitespace-delimited words
and 1,000 Unicode code points; longer witnesses fail closed rather than truncate
replay evidence.
After exclusions:
1. duplicate witnesses with the same `(rule_id, segment_id, byte_start, byte_end)`
collapse;
2. `rule_match_count` counts the remaining unique rule witnesses;
3. overlapping or identical witnesses within one segment form a connected overlap
component; and
4. `unique_instance_count` counts those components. Its representative is selected by
earliest byte start, then longest byte span, then lexical rule id.
Adjacent non-overlapping spans are separate instances. Repeated occurrences at
different spans are separate instances. Reports always include both counts so a
single phrase matched by multiple rules cannot masquerade as repeated prose, while
genuine repetition remains visible. Counts are not severity tiers.
## 5. Text normalization and token boundaries
The matcher consumes exact raw UTF-8 bytes and retains their SHA-256 before any
transformation. It decodes strictly, rejects a BOM and isolated carriage returns,
and recognizes LF and CRLF line endings without rewriting the stored input.
Normalization is segment-local and read-only:
1. recognize the explicit Markdown or LaTeX structure under §6;
2. join a discretionary line-wrap hyphen only for `letter + '-' + newline + letter`
inside the same scannable segment, retaining a map to the full raw span;
3. remove U+00AD SOFT HYPHEN only when it joins word characters, joining those
characters into one token while retaining the full raw-span mapping; otherwise it
is a separator;
4. split at whitespace, punctuation, symbols, controls, and format characters;
dash-punctuation characters are separators except for the line-wrap rule above;
5. form tokens from maximal runs beginning with a Unicode Letter or Number and
followed by Letters, Numbers, or Marks; and
6. apply Unicode NFKC and Unicode casefold independently to each token.
The same function normalizes AST literals and input tokens. Empty normalized tokens
are rejected in the AST and ignored as separators in input. Ordinary hyphens do not
join words, zero-width format characters do not fuse tokens, and matching never uses
locale, filesystem encoding, regular-expression locale state, or an ambient clock.
Token equality is exact after this pipeline. Substring-inside-token matching,
stemming, lemmatization, edit distance, semantic similarity, and language-model
classification are out of scope.
## 6. Markdown and LaTeX segmentation
The caller must pass `--format markdown` or `--format latex`; format guessing is
forbidden. Parsing produces stable, ordered segments with one of these closed
contexts:
```text
author_prose
quote
cited_title
reference_entry
code_or_verbatim
unknown
cited_abstract
```
Every context remains scannable so a parser classification does not silently erase a
list observation. Context changes only the disposition: `author_prose` asks for
review with no automatic rewrite; quote, cited-title, reference and code/verbatim
contexts are preserved verbatim; a cited abstract routes to cited-source review.
An `unknown` context is reported explicitly and never gains a rewrite suggestion.
Malformed structure that prevents complete segmentation makes the artifact-level
check `degraded/unresolved`; a degraded artifact cannot emit a zero-match statement.
The Markdown subset recognizes fenced and inline code, block-quote lines,
HTML-comment spans, DOI-link title text, and a reference section introduced by the
closed ATX heading vocabulary `References`, `Bibliography`, or `Works Cited`.
Dollar and inline-code delimiters preceded by an odd run of backslashes are literal;
an even run leaves the delimiter active. The same parity rule applies to LaTeX dollar
math and percent comments, so escaped literal syntax never creates a false
`unknown` segment.
Within that recognized section, every non-empty physical reference line begins a new
exclusion-scope segment. The LaTeX `thebibliography` context similarly begins a new
exclusion-scope segment at each `\bibitem`. This conservative rule prevents an
exclusion phrase in one reference entry from suppressing a hit in another; it does not
claim full citation parsing.
Everything else—including ordinary paragraphs, ATX/Setext headings, YAML front
matter, ordinary links, reference definitions outside the recognized section, and
raw HTML—is scanned conservatively as `author_prose`; V1 does not claim to parse
those constructs. An unclosed fence, inline-code delimiter, or HTML comment becomes
`unknown` and degrades the artifact-level result.
The LaTeX subset recognizes comments; `verbatim`, `lstlisting`, and `minted`;
`quote`/`quotation`; `thebibliography`; `\verb`; `\(...\)`, `\[...\]`, and dollar
math. Text commands and other macro bytes remain scannable as `author_prose` rather
than being interpreted. An unclosed recognized environment, delimiter, or `\verb`
becomes `unknown`; V1 does not claim complete TeX expansion or brace validation.
External `.bib` content is never dereferenced by this parser; cited titles come from
the structured corpus surface instead.
Opaque constructs are lexed strictly in source order: after the earliest eligible
opener is selected, opener-like bytes inside that interval cannot consume a closer
belonging to a later construct. Contextual quote and DOI-title recognizers likewise
cannot start inside, or pair across, an opaque interval. Fixed-cost opener scans and
monotonic delimiter cursors replace retrying backreference searches; candidate
collection fails at the parser-interval ceiling. `\verb*` is an explicit recognized
variant, while longer control words such as `\verbose` and `\verbatim` remain prose;
a bare `\verb` without a legal delimiter becomes `unknown`. Same-name nested
`quote`/`quotation` environments remain one protected outer quote interval.
Opaque membership uses binary search over sorted non-overlapping intervals; ignored
or unmatched context tokens still consume the same document-wide parser-work budget,
so exclusion checks cannot multiply candidate count by interval count without hitting
a declared ceiling.
Matches in manuscript prose receive
`review_author_prose_no_automatic_rewrite`. Matches in quotes, cited-title text,
reference entries, code and verbatim contexts receive
`preserve_verbatim_review_context`; the separate corpus title row supplies the
cited-source route. Unknown context receives `review_unknown_no_automatic_rewrite`.
No disposition contains replacement text. A cited title remains byte-for-byte
unchanged even when its corpus title row is `detected`.
## 7. Phase 1 — matcher, provenance, and synthetic seeds
Phase 1 implements only the pure local substrate:
- closed detached-manifest and AST snapshot schemas;
- exact-byte manifest validation;
- Markdown/LaTeX segmentation;
- normalization, matching, exclusion, overlap, and count functions;
- a CLI requiring explicit input, format, snapshot, manifest, and timestamps; and
- public, invented positive and negative conformance fixtures.
The seed set covers literal boundaries, Unicode compatibility/casefold behavior,
ordinary and line-wrap hyphens, soft hyphens, `all`, `any`, bounded `near`,
gap edges, segment-scoped exclusions, overlapping rules, repeated instances,
Markdown/LaTeX classification, malformed input, manifest mismatch, and unsupported
AST rejection.
These fixtures measure implementation conformance only. They do not estimate
real-world false-positive rate, false-negative rate, precision, recall, list
coverage, contextual validity, or publisher screening behavior. Their expected
results are hand-authored mechanical expectations, not judge labels.
Phase 1 neither emits the own-draft advisory nor writes corpus rows. A Phase-1-only
change references #660 but cannot close it.
## 8. Phase 2 — own-draft advisory
### 8.1 Machine artifact
Phase 2 adds a recursively closed `tortured-phrase-advisory/1.0` artifact. Its
normative schema groups exact raw input/snapshot/manifest/timestamp bindings under
`input_binding`, keeps the closed status/finding/reason-code state, retains all match
records and context counts, and includes a self-hash. The following semantic values
are immutable:
```text
schema_version = tortured-phrase-advisory/1.0
layer = HEURISTIC-ADVISORY
evaluation_status = UNMEASURED
check_status
finding
matches
counts (including every match record and unique textual instances)
boundary
```
The complete machine artifact retains every match and the closed segment/context
coverage counts used by the reducer. Exact source and snapshot replay reconstructs
the underlying segment partition. It is an
advisory transcript, not a score or submission gate. `evaluation_status` cannot be
promoted by passing synthetic tests.
An empty or whitespace-only draft is `degraded/unresolved` with `DOCUMENT_EMPTY`;
it is never represented as a checked zero-match surface.
### 8.2 Pipeline placement
In the full pipeline, the orchestrator runs the checker on the exact accepted working
draft after the final integrity pass and immediately before Stage 5 conversion. In
standalone `academic-paper` formatting, it runs on the exact format-conversion input.
The formatter receives the already validated artifact and only renders it; it does
not re-run matching, infer context, alter counts, or rewrite prose.
A detected, not-checked, or degraded result remains advisory and does not add a
terminal policy or bypass the existing mandatory Stage-5 confirmation. The user may
choose to revise, preserve, or proceed. Any revision creates new input bytes and
requires a new check with new explicit timestamps before the result may be called
current.
Phase 2 does not write cited-source rows. A Phase-2-only change references #660 but
cannot close it.
## 9. One-page bounded human renderer
The own-draft #660 match-detail renderer is fixed to at most 25 match rows. There is
no `--all`, limit override, environment variable, configuration key, or alternate
unbounded human-rendering path. There is exactly one rendered page: no `--page`,
`--page-size`, cursor, next-page token, or repeated pagination path may expose the
omitted rows. CLI parsing must reject `--all`, `--page`, and `--page-size` as unknown
options.
Rows are selected from the canonical machine order: raw code-point start, raw
code-point end, lexical rule id, then segment id. The renderer reports total
`unique_instance_count`, `rule_match_count`, shown count, and omitted count, and
points to the complete machine JSON when rows are
omitted. It escapes Markdown table delimiters, line breaks, and control characters;
it never interpolates a match as executable Markdown or LaTeX.
The cap applies to own-draft detail, not to the canonical #678 advisory table: every
corpus signal row, including `not_checked` abstract rows, remains visible there. The
aggregate table projects at most three bounded match witnesses inside each v1.2 row
and reports the number of omitted machine witnesses; that per-row projection cannot
hide the row's status or counts.
## 10. Phase 3 — cited-source integration
Phase 3 adds the v1.2 carrier profile and a pure producer that scans only the exact
local `literature_corpus[].title` and optional `.abstract` strings. It makes no API
call and does not dereference `source_pointer`.
The required `tortured_phrase_context` records the exact closed shape:
```text
layer = HEURISTIC-ADVISORY
evaluation_status = UNMEASURED
surface = cited_title | cited_abstract
surface_binding { content_sha256, content_utf8_bytes }
snapshot {
status, reason_code, snapshot_sha256, manifest_sha256, snapshot_id,
source, supply_mode, snapshot_schema_version, grammar_profile,
normalizer_profile, unicode_data_version, rule_count,
unsupported_rule_count, rights
}
reason_code
counts { rules_evaluated, matched_rule_count, rule_match_count,
unique_instance_count, segments_total, unknown_segments,
matches_by_context }
matches
boundary
```
The two `surface_binding` members are non-null only for a present, non-whitespace
surface. Checked rows bind the versioned grammar/normalizer profiles, Unicode data
version, exact snapshot and manifest hashes, and explicit timestamps; the integration
guard additionally freezes the shipped runtime bytes. Detected rows carry structured
match witnesses; checked
not-detected rows carry a list-observation evidence item; absent/empty/not-supplied
and degraded rows carry a typed degradation item. Evidence never embeds a full
abstract: a witness contains its bounded raw span/hash and rule id, while the local
machine artifact remains the complete replay authority.
`unicode_data_version` records the producer interpreter's Unicode database; it is
not rewritten to the validator's local version. The shipped v1.2 fixtures freeze
their generation provenance at `14.0.0`. A validator running another Unicode version
must validate that frozen value and compare the remaining replay projection exactly,
not relabel the stored artifact with its ambient version. A newly produced row still
records its actual runtime version, and byte-exact reproduction of that row requires
the same recorded Unicode data version.
The producer is idempotent by the citation×surface current-row rule in §2.2.
Re-running on unchanged bytes and identical explicit timestamps yields byte-identical
output. Changed surface or snapshot bytes change the hash-bound id and atomically
supersede the one prior current row in the returned copy. A timestamp-only change
updates the row but does not create history. Before supersession, an existing current
row must pass its entry-independent schema, id/hash, count, state, provenance, and
match self-binding checks and still join the containing citation key and source
pointer. The old surface-content hash may differ from newly supplied surface bytes;
that difference is the legitimate supersession case. Duplicate current rows,
duplicate ids, internally stale bindings, or a self-bound surface/id mismatch fail
closed rather than using first- or last-write-wins.
Formatter integration renders the fixed vocabulary from §1, exact per-surface
status, both counts, snapshot version/date/hash, explicit times, and unresolved
coverage. It composes with retraction and every other #678 row without minting a
marker or policy result.
Only Phase 3 completes the two-surface implementation. That implementation PR keeps
#660 open; after the exact accepted commit reaches `main`, the separate preregistered
mechanical-conformance PR may close #660 only when every row of §13 is satisfied.
## 11. Explicit time and reproducibility contract
Every producer requires schema-valid RFC 3339 timestamps as explicit arguments.
Fractional seconds, when present, contain one through six digits so ordering is
lossless in the shared runtime. Required timestamp fields have no default. The runtime must not read system time,
monotonic time as a substitute, timezone, file modification time, Git author/commit
time, network time, or UUID time.
At minimum the invocation supplies `--checked-at` and `--recorded-at`.
`source.as_of` (an ISO date) and source version are explicit fields in the already
supplied, validated detached manifest; the runtime never synthesizes either. Missing
or malformed timestamps fail before a checked result is emitted.
Hermetic clock-mutation tests monkeypatch or deny common clock functions and verify
that output identity depends only on explicit inputs. Two runs over identical raw
bytes, runtime version, and explicit timestamps must produce identical canonical
JSON bytes.
## 12. Measurement and claim ceiling
No model, API, human judge, model judge, or contextual classifier is run for #660.
The repository seed set is public synthetic conformance material, not a held-out
effectiveness evaluation. Phase acceptance uses ordinary deterministic tests and
keeps both advisory surfaces `UNMEASURED` with respect to contextual validity.
No #654 measurement row may be manufactured in the implementation PR. That PR lands
the synthetic suite, frozen expectations, measurement plan, and registry entry, but
publishes no scored envelope. The closure PR adds a `mechanical_match` report only
after the exact suite and matcher commit exists on `main`, so
`subject.config.suite_commit` and `preregistration.frozen_commit` can name a real
40-hex main-history object and all plan, execution-manifest, and raw-output
references can resolve. That closure row uses zero judges,
`judge_plan.exception: mechanical_suite`, no adjudication, and a headline explicitly
limited to synthetic conformance. It cannot retrospectively authorize real-world
precision/recall, contextual validity, source cleanliness, or publisher acceptance.
The mechanical #654 row is required before #660 closes. It scores the public
synthetic positive-match and negative-non-match expectations only; those labels are
grammar-conformance fixtures, not measured contextual false-positive or
false-negative rates. Introducing a judged contextual layer would reopen design and
invoke the applicable #654 judge, blinding, preregistration, and raw-output
requirements.
## 13. Closure and acceptance matrix
The design document itself removes the design ambiguity but does not close the issue.
The `status/needs-design` label may be removed after this frozen design is accepted.
Closing keywords are reserved for the change that can demonstrate every row below.
| Requirement | Frozen proof required before closure | Phase |
|---|---|---|
| Shared authority | v1.2 profile extends #678; v1.0/v1.1 compatibility and retraction identity tests pass; no new corpus carrier | 3 |
| Snapshot rights | no PPS content/importer/fetcher; exact-byte detached manifest; only user-supplied or synthetic modes | 1 |
| Grammar | closed `literal`/`all`/`any`/`near` AST and rule-level segment-scoped `exclude_if`; unsupported input fails closed | 1 |
| Normalization | UTF-8/BOM, NFKC/casefold, token, Unicode, hyphenation, raw-span and overlap/count semantics are mutation-tested | 1 |
| Parsing | explicit Markdown/LaTeX modes classify prose, quotes, cited titles, references and unsupported coverage | 1 |
| Public seed | invented positive/negative fixtures cover all grammar and normalization branches; only synthetic conformance is claimed | 1 |
| Mechanical measurement | a post-merge `heldout-measurement/1.1` `mechanical_match` row resolves the precommitted plan, exact main-history suite commit, write-once execution manifest, and retained raw transcript; zero judges and no contextual-accuracy claim | closure |
| Own-draft surface | closed HEURISTIC-ADVISORY/UNMEASURED artifact scans exact final input and never rewrites or gates | 2 |
| Preservation/routing | quotes and cited/reference titles remain verbatim; cited-title matches route to the corpus advisory surface | 2/3 |
| Cited-source surface | independent title and abstract rows; absent/empty abstract is explicit `not_checked/unresolved`; manual entries are covered | 3 |
| Epistemic separation | deterministic runtime is not relabelled deterministic fact; every row is HEURISTIC-INDICATOR; contextual judgment is `not_performed` | 1/2/3 |
| Counts and tiers | `unique_instance_count` and `rule_match_count` always render; no severity tier exists | 1/2/3 |
| Provenance | every checked result carries snapshot version/date/raw hash, manifest hash, exact surface/input hash, and explicit times | 1/2/3 |
| No ambient execution | no network/model/API/judge/clock path; hermetic integration guards and clock mutations pass | 1/2/3 |
| Bounded rendering | own-draft detail has one fixed page with at most 25 matches and no traversal flags; the complete canonical corpus table retains every signal row and bounds evidence per row | 2/3 |
| Composition | one complete Bibliographic Integrity Advisories section; lexical rows; no advisory marker, terminal policy, or formatter re-judgment | 3 |
| Claim vocabulary | positive, zero, missing and degraded wording matches §1; banned origin/contamination/clean claims are mutation-tested | 2/3 |
| Consumer wiring | bibliography producer, pipeline orchestrator, formatter, protocol docs, fixtures, integration checker and CI manifest agree | 3 |
| Validation | focused tests, adjacent #651/#678 carrier regressions, schemas, integration guards, mirror equality where applicable, lint, compile and diff checks are green | 3 |
Phase 1, Phase 2, and Phase 3 implementation changes use `Refs #660`. The subsequent
mechanical-measurement PR may use `Closes #660` only if its exact implementation
commit is already reachable on `main` and the complete matrix is green in that exact
branch. The three implementation phases may be separate PRs or separately reviewable
commits in one PR, but no phase may be omitted or treated as implied by the #678
scaffold. After merge, closure is confirmed from GitHub rather than inferred from a
commit message.
## 14. Non-goals and future changes
V1 explicitly does not provide:
- AI-text detection or authorship/origin inference;
- misconduct, paper-mill, contamination, quality, or publishability findings;
- a clean-text or clean-source certificate;
- contextual false-positive classification;
- automatic replacement or rewrite suggestions;
- severity tiers or terminal policy;
- a native PPS parser, list downloader, redistributed PPS content, or list-completeness
claim;
- full-text cited-source screening;
- a model/API/judge run or real-world accuracy estimate; or
- an unbounded human renderer.
Changing the AST operators, exclusion scope, normalization, overlap/count rule,
surface unit, epistemic label, output vocabulary, timestamp source, renderer cap, or
claim ceiling requires a new contract version and updated mutation tests. It may not
be introduced as an undocumented implementation detail.
+1
View File
@@ -78,6 +78,7 @@ table below is an informative mirror:
| `pipeline_behavior_robustness` | `mechanical_match` | full-expectation mechanical match; judge only transcribes |
| `reviewer_seeded_defects` | `seeded_manifest_adjudicated` | E4 machinery remains normative and unchanged; see adoption surface below |
| `re_review_persuasion_invariance` | `paired_controls` | reuses E4 machinery per its README (SD-11) |
| `tortured_phrase_conformance` | `mechanical_match` | synthetic grammar, normalization, parsing, replay, and fail-safe conformance only; no contextual-accuracy claim |
Class semantics (schema branches B1-B3 + checker):
+2 -1
View File
@@ -4,5 +4,6 @@
"rq_framing_offlist": "llm_judged",
"pipeline_behavior_robustness": "mechanical_match",
"reviewer_seeded_defects": "seeded_manifest_adjudicated",
"re_review_persuasion_invariance": "paired_controls"
"re_review_persuasion_invariance": "paired_controls",
"tortured_phrase_conformance": "mechanical_match"
}
@@ -0,0 +1,23 @@
# Tortured-Phrase Mechanical Conformance
Issue: #660. Suite class: `mechanical_match`.
This suite measures only whether the deterministic #660 runtime matches the public,
repository-owned synthetic expectations in
`scripts/fixtures/tortured_phrase_screening/seed_expectations.json`. It does not
contain Problematic Paper Screener content, a native PPS importer, real manuscripts,
or contextual false-positive/false-negative labels. A passing row therefore supports
only a synthetic grammar/normalization/parser/replay conformance statement. It does
not support claims about real-world accuracy, paper-mill or AI origin, source quality,
source cleanliness, publisher screening, or contextual validity.
The implementation PR registers and freezes this suite but remains `UNMEASURED`.
Measurement occurs only after that exact implementation is reachable on `main`, under
the precommitted [measurement plan](measurement_plan.md). The resulting
`heldout-measurement/1.1` row uses zero judges,
`judge_plan.exception: mechanical_suite`, and `adjudication.applies: false`. Raw
command/output bytes and a write-once execution manifest remain beside the row.
The synthetic positive cases mean “the frozen matcher must emit this match.” The
synthetic negative cases mean “the frozen matcher must not emit this match.” They are
not empirical false-positive or false-negative rates and are never relabelled as such.
@@ -0,0 +1,133 @@
# Frozen measurement plan — tortured_phrase_conformance/1.0
Status: PRE-REGISTERED / NOT RUN. Issue: #660. Contract:
`heldout-measurement/1.1`; suite class: `mechanical_match`.
## Frozen subject and inputs
The subject is the exact accepted #660 implementation at the first commit reachable
on `main` that contains the complete implementation, public suite, and this plan. The
measurement row must use that same 40-hex main-history commit for both
`subject.config.suite_commit` and `preregistration.frozen_commit`; no working-tree,
plan-only, or later runtime is eligible.
The six seed/data inputs are:
- `scripts/fixtures/tortured_phrase_screening/snapshot.json`;
- `scripts/fixtures/tortured_phrase_screening/snapshot_manifest.json`;
- `scripts/fixtures/tortured_phrase_screening/seed_expectations.json`;
- `scripts/fixtures/tortured_phrase_screening/own_draft.md`;
- `scripts/fixtures/tortured_phrase_screening/own_draft.tex`; and
- `scripts/fixtures/tortured_phrase_screening/corpus_input.yaml`.
The full measurement subject also includes the test harness, runtime, every schema
and carrier fixture it loads, and the Python dependency environment disclosed in the
raw transcript. Every repository path is read from the same frozen commit. The six
items above are not an exhaustive dependency list and may not be substituted from a
later tree.
Only repository-owned synthetic phrases are in scope. No native PPS bytes, network,
model, API, human judge, model judge, or contextual classifier may be introduced.
## Frozen execution
From a clean checkout of the pinned commit, execute exactly once:
```text
python -m pytest -q scripts/test_tortured_phrase_screening.py
```
The exact command is encoded as UTF-8 with no trailing newline;
`execution_manifest.calls[0].prompt_sha256` is SHA-256 over those exact bytes. The
single retained transcript path is
`evals/heldout/tortured_phrase_conformance/runs/2026-08-10/raw/pytest-transcript.json`.
It is strict UTF-8 JSON without a BOM, serialized with sorted keys, no insignificant
whitespace, no non-finite values, and one terminal LF. Its closed fields are:
```text
schema_version = tortured-phrase-conformance-transcript/1.0
command_utf8
started_at
completed_at
exit_code
environment {
python_implementation
python_version
pytest_version
jsonschema_version
ruamel_yaml_version
}
stdout_utf8
stderr_utf8
```
`command_utf8` equals the command above exactly. Start/completion are explicit RFC
3339 values; `environment` records the exact interpreter and required package
versions; stdout and stderr preserve the complete decoded strict-UTF-8 streams
separately; and exit code is the process exit status. The execution manifest's
`output_sha256` is SHA-256 over the exact transcript JSON bytes, so stdout, stderr,
exit status, command, and timestamps share one unambiguous framing. The report lists
that exact file under `raw_outputs.paths`.
The execution manifest is
`evals/heldout/tortured_phrase_conformance/runs/2026-08-10/execution-manifest.json`.
It has `created_at == calls[0].completed_at`, one call with
`call_id: tpc-2026-08-10-pytest` and `sequence_index: 1`, and makes no
`same_window`, `ordering`, or `concurrency` claim. No retry is permitted. A non-zero
exit, partial transcript, changed fixture/runtime bytes, or unavailable dependency is
a failed/blocked run and may not be replaced by a hand-authored success record. The
transcript and execution manifest are emitted once, hashed, and thereafter
write-once.
## Frozen metric and verdict
Primary metric: `synthetic_conformance_test_pass_rate`.
- Numerator: pytest test cases reported passed by the exact command.
- Denominator: all collected test cases in that exact command.
- A success verdict requires exit status 0, zero failed/error/skipped/xfailed/xpassed
cases, and numerator equal to denominator.
- Any collection drift is visible in the raw transcript and reported count. There is
no outcome-dependent exclusion, rerun, adjudication, or threshold adjustment.
- `replicates.per_item` is 1 with the explicit mechanical determinism exception; the
same bytes are not rerun to manufacture variance.
The report publishes the exact passed/collected counts, `1.0` only if all cases pass,
and a `point_estimate` label strictly for this finite synthetic suite. It uses zero
judges, no agreement statistic, no adjudication, and no experimental arms.
The report freezes these envelope values:
- `decision_relevant: true`, `suite_class: mechanical_match`,
`subject.model_id: deterministic-runtime/tortured_phrase_screening.py`, and the
common main-history commit binding above;
- `judge_plan.exception: mechanical_suite`, `judges: []`, and
`adjudication: {applies: false}`;
- `aggregate.agreement` is exactly `rate: null`, `divergent_items: []`, with a note
saying there are no judges; the headline construction rule is the numerator divided
by denominator under the success rule above;
- `replicates.per_item: 1` and an explicit exception that deterministic exact-byte
replay is not rerun to manufacture variance;
- preregistration has the exact plan path/hash, the shared frozen commit,
`frozen_before_dispatch: true`, `rubric_and_plan_frozen_together: true`, no rubric
or judge-template fields, `amendments_append_only: true`, and an empty amendment
ledger;
- `results.design: single-arm deterministic synthetic conformance`, both reserved arm
arrays empty, and suite-specific `passed_tests`, `collected_tests`, `exit_status`,
and `synthetic_conformance_test_pass_rate` fields; and
- on success, `attempts.partial_published: false` with no blocked runs; otherwise it is
true and names the single failed/blocked call while retaining the raw transcript.
## Claim ceiling
A passing row may say only: “the pinned deterministic runtime passed the pinned
repository-owned synthetic conformance suite.” It remains `UNMEASURED` for contextual
validity and real-world false-positive/false-negative performance. It cannot certify
clean text or infer paper-mill, AI, misconduct, contamination, quality, acceptance, or
origin. Any wording beyond this ceiling invalidates the row.
## Amendments
The amendment ledger starts empty and is append-only. Any change to subject bytes,
fixtures, command, metric, success rule, claim ceiling, or retry policy requires a new
plan version and a new post-freeze run; it cannot be pooled with this plan.
+8
View File
@@ -440,6 +440,14 @@ path = "scripts/test_check_human_subjects_reference_migration.py"
id = "678-bibliographic-integrity-signals"
path = "scripts/test_check_bibliographic_integrity_signals.py"
[[pytest]]
id = "660-tortured-phrase-screening"
path = "scripts/test_tortured_phrase_screening.py"
[[pytest]]
id = "660-tortured-phrase-integration"
path = "scripts/test_check_tortured_phrase_screening_integration.py"
[[pytest]]
id = "651-retraction-status"
path = "scripts/test_retraction_status.py"
@@ -2,10 +2,12 @@
against their schemas and enforces citation_key uniqueness."""
import copy
import json
import os
from pathlib import Path
import subprocess
import sys
import yaml
import pytest
REPO_ROOT = Path(__file__).resolve().parents[3]
@@ -13,6 +15,10 @@ SCRIPT = REPO_ROOT / "scripts/check_literature_corpus_schema.py"
SIGNAL_FIXTURE = (
REPO_ROOT / "scripts/fixtures/bibliographic_integrity_signals/retraction.json"
)
TORTURED_PHRASE_FIXTURES = REPO_ROOT / "scripts/fixtures/tortured_phrase_screening"
sys.path.insert(0, str(REPO_ROOT / "scripts"))
import tortured_phrase_screening as screening # noqa: E402
def _write_yaml(tmp_path, name, data):
@@ -35,6 +41,42 @@ def test_script_exists():
assert SCRIPT.exists()
def test_package_import_cannot_be_shadowed_by_pythonpath(tmp_path):
for module_name in ("check_v3_10_policy", "tortured_phrase_screening"):
(tmp_path / f"{module_name}.py").write_text(
f'raise RuntimeError("shadow {module_name} imported")\n',
encoding="utf-8",
)
env = os.environ.copy()
env["PYTHONPATH"] = os.pathsep.join((str(tmp_path), str(REPO_ROOT)))
probe = "\n".join(
(
"from pathlib import Path",
"import sys",
"import scripts.check_literature_corpus_schema as checker",
f"repo = Path({str(REPO_ROOT)!r})",
"assert checker.assert_venue_type_source_clean.__module__ "
"== 'scripts.check_v3_10_policy'",
"assert checker.validate_cited_signal_binding.__module__ "
"== 'scripts.tortured_phrase_screening'",
"assert Path(sys.modules['scripts.check_v3_10_policy'].__file__).resolve() "
"== repo / 'scripts/check_v3_10_policy.py'",
"assert Path(sys.modules['scripts.tortured_phrase_screening'].__file__).resolve() "
"== repo / 'scripts/tortured_phrase_screening.py'",
)
)
result = subprocess.run(
[sys.executable, "-c", probe],
cwd=tmp_path,
env=env,
capture_output=True,
text=True,
check=False,
)
assert result.returncode == 0, result.stderr
def test_passes_on_valid_passport(tmp_path):
passport = {
"literature_corpus": [
@@ -63,6 +105,37 @@ def _entry_with_signals(signals):
}
def _runtime_enriched_passport(citation_key="fixture_complete_2026"):
source = yaml.safe_load(
(TORTURED_PHRASE_FIXTURES / "corpus_input.yaml").read_text(encoding="utf-8")
)
entry = copy.deepcopy(
next(
item
for item in source["literature_corpus"]
if item["citation_key"] == citation_key
)
)
bundle = screening.load_snapshot(
TORTURED_PHRASE_FIXTURES / "snapshot.json",
TORTURED_PHRASE_FIXTURES / "snapshot_manifest.json",
)
state = screening.SnapshotState(
status="loaded",
reason_code="CHECK_COMPLETED",
bundle=bundle,
snapshot_sha256=bundle.snapshot_sha256,
manifest_sha256=bundle.manifest_sha256,
detail=None,
)
return screening.enrich_passport(
{"literature_corpus": [entry]},
state=state,
checked_at="2026-08-10T01:00:00Z",
recorded_at="2026-08-10T01:00:01Z",
)
def test_passport_cross_validates_canonical_signal(tmp_path):
signal = json.loads(SIGNAL_FIXTURE.read_text(encoding="utf-8"))
signal["subject"]["citation_key"] = "chen2024"
@@ -72,6 +145,109 @@ def test_passport_cross_validates_canonical_signal(tmp_path):
assert result.returncode == 0, result.stderr
def test_passport_accepts_runtime_generated_current_tortured_phrase_pair(tmp_path):
passport = _runtime_enriched_passport()
signals = passport["literature_corpus"][0]["bibliographic_integrity_signals"]
surfaces = {
signal["tortured_phrase_context"]["surface"]
for signal in signals
if signal.get("schema_version") == "bibliographic-integrity-signal/1.2"
}
assert surfaces == {"cited_title", "cited_abstract"}
path = _write_yaml(tmp_path, "passport.yaml", passport)
result = _run(["--passport", str(path)])
assert result.returncode == 0, result.stderr
def test_passport_accepts_missing_abstract_current_pair_with_null_checked_at(tmp_path):
passport = _runtime_enriched_passport("fixture_missing_2026")
signals = passport["literature_corpus"][0]["bibliographic_integrity_signals"]
title = next(
signal
for signal in signals
if signal["tortured_phrase_context"]["surface"] == "cited_title"
)
abstract = next(
signal
for signal in signals
if signal["tortured_phrase_context"]["surface"] == "cited_abstract"
)
assert title["provenance"]["checked_at"] == "2026-08-10T01:00:00Z"
assert abstract["provenance"]["checked_at"] is None
assert (
title["tortured_phrase_context"]["snapshot"]
== abstract["tortured_phrase_context"]["snapshot"]
)
assert title["provenance"]["recorded_at"] == abstract["provenance"]["recorded_at"]
path = _write_yaml(tmp_path, "passport.yaml", passport)
result = _run(["--passport", str(path)])
assert result.returncode == 0, result.stderr
def test_passport_rejects_current_tortured_phrase_pair_with_one_row_deleted(
tmp_path,
):
passport = _runtime_enriched_passport()
entry = passport["literature_corpus"][0]
entry["bibliographic_integrity_signals"] = [
signal
for signal in entry["bibliographic_integrity_signals"]
if signal["tortured_phrase_context"]["surface"] != "cited_abstract"
]
path = _write_yaml(tmp_path, "passport.yaml", passport)
result = _run(["--passport", str(path)])
assert result.returncode != 0
assert "require exactly one 'cited_abstract' row" in result.stderr
def test_passport_rejects_duplicate_current_tortured_phrase_surface(tmp_path):
passport = _runtime_enriched_passport()
entry = passport["literature_corpus"][0]
title = next(
signal
for signal in entry["bibliographic_integrity_signals"]
if signal["tortured_phrase_context"]["surface"] == "cited_title"
)
entry["bibliographic_integrity_signals"].append(copy.deepcopy(title))
path = _write_yaml(tmp_path, "passport.yaml", passport)
result = _run(["--passport", str(path)])
assert result.returncode != 0
assert (
"multiple current tortured-phrase rows for surface 'cited_title'"
in result.stderr
)
@pytest.mark.parametrize(
("mutation", "message"),
[
("snapshot", "same snapshot and detached manifest"),
("recorded_at", "same recorded run"),
("checked_at", "exact checked_at value"),
],
)
def test_passport_rejects_incoherent_current_tortured_phrase_pair(
tmp_path, mutation, message
):
passport = _runtime_enriched_passport()
signals = passport["literature_corpus"][0]["bibliographic_integrity_signals"]
abstract = next(
signal
for signal in signals
if signal["tortured_phrase_context"]["surface"] == "cited_abstract"
)
if mutation == "snapshot":
abstract["tortured_phrase_context"]["snapshot"]["manifest_sha256"] = "0" * 64
elif mutation == "recorded_at":
abstract["provenance"]["recorded_at"] = "2026-08-10T01:00:02Z"
else:
abstract["provenance"]["checked_at"] = "2026-08-10T01:00:00.5Z"
path = _write_yaml(tmp_path, "passport.yaml", passport)
result = _run(["--passport", str(path)])
assert result.returncode != 0
assert message in result.stderr
def test_passport_rejects_schema_invalid_canonical_signal(tmp_path):
signal = json.loads(SIGNAL_FIXTURE.read_text(encoding="utf-8"))
signal["subject"]["citation_key"] = "chen2024"
+406 -7
View File
@@ -9,10 +9,13 @@ from __future__ import annotations
import argparse
import copy
import datetime as dt
import hashlib
import html
import json
import re
import sys
import unicodedata
from pathlib import Path
from typing import Any
@@ -44,6 +47,35 @@ _SUMMARY_LABELS = {
"tortured_phrase_match": "Tortured-phrase heuristic match",
}
_TORTURED_PHRASE_CONTEXTS = (
"author_prose",
"quote",
"cited_title",
"reference_entry",
"code_or_verbatim",
"unknown",
"cited_abstract",
)
_TORTURED_PHRASE_RFC3339_RE = re.compile(
r"^[0-9]{4}-[0-9]{2}-[0-9]{2}[Tt]"
r"(?:[01][0-9]|2[0-3]):[0-5][0-9]:[0-5][0-9]"
r"(?:\.[0-9]{1,6})?(?:[Zz]|[+-](?:[01][0-9]|2[0-3]):[0-5][0-9])$"
)
def _empty_tortured_phrase_counts() -> dict[str, Any]:
return {
"rules_evaluated": 0,
"matched_rule_count": 0,
"rule_match_count": 0,
"unique_instance_count": 0,
"segments_total": 0,
"unknown_segments": 0,
"matches_by_context": {
context: 0 for context in _TORTURED_PHRASE_CONTEXTS
},
}
def _signal_id(citation_key: str, signal_type: str) -> str:
if re.fullmatch(r"[A-Za-z0-9_:-]+", citation_key):
@@ -65,6 +97,282 @@ def validation_errors(signal: dict[str, Any]) -> list[str]:
return sorted(error.message for error in validator.iter_errors(signal))
def _canonical_json(value: Any) -> str:
return json.dumps(
value,
ensure_ascii=False,
sort_keys=True,
separators=(",", ":"),
allow_nan=False,
)
def _validate_tortured_phrase_projection(signal: dict[str, Any]) -> None:
"""Validate v1.2 invariants that do not require the source corpus entry."""
if signal.get("schema_version") != "bibliographic-integrity-signal/1.2":
return
context = signal["tortured_phrase_context"]
surface = context["surface"]
binding = context["surface_binding"]
snapshot = context["snapshot"]
content_utf8_bytes = binding["content_utf8_bytes"]
if binding["content_sha256"] is not None and (
isinstance(content_utf8_bytes, bool)
or not isinstance(content_utf8_bytes, int)
or content_utf8_bytes < 1
):
raise ValueError(
"present tortured-phrase surface requires positive content_utf8_bytes"
)
citation_key = signal["subject"]["citation_key"]
payload = {
"citation_key": citation_key,
"surface": surface,
"snapshot_sha256": snapshot["snapshot_sha256"],
"content_sha256": binding["content_sha256"],
}
suffix = hashlib.sha256(_canonical_json(payload).encode("utf-8")).hexdigest()[:20]
surface_slug = "title" if surface == "cited_title" else "abstract"
expected_signal_id = f"bis:{citation_key}:tpm_{surface_slug}_{suffix}"
if signal["signal_id"] != expected_signal_id:
raise ValueError("tortured-phrase signal_id binding mismatch")
matches = context["matches"]
counts = context["counts"]
if counts["rule_match_count"] != len(matches):
raise ValueError("tortured-phrase rule_match_count mismatch")
if counts["matched_rule_count"] != len({item["pattern_id"] for item in matches}):
raise ValueError("tortured-phrase matched_rule_count mismatch")
intervals: dict[str, list[tuple[int, int]]] = {}
seen_match_ids: set[str] = set()
artifact_sha = binding["content_sha256"] or hashlib.sha256(b"").hexdigest()
expected_context = surface
expected_disposition = (
"preserve_verbatim_review_context"
if surface == "cited_title"
else "review_cited_source_no_automatic_rewrite"
)
for index, match in enumerate(matches):
if match["segment_id"] != "SEG-000001":
raise ValueError(
f"tortured-phrase matches[{index}] segment_id mismatch"
)
if match["context"] != expected_context:
raise ValueError(f"tortured-phrase matches[{index}] context mismatch")
if match["disposition"] != expected_disposition:
raise ValueError(
f"tortured-phrase matches[{index}] disposition mismatch"
)
if hashlib.sha256(match["matched_text"].encode("utf-8")).hexdigest() != match[
"matched_text_sha256"
]:
raise ValueError(f"tortured-phrase matches[{index}] text hash mismatch")
span = match["source_span"]
codepoint_start = span["codepoint_start"]
codepoint_end = span["codepoint_end"]
utf8_start = span["utf8_start"]
utf8_end = span["utf8_end"]
matched_text = match["matched_text"]
if not codepoint_start < codepoint_end or not utf8_start < utf8_end:
raise ValueError(f"tortured-phrase matches[{index}] has an empty/reversed span")
if codepoint_end - codepoint_start != len(matched_text):
raise ValueError(
f"tortured-phrase matches[{index}] codepoint span length mismatch"
)
if utf8_end - utf8_start != len(matched_text.encode("utf-8")):
raise ValueError(
f"tortured-phrase matches[{index}] UTF-8 span length mismatch"
)
if not (
codepoint_start <= utf8_start <= 4 * codepoint_start
and codepoint_end <= utf8_end <= 4 * codepoint_end
):
raise ValueError(
f"tortured-phrase matches[{index}] has an impossible UTF-8 prefix offset"
)
if (
not isinstance(content_utf8_bytes, int)
or codepoint_end > content_utf8_bytes
or utf8_end > content_utf8_bytes
):
raise ValueError(
f"tortured-phrase matches[{index}] exceeds the bound surface bytes"
)
if len(matched_text) > 1000 or len(matched_text.split()) > 25:
raise ValueError(f"tortured-phrase matches[{index}] exceeds evidence bounds")
match_payload = {
"artifact_sha256": artifact_sha,
"snapshot_sha256": snapshot["snapshot_sha256"],
"surface": surface,
"segment_id": match["segment_id"],
"context": match["context"],
"rule_id": match["pattern_id"],
"codepoint_start": span["codepoint_start"],
"codepoint_end": span["codepoint_end"],
}
expected_match_id = "tpm-" + hashlib.sha256(
_canonical_json(match_payload).encode("utf-8")
).hexdigest()[:24]
if match["match_id"] != expected_match_id:
raise ValueError(f"tortured-phrase matches[{index}] match_id mismatch")
if match["match_id"] in seen_match_ids:
raise ValueError(f"tortured-phrase matches[{index}] duplicates match_id")
seen_match_ids.add(match["match_id"])
intervals.setdefault(match["segment_id"], []).append(
(span["codepoint_start"], span["codepoint_end"])
)
unique_instances = 0
for values in intervals.values():
current_end: int | None = None
for start, end in sorted(set(values)):
if current_end is None or start >= current_end:
unique_instances += 1
current_end = end
else:
current_end = max(current_end, end)
if counts["unique_instance_count"] != unique_instances:
raise ValueError("tortured-phrase unique_instance_count mismatch")
if counts["matched_rule_count"] > counts["rules_evaluated"]:
raise ValueError("tortured-phrase matched_rule_count exceeds rules_evaluated")
matches_by_context = counts["matches_by_context"]
if any(
matches_by_context[name]
!= (len(matches) if name == expected_context else 0)
for name in _TORTURED_PHRASE_CONTEXTS
):
raise ValueError("tortured-phrase matches_by_context mismatch")
status = signal["check_status"]
reason_code = context["reason_code"]
if binding["content_sha256"] is None:
if (
surface != "cited_abstract"
or reason_code not in {"ABSTRACT_MISSING", "ABSTRACT_EMPTY"}
or status != "not_checked"
or signal["finding"] != "unresolved"
or matches
or counts != _empty_tortured_phrase_counts()
):
raise ValueError(
"unbound cited surface must be an explicit missing/empty abstract"
)
elif status == "checked":
if (
reason_code != "CHECK_COMPLETED"
or snapshot["status"] != "loaded"
or snapshot["reason_code"] != "CHECK_COMPLETED"
or counts["rules_evaluated"] != snapshot["rule_count"]
or counts["segments_total"] != 1
or counts["unknown_segments"] != 0
):
raise ValueError(
"checked cited surface lacks complete loaded-snapshot coverage"
)
else:
if matches or counts != _empty_tortured_phrase_counts():
raise ValueError(
"not-checked/degraded cited surface must discard partial match state"
)
if reason_code == "MATCH_RESOURCE_LIMIT":
if (
status != "degraded"
or signal["finding"] != "unresolved"
or snapshot["status"] != "loaded"
or snapshot["reason_code"] != "CHECK_COMPLETED"
):
raise ValueError(
"match resource failure requires a loaded snapshot and degraded output"
)
elif (
reason_code != snapshot["reason_code"]
or status != snapshot["status"]
or signal["finding"] != "unresolved"
):
raise ValueError(
"cited surface status/reason does not replay snapshot availability"
)
evidence = signal["evidence"]
if len(evidence) != 1:
raise ValueError("tortured-phrase signal requires exactly one evidence row")
expected_type = (
"phrase_match"
if signal["check_status"] == "checked" and signal["finding"] == "detected"
else "list_record"
if signal["check_status"] == "checked"
else "degradation_record"
)
expected_value: Any = (
counts["rule_match_count"]
if signal["check_status"] == "checked"
else context["reason_code"]
)
evidence_row = evidence[0]
if evidence_row["evidence_type"] != expected_type:
raise ValueError("tortured-phrase evidence_type mismatch")
if evidence_row.get("record_locator") != (
"title" if surface == "cited_title" else "abstract"
):
raise ValueError("tortured-phrase evidence locator mismatch")
if signal["check_status"] == "checked" and (
isinstance(evidence_row["observed_value"], bool)
or not isinstance(evidence_row["observed_value"], int)
):
raise ValueError("checked tortured-phrase evidence count must be an integer")
if evidence_row["observed_value"] != expected_value:
raise ValueError("tortured-phrase evidence observed_value mismatch")
if evidence_row.get("evidence_sha256") != binding["content_sha256"]:
raise ValueError("tortured-phrase evidence hash mismatch")
provenance = signal["provenance"]
if provenance["source_sha256"] != snapshot["snapshot_sha256"]:
raise ValueError("tortured-phrase provenance snapshot hash mismatch")
snapshot_source = snapshot["source"]
expected_source_name = (
snapshot_source["name"]
if isinstance(snapshot_source, dict)
else "tortured-phrase snapshot unavailable"
)
expected_source_version = (
snapshot_source["version"] if isinstance(snapshot_source, dict) else None
)
if provenance["source_name"] != expected_source_name:
raise ValueError("tortured-phrase provenance source_name mismatch")
if provenance["source_version"] != expected_source_version:
raise ValueError("tortured-phrase provenance source_version mismatch")
if evidence_row["source_name"] != expected_source_name:
raise ValueError("tortured-phrase evidence source_name mismatch")
checked_at = provenance["checked_at"]
recorded_at = provenance["recorded_at"]
if signal["check_status"] == "not_checked":
if checked_at is not None:
raise ValueError("not-checked tortured-phrase row must have null checked_at")
else:
if not isinstance(checked_at, str) or not isinstance(recorded_at, str):
raise ValueError("checked/degraded tortured-phrase row requires checked_at")
if (
_TORTURED_PHRASE_RFC3339_RE.fullmatch(checked_at) is None
or _TORTURED_PHRASE_RFC3339_RE.fullmatch(recorded_at) is None
):
raise ValueError(
"tortured-phrase timestamps require RFC 3339 with at most 6 fraction digits"
)
checked_value = (
checked_at[:-1] + "+00:00"
if checked_at[-1] in {"Z", "z"}
else checked_at
)
recorded_value = (
recorded_at[:-1] + "+00:00"
if recorded_at[-1] in {"Z", "z"}
else recorded_at
)
checked = dt.datetime.fromisoformat(checked_value)
recorded = dt.datetime.fromisoformat(recorded_value)
if recorded < checked:
raise ValueError("tortured-phrase recorded_at precedes checked_at")
def _signal(
*,
citation_key: str,
@@ -283,22 +591,44 @@ def migrate_legacy_entry(
def render_advisory_section(signals: list[dict[str, Any]]) -> str:
"""Compose any number of signals in one provenance-summary section."""
"""Compose one complete, injection-safe canonical provenance summary."""
if not signals:
return ""
def cell(value: Any) -> str:
if value is None or value == "":
return ""
return str(value).replace("|", "\\|").replace("\n", " ")
rendered = str(value).replace("\r", " ").replace("\n", " ")
rendered = "".join(
char
if unicodedata.category(char) not in {"Cc", "Cf", "Zl", "Zp"}
else f"U+{ord(char):04X}"
for char in rendered
)
if len(rendered) > 1000:
rendered = rendered[:999] + ""
rendered = html.escape(rendered, quote=False)
for char in ("\\", "|", "`", "[", "]", "!", "*"):
rendered = rendered.replace(char, "\\" + char)
return rendered
ordered = sorted(signals, key=lambda item: item["signal_id"])
for signal in ordered:
problems = validation_errors(signal)
if problems:
raise ValueError(
f"invalid bibliographic-integrity signal {signal.get('signal_id')!r}: "
+ "; ".join(problems)
)
_validate_tortured_phrase_projection(signal)
lines = [
"## Bibliographic Integrity Advisories",
"",
"| signal_id | signal type | citation | label | status | finding | source | source version | source sha256 | checked at | recorded at | stale after | freshness | source pointer | claims | context |",
"|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|",
"| signal_id | signal type | citation | label | summary label | layer | evaluation status | status | finding | reason code | surface | rule matches | unique instances | evidence | source | source version | snapshot as of | source sha256 | manifest sha256 | checked at | recorded at | stale after | freshness | source pointer | claims | context |",
"|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|",
]
for signal in sorted(signals, key=lambda item: item["signal_id"]):
for signal in ordered:
status = signal["check_status"]
finding = signal["finding"]
if finding == "unresolved" or status in {"not_checked", "unknown", "degraded"}:
@@ -306,6 +636,16 @@ def render_advisory_section(signals: list[dict[str, Any]]) -> str:
claims = ", ".join(signal["subject"]["affected_claims"]) or ""
provenance = signal["provenance"]
context = ""
summary_label = ""
layer = ""
evaluation_status = ""
reason_code = ""
surface = ""
rule_matches: Any = ""
unique_instances: Any = ""
evidence_projection = ""
snapshot_as_of = ""
manifest_sha256 = ""
retraction = signal.get("retraction_context")
if isinstance(retraction, dict):
reasons = ", ".join(retraction.get("retraction_reasons", [])) or "not served"
@@ -320,20 +660,79 @@ def render_advisory_section(signals: list[dict[str, Any]]) -> str:
f"timing={retraction.get('timing_vs_acquisition', 'not_decidable')}; "
f"legitimate_exception={str(bool(legitimate)).lower()}"
)
phrase = signal.get("tortured_phrase_context")
if isinstance(phrase, dict):
summary_label = signal["display"]["summary_label"]
layer = phrase["layer"]
evaluation_status = phrase["evaluation_status"]
reason_code = phrase["reason_code"]
surface = phrase.get("surface", "")
snapshot = phrase["snapshot"]
snapshot_source = snapshot["source"]
if isinstance(snapshot_source, dict):
snapshot_as_of = snapshot_source["as_of"]
manifest_sha256 = snapshot["manifest_sha256"] or ""
counts = phrase.get("counts", {})
rule_matches = counts.get("rule_match_count", "")
unique_instances = counts.get("unique_instance_count", "")
projected: list[str] = []
for match in phrase.get("matches", [])[:3]:
span = match.get("source_span", {})
projected.append(
"{pattern}@{start}:{end}={text} [{disposition}]".format(
pattern=match.get("pattern_id", "?"),
start=span.get("utf8_start", "?"),
end=span.get("utf8_end", "?"),
text=match.get("matched_text", "?"),
disposition=match.get("disposition", "?"),
)
)
omitted = len(phrase.get("matches", [])) - len(projected)
if omitted > 0:
projected.append(f"+{omitted} more machine rows")
evidence_projection = "; ".join(projected) or phrase.get(
"reason_code", ""
)
if finding == "detected":
outcome = "phrase-list match requiring review"
elif finding == "not_detected":
outcome = (
"no phrase-list match observed on the checked surface; "
"absence is not a clean certificate"
)
else:
outcome = "phrase-list screening unresolved; no clean conclusion"
context = (
f"{outcome}; list-match only; origin not inferred; contextual judgment "
"not performed; no automatic rewrite"
)
lines.append(
"| {signal_id} | {signal_type} | {citation} | {label} | {status} | {finding} | "
"{source} | {source_version} | {source_sha256} | {checked_at} | "
"| {signal_id} | {signal_type} | {citation} | {label} | {summary_label} | "
"{layer} | {evaluation_status} | {status} | {finding} | {reason_code} | "
"{surface} | {rule_matches} | {unique_instances} | {evidence} | "
"{source} | {source_version} | {snapshot_as_of} | {source_sha256} | "
"{manifest_sha256} | {checked_at} | "
"{recorded_at} | {stale_after} | {freshness} | {source_pointer} | "
"{claims} | {context} |".format(
signal_id=cell(signal["signal_id"]),
signal_type=cell(signal["signal_type"]),
citation=cell(signal["subject"]["citation_key"]),
label=cell(signal["epistemic_label"]),
summary_label=cell(summary_label),
layer=cell(layer),
evaluation_status=cell(evaluation_status),
status=cell(status),
finding=cell(finding),
reason_code=cell(reason_code),
surface=cell(surface),
rule_matches=cell(rule_matches),
unique_instances=cell(unique_instances),
evidence=cell(evidence_projection),
source=cell(provenance["source_name"]),
source_version=cell(provenance["source_version"]),
snapshot_as_of=cell(snapshot_as_of),
source_sha256=cell(provenance["source_sha256"]),
manifest_sha256=cell(manifest_sha256),
checked_at=cell(provenance["checked_at"]),
recorded_at=cell(provenance["recorded_at"]),
stale_after=cell(provenance["stale_after"]),
@@ -3,16 +3,128 @@
from __future__ import annotations
import argparse
import copy
import hashlib
import importlib.util
import json
import sys
from pathlib import Path
from typing import Any
import yaml
from jsonschema import Draft202012Validator
DEFAULT_REPO_ROOT = Path(__file__).resolve().parent.parent
LEGACY_FIXTURE_SHA256 = {
"retraction.json": "18b84b2988c4460e46c02de135b6b985b51251ca2c86ab33fb0f38b21654e93e",
"retraction_check_attestation.json": "7d8acc7458128e3d58cb43fe7420bd7981d721437f1607de79ca9c2bffd428bf",
"tortured_phrase.json": "162bcb749fb944dfeb9d51b9e81a166a201bebe9a289e93a9a07bec46dee113f",
}
V12_DETECTED_FIXTURE = "tortured_phrase_v1_2_detected.json"
V12_ABSTRACT_MISSING_FIXTURE = "tortured_phrase_v1_2_abstract_missing.json"
V12_FIXTURE_UNICODE_DATA_VERSION = "14.0.0"
CONTEXTS = (
"author_prose",
"quote",
"cited_title",
"reference_entry",
"code_or_verbatim",
"unknown",
"cited_abstract",
)
BOUNDARY = {
"list_match_only": True,
"origin_inference": "not_performed",
"contextual_judgment": "not_performed",
"automatic_rewrite": False,
"absence_is_clean_certificate": False,
"native_pps_compatibility": "not_claimed",
"sharing_scope": "local_only",
}
TERMINAL_LOCK = {
"eligible": False,
"owner": "none",
"policy_key": None,
"current_effect": "advisory_only",
}
def _sha256_bytes(value: bytes) -> str:
return hashlib.sha256(value).hexdigest()
def _sha256_text(value: str) -> str:
return _sha256_bytes(value.encode("utf-8", errors="strict"))
def _canonical_json(value: Any) -> str:
return json.dumps(
value,
ensure_ascii=False,
sort_keys=True,
separators=(",", ":"),
allow_nan=False,
)
def _signal_id(
citation_key: str,
surface: str,
snapshot_sha256: str | None,
content_sha256: str | None,
) -> str:
payload = {
"citation_key": citation_key,
"surface": surface,
"snapshot_sha256": snapshot_sha256,
"content_sha256": content_sha256,
}
suffix = _sha256_text(_canonical_json(payload))[:20]
surface_slug = "title" if surface == "cited_title" else "abstract"
return f"bis:{citation_key}:tpm_{surface_slug}_{suffix}"
def _byte_offsets(value: str) -> list[int]:
offsets = [0]
total = 0
for char in value:
total += len(char.encode("utf-8", errors="strict"))
offsets.append(total)
return offsets
def _unique_instance_count(matches: list[dict[str, Any]]) -> int | None:
by_segment: dict[str, list[tuple[int, int]]] = {}
for match in matches:
if not isinstance(match, dict):
return None
segment_id = match.get("segment_id")
span = match.get("source_span")
if not isinstance(segment_id, str) or not isinstance(span, dict):
return None
start = span.get("codepoint_start")
end = span.get("codepoint_end")
if (
isinstance(start, bool)
or isinstance(end, bool)
or not isinstance(start, int)
or not isinstance(end, int)
):
return None
by_segment.setdefault(segment_id, []).append((start, end))
count = 0
for intervals in by_segment.values():
current_end: int | None = None
for start, end in sorted(set(intervals)):
if current_end is None or start >= current_end:
count += 1
current_end = end
else:
current_end = max(current_end, end)
return count
def _load(path: Path, errors: list[str]) -> dict[str, Any] | None:
try:
@@ -26,6 +138,508 @@ def _load(path: Path, errors: list[str]) -> dict[str, Any] | None:
return value
def _load_fixture_material(
repo_root: Path, errors: list[str]
) -> tuple[
dict[str, Any] | None,
dict[str, Any] | None,
dict[str, dict[str, Any]],
str | None,
str | None,
]:
fixture_root = repo_root / "scripts/fixtures/tortured_phrase_screening"
snapshot_path = fixture_root / "snapshot.json"
manifest_path = fixture_root / "snapshot_manifest.json"
snapshot = _load(snapshot_path, errors)
manifest = _load(manifest_path, errors)
snapshot_sha256: str | None = None
manifest_sha256: str | None = None
try:
snapshot_sha256 = _sha256_bytes(snapshot_path.read_bytes())
except OSError as exc:
errors.append(f"cannot hash {snapshot_path}: {exc}")
try:
manifest_sha256 = _sha256_bytes(manifest_path.read_bytes())
except OSError as exc:
errors.append(f"cannot hash {manifest_path}: {exc}")
if (
manifest is not None
and snapshot_sha256 is not None
and manifest.get("snapshot_sha256") != snapshot_sha256
):
errors.append(
"synthetic snapshot manifest does not bind the exact snapshot fixture bytes"
)
if snapshot is not None and manifest is not None:
rules = snapshot.get("rules")
if not isinstance(rules, list) or manifest.get("rule_count") != len(rules):
errors.append("synthetic snapshot manifest rule_count drifted")
for field in ("snapshot_id", "grammar_profile", "normalizer_profile"):
if snapshot.get(field) != manifest.get(field):
errors.append(
f"synthetic snapshot and manifest disagree on {field}"
)
if manifest.get("snapshot_schema_version") != snapshot.get("schema_version"):
errors.append(
"synthetic snapshot manifest snapshot_schema_version drifted"
)
corpus_path = fixture_root / "corpus_input.yaml"
corpus_by_key: dict[str, dict[str, Any]] = {}
try:
document = yaml.safe_load(corpus_path.read_text(encoding="utf-8"))
except (OSError, yaml.YAMLError) as exc:
errors.append(f"cannot load {corpus_path}: {exc}")
else:
corpus = document.get("literature_corpus") if isinstance(document, dict) else None
if not isinstance(corpus, list):
errors.append(f"{corpus_path}: expected literature_corpus[]")
else:
for index, entry in enumerate(corpus):
citation_key = entry.get("citation_key") if isinstance(entry, dict) else None
if not isinstance(citation_key, str) or not citation_key:
errors.append(
f"{corpus_path}: literature_corpus[{index}] lacks citation_key"
)
continue
if citation_key in corpus_by_key:
errors.append(
f"{corpus_path}: duplicate citation_key {citation_key!r}"
)
continue
corpus_by_key[citation_key] = entry
return snapshot, manifest, corpus_by_key, snapshot_sha256, manifest_sha256
def _check_v12_fixture(
fixture_path: Path,
fixture: dict[str, Any],
*,
snapshot: dict[str, Any] | None,
manifest: dict[str, Any] | None,
corpus_by_key: dict[str, dict[str, Any]],
snapshot_sha256: str | None,
manifest_sha256: str | None,
errors: list[str],
) -> None:
label = str(fixture_path)
def problem(message: str) -> None:
errors.append(f"{label}: {message}")
locks = {
"signal_type": "tortured_phrase_match",
"epistemic_class": "heuristic_advisory",
"epistemic_label": "HEURISTIC-INDICATOR",
}
for field, expected in locks.items():
if fixture.get(field) != expected:
problem(f"v1.2 advisory lock {field} drifted")
if fixture.get("terminal_policy") != TERMINAL_LOCK:
problem("v1.2 tortured-phrase signal escaped the non-terminal lock")
if fixture.get("display") != {
"carrier": "provenance_summary",
"section": "Bibliographic Integrity Advisories",
"summary_label": "Phrase-list screening advisory",
"marker_token": None,
}:
problem("v1.2 advisory display lock drifted")
context = fixture.get("tortured_phrase_context")
if not isinstance(context, dict):
problem("missing tortured_phrase_context object")
return
if context.get("layer") != "HEURISTIC-ADVISORY":
problem("tortured_phrase_context.layer drifted")
if context.get("evaluation_status") != "UNMEASURED":
problem("tortured_phrase_context.evaluation_status drifted")
if context.get("boundary") != BOUNDARY:
problem("tortured_phrase_context advisory boundary drifted")
snapshot_binding = context.get("snapshot")
if not isinstance(snapshot_binding, dict):
problem("missing snapshot binding")
return
if snapshot is not None and manifest is not None:
expected_snapshot = {
"status": "loaded",
"reason_code": "CHECK_COMPLETED",
"snapshot_sha256": snapshot_sha256,
"manifest_sha256": manifest_sha256,
"snapshot_id": manifest.get("snapshot_id"),
"source": manifest.get("source"),
"supply_mode": manifest.get("supply_mode"),
"snapshot_schema_version": manifest.get("snapshot_schema_version"),
"grammar_profile": manifest.get("grammar_profile"),
"normalizer_profile": manifest.get("normalizer_profile"),
"unicode_data_version": V12_FIXTURE_UNICODE_DATA_VERSION,
"rule_count": manifest.get("rule_count"),
"unsupported_rule_count": manifest.get("unsupported_rule_count"),
"rights": manifest.get("rights"),
}
for field, expected in expected_snapshot.items():
if snapshot_binding.get(field) != expected:
problem(f"snapshot binding {field} does not replay fixture provenance")
subject = fixture.get("subject")
if not isinstance(subject, dict):
problem("subject must be an object")
return
citation_key = subject.get("citation_key")
if not isinstance(citation_key, str) or citation_key not in corpus_by_key:
problem("subject citation_key does not join the synthetic corpus fixture")
return
entry = corpus_by_key[citation_key]
if subject.get("source_pointer") != entry.get("source_pointer"):
problem("subject source_pointer does not join the synthetic corpus fixture")
surface = context.get("surface")
if surface == "cited_title":
value = entry.get("title")
locator = "title"
expected_disposition = "preserve_verbatim_review_context"
elif surface == "cited_abstract":
value = entry.get("abstract")
locator = "abstract"
expected_disposition = "review_cited_source_no_automatic_rewrite"
else:
problem("surface is not cited_title or cited_abstract")
return
if value is None or value == "":
expected_content_sha256 = None
expected_content_bytes = None
source_text = ""
elif isinstance(value, str):
raw = value.encode("utf-8", errors="strict")
expected_content_sha256 = _sha256_bytes(raw)
expected_content_bytes = len(raw)
source_text = value
else:
problem("bound synthetic corpus surface is not a string")
return
surface_binding = context.get("surface_binding")
if not isinstance(surface_binding, dict):
problem("surface_binding must be an object")
else:
if surface_binding.get("content_sha256") != expected_content_sha256:
problem("surface content_sha256 is stale")
if surface_binding.get("content_utf8_bytes") != expected_content_bytes:
problem("surface content_utf8_bytes is stale")
expected_id = _signal_id(
citation_key,
surface,
snapshot_binding.get("snapshot_sha256"),
expected_content_sha256,
)
if fixture.get("signal_id") != expected_id:
problem("signal_id does not replay the exact surface/snapshot binding")
provenance = fixture.get("provenance")
if not isinstance(provenance, dict):
problem("provenance must be an object")
else:
source = snapshot_binding.get("source")
source = source if isinstance(source, dict) else {}
expected_provenance = {
"source_name": source.get("name"),
"source_version": source.get("version"),
"source_sha256": snapshot_binding.get("snapshot_sha256"),
"stale_after": None,
"freshness": "unknown",
}
for field, expected in expected_provenance.items():
if provenance.get(field) != expected:
problem(f"provenance {field} drifted from the exact snapshot binding")
counts = context.get("counts")
matches = context.get("matches")
if not isinstance(counts, dict) or not isinstance(matches, list):
problem("counts and matches must be present")
return
context_counts = counts.get("matches_by_context")
if not isinstance(context_counts, dict) or set(context_counts) != set(CONTEXTS):
problem("matches_by_context key set drifted")
context_counts = {}
if counts.get("rule_match_count") != len(matches):
problem("rule_match_count does not equal the complete match array")
pattern_ids = {
match.get("pattern_id")
for match in matches
if isinstance(match, dict) and isinstance(match.get("pattern_id"), str)
}
if counts.get("matched_rule_count") != len(pattern_ids):
problem("matched_rule_count does not equal unique pattern ids")
instance_count = _unique_instance_count(matches)
if counts.get("unique_instance_count") != instance_count:
problem("unique_instance_count does not replay overlap components")
expected_rules_evaluated = (
snapshot_binding.get("rule_count")
if context.get("reason_code") == "CHECK_COMPLETED"
else 0
)
if counts.get("rules_evaluated") != expected_rules_evaluated:
problem("rules_evaluated does not match the completed-check state")
if counts.get("unknown_segments") != 0:
problem("a structured cited surface reported unknown segments")
expected_segments = 0 if expected_content_sha256 is None else 1
if counts.get("segments_total") != expected_segments:
problem("segments_total does not match surface availability")
for context_name in CONTEXTS:
expected_count = len(matches) if context_name == surface else 0
if context_counts.get(context_name) != expected_count:
problem(f"matches_by_context.{context_name} drifted")
rules_by_id: dict[str, dict[str, Any]] = {}
if snapshot is not None and isinstance(snapshot.get("rules"), list):
rules_by_id = {
rule["rule_id"]: rule
for rule in snapshot["rules"]
if isinstance(rule, dict) and isinstance(rule.get("rule_id"), str)
}
offsets = _byte_offsets(source_text)
for index, match in enumerate(matches):
if not isinstance(match, dict):
problem(f"matches[{index}] is not an object")
continue
if match.get("context") != surface:
problem(f"matches[{index}] context escaped its declared surface")
if match.get("disposition") != expected_disposition:
problem(f"matches[{index}] disposition drifted")
span = match.get("source_span")
if not isinstance(span, dict):
problem(f"matches[{index}] lacks source_span")
continue
cp_start = span.get("codepoint_start")
cp_end = span.get("codepoint_end")
if (
isinstance(cp_start, bool)
or isinstance(cp_end, bool)
or not isinstance(cp_start, int)
or not isinstance(cp_end, int)
or not 0 <= cp_start < cp_end <= len(source_text)
):
problem(f"matches[{index}] codepoint span is stale")
continue
matched_text = source_text[cp_start:cp_end]
if span.get("utf8_start") != offsets[cp_start] or span.get(
"utf8_end"
) != offsets[cp_end]:
problem(f"matches[{index}] UTF-8 span is stale")
if match.get("matched_text") != matched_text:
problem(f"matches[{index}] exact text is stale")
if match.get("matched_text_sha256") != _sha256_text(matched_text):
problem(f"matches[{index}] exact text hash is stale")
rule = rules_by_id.get(match.get("pattern_id"))
expected_pattern_sha = (
_sha256_text(_canonical_json(rule)) if rule is not None else None
)
if match.get("pattern_sha256") != expected_pattern_sha:
problem(f"matches[{index}] pattern hash does not replay the snapshot")
match_key = {
"artifact_sha256": expected_content_sha256,
"snapshot_sha256": snapshot_binding.get("snapshot_sha256"),
"surface": surface,
"segment_id": match.get("segment_id"),
"context": match.get("context"),
"rule_id": match.get("pattern_id"),
"codepoint_start": cp_start,
"codepoint_end": cp_end,
}
expected_match_id = "tpm-" + _sha256_text(_canonical_json(match_key))[:24]
if match.get("match_id") != expected_match_id:
problem(f"matches[{index}] match_id binding drifted")
evidence = fixture.get("evidence")
if not isinstance(evidence, list) or len(evidence) != 1 or not isinstance(
evidence[0], dict
):
problem("v1.2 tortured-phrase carrier must contain one evidence row")
evidence_row: dict[str, Any] = {}
else:
evidence_row = evidence[0]
source = snapshot_binding.get("source")
expected_source_name = (
source.get("name")
if isinstance(source, dict)
else "tortured-phrase snapshot unavailable"
)
if evidence_row.get("source_name") != expected_source_name:
problem("evidence source_name drifted from snapshot provenance")
if evidence_row.get("record_locator") != locator:
problem("evidence record_locator drifted from the declared surface")
if evidence_row.get("evidence_sha256") != expected_content_sha256:
problem("evidence_sha256 drifted from the exact surface binding")
reason_code = context.get("reason_code")
if reason_code == "CHECK_COMPLETED":
expected_finding = "detected" if matches else "not_detected"
expected_evidence_type = "phrase_match" if matches else "list_record"
if expected_content_sha256 is None:
problem("CHECK_COMPLETED cannot describe an absent cited surface")
if fixture.get("check_status") != "checked":
problem("CHECK_COMPLETED carrier is not checked")
if fixture.get("finding") != expected_finding:
problem("CHECK_COMPLETED finding does not follow match count")
if evidence_row.get("evidence_type") != expected_evidence_type:
problem("checked evidence_type does not follow finding")
if evidence_row.get("observed_value") != len(matches):
problem("checked evidence observed_value does not equal match count")
if not isinstance(provenance, dict) or provenance.get("checked_at") is None:
problem("checked carrier lacks checked_at")
elif reason_code == "ABSTRACT_MISSING":
if surface != "cited_abstract" or expected_content_sha256 is not None:
problem("ABSTRACT_MISSING is not bound to an absent abstract")
if fixture.get("check_status") != "not_checked" or fixture.get(
"finding"
) != "unresolved":
problem("ABSTRACT_MISSING was promoted to a clean result")
if matches or any(
counts.get(field) != 0
for field in (
"matched_rule_count",
"rule_match_count",
"unique_instance_count",
"segments_total",
"unknown_segments",
)
):
problem("ABSTRACT_MISSING carries fabricated match or segment counts")
if evidence_row.get("evidence_type") != "degradation_record" or evidence_row.get(
"observed_value"
) != "ABSTRACT_MISSING":
problem("ABSTRACT_MISSING evidence row drifted")
if not isinstance(provenance, dict) or provenance.get("checked_at") is not None:
problem("ABSTRACT_MISSING must not claim checked_at")
else:
problem("synthetic v1.2 fixture has an unexpected reason_code")
def _check_renderer(
repo_root: Path,
*,
legacy_fixture: dict[str, Any] | None,
detected_fixture: dict[str, Any] | None,
abstract_missing_fixture: dict[str, Any] | None,
errors: list[str],
) -> None:
runtime_path = repo_root / "scripts/bibliographic_integrity_signals.py"
try:
spec = importlib.util.spec_from_file_location(
"_bibliographic_integrity_renderer_guard", runtime_path
)
if spec is None or spec.loader is None:
raise ImportError("could not create module spec")
runtime = importlib.util.module_from_spec(spec)
spec.loader.exec_module(runtime)
except (OSError, ImportError, ModuleNotFoundError) as exc:
errors.append(f"cannot load bounded advisory renderer {runtime_path}: {exc}")
return
if (
legacy_fixture is None
or detected_fixture is None
or abstract_missing_fixture is None
):
return
malicious = "fixture://safe|`[link](https://invalid)\n<!--ref:INJECTED-->"
rows: list[dict[str, Any]] = []
for index in range(26):
row = copy.deepcopy(legacy_fixture)
row["signal_id"] = f"bis:render{index:02d}:retraction_status"
row["subject"]["citation_key"] = f"render{index:02d}"
row["subject"]["source_pointer"] = malicious if index == 0 else None
rows.append(row)
try:
rendered = runtime.render_advisory_section(rows)
except (OSError, TypeError, ValueError, KeyError) as exc:
errors.append(f"bounded advisory renderer probe failed: {exc}")
else:
if rendered.count("## Bibliographic Integrity Advisories") != 1:
errors.append("advisory renderer no longer emits exactly one section")
if "<!--ref:INJECTED-->" in rendered or "[link](https://invalid)" in rendered:
errors.append("advisory renderer emitted unescaped injected markup")
if "CONTAMINATED-" in rendered:
errors.append("advisory renderer minted a terminal marker token")
if "bis:render25:retraction_status" not in rendered:
errors.append("canonical advisory renderer dropped a complete signal row")
phrase = copy.deepcopy(detected_fixture)
context = phrase.get("tortured_phrase_context")
if not isinstance(context, dict) or not context.get("matches"):
errors.append("detected v1.2 fixture cannot probe bounded match rendering")
return
seed_match = context["matches"][0]
projected_matches: list[dict[str, Any]] = []
for index in range(4):
match = copy.deepcopy(seed_match)
match["pattern_id"] = f"render_pattern_{index}"
span = match["source_span"]
match_payload = {
"artifact_sha256": context["surface_binding"]["content_sha256"],
"snapshot_sha256": context["snapshot"]["snapshot_sha256"],
"surface": context["surface"],
"segment_id": match["segment_id"],
"context": match["context"],
"rule_id": match["pattern_id"],
"codepoint_start": span["codepoint_start"],
"codepoint_end": span["codepoint_end"],
}
match["match_id"] = "tpm-" + _sha256_text(
_canonical_json(match_payload)
)[:24]
projected_matches.append(match)
context["matches"] = projected_matches
context["counts"]["rule_match_count"] = 4
context["counts"]["matched_rule_count"] = 4
context["counts"]["matches_by_context"][context["surface"]] = 4
phrase["evidence"][0]["observed_value"] = 4
try:
phrase_rendered = runtime.render_advisory_section([phrase])
except (OSError, TypeError, ValueError, KeyError) as exc:
errors.append(f"bounded phrase-match renderer probe failed: {exc}")
else:
if "+1 more machine rows" not in phrase_rendered:
errors.append("phrase evidence projection exceeded its three-match detail cap")
if "phrase-list match requiring review" not in phrase_rendered:
errors.append("detected phrase renderer outcome wording drifted")
no_match = copy.deepcopy(detected_fixture)
no_match_context = no_match["tortured_phrase_context"]
no_match["finding"] = "not_detected"
no_match_context["matches"] = []
no_match_context["counts"]["matched_rule_count"] = 0
no_match_context["counts"]["rule_match_count"] = 0
no_match_context["counts"]["unique_instance_count"] = 0
no_match_context["counts"]["matches_by_context"][
no_match_context["surface"]
] = 0
no_match["evidence"][0]["evidence_type"] = "list_record"
no_match["evidence"][0]["observed_value"] = 0
try:
no_match_rendered = runtime.render_advisory_section([no_match])
except (OSError, TypeError, ValueError, KeyError) as exc:
errors.append(f"zero-match renderer probe failed: {exc}")
else:
if (
"no phrase-list match observed on the checked surface" not in no_match_rendered
or "absence is not a clean certificate" not in no_match_rendered
):
errors.append("zero-match renderer clean-claim boundary drifted")
try:
unresolved_rendered = runtime.render_advisory_section(
[abstract_missing_fixture]
)
except (OSError, TypeError, ValueError, KeyError) as exc:
errors.append(f"unresolved renderer probe failed: {exc}")
else:
if (
"phrase-list screening unresolved; no clean conclusion"
not in unresolved_rendered
):
errors.append("unresolved renderer clean-claim boundary drifted")
def run_checks(repo_root: Path) -> list[str]:
errors: list[str] = []
schema_path = (
@@ -64,6 +678,7 @@ def run_checks(repo_root: Path) -> list[str]:
fixtures_dir = repo_root / "scripts/fixtures/bibliographic_integrity_signals"
fixtures: list[dict[str, Any]] = []
fixtures_by_name: dict[str, dict[str, Any]] = {}
for fixture_path in sorted(fixtures_dir.glob("*.json")):
fixture = _load(fixture_path, errors)
if fixture is None:
@@ -74,6 +689,16 @@ def run_checks(repo_root: Path) -> list[str]:
if json.loads(json.dumps(fixture, sort_keys=True)) != fixture:
errors.append(f"{fixture_path}: JSON round-trip changed the fixture")
fixtures.append(fixture)
fixtures_by_name[fixture_path.name] = fixture
for name, expected_sha256 in LEGACY_FIXTURE_SHA256.items():
path = fixtures_dir / name
try:
actual_sha256 = _sha256_bytes(path.read_bytes())
except OSError as exc:
errors.append(f"cannot hash legacy fixture {path}: {exc}")
continue
if actual_sha256 != expected_sha256:
errors.append(f"{path}: legacy fixture byte identity drifted")
fixture_types = {fixture.get("signal_type") for fixture in fixtures}
if not {"retraction_status", "tortured_phrase_match"}.issubset(fixture_types):
errors.append("fixtures must cover both #651 retraction and #660 tortured phrase")
@@ -91,7 +716,54 @@ def run_checks(repo_root: Path) -> list[str]:
if labels.get("deterministic_fact") == labels.get("heuristic_advisory"):
errors.append("deterministic facts and heuristics share an epistemic label")
if any(fixture.get("display", {}).get("marker_token") is not None for fixture in fixtures):
errors.append("v1 fixture minted a ref-marker advisory token")
errors.append("fixture minted a ref-marker advisory token")
snapshot, manifest, corpus_by_key, snapshot_sha256, manifest_sha256 = (
_load_fixture_material(repo_root, errors)
)
required_v12 = {
V12_DETECTED_FIXTURE: ("checked", "detected", "CHECK_COMPLETED"),
V12_ABSTRACT_MISSING_FIXTURE: (
"not_checked",
"unresolved",
"ABSTRACT_MISSING",
),
}
for name, expected_state in required_v12.items():
fixture = fixtures_by_name.get(name)
if fixture is None:
errors.append(f"missing required v1.2 tortured-phrase fixture {name}")
continue
context = fixture.get("tortured_phrase_context")
actual_state = (
fixture.get("check_status"),
fixture.get("finding"),
context.get("reason_code") if isinstance(context, dict) else None,
)
if actual_state != expected_state:
errors.append(f"{fixtures_dir / name}: required v1.2 state drifted")
for name, fixture in fixtures_by_name.items():
if fixture.get("schema_version") != "bibliographic-integrity-signal/1.2":
continue
_check_v12_fixture(
fixtures_dir / name,
fixture,
snapshot=snapshot,
manifest=manifest,
corpus_by_key=corpus_by_key,
snapshot_sha256=snapshot_sha256,
manifest_sha256=manifest_sha256,
errors=errors,
)
_check_renderer(
repo_root,
legacy_fixture=fixtures_by_name.get("retraction.json"),
detected_fixture=fixtures_by_name.get(V12_DETECTED_FIXTURE),
abstract_missing_fixture=fixtures_by_name.get(
V12_ABSTRACT_MISSING_FIXTURE
),
errors=errors,
)
entry_schema = _load(
repo_root / "shared/contracts/passport/literature_corpus_entry.schema.json",
+102 -7
View File
@@ -37,13 +37,22 @@ except ImportError as e:
)
sys.exit(2)
# Dual-path import (mirrors arxiv_client.py): the v3.10 laundering guard
# (#329) lives in check_v3_10_policy and is wired here so it runs over REAL
# passport entries, not just fixtures.
try:
# Dual-path imports: package mode must stay anchored to this repository's
# ``scripts`` namespace so an earlier PYTHONPATH entry cannot shadow either
# policy dependency. Direct-script mode has no package context and therefore
# uses the sibling modules exposed by the script directory on sys.path.
if __package__:
from .check_v3_10_policy import assert_venue_type_source_clean
from .tortured_phrase_screening import (
ScreeningError as TorturedPhraseScreeningError,
validate_cited_signal_binding,
)
else:
from check_v3_10_policy import assert_venue_type_source_clean
except ImportError:
from scripts.check_v3_10_policy import assert_venue_type_source_clean
from tortured_phrase_screening import (
ScreeningError as TorturedPhraseScreeningError,
validate_cited_signal_binding,
)
REPO_ROOT = Path(__file__).resolve().parent.parent
ENTRY_SCHEMA_PATH = REPO_ROOT / "shared/contracts/passport/literature_corpus_entry.schema.json"
@@ -162,8 +171,10 @@ def validate_passport(
signals = entry.get("bibliographic_integrity_signals", [])
if isinstance(signals, list):
signal_ids: dict[str, int] = {}
current_phrase_surfaces: dict[str, list[dict]] = {}
for signal_i, signal in enumerate(signals):
for err in signal_validator.iter_errors(signal):
signal_problems = list(signal_validator.iter_errors(signal))
for err in signal_problems:
errors.append(
f"{path}: literature_corpus[{i}]."
f"bibliographic_integrity_signals[{signal_i}] "
@@ -188,6 +199,30 @@ def validate_passport(
f"targets citation_key {signal_citation!r}, not "
f"its containing entry {entry_citation!r}"
)
if (
not signal_problems
and signal.get("signal_type")
== "tortured_phrase_match"
):
try:
validate_cited_signal_binding(signal, entry)
except TorturedPhraseScreeningError as exc:
errors.append(
f"{path}: literature_corpus[{i}]."
f"bibliographic_integrity_signals[{signal_i}] "
f"tortured-phrase binding error: {exc}"
)
context = signal.get("tortured_phrase_context")
if (
signal.get("schema_version")
== "bibliographic-integrity-signal/1.2"
and isinstance(context, dict)
and isinstance(context.get("surface"), str)
):
surface = context["surface"]
current_phrase_surfaces.setdefault(surface, []).append(
signal
)
for signal_id, count in signal_ids.items():
if count > 1:
errors.append(
@@ -195,6 +230,66 @@ def validate_passport(
"bibliographic_integrity_signals: duplicate "
f"signal_id {signal_id!r} appears {count} times"
)
for surface, rows in current_phrase_surfaces.items():
if len(rows) > 1:
errors.append(
f"{path}: literature_corpus[{i}]."
"bibliographic_integrity_signals: multiple current "
f"tortured-phrase rows for surface {surface!r}"
)
if current_phrase_surfaces:
for required_surface in (
"cited_title",
"cited_abstract",
):
if not current_phrase_surfaces.get(required_surface):
errors.append(
f"{path}: literature_corpus[{i}]."
"bibliographic_integrity_signals: current "
"tortured-phrase v1.2 rows require exactly one "
f"{required_surface!r} row; none was found"
)
title_rows = current_phrase_surfaces.get("cited_title", [])
abstract_rows = current_phrase_surfaces.get(
"cited_abstract", []
)
if len(title_rows) == len(abstract_rows) == 1:
title_row = title_rows[0]
abstract_row = abstract_rows[0]
title_context = title_row["tortured_phrase_context"]
abstract_context = abstract_row["tortured_phrase_context"]
if title_context["snapshot"] != abstract_context["snapshot"]:
errors.append(
f"{path}: literature_corpus[{i}]."
"bibliographic_integrity_signals: current "
"tortured-phrase title/abstract rows must bind the "
"same snapshot and detached manifest"
)
title_provenance = title_row["provenance"]
abstract_provenance = abstract_row["provenance"]
if (
title_provenance["recorded_at"]
!= abstract_provenance["recorded_at"]
):
errors.append(
f"{path}: literature_corpus[{i}]."
"bibliographic_integrity_signals: current "
"tortured-phrase title/abstract rows must come "
"from the same recorded run"
)
title_checked = title_provenance["checked_at"]
abstract_checked = abstract_provenance["checked_at"]
if (
title_checked is not None
and abstract_checked is not None
and title_checked != abstract_checked
):
errors.append(
f"{path}: literature_corpus[{i}]."
"bibliographic_integrity_signals: checked "
"tortured-phrase title/abstract rows must share "
"the exact checked_at value"
)
vts = entry.get("venue_type_source", "")
vtp = entry.get("venue_type_provenance", "")
if isinstance(vts, str) and isinstance(vtp, str):
+2 -2
View File
@@ -70,8 +70,8 @@ REPO_ROOT = Path(__file__).resolve().parent.parent
# reviewed against the #528 resolutions.
# ---------------------------------------------------------------------------
CONTENT_LOCKS = {
"academic-pipeline/SKILL.md": "2ec9302ce877f5126d2ef0d69ef306009e478a0c0dadc2343b18c15692b13bba",
"academic-pipeline/agents/pipeline_orchestrator_agent.md": "a808d34ea88aee360b1970e0894270d56ef81a0842c7c885f9930ec51bd42dc7",
"academic-pipeline/SKILL.md": "dfb1829a7389fb5b727ce75027a7e4da287db761ae0cb32a4b85bd365532be64",
"academic-pipeline/agents/pipeline_orchestrator_agent.md": "60b9a55f62836a78a55e85c00d987800881de29e64b10857182257b3982006e1",
"academic-pipeline/agents/state_tracker_agent.md": "1222cf0ca75646cb5e0a5e27fba344e15f45ea45df4ed5ce10301cbf5bf265d8",
"academic-pipeline/references/pipeline_state_machine.md": "34d6baa2af3dddc0e8692584195405153f5fcfe0f4794786a6c19b959b16e404",
"academic-pipeline/references/process_summary_protocol.md": "5c7053230d73b39d0a5d9d6f5e9f339c12570ae6d3aa2eae2eaf74f51d571e94",
File diff suppressed because it is too large Load Diff
@@ -39,7 +39,7 @@ REPO_ROOT = Path(__file__).resolve().parents[1]
SCHEMAS = REPO_ROOT / "shared/contracts/passport"
BIBLIOGRAPHY_AGENT_PATH = REPO_ROOT / "deep-research/agents/bibliography_agent.md"
BIBLIOGRAPHY_AGENT_SHA256 = "1e4f4bed354fdacdf11d36cda4f69b135477af876f77f88e3a9104971d3d4933" # #651 baseline: added DOI-keyed retraction-status production only; NO temporal/M6/M5 logic touched, ownership invariant intact. Previous accepted #548/#511 baseline recorded the Last Searched and omission-provenance additions under the same rule.
BIBLIOGRAPHY_AGENT_SHA256 = "19fdf3cc299d230a4cec44ab3a5bee2030f6293c4b60be9e8a9355cdd05adb93" # #660 baseline: added local, non-retrieving tortured-phrase metadata enrichment only; NO temporal/M6/M5 logic touched, ownership invariant intact. Previous accepted #651/#548/#511 additions remain covered by the same rule.
def _validate(yaml_path: Path, schema_path: Path) -> list[str]:
@@ -0,0 +1,107 @@
{
"schema_version": "bibliographic-integrity-signal/1.2",
"signal_id": "bis:fixture_missing_2026:tpm_abstract_cdbd5c77f3219d7822d3",
"signal_type": "tortured_phrase_match",
"epistemic_class": "heuristic_advisory",
"epistemic_label": "HEURISTIC-INDICATOR",
"check_status": "not_checked",
"finding": "unresolved",
"evidence": [
{
"evidence_type": "degradation_record",
"source_name": "ARS synthetic matcher conformance fixture",
"record_locator": "abstract",
"observed_value": "ABSTRACT_MISSING",
"evidence_sha256": null
}
],
"provenance": {
"source_name": "ARS synthetic matcher conformance fixture",
"source_version": "synthetic/1.0",
"source_sha256": "962879909bfdd338047dc4569ff42188c4200fcd56a0a459b0e0f9169f2446c4",
"checked_at": null,
"recorded_at": "2026-08-10T01:00:01Z",
"stale_after": null,
"freshness": "unknown"
},
"subject": {
"citation_key": "fixture_missing_2026",
"source_pointer": "fixture://tortured-phrase/missing-abstract",
"affected_claims": []
},
"terminal_policy": {
"eligible": false,
"owner": "none",
"policy_key": null,
"current_effect": "advisory_only"
},
"display": {
"carrier": "provenance_summary",
"section": "Bibliographic Integrity Advisories",
"summary_label": "Phrase-list screening advisory",
"marker_token": null
},
"tortured_phrase_context": {
"layer": "HEURISTIC-ADVISORY",
"evaluation_status": "UNMEASURED",
"surface": "cited_abstract",
"surface_binding": {
"content_sha256": null,
"content_utf8_bytes": null
},
"snapshot": {
"status": "loaded",
"reason_code": "CHECK_COMPLETED",
"snapshot_sha256": "962879909bfdd338047dc4569ff42188c4200fcd56a0a459b0e0f9169f2446c4",
"manifest_sha256": "34bfa9f92a612cd7e794024dded98ed96b0968dbc0f2c924d68cafbe3f280924",
"snapshot_id": "ars-synthetic-conformance-2026-08-10",
"source": {
"name": "ARS synthetic matcher conformance fixture",
"version": "synthetic/1.0",
"as_of": "2026-08-10",
"locator": "fixture://tortured-phrase/snapshot"
},
"supply_mode": "synthetic_fixture",
"snapshot_schema_version": "tortured-phrase-snapshot/1.0",
"grammar_profile": "ars-tortured-phrase-canonical-ast/1.0",
"normalizer_profile": "ars-nfkc-casefold-token/1.0",
"unicode_data_version": "14.0.0",
"rule_count": 17,
"unsupported_rule_count": 0,
"rights": {
"basis": "synthetic_fixture",
"redistribution_status": "permitted",
"reference": null,
"user_declaration": null
}
},
"reason_code": "ABSTRACT_MISSING",
"counts": {
"rules_evaluated": 0,
"matched_rule_count": 0,
"rule_match_count": 0,
"unique_instance_count": 0,
"segments_total": 0,
"unknown_segments": 0,
"matches_by_context": {
"author_prose": 0,
"quote": 0,
"cited_title": 0,
"reference_entry": 0,
"code_or_verbatim": 0,
"unknown": 0,
"cited_abstract": 0
}
},
"matches": [],
"boundary": {
"list_match_only": true,
"origin_inference": "not_performed",
"contextual_judgment": "not_performed",
"automatic_rewrite": false,
"absence_is_clean_certificate": false,
"native_pps_compatibility": "not_claimed",
"sharing_scope": "local_only"
}
}
}
@@ -0,0 +1,124 @@
{
"schema_version": "bibliographic-integrity-signal/1.2",
"signal_id": "bis:fixture_complete_2026:tpm_title_bd2b4979a30f74c7fd4a",
"signal_type": "tortured_phrase_match",
"epistemic_class": "heuristic_advisory",
"epistemic_label": "HEURISTIC-INDICATOR",
"check_status": "checked",
"finding": "detected",
"evidence": [
{
"evidence_type": "phrase_match",
"source_name": "ARS synthetic matcher conformance fixture",
"record_locator": "title",
"observed_value": 1,
"evidence_sha256": "ac8cea1f2d462d7b9d80ed20a9be321dd2b928a4d3672ed8b716357f41c9f278"
}
],
"provenance": {
"source_name": "ARS synthetic matcher conformance fixture",
"source_version": "synthetic/1.0",
"source_sha256": "962879909bfdd338047dc4569ff42188c4200fcd56a0a459b0e0f9169f2446c4",
"checked_at": "2026-08-10T01:00:00Z",
"recorded_at": "2026-08-10T01:00:01Z",
"stale_after": null,
"freshness": "unknown"
},
"subject": {
"citation_key": "fixture_complete_2026",
"source_pointer": "fixture://tortured-phrase/complete",
"affected_claims": []
},
"terminal_policy": {
"eligible": false,
"owner": "none",
"policy_key": null,
"current_effect": "advisory_only"
},
"display": {
"carrier": "provenance_summary",
"section": "Bibliographic Integrity Advisories",
"summary_label": "Phrase-list screening advisory",
"marker_token": null
},
"tortured_phrase_context": {
"layer": "HEURISTIC-ADVISORY",
"evaluation_status": "UNMEASURED",
"surface": "cited_title",
"surface_binding": {
"content_sha256": "ac8cea1f2d462d7b9d80ed20a9be321dd2b928a4d3672ed8b716357f41c9f278",
"content_utf8_bytes": 41
},
"snapshot": {
"status": "loaded",
"reason_code": "CHECK_COMPLETED",
"snapshot_sha256": "962879909bfdd338047dc4569ff42188c4200fcd56a0a459b0e0f9169f2446c4",
"manifest_sha256": "34bfa9f92a612cd7e794024dded98ed96b0968dbc0f2c924d68cafbe3f280924",
"snapshot_id": "ars-synthetic-conformance-2026-08-10",
"source": {
"name": "ARS synthetic matcher conformance fixture",
"version": "synthetic/1.0",
"as_of": "2026-08-10",
"locator": "fixture://tortured-phrase/snapshot"
},
"supply_mode": "synthetic_fixture",
"snapshot_schema_version": "tortured-phrase-snapshot/1.0",
"grammar_profile": "ars-tortured-phrase-canonical-ast/1.0",
"normalizer_profile": "ars-nfkc-casefold-token/1.0",
"unicode_data_version": "14.0.0",
"rule_count": 17,
"unsupported_rule_count": 0,
"rights": {
"basis": "synthetic_fixture",
"redistribution_status": "permitted",
"reference": null,
"user_declaration": null
}
},
"reason_code": "CHECK_COMPLETED",
"counts": {
"rules_evaluated": 17,
"matched_rule_count": 1,
"rule_match_count": 1,
"unique_instance_count": 1,
"segments_total": 1,
"unknown_segments": 0,
"matches_by_context": {
"author_prose": 0,
"quote": 0,
"cited_title": 1,
"reference_entry": 0,
"code_or_verbatim": 0,
"unknown": 0,
"cited_abstract": 0
}
},
"matches": [
{
"match_id": "tpm-04fd3e70898ad5db821c74aa",
"pattern_id": "syn_context_reference",
"pattern_sha256": "2d027178f5c464a221e29031a8a681d7de16ed35371d0d8a2aaadc59dc114230",
"segment_id": "SEG-000001",
"context": "cited_title",
"disposition": "preserve_verbatim_review_context",
"source_span": {
"codepoint_start": 0,
"codepoint_end": 18,
"utf8_start": 0,
"utf8_end": 18
},
"matched_text": "Cited Turnip Atlas",
"matched_text_sha256": "d4376a7af4908abd3f1b9b7af94af9d75d6bf30da78ee43590e12bd9e15e8dff"
}
],
"boundary": {
"list_match_only": true,
"origin_inference": "not_performed",
"contextual_judgment": "not_performed",
"automatic_rewrite": false,
"absence_is_clean_certificate": false,
"native_pps_compatibility": "not_claimed",
"sharing_scope": "local_only"
}
}
}
@@ -0,0 +1,30 @@
literature_corpus:
- citation_key: fixture_complete_2026
title: Cited Turnip Atlas for Imaginary Orchards
authors:
- family: Fixture
given: Ada
year: 2026
source_pointer: fixture://tortured-phrase/complete
obtained_via: manual
obtained_at: '2026-08-10T00:00:00Z'
abstract: An abstract saffron circuit appears in this wholly synthetic summary.
- citation_key: fixture_missing_2026
title: Missing Indigo Compass in a Paper Orchard
authors:
- family: Fixture
given: Bea
year: 2026
source_pointer: fixture://tortured-phrase/missing-abstract
obtained_via: manual
obtained_at: '2026-08-10T00:00:00Z'
- citation_key: fixture_negative_2026
title: Orbital Greenhouses and Ordinary Seed Catalogues
authors:
- family: Fixture
given: Cy
year: 2026
source_pointer: fixture://tortured-phrase/legitimate-negative
obtained_via: manual
obtained_at: '2026-08-10T00:00:00Z'
abstract: This recipe glossary explains why the harmless moonlit pickle is expected here.
@@ -0,0 +1,36 @@
# Synthetic Orchard Notes
This repository-authored conformance draft contains a luminous turnip in ordinary author prose.
The velvet algorithm shares this single paragraph with the meadow protocol for a conjunction case.
A quartz llama appears as the first of two disjunctive alternatives.
The ordered window places a nebula pebble cabbage sequence in one sentence.
This recipe glossary records the harmless moonlit pickle used by the fixture.
One span contains the prism otter engine so that two synthetic rules overlap.
The orbital greenhouse is a legitimate token-boundary negative.
Compatibility text writes with compatibility characters and capitals.
A same-line separator divides SILVER-LEAF into two tokens.
The format-character case is crystal­fern inside author prose.
The line-break case is silver-
leaf inside the same paragraph.
> An attributed quotation preserves the quoted acorn engine and a moonlit pickle verbatim.
The inline example `coded radish loop` is mechanically visible but remains code context.
```text
coded radish loop
```
## References
Fixture, A. (2099). *Cited turnip atlas for imaginary orchards*. Synthetic Fixture Press.
@@ -0,0 +1,21 @@
\documentclass{article}
\begin{document}
Ordinary author prose mentions a luminous turnip once.
The same-line TeX dash case is SILVER-LEAF.
\begin{quote}
An attributed quotation preserves the quoted acorn engine verbatim.
\end{quote}
\begin{verbatim}
coded radish loop
\end{verbatim}
\begin{thebibliography}{1}
\bibitem{synthetic2099}
A. Fixture, \emph{Cited turnip atlas for imaginary orchards}, Synthetic Fixture Press, 2099.
\end{thebibliography}
\end{document}
@@ -0,0 +1,218 @@
{
"schema_version": "tortured-phrase-seed-expectations/1.0",
"snapshot_id": "ars-synthetic-conformance-2026-08-10",
"snapshot_sha256": "962879909bfdd338047dc4569ff42188c4200fcd56a0a459b0e0f9169f2446c4",
"label_scope": "mechanical_matcher_conformance_only",
"empirical_accuracy_claimed": false,
"contextual_false_positive_labels_provided": false,
"contextual_false_negative_labels_provided": false,
"inputs": [
{
"path": "own_draft.md",
"expected_counts": {
"matched_rule_count": 14,
"rule_match_count": 15,
"unique_instance_count": 14
},
"expected_matches": [
{
"rule_id": "syn_literal_author",
"count": 1,
"context": "author_prose"
},
{
"rule_id": "syn_all_segment",
"count": 1,
"context": "author_prose"
},
{
"rule_id": "syn_any_segment",
"count": 1,
"context": "author_prose",
"matched_alternative_index": 0
},
{
"rule_id": "syn_near_ordered",
"count": 1,
"context": "author_prose"
},
{
"rule_id": "syn_exclude_segment",
"count": 1,
"context": "quote"
},
{
"rule_id": "syn_overlap_long",
"count": 1,
"context": "author_prose",
"overlap_group": "md-overlap-1"
},
{
"rule_id": "syn_overlap_short",
"count": 1,
"context": "author_prose",
"overlap_group": "md-overlap-1"
},
{
"rule_id": "syn_nfkc_casefold",
"count": 1,
"context": "author_prose",
"normalization_case": "nfkc_casefold"
},
{
"rule_id": "syn_dash_same_line",
"count": 1,
"context": "author_prose",
"normalization_case": "same_line_dash_separator"
},
{
"rule_id": "syn_soft_hyphen",
"count": 1,
"context": "author_prose",
"normalization_case": "soft_hyphen_join"
},
{
"rule_id": "syn_hyphen_line_break",
"count": 1,
"context": "author_prose",
"normalization_case": "line_break_hyphen_join"
},
{
"rule_id": "syn_context_quote",
"count": 1,
"context": "quote"
},
{
"rule_id": "syn_context_code",
"count": 2,
"context": "code_or_verbatim"
},
{
"rule_id": "syn_context_reference",
"count": 1,
"context": "reference_entry"
}
],
"expected_suppressions": [
{
"rule_id": "syn_exclude_segment",
"count": 1,
"context": "author_prose",
"reason": "exclude_if_within_token_window"
}
],
"expected_non_matches": [
{
"rule_id": "syn_boundary_orb",
"reason": "token_boundary"
}
]
},
{
"path": "own_draft.tex",
"expected_counts": {
"matched_rule_count": 5,
"rule_match_count": 5,
"unique_instance_count": 5
},
"expected_matches": [
{
"rule_id": "syn_literal_author",
"count": 1,
"context": "author_prose"
},
{
"rule_id": "syn_dash_same_line",
"count": 1,
"context": "author_prose",
"normalization_case": "same_line_dash_separator"
},
{
"rule_id": "syn_context_quote",
"count": 1,
"context": "quote"
},
{
"rule_id": "syn_context_code",
"count": 1,
"context": "code_or_verbatim"
},
{
"rule_id": "syn_context_reference",
"count": 1,
"context": "reference_entry"
}
],
"expected_suppressions": [],
"expected_non_matches": []
},
{
"path": "corpus_input.yaml",
"records": [
{
"citation_key": "fixture_complete_2026",
"surfaces": {
"title": {
"check_status": "checked",
"reason": "CHECK_COMPLETED",
"expected_rule_ids": [
"syn_context_reference"
]
},
"abstract": {
"check_status": "checked",
"reason": "CHECK_COMPLETED",
"expected_rule_ids": [
"syn_corpus_abstract"
]
}
}
},
{
"citation_key": "fixture_missing_2026",
"surfaces": {
"title": {
"check_status": "checked",
"reason": "CHECK_COMPLETED",
"expected_rule_ids": [
"syn_corpus_missing_abstract_title"
]
},
"abstract": {
"check_status": "not_checked",
"reason": "ABSTRACT_MISSING",
"expected_rule_ids": []
}
}
},
{
"citation_key": "fixture_negative_2026",
"surfaces": {
"title": {
"check_status": "checked",
"reason": "CHECK_COMPLETED",
"expected_rule_ids": [],
"non_match_reasons": [
{
"rule_id": "syn_boundary_orb",
"reason": "token_boundary"
}
]
},
"abstract": {
"check_status": "checked",
"reason": "CHECK_COMPLETED",
"expected_rule_ids": [],
"non_match_reasons": [
{
"rule_id": "syn_exclude_segment",
"reason": "exclude_if_in_same_segment"
}
]
}
}
}
]
}
]
}
@@ -0,0 +1,164 @@
{
"schema_version": "tortured-phrase-snapshot/1.0",
"snapshot_id": "ars-synthetic-conformance-2026-08-10",
"grammar_profile": "ars-tortured-phrase-canonical-ast/1.0",
"normalizer_profile": "ars-nfkc-casefold-token/1.0",
"rules": [
{
"rule_id": "syn_literal_author",
"expression": {
"op": "literal",
"value": "luminous turnip"
}
},
{
"rule_id": "syn_all_segment",
"expression": {
"op": "all",
"terms": [
{
"op": "literal",
"value": "velvet algorithm"
},
{
"op": "literal",
"value": "meadow protocol"
}
],
"max_span_tokens": 12
}
},
{
"rule_id": "syn_any_segment",
"expression": {
"op": "any",
"alternatives": [
{
"op": "literal",
"value": "quartz llama"
},
{
"op": "literal",
"value": "cobalt radish"
}
]
}
},
{
"rule_id": "syn_near_ordered",
"expression": {
"op": "near",
"left": {
"op": "literal",
"value": "nebula"
},
"right": {
"op": "literal",
"value": "cabbage"
},
"max_gap_tokens": 1,
"ordered": true
}
},
{
"rule_id": "syn_exclude_segment",
"expression": {
"op": "literal",
"value": "moonlit pickle"
},
"exclude_if": [
{
"expression": {
"op": "literal",
"value": "recipe glossary"
},
"within_tokens": 8
}
]
},
{
"rule_id": "syn_overlap_long",
"expression": {
"op": "literal",
"value": "prism otter engine"
}
},
{
"rule_id": "syn_overlap_short",
"expression": {
"op": "literal",
"value": "otter engine"
}
},
{
"rule_id": "syn_boundary_orb",
"expression": {
"op": "literal",
"value": "orb"
}
},
{
"rule_id": "syn_nfkc_casefold",
"expression": {
"op": "literal",
"value": "fullwidth comet"
}
},
{
"rule_id": "syn_dash_same_line",
"expression": {
"op": "literal",
"value": "silver leaf"
}
},
{
"rule_id": "syn_soft_hyphen",
"expression": {
"op": "literal",
"value": "crystalfern"
}
},
{
"rule_id": "syn_hyphen_line_break",
"expression": {
"op": "literal",
"value": "silverleaf"
}
},
{
"rule_id": "syn_context_quote",
"expression": {
"op": "literal",
"value": "quoted acorn engine"
}
},
{
"rule_id": "syn_context_reference",
"expression": {
"op": "literal",
"value": "cited turnip atlas"
}
},
{
"rule_id": "syn_context_code",
"expression": {
"op": "literal",
"value": "coded radish loop"
}
},
{
"rule_id": "syn_corpus_abstract",
"expression": {
"op": "literal",
"value": "abstract saffron circuit"
}
},
{
"rule_id": "syn_corpus_missing_abstract_title",
"expression": {
"op": "literal",
"value": "missing indigo compass"
}
}
]
}
@@ -0,0 +1,29 @@
{
"schema_version": "tortured-phrase-snapshot-manifest/1.0",
"snapshot_id": "ars-synthetic-conformance-2026-08-10",
"source": {
"name": "ARS synthetic matcher conformance fixture",
"version": "synthetic/1.0",
"as_of": "2026-08-10",
"locator": "fixture://tortured-phrase/snapshot"
},
"supply_mode": "synthetic_fixture",
"snapshot_schema_version": "tortured-phrase-snapshot/1.0",
"snapshot_sha256": "962879909bfdd338047dc4569ff42188c4200fcd56a0a459b0e0f9169f2446c4",
"grammar_profile": "ars-tortured-phrase-canonical-ast/1.0",
"normalizer_profile": "ars-nfkc-casefold-token/1.0",
"preprocessor": {
"name": "ars-synthetic-fixture-authoring",
"version": "1.0",
"native_grammar": "synthetic-canonical-ast/1.0",
"reduction_notes": []
},
"unsupported_rule_count": 0,
"rule_count": 17,
"rights": {
"basis": "synthetic_fixture",
"redistribution_status": "permitted",
"reference": null,
"user_declaration": null
}
}
@@ -2,12 +2,14 @@
from __future__ import annotations
import copy
import hashlib
import json
import shutil
import sys
from pathlib import Path
import pytest
import yaml
from jsonschema import Draft202012Validator
@@ -17,6 +19,7 @@ sys.path.insert(0, str(SCRIPTS))
import bibliographic_integrity_signals as signals # noqa: E402
import check_bibliographic_integrity_signals as checker # noqa: E402
import tortured_phrase_screening as screening # noqa: E402
def _fixture(name: str) -> dict:
@@ -30,17 +33,91 @@ def _validator() -> Draft202012Validator:
)
def _screening_state() -> screening.SnapshotState:
fixture_root = SCRIPTS / "fixtures/tortured_phrase_screening"
bundle = screening.load_snapshot(
fixture_root / "snapshot.json", fixture_root / "snapshot_manifest.json"
)
return screening.SnapshotState(
status="loaded",
reason_code="CHECK_COMPLETED",
bundle=bundle,
snapshot_sha256=bundle.snapshot_sha256,
manifest_sha256=bundle.manifest_sha256,
detail=None,
)
def _rendered_row(signal: dict) -> dict[str, str]:
lines = signals.render_advisory_section([signal]).splitlines()
headings = lines[2].strip("|").split(" | ")
values = lines[4].strip("|").split(" | ")
assert len(values) == len(headings)
return dict(zip(headings, values, strict=True))
def test_all_epistemic_class_fixtures_round_trip() -> None:
for name in (
"retraction.json",
"retraction_check_attestation.json",
"tortured_phrase.json",
"tortured_phrase_v1_2_detected.json",
"tortured_phrase_v1_2_abstract_missing.json",
):
fixture = _fixture(name)
assert not list(_validator().iter_errors(fixture))
assert json.loads(json.dumps(fixture)) == fixture
def test_v1_2_fixtures_replay_except_frozen_unicode_provenance() -> None:
corpus_path = SCRIPTS / "fixtures/tortured_phrase_screening/corpus_input.yaml"
corpus = yaml.safe_load(corpus_path.read_text(encoding="utf-8"))[
"literature_corpus"
]
by_key = {entry["citation_key"]: entry for entry in corpus}
cases = (
(
"tortured_phrase_v1_2_detected.json",
"fixture_complete_2026",
"cited_title",
),
(
"tortured_phrase_v1_2_abstract_missing.json",
"fixture_missing_2026",
"cited_abstract",
),
)
state = _screening_state()
for fixture_name, citation_key, surface in cases:
expected = screening.build_cited_signal(
by_key[citation_key],
surface=surface,
state=state,
checked_at="2026-08-10T01:00:00Z",
recorded_at="2026-08-10T01:00:01Z",
)
fixture = _fixture(fixture_name)
runtime_version = expected["tortured_phrase_context"]["snapshot"][
"unicode_data_version"
]
assert runtime_version == state.bundle.unicode_data_version
fixture_version = fixture["tortured_phrase_context"]["snapshot"][
"unicode_data_version"
]
assert fixture_version == checker.V12_FIXTURE_UNICODE_DATA_VERSION
expected["tortured_phrase_context"]["snapshot"][
"unicode_data_version"
] = fixture_version
assert fixture == expected
screening.validate_cited_signal_binding(fixture, by_key[citation_key])
def test_legacy_fixture_byte_identity_is_frozen() -> None:
fixture_root = SCRIPTS / "fixtures/bibliographic_integrity_signals"
for name, expected_sha256 in checker.LEGACY_FIXTURE_SHA256.items():
assert hashlib.sha256((fixture_root / name).read_bytes()).hexdigest() == expected_sha256
def test_closed_schema_rejects_undeclared_field() -> None:
fixture = _fixture("retraction.json")
fixture["undeclared"] = True
@@ -106,6 +183,251 @@ def test_multiple_signals_compose_in_one_section_without_marker_token() -> None:
assert heading in rendered
def test_v1_2_renderer_projects_complete_neutral_metadata_for_all_states() -> None:
corpus = yaml.safe_load(
(SCRIPTS / "fixtures/tortured_phrase_screening/corpus_input.yaml").read_text(
encoding="utf-8"
)
)["literature_corpus"]
by_key = {entry["citation_key"]: entry for entry in corpus}
loaded = _screening_state()
detected = _fixture("tortured_phrase_v1_2_detected.json")
zero = screening.build_cited_signal(
by_key["fixture_negative_2026"],
surface="cited_title",
state=loaded,
checked_at="2026-08-10T01:00:00Z",
recorded_at="2026-08-10T01:00:01Z",
)
missing = _fixture("tortured_phrase_v1_2_abstract_missing.json")
degraded = screening.build_cited_signal(
by_key["fixture_negative_2026"],
surface="cited_title",
state=screening.SnapshotState(
status="degraded",
reason_code="SNAPSHOT_HASH_MISMATCH",
bundle=None,
snapshot_sha256="0" * 64,
manifest_sha256="1" * 64,
detail="synthetic degraded renderer probe",
),
checked_at="2026-08-10T01:00:00Z",
recorded_at="2026-08-10T01:00:01Z",
)
loaded_manifest = "34bfa9f92a612cd7e794024dded98ed96b0968dbc0f2c924d68cafbe3f280924"
shared = {
"summary label": "Phrase-list screening advisory",
"layer": "HEURISTIC-ADVISORY",
"evaluation status": "UNMEASURED",
}
expectations = (
(
detected,
{
**shared,
"status": "checked",
"finding": "detected",
"reason code": "CHECK_COMPLETED",
"snapshot as of": "2026-08-10",
"manifest sha256": loaded_manifest,
},
),
(
zero,
{
**shared,
"status": "checked",
"finding": "not_detected",
"reason code": "CHECK_COMPLETED",
"snapshot as of": "2026-08-10",
"manifest sha256": loaded_manifest,
},
),
(
missing,
{
**shared,
"status": "not_checked",
"finding": "NOT CLEAN — UNRESOLVED",
"reason code": "ABSTRACT_MISSING",
"snapshot as of": "2026-08-10",
"manifest sha256": loaded_manifest,
},
),
(
degraded,
{
**shared,
"status": "degraded",
"finding": "NOT CLEAN — UNRESOLVED",
"reason code": "SNAPSHOT_HASH_MISMATCH",
"snapshot as of": "",
"manifest sha256": "1" * 64,
},
),
)
for signal, expected in expectations:
row = _rendered_row(signal)
assert {key: row[key] for key in expected} == expected
for signal in (zero, missing, degraded):
assert "phrase-list match requiring review" not in signals.render_advisory_section(
[signal]
)
legacy = _rendered_row(_fixture("retraction.json"))
for heading in (
"summary label",
"layer",
"evaluation status",
"reason code",
"snapshot as of",
"manifest sha256",
):
assert legacy[heading] == ""
def test_renderer_is_complete_injection_safe_and_bounds_match_projection() -> None:
malicious = "fixture://safe|`[link](https://invalid)\n<!--ref:INJECTED-->"
rows = []
for index in range(26):
row = copy.deepcopy(_fixture("retraction.json"))
row["signal_id"] = f"bis:render{index:02d}:retraction_status"
row["subject"]["citation_key"] = f"render{index:02d}"
row["subject"]["source_pointer"] = malicious if index == 0 else None
rows.append(row)
rendered = signals.render_advisory_section(rows)
assert rendered.count("## Bibliographic Integrity Advisories") == 1
assert "bis:render25:retraction_status" in rendered
assert "<!--ref:INJECTED-->" not in rendered
assert "[link](https://invalid)" not in rendered
assert "CONTAMINATED-" not in rendered
phrase = copy.deepcopy(_fixture("tortured_phrase_v1_2_detected.json"))
context = phrase["tortured_phrase_context"]
seed_match = context["matches"][0]
projected_matches = []
for index in range(4):
match = copy.deepcopy(seed_match)
match["pattern_id"] = f"render_pattern_{index}"
span = match["source_span"]
payload = {
"artifact_sha256": context["surface_binding"]["content_sha256"],
"snapshot_sha256": context["snapshot"]["snapshot_sha256"],
"surface": context["surface"],
"segment_id": match["segment_id"],
"context": match["context"],
"rule_id": match["pattern_id"],
"codepoint_start": span["codepoint_start"],
"codepoint_end": span["codepoint_end"],
}
match["match_id"] = "tpm-" + hashlib.sha256(
checker._canonical_json(payload).encode("utf-8")
).hexdigest()[:24]
projected_matches.append(match)
context["matches"] = projected_matches
context["counts"]["rule_match_count"] = 4
context["counts"]["matched_rule_count"] = 4
context["counts"]["matches_by_context"]["cited_title"] = 4
phrase["evidence"][0]["observed_value"] = 4
phrase_rendered = signals.render_advisory_section([phrase])
assert phrase_rendered.count("render_pattern_") == 3
assert "+1 more machine rows" in phrase_rendered
assert "phrase-list match requiring review" in phrase_rendered
injected = copy.deepcopy(_fixture("tortured_phrase_v1_2_detected.json"))
injected_match = injected["tortured_phrase_context"]["matches"][0]
malicious_match = "![x](y)|`<tag>`!!!"
assert len(malicious_match) == len(injected_match["matched_text"])
injected_match["matched_text"] = malicious_match
injected_match["matched_text_sha256"] = hashlib.sha256(
malicious_match.encode("utf-8")
).hexdigest()
injected_rendered = signals.render_advisory_section([injected])
assert malicious_match not in injected_rendered
assert "<tag>" not in injected_rendered
assert "&lt;tag&gt;" in injected_rendered
assert "\\|" in injected_rendered
assert "\\`" in injected_rendered
assert "\\[x\\]" in injected_rendered
corpus_path = SCRIPTS / "fixtures/tortured_phrase_screening/corpus_input.yaml"
negative_entry = yaml.safe_load(corpus_path.read_text(encoding="utf-8"))[
"literature_corpus"
][2]
no_match = screening.build_cited_signal(
negative_entry,
surface="cited_title",
state=_screening_state(),
checked_at="2026-08-10T01:00:00Z",
recorded_at="2026-08-10T01:00:01Z",
)
no_match_rendered = signals.render_advisory_section([no_match])
assert "no phrase-list match observed on the checked surface" in no_match_rendered
assert "absence is not a clean certificate" in no_match_rendered
unresolved = _fixture("tortured_phrase_v1_2_abstract_missing.json")
unresolved_rendered = signals.render_advisory_section([unresolved])
assert "phrase-list screening unresolved; no clean conclusion" in unresolved_rendered
@pytest.mark.parametrize(
"mutation",
[
"reversed_codepoint",
"zero_utf8_span",
"codepoint_length",
"too_many_words",
"outside_surface",
"impossible_utf8_prefix",
"zero_surface_bytes",
"matched_rules_exceed_evaluated",
"submicrosecond_order",
],
)
def test_source_independent_phrase_projection_mutations_fail(mutation: str) -> None:
phrase = copy.deepcopy(_fixture("tortured_phrase_v1_2_detected.json"))
context = phrase["tortured_phrase_context"]
match = context["matches"][0]
span = match["source_span"]
if mutation == "reversed_codepoint":
span["codepoint_end"] = span["codepoint_start"]
elif mutation == "zero_utf8_span":
span["utf8_end"] = span["utf8_start"]
elif mutation == "codepoint_length":
span["codepoint_end"] += 1
elif mutation == "too_many_words":
text = " ".join(["word"] * 26)
match["matched_text"] = text
match["matched_text_sha256"] = hashlib.sha256(text.encode()).hexdigest()
span["codepoint_start"] = 0
span["codepoint_end"] = len(text)
span["utf8_start"] = 0
span["utf8_end"] = len(text.encode())
elif mutation == "outside_surface":
width = span["codepoint_end"] - span["codepoint_start"]
start = context["surface_binding"]["content_utf8_bytes"] + 1
span["codepoint_start"] = start
span["codepoint_end"] = start + width
span["utf8_start"] = start
span["utf8_end"] = start + len(match["matched_text"].encode())
elif mutation == "impossible_utf8_prefix":
span["utf8_start"] += 1
span["utf8_end"] += 1
elif mutation == "zero_surface_bytes":
context["surface_binding"]["content_utf8_bytes"] = 0
elif mutation == "submicrosecond_order":
phrase["provenance"]["checked_at"] = "2026-08-10T01:00:00.0000009Z"
phrase["provenance"]["recorded_at"] = "2026-08-10T01:00:00.0000001Z"
else:
context["counts"]["rules_evaluated"] = 0
if mutation == "submicrosecond_order":
assert signals.validation_errors(phrase)
with pytest.raises(ValueError):
signals._validate_tortured_phrase_projection(phrase)
def test_checked_but_unresolved_attestation_is_visibly_not_clean() -> None:
rendered = signals.render_advisory_section(
[_fixture("retraction_check_attestation.json")]
@@ -122,9 +444,15 @@ def _minimal_repo(tmp_path: Path) -> Path:
"shared/handoff_schemas.md",
"academic-pipeline/agents/pipeline_orchestrator_agent.md",
"academic-paper/agents/formatter_agent.md",
"scripts/bibliographic_integrity_signals.py",
"scripts/fixtures/bibliographic_integrity_signals/retraction.json",
"scripts/fixtures/bibliographic_integrity_signals/retraction_check_attestation.json",
"scripts/fixtures/bibliographic_integrity_signals/tortured_phrase.json",
"scripts/fixtures/bibliographic_integrity_signals/tortured_phrase_v1_2_detected.json",
"scripts/fixtures/bibliographic_integrity_signals/tortured_phrase_v1_2_abstract_missing.json",
"scripts/fixtures/tortured_phrase_screening/corpus_input.yaml",
"scripts/fixtures/tortured_phrase_screening/snapshot.json",
"scripts/fixtures/tortured_phrase_screening/snapshot_manifest.json",
]
for relative in paths:
source = REPO_ROOT / relative
@@ -149,6 +477,85 @@ def test_sync_checker_detects_formatter_drift(tmp_path: Path) -> None:
assert any("formatter_agent.md" in error for error in found)
def test_sync_checker_detects_legacy_fixture_byte_drift(tmp_path: Path) -> None:
root = _minimal_repo(tmp_path)
path = root / "scripts/fixtures/bibliographic_integrity_signals/retraction.json"
path.write_bytes(path.read_bytes() + b"\n")
found = checker.run_checks(root)
assert any("legacy fixture byte identity drifted" in error for error in found)
@pytest.mark.parametrize(
("mutation", "expected_error"),
[
("deterministic_fact", "v1.2 advisory lock epistemic_class drifted"),
("terminal", "escaped the non-terminal lock"),
("stale", "surface content_sha256 is stale"),
("hash", "evidence_sha256 drifted"),
("count", "rule_match_count does not equal the complete match array"),
("provenance", "provenance stale_after drifted"),
("unicode_version", "unicode_data_version does not replay fixture provenance"),
("missing_abstract", "CHECK_COMPLETED cannot describe an absent cited surface"),
],
)
def test_sync_checker_rejects_v1_2_mutations(
tmp_path: Path, mutation: str, expected_error: str
) -> None:
root = _minimal_repo(tmp_path)
fixtures = root / "scripts/fixtures/bibliographic_integrity_signals"
detected_path = fixtures / "tortured_phrase_v1_2_detected.json"
missing_path = fixtures / "tortured_phrase_v1_2_abstract_missing.json"
if mutation == "stale":
corpus_path = root / "scripts/fixtures/tortured_phrase_screening/corpus_input.yaml"
corpus = yaml.safe_load(corpus_path.read_text(encoding="utf-8"))
corpus["literature_corpus"][0]["title"] += " stale"
corpus_path.write_text(
yaml.safe_dump(corpus, sort_keys=False, allow_unicode=True),
encoding="utf-8",
)
elif mutation == "missing_abstract":
fixture = json.loads(missing_path.read_text(encoding="utf-8"))
fixture["check_status"] = "checked"
fixture["finding"] = "not_detected"
fixture["provenance"]["checked_at"] = "2026-08-10T01:00:00Z"
fixture["evidence"][0]["evidence_type"] = "list_record"
fixture["evidence"][0]["observed_value"] = 0
fixture["tortured_phrase_context"]["reason_code"] = "CHECK_COMPLETED"
missing_path.write_text(
json.dumps(fixture, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
else:
fixture = json.loads(detected_path.read_text(encoding="utf-8"))
if mutation == "deterministic_fact":
fixture["epistemic_class"] = "deterministic_fact"
fixture["epistemic_label"] = "RESOLVER-OR-LIST-OBSERVATION"
elif mutation == "terminal":
fixture["terminal_policy"] = {
"eligible": True,
"owner": "citation_finalizer",
"policy_key": "tortured_phrase",
"current_effect": "policy_gated",
}
elif mutation == "count":
fixture["tortured_phrase_context"]["counts"]["rule_match_count"] += 1
elif mutation == "provenance":
fixture["provenance"]["stale_after"] = "2026-08-11T01:00:00Z"
fixture["provenance"]["freshness"] = "stale"
elif mutation == "unicode_version":
fixture["tortured_phrase_context"]["snapshot"][
"unicode_data_version"
] = "15.0.0"
else:
fixture["evidence"][0]["evidence_sha256"] = "0" * 64
detected_path.write_text(
json.dumps(fixture, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
found = checker.run_checks(root)
assert any(expected_error in error for error in found), found
def test_existing_canonical_record_wins_idempotently() -> None:
fixture = _fixture("retraction.json")
entry = {
@@ -0,0 +1,964 @@
"""Mutation tests for the hermetic #660 integration guard."""
from __future__ import annotations
import ast
import json
import shutil
import sys
from pathlib import Path
from typing import Any, Callable
import pytest
SCRIPTS = Path(__file__).resolve().parent
REPO_ROOT = SCRIPTS.parent
sys.path.insert(0, str(SCRIPTS))
import check_tortured_phrase_screening_integration as integration # noqa: E402
COPIED_FILES = (
integration.SNAPSHOT_SCHEMA,
integration.MANIFEST_SCHEMA,
integration.ADVISORY_SCHEMA,
integration.SIGNAL_SCHEMA,
integration.RUNTIME,
integration.RUNTIME_TEST,
Path("scripts/check_tortured_phrase_screening_integration.py"),
Path("scripts/test_check_tortured_phrase_screening_integration.py"),
integration.SNAPSHOT_FIXTURE,
integration.MANIFEST_FIXTURE,
integration.EXPECTATIONS_FIXTURE,
integration.CITED_DETECTED_FIXTURE,
integration.CITED_MISSING_FIXTURE,
integration.FIXTURE_ROOT / "own_draft.md",
integration.FIXTURE_ROOT / "own_draft.tex",
integration.FIXTURE_ROOT / "corpus_input.yaml",
integration.DESIGN_SPEC,
integration.SIGNAL_PROTOCOL,
integration.CONTRACT_README,
integration.INTEGRITY_PROTOCOL,
integration.CORPUS_CONSUMER_PROTOCOL,
integration.BIBLIOGRAPHY_AGENT,
integration.PIPELINE_SKILL,
integration.PIPELINE_ORCHESTRATOR,
integration.FORMATTER_AGENT,
integration.HANDOFF_SCHEMAS,
integration.SIGNAL_RUNTIME,
integration.CORPUS_CHECKER,
integration.CI_MANIFEST,
integration.WORKFLOW,
integration.CONFORMANCE_README,
integration.MEASUREMENT_PLAN,
integration.SUITE_REGISTRY,
integration.MEASUREMENT_CONTRACT,
)
def _copy_repo(tmp_path: Path) -> Path:
root = tmp_path / "repo"
for relative in COPIED_FILES:
target = root / relative
target.parent.mkdir(parents=True, exist_ok=True)
shutil.copy2(REPO_ROOT / relative, target)
return root
def _rewrite_json(
root: Path, relative: Path, mutation: Callable[[dict[str, Any]], None]
) -> None:
path = root / relative
value = json.loads(path.read_text(encoding="utf-8"))
mutation(value)
path.write_text(
json.dumps(value, ensure_ascii=False, indent=2, allow_nan=False) + "\n",
encoding="utf-8",
)
def _replace(root: Path, relative: Path, old: str, new: str) -> None:
path = root / relative
text = path.read_text(encoding="utf-8")
assert old in text
path.write_text(text.replace(old, new, 1), encoding="utf-8")
def _replace_nth(
root: Path, relative: Path, old: str, new: str, occurrence: int
) -> None:
path = root / relative
text = path.read_text(encoding="utf-8")
start = -len(old)
for _ in range(occurrence):
start = text.find(old, start + len(old))
assert start >= 0
path.write_text(text[:start] + new + text[start + len(old) :], encoding="utf-8")
def _assert_error(root: Path, expected: str) -> None:
errors = integration.run_checks(root)
assert any(expected in error for error in errors), errors
def test_repository_integration_is_currently_green() -> None:
assert integration.run_checks(REPO_ROOT) == []
@pytest.mark.parametrize(
"relative",
(
integration.RUNTIME,
integration.SIGNAL_RUNTIME,
integration.CORPUS_CHECKER,
),
)
def test_reviewed_runtime_byte_drift_fails(tmp_path: Path, relative: Path) -> None:
root = _copy_repo(tmp_path)
path = root / relative
path.write_bytes(path.read_bytes() + b"\n# frozen-byte drift\n")
_assert_error(root, "frozen runtime sha256 drifted")
def test_closed_snapshot_schema_operator_drift_fails(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(schema: dict[str, Any]) -> None:
all_expression = schema["$defs"]["all_expression"]
all_expression["required"][1] = "operands"
all_expression["properties"]["operands"] = all_expression["properties"].pop(
"terms"
)
_rewrite_json(root, integration.SNAPSHOT_SCHEMA, mutate)
_assert_error(root, "$defs.all_expression properties drifted")
def test_rule_exclusion_shape_drift_fails(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(schema: dict[str, Any]) -> None:
schema["$defs"]["rule"]["properties"]["exclude_if"]["maxItems"] = 9
_rewrite_json(root, integration.SNAPSHOT_SCHEMA, mutate)
_assert_error(root, "rule.exclude_if must contain 1..8")
def test_manifest_shape_drift_fails(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(schema: dict[str, Any]) -> None:
source = schema["$defs"]["source"]
source["required"].remove("locator")
source["properties"].pop("locator")
_rewrite_json(root, integration.MANIFEST_SCHEMA, mutate)
_assert_error(root, "$defs.source: required field set drifted")
@pytest.mark.parametrize(
("field", "value"),
[("layer", "DETERMINISTIC"), ("evaluation_status", "MEASURED")],
)
def test_own_advisory_claim_const_drift_fails(
tmp_path: Path, field: str, value: str
) -> None:
root = _copy_repo(tmp_path)
def mutate(schema: dict[str, Any]) -> None:
schema["properties"][field]["const"] = value
_rewrite_json(root, integration.ADVISORY_SCHEMA, mutate)
_assert_error(root, f"{field} must equal")
def test_own_advisory_schema_requires_document_empty(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(schema: dict[str, Any]) -> None:
schema["$defs"]["reason_code"]["enum"].remove("DOCUMENT_EMPTY")
_rewrite_json(root, integration.ADVISORY_SCHEMA, mutate)
_assert_error(root, "reason_code must include DOCUMENT_EMPTY")
def test_cited_schema_requires_abstract_empty(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(schema: dict[str, Any]) -> None:
schema["properties"]["tortured_phrase_context"]["properties"]["reason_code"][
"enum"
].remove("ABSTRACT_EMPTY")
_rewrite_json(root, integration.SIGNAL_SCHEMA, mutate)
_assert_error(root, "reason_code must include ABSTRACT_EMPTY")
def test_abstract_empty_must_remain_not_checked_unresolved(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(schema: dict[str, Any]) -> None:
branch = next(
item
for item in schema["allOf"]
if "ABSTRACT_EMPTY"
in item.get("if", {})
.get("properties", {})
.get("tortured_phrase_context", {})
.get("properties", {})
.get("reason_code", {})
.get("enum", ())
)
branch["if"]["properties"]["tortured_phrase_context"]["properties"][
"reason_code"
]["enum"].remove("ABSTRACT_EMPTY")
_rewrite_json(root, integration.SIGNAL_SCHEMA, mutate)
_assert_error(root, "ABSTRACT_EMPTY must map to not_checked/unresolved")
def test_v12_signal_epistemic_profile_drift_fails(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(schema: dict[str, Any]) -> None:
branch = next(
item
for item in schema["allOf"]
if item.get("if", {})
.get("properties", {})
.get("schema_version", {})
.get("const")
== integration.SIGNAL_VERSION
)
branch["then"]["properties"]["epistemic_label"]["const"] = (
"RESOLVER-OR-LIST-OBSERVATION"
)
_rewrite_json(root, integration.SIGNAL_SCHEMA, mutate)
_assert_error(root, "v1.2 tortured-phrase profile drifted")
def test_v12_signal_neutral_label_drift_fails(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(schema: dict[str, Any]) -> None:
branch = next(
item
for item in schema["allOf"]
if item.get("if", {})
.get("properties", {})
.get("schema_version", {})
.get("const")
== integration.SIGNAL_VERSION
)
branch["then"]["properties"]["display"]["properties"]["summary_label"][
"const"
] = "Origin finding"
_rewrite_json(root, integration.SIGNAL_SCHEMA, mutate)
_assert_error(root, "v1.2 tortured-phrase profile drifted")
def test_signal_terminal_promotion_fails(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(schema: dict[str, Any]) -> None:
branch = next(
item
for item in schema["allOf"]
if item.get("if", {})
.get("properties", {})
.get("signal_type", {})
.get("const")
== "tortured_phrase_match"
)
branch["then"]["properties"]["terminal_policy"]["properties"]["eligible"] = {
"const": True
}
_rewrite_json(root, integration.SIGNAL_SCHEMA, mutate)
_assert_error(root, "heuristic/advisory-only policy boundary drifted")
def test_strict_json_duplicate_schema_key_fails(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
path = root / integration.SNAPSHOT_SCHEMA
text = path.read_text(encoding="utf-8")
path.write_text(
text.replace(
'"title": "Tortured-phrase canonical snapshot",',
'"title": "Tortured-phrase canonical snapshot",\n "title": "duplicate",',
1,
),
encoding="utf-8",
)
_assert_error(root, "duplicate JSON key")
def test_fixture_snapshot_exact_byte_hash_drift_fails(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
path = root / integration.SNAPSHOT_FIXTURE
path.write_bytes(path.read_bytes() + b" ")
_assert_error(root, "snapshot_sha256 does not bind exact bytes")
def test_fixture_must_remain_synthetic_only(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(manifest: dict[str, Any]) -> None:
manifest["supply_mode"] = "user_supplied"
_rewrite_json(root, integration.MANIFEST_FIXTURE, mutate)
_assert_error(root, "synthetic-only provenance/rights boundary drifted")
def test_cited_fixture_must_bind_the_detached_manifest_bytes(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(fixture: dict[str, Any]) -> None:
fixture["tortured_phrase_context"]["snapshot"]["manifest_sha256"] = "0" * 64
_rewrite_json(root, integration.CITED_DETECTED_FIXTURE, mutate)
_assert_error(root, "cited v1.2 fixtures: carrier/replay/claim binding drifted")
@pytest.mark.parametrize(
"field",
(
"empirical_accuracy_claimed",
"contextual_false_positive_labels_provided",
"contextual_false_negative_labels_provided",
),
)
def test_seed_cannot_claim_contextual_measurement(tmp_path: Path, field: str) -> None:
root = _copy_repo(tmp_path)
def mutate(expectations: dict[str, Any]) -> None:
expectations[field] = True
_rewrite_json(root, integration.EXPECTATIONS_FIXTURE, mutate)
_assert_error(root, "synthetic UNMEASURED claim ceiling drifted")
@pytest.mark.parametrize(
"module", ("requests", "openai", "subprocess", "urllib.request")
)
def test_runtime_forbids_network_model_and_process_imports(
tmp_path: Path, module: str
) -> None:
root = _copy_repo(tmp_path)
path = root / integration.RUNTIME
path.write_text(
path.read_text(encoding="utf-8") + f"\nimport {module}\n",
encoding="utf-8",
)
_assert_error(root, f"forbidden model/network/process import {module}")
def test_runtime_forbids_process_calls_without_subprocess_import(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
path = root / integration.RUNTIME
path.write_text(
path.read_text(encoding="utf-8")
+ "\n\ndef _process_mutation():\n return os.system('forbidden')\n",
encoding="utf-8",
)
_assert_error(root, "forbidden model/network/process call os.system")
@pytest.mark.parametrize(
"body",
(
"return dt.datetime.now()",
"return dt.datetime.utcnow()",
"import time\n return time.time()",
),
)
def test_runtime_forbids_ambient_clock(tmp_path: Path, body: str) -> None:
root = _copy_repo(tmp_path)
path = root / integration.RUNTIME
path.write_text(
path.read_text(encoding="utf-8") + f"\n\ndef _ambient_clock_mutation():\n {body}\n",
encoding="utf-8",
)
_assert_error(root, "forbidden ambient clock call")
def test_runtime_forbids_source_file_time(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
path = root / integration.RUNTIME
path.write_text(
path.read_text(encoding="utf-8")
+ "\n\ndef _file_time_mutation(path):\n return path.stat().st_mtime\n",
encoding="utf-8",
)
_assert_error(root, "forbidden ambient file-time call")
def test_runtime_forbids_ambient_environment(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
path = root / integration.RUNTIME
path.write_text(
path.read_text(encoding="utf-8")
+ "\n\ndef _ambient_environment_mutation():\n return os.environ.get('API_KEY')\n",
encoding="utf-8",
)
_assert_error(root, "ambient environment access is forbidden")
def test_runtime_forbids_ambient_file_discovery(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
path = root / integration.RUNTIME
path.write_text(
path.read_text(encoding="utf-8")
+ "\n\ndef _ambient_scan_mutation(path):\n return path.rglob('*.json')\n",
encoding="utf-8",
)
_assert_error(root, "forbidden ambient discovery call")
@pytest.mark.parametrize(
("body", "expected"),
(
("import requests\n", "forbidden transitive model/network/process import"),
(
"\n\ndef _clock_mutation():\n return dt.datetime.now()\n",
"forbidden transitive ambient clock call",
),
(
"\nimport os\n\ndef _environment_mutation():\n"
" return os.environ.get('TOKEN')\n",
"transitive ambient environment access is forbidden",
),
(
"\nimport subprocess\n\ndef _process_mutation():\n"
" return subprocess.run([])\n",
"forbidden transitive model/network/process import",
),
),
)
def test_imported_signal_helper_shares_the_hermetic_boundary(
tmp_path: Path, body: str, expected: str
) -> None:
root = _copy_repo(tmp_path)
path = root / integration.SIGNAL_RUNTIME
path.write_text(path.read_text(encoding="utf-8") + body, encoding="utf-8")
_assert_error(root, expected)
@pytest.mark.parametrize("option", ("--all", "--page", "--page-size"))
def test_renderer_traversal_flags_are_forbidden(tmp_path: Path, option: str) -> None:
root = _copy_repo(tmp_path)
path = root / integration.RUNTIME
path.write_text(
path.read_text(encoding="utf-8")
+ f"\n\ndef _renderer_option_mutation(parser):\n parser.add_argument({option!r})\n",
encoding="utf-8",
)
_assert_error(root, "forbidden renderer options")
@pytest.mark.parametrize(
("old", "new", "constant"),
(
(
'ADVISORY_LABEL = "Phrase-list screening advisory"',
'ADVISORY_LABEL = "Origin classifier"',
"ADVISORY_LABEL",
),
("MAX_NODE_WITNESSES = 512", "MAX_NODE_WITNESSES = 513", "MAX_NODE_WITNESSES"),
(
"MAX_NODE_COMBINATIONS = 100_000",
"MAX_NODE_COMBINATIONS = 100_001",
"MAX_NODE_COMBINATIONS",
),
("MAX_PARSE_INTERVALS = 4096", "MAX_PARSE_INTERVALS = 4097", "MAX_PARSE_INTERVALS"),
("MAX_SEGMENTS = 4096", "MAX_SEGMENTS = 4097", "MAX_SEGMENTS"),
(
"MAX_PARSE_WORK_UNITS = 100_000",
"MAX_PARSE_WORK_UNITS = 100_001",
"MAX_PARSE_WORK_UNITS",
),
(
"MAX_RAW_TOKEN_CODEPOINTS = 4096",
"MAX_RAW_TOKEN_CODEPOINTS = 4097",
"MAX_RAW_TOKEN_CODEPOINTS",
),
("MAX_CORPUS_ENTRIES = 512", "MAX_CORPUS_ENTRIES = 513", "MAX_CORPUS_ENTRIES"),
(
"MAX_CORPUS_EXISTING_SIGNALS = 8192",
"MAX_CORPUS_EXISTING_SIGNALS = 8193",
"MAX_CORPUS_EXISTING_SIGNALS",
),
(
"MAX_STRUCTURE_DEPTH = 64",
"MAX_STRUCTURE_DEPTH = 65",
"MAX_STRUCTURE_DEPTH",
),
(
"MAX_STRUCTURE_NODES = 200_000",
"MAX_STRUCTURE_NODES = 200_001",
"MAX_STRUCTURE_NODES",
),
),
)
def test_runtime_frozen_constant_drift_fails(
tmp_path: Path, old: str, new: str, constant: str
) -> None:
root = _copy_repo(tmp_path)
_replace(root, integration.RUNTIME, old, new)
_assert_error(root, f"{constant} must equal")
@pytest.mark.parametrize(
("old", "new", "expected"),
(
(
"if len(corpus) > MAX_CORPUS_ENTRIES:",
"if len(corpus) >= MAX_CORPUS_ENTRIES:",
"MAX_CORPUS_ENTRIES with the strict N/N+1 guard",
),
(
"if existing_signal_count > MAX_CORPUS_EXISTING_SIGNALS:",
"if existing_signal_count >= MAX_CORPUS_EXISTING_SIGNALS:",
"MAX_CORPUS_EXISTING_SIGNALS with the strict N/N+1 guard",
),
),
)
def test_corpus_cardinality_boundary_drift_fails(
tmp_path: Path,
old: str,
new: str,
expected: str,
) -> None:
root = _copy_repo(tmp_path)
_replace(root, integration.RUNTIME, old, new)
_assert_error(root, expected)
def test_raw_token_boundary_must_remain_strict_n_plus_one(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.RUNTIME,
"if len(buffer) > MAX_RAW_TOKEN_CODEPOINTS:",
"if len(buffer) >= MAX_RAW_TOKEN_CODEPOINTS:",
)
_assert_error(root, "MAX_RAW_TOKEN_CODEPOINTS with the strict N/N+1 guard")
def test_existing_signal_limit_must_count_the_aggregate(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.RUNTIME,
"existing_signal_count = sum(",
"existing_signal_count = max(",
)
_assert_error(root, "existing_signal_count must aggregate every pre-existing")
def test_passport_write_cannot_precede_successful_enrichment(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.RUNTIME,
" document, kind = _load_passport(args.input)\n"
" output = enrich_passport(",
" document, kind = _load_passport(args.input)\n"
" _atomic_write_passport(args.output, document, kind)\n"
" output = enrich_passport(",
)
_assert_error(root, "write the passport only after enrich_passport returns")
@pytest.mark.parametrize(
("old", "new", "expected"),
(
(
"if len(unique) > MAX_NODE_WITNESSES:",
"if False: # removed witness guard",
"_minimal_witnesses must enforce MAX_NODE_WITNESSES",
),
(
"if len(result) > MAX_NODE_WITNESSES:",
"if False: # removed literal witness guard",
"_literal_witnesses must enforce MAX_NODE_WITNESSES",
),
(
"if attempts > MAX_NODE_COMBINATIONS:",
"if False: # removed combination guard",
"evaluate_expression must enforce MAX_NODE_COMBINATIONS",
),
(
"if len(intervals) > MAX_PARSE_INTERVALS:",
"if False: # removed interval guard",
"_append_parse_interval must enforce MAX_PARSE_INTERVALS",
),
(
"if len(openers) > MAX_PARSE_INTERVALS:",
"if False: # removed opaque candidate guard",
"_append_opaque_opener must enforce MAX_PARSE_INTERVALS",
),
(
"if len(buffer) > MAX_RAW_TOKEN_CODEPOINTS:",
"if False: # removed raw-token guard",
"tokenize must enforce MAX_RAW_TOKEN_CODEPOINTS",
),
(
"if len(segments) > MAX_SEGMENTS:",
"if False: # removed segment guard",
"build_own_draft_report must enforce MAX_SEGMENTS",
),
(
"if nodes_seen > MAX_STRUCTURE_NODES:",
"if False: # removed structure-node guard",
"_reject_nonfinite_recursive must enforce MAX_STRUCTURE_NODES",
),
(
'"summary_label": ADVISORY_LABEL,',
'"summary_label": "Phrase-list screening advisory",',
"build_cited_signal must bind summary_label to ADVISORY_LABEL",
),
),
)
def test_runtime_limit_and_label_structure_drift_fails(
tmp_path: Path, old: str, new: str, expected: str
) -> None:
root = _copy_repo(tmp_path)
_replace(root, integration.RUNTIME, old, new)
_assert_error(root, expected)
def test_iterative_structure_depth_guard_cannot_be_removed(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace_nth(
root,
integration.RUNTIME,
"if depth > MAX_STRUCTURE_DEPTH:",
"if False: # removed iterative structure-depth guard",
3,
)
_assert_error(
root,
"_reject_nonfinite_recursive must enforce MAX_STRUCTURE_DEPTH",
)
def test_segment_document_retains_its_aggregate_interval_cap(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace_nth(
root,
integration.RUNTIME,
"if len(intervals) > MAX_PARSE_INTERVALS:",
"if False: # removed aggregate interval guard",
2,
)
_assert_error(root, "segment_document must enforce MAX_PARSE_INTERVALS")
@pytest.mark.parametrize(
("old", "expected"),
(
(
" _reject_nonfinite_recursive(value, path=label)\n",
"strict JSON loader must apply the structure guard",
),
(
' _reject_nonfinite_recursive(document, path="passport")\n',
"direct passport enricher must apply the structure guard",
),
),
)
def test_structure_guard_call_cannot_be_removed(
tmp_path: Path, old: str, expected: str
) -> None:
root = _copy_repo(tmp_path)
_replace(root, integration.RUNTIME, old, "")
_assert_error(root, expected)
def test_renderer_cap_drift_fails(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(root, integration.RUNTIME, "MAX_RENDER_PAGE_SIZE = 25", "MAX_RENDER_PAGE_SIZE = 26")
_assert_error(root, "MAX_RENDER_PAGE_SIZE must equal 25")
def test_renderer_unbounded_slice_fails(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.RUNTIME,
"selected = matches[:MAX_RENDER_PAGE_SIZE]",
"selected = matches[:]",
)
_assert_error(root, "renderer must slice matches at the fixed cap")
def test_renderer_must_iterate_the_bounded_selection(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.RUNTIME,
"for match in selected:",
"for match in matches:",
)
_assert_error(root, "renderer must iterate only the fixed slice")
@pytest.mark.parametrize(
("old", "new", "expected"),
(
(
'reason_code = "DOCUMENT_EMPTY"',
'reason_code = "DOCUMENT_BLANK"',
"empty draft must emit DOCUMENT_EMPTY",
),
(
' "ABSTRACT_EMPTY",',
' "ABSTRACT_BLANK",',
"empty abstract must emit ABSTRACT_EMPTY",
),
),
)
def test_runtime_empty_reason_drift_fails(
tmp_path: Path, old: str, new: str, expected: str
) -> None:
root = _copy_repo(tmp_path)
_replace(root, integration.RUNTIME, old, new)
_assert_error(root, expected)
def test_explicit_timestamp_parameters_are_required(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(root, integration.RUNTIME, " checked_at: str,", " captured_at: str,")
_assert_error(root, "requires explicit checked_at/recorded_at")
def test_design_claim_ceiling_binding_fails_closed(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.DESIGN_SPEC,
"No #654 measurement row may be manufactured",
"A #654 measurement row may be manufactured",
)
_assert_error(root, "missing #660 integration tokens")
@pytest.mark.parametrize(
"token",
(
"MAX_CORPUS_ENTRIES = 512",
"MAX_CORPUS_EXISTING_SIGNALS = 8192",
"MAX_STRUCTURE_DEPTH = 64",
"MAX_STRUCTURE_NODES = 200_000",
"MAX_PARSE_WORK_UNITS = 100_000",
"pre-normalization maximum of 4,096 raw code points per token",
"Opaque constructs are lexed strictly in source order",
"preceded by an odd run of backslashes are literal",
"`N+1` raises the resource-limit failure",
"leaves a pre-existing output byte-identical",
),
)
def test_design_corpus_cardinality_boundary_is_bound(
tmp_path: Path,
token: str,
) -> None:
root = _copy_repo(tmp_path)
_replace(root, integration.DESIGN_SPEC, token, "REMOVED_CORPUS_BOUNDARY")
_assert_error(root, "missing #660 integration tokens")
@pytest.mark.parametrize("token", ("ABSTRACT_EMPTY", "DOCUMENT_EMPTY"))
def test_design_empty_reason_vocabulary_is_bound(tmp_path: Path, token: str) -> None:
root = _copy_repo(tmp_path)
_replace(root, integration.DESIGN_SPEC, token, "REMOVED_EMPTY_REASON")
_assert_error(root, "missing #660 integration tokens")
def test_formatter_neutral_advisory_label_is_bound(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.FORMATTER_AGENT,
"neutral **Phrase-list screening advisory**",
"categorical **Origin classifier**",
)
_assert_error(root, "missing #660 integration tokens")
def test_orchestrator_preserves_degraded_artifact_after_exit_one(
tmp_path: Path,
) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.PIPELINE_ORCHESTRATOR,
"schema-valid degraded advisory and then exit 1",
"degraded output and then exits",
)
_assert_error(root, "missing #660 integration tokens")
def test_design_neutral_advisory_label_is_bound(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.DESIGN_SPEC,
"category label is the neutral **“Phrase-list screening",
"category label is the categorical **“Origin classifier",
)
_assert_error(root, "missing #660 integration tokens")
def test_consumer_binding_drift_fails(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.CORPUS_CHECKER,
" validate_cited_signal_binding(signal, entry)",
" validate_removed_signal_binding(signal, entry)",
)
_assert_error(root, "must call validate_cited_signal_binding exactly once")
def test_runtime_manifest_entry_is_required(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
block = (
'[[pytest]]\n'
'id = "660-tortured-phrase-screening"\n'
'path = "scripts/test_tortured_phrase_screening.py"\n\n'
)
_replace(root, integration.CI_MANIFEST, block, "")
_assert_error(root, "660-tortured-phrase-screening")
def test_integration_manifest_entry_is_required(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
block = (
'[[pytest]]\n'
'id = "660-tortured-phrase-integration"\n'
'path = "scripts/test_check_tortured_phrase_screening_integration.py"\n\n'
)
_replace(root, integration.CI_MANIFEST, block, "")
_assert_error(root, "660-tortured-phrase-integration")
def test_direct_workflow_checker_is_required(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.WORKFLOW,
"python3 scripts/check_tortured_phrase_screening_integration.py",
"python3 scripts/check_bibliographic_integrity_signals.py",
)
_assert_error(root, "direct #660 integration checker must run exactly once")
def test_conformance_suite_registry_entry_is_required(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
def mutate(registry: dict[str, Any]) -> None:
registry.pop("tortured_phrase_conformance")
_rewrite_json(root, integration.SUITE_REGISTRY, mutate)
_assert_error(root, "tortured_phrase_conformance must map exactly to mechanical_match")
def test_measurement_contract_registry_mirror_is_required(tmp_path: Path) -> None:
root = _copy_repo(tmp_path)
_replace(
root,
integration.MEASUREMENT_CONTRACT,
"| `tortured_phrase_conformance` | `mechanical_match` |",
"| `tortured_phrase_conformance` | `llm_judged` |",
)
_assert_error(root, "registry mirror must contain exactly one")
@pytest.mark.parametrize(
("old", "new"),
(
("Status: PRE-REGISTERED / NOT RUN", "Status: DRAFT / MAY RUN"),
(
"first commit reachable",
"working-tree state not reachable on the default branch",
),
(
"use that same 40-hex main-history commit for both",
"different commits may be used for",
),
("No retry is permitted", "One retry is permitted"),
("synthetic_conformance_test_pass_rate", "contextual_accuracy"),
("It uses zero", "It uses two"),
("A passing row may say only", "A passing row may broadly claim"),
),
)
def test_frozen_measurement_plan_token_drift_fails(
tmp_path: Path, old: str, new: str
) -> None:
root = _copy_repo(tmp_path)
_replace(root, integration.MEASUREMENT_PLAN, old, new)
_assert_error(root, "frozen plan tokens drifted")
@pytest.mark.parametrize(
("old", "new"),
(
("remains `UNMEASURED`", "becomes `MEASURED`"),
(
"contextual false-positive/false-negative labels",
"contextual accuracy labels",
),
),
)
def test_conformance_readme_claim_ceiling_drift_fails(
tmp_path: Path, old: str, new: str
) -> None:
root = _copy_repo(tmp_path)
_replace(root, integration.CONFORMANCE_README, old, new)
_assert_error(root, "UNMEASURED/contextual claim ceiling drifted")
def test_scored_heldout_row_is_forbidden_on_implementation_branch(
tmp_path: Path,
) -> None:
root = _copy_repo(tmp_path)
path = root / integration.HELDOUT_ROOT / "tortured_phrase" / "measurement.json"
path.parent.mkdir(parents=True)
path.write_text(
json.dumps(
{
"measurement_contract": "heldout-measurement/1.0",
"surface": "tortured_phrase",
}
),
encoding="utf-8",
)
_assert_error(root, "must remain UNMEASURED")
def test_checker_imports_no_network_model_or_process_modules() -> None:
source = Path(integration.__file__).read_text(encoding="utf-8")
tree = ast.parse(source)
imports = {
alias.name.split(".", 1)[0]
for node in ast.walk(tree)
if isinstance(node, ast.Import)
for alias in node.names
} | {
node.module.split(".", 1)[0]
for node in ast.walk(tree)
if isinstance(node, ast.ImportFrom) and node.module
}
assert not imports & integration.FORBIDDEN_IMPORT_ROOTS
def test_cli_main_reports_success_and_invocation_error(
tmp_path: Path, capsys: pytest.CaptureFixture[str]
) -> None:
root = _copy_repo(tmp_path)
assert integration.main(["--root", str(root)]) == 0
assert "integration: ok" in capsys.readouterr().out
assert integration.main(["--unknown"]) == 2
File diff suppressed because it is too large Load Diff
+53 -3
View File
@@ -161,6 +161,14 @@ LINE_BUDGET_576_STAGE3P_DISPATCH = 30
# landing: 54 lines; budget 60 leaves 6 lines of headroom.
LINE_BUDGET_656_EVIDENCE_RENDERING = 60
# #660 ships the `## Tortured-Phrase Advisory Dispatch (#660)` H2 block.
# It carries the exact pre-format dispatch, explicit-time/no-network boundary,
# degraded-artifact handoff, advisory-only claim ceiling, and read-only cited
# carrier rules. This is an independent 2026-08 extension, so it is subtracted
# from the historical v3.6.7 budget and receives its own bounded test. Measured
# at landing: 55 lines; budget 60 leaves 5 lines of headroom.
LINE_BUDGET_660_ADVISORY_DISPATCH = 60
# All 24 failure phase IDs from spec §5.6 inventory (7 P-PA-* + 17 P-PB-*).
# These must each appear at least once in the orchestrator prompt as
# cross-references to spec §5.6 (NOT inline procedural definitions —
@@ -613,6 +621,45 @@ def _measure_656_evidence_rendering_block_lines(text: str) -> int:
return policy_lines + insertion_end - insertion_start + 1
def _measure_660_advisory_dispatch_block_lines(text: str) -> int:
"""Return lines in the #660 tortured-phrase dispatch H2 block."""
import re as _re
anchor = _re.compile(
r"(?m)^[ \t]*##[ \t]+Tortured-Phrase Advisory Dispatch "
r"\(#660\)[ \t]*$"
)
m = anchor.search(text)
if m is None:
return 0
next_h = _re.compile(r"(?m)^[ \t]*#{1,2}[ \t]+")
head_eol = text.find("\n", m.end())
search_start = (head_eol + 1) if head_eol >= 0 else len(text)
nm = next_h.search(text, search_start)
end = nm.start() if nm else len(text)
return len(text[m.start():end].splitlines())
class Advisory660LineBudgetTest(unittest.TestCase):
"""#660 tortured-phrase dispatch block stays independently bounded."""
def test_660_advisory_dispatch_block_within_budget(self) -> None:
text = _read_prompt()
block_lines = _measure_660_advisory_dispatch_block_lines(text)
self.assertGreater(
block_lines,
0,
"#660 tortured-phrase advisory dispatch block is missing from "
"pipeline_orchestrator_agent.md",
)
self.assertLessEqual(
block_lines,
LINE_BUDGET_660_ADVISORY_DISPATCH,
f"#660 tortured-phrase advisory dispatch block is {block_lines} "
f"lines, over its {LINE_BUDGET_660_ADVISORY_DISPATCH}-line budget",
)
class Dispatch576LineBudgetTest(unittest.TestCase):
"""#576 Spec B Stage 3' contract-dispatch block within
`LINE_BUDGET_576_STAGE3P_DISPATCH` line budget.
@@ -693,18 +740,20 @@ class Phase66LineBudgetTest(unittest.TestCase):
authority_670_lines = _measure_670_authority_extension_lines(text)
dispatch_576_lines = _measure_576_stage3p_dispatch_block_lines(text)
evidence_656_lines = _measure_656_evidence_rendering_block_lines(text)
advisory_660_lines = _measure_660_advisory_dispatch_block_lines(text)
# v3.6.7-only line count: total minus v3.7.1 Step 3b, v3.7.3
# finalizer extension, v3.8 §3.6 audit-gate, v3.9.0 triangulation
# extension, v3.10 terminal-policy extension, the #394 slice-4
# submission-package gate, the #390 Slice B revision-patch
# sequencing, the #670 authority/bundle extension, the #576 Spec B
# Stage 3' contract-dispatch, AND the #656 Phase E evidence-row
# checkpoint-rendering
# checkpoint-rendering, AND the #660 tortured-phrase advisory dispatch
# subsections (each has its own dedicated budget test).
v367_line_count = (
total_lines - step_3b_lines - v3_7_3_lines - v3_8_lines
- v3_9_0_lines - v3_10_lines - gate_394_lines - seq_390_lines
- authority_670_lines - dispatch_576_lines - evidence_656_lines
- advisory_660_lines
)
ceiling = BASELINE_LINE_COUNT + LINE_BUDGET_OVER_BASELINE
self.assertLessEqual(
@@ -721,8 +770,9 @@ class Phase66LineBudgetTest(unittest.TestCase):
f"subsection, {gate_394_lines} are in the #394 submission-"
f"package gate, {seq_390_lines} are in the #390 revision-patch "
f"sequencing subsection, {dispatch_576_lines} are in the #576 "
f"dispatch subsection, and {evidence_656_lines} are in the #656 "
f"evidence-rendering subsection; v3.6.7-attributed lines = "
f"dispatch subsection, {evidence_656_lines} are in the #656 "
f"evidence-rendering subsection, and {advisory_660_lines} are in "
f"the #660 advisory-dispatch subsection; v3.6.7-attributed lines = "
f"{v367_line_count} exceeds {ceiling} (baseline "
f"{BASELINE_LINE_COUNT} + Phase 6.6 budget "
f"{LINE_BUDGET_OVER_BASELINE}). Tighten the §3.5 Audit "
File diff suppressed because it is too large Load Diff
+46
View File
@@ -5,6 +5,11 @@ single schema authority for bibliographic-integrity observations carried in
`literature_corpus[].bibliographic_integrity_signals[]`. Version 1.0 is the
additive migration carrier. Version 1.1 is the #651 retraction-status policy
cutover. Both preserve the one-advisory-token reference-marker grammar.
Version 1.2 is the #660 tortured-phrase advisory profile; it is additive and
does not change the validity of shipped canonical v1.0/v1.1 fixtures or producer
outputs, nor their policy meaning. Its signal-type invariant intentionally rejects
noncanonical, previously underconstrained tortured-phrase mutations that claimed
deterministic or terminal authority.
## Epistemic boundary
@@ -50,6 +55,43 @@ declared-legitimate exception did not fire. The finalizer then evaluates the
explicit `terminal_policies.retraction` choice. Adding a signal never silently
promotes it to `HIGH-BLOCK`.
## Tortured-phrase advisory profile (v1.2 / #660)
Version 1.2 carries one current `tortured_phrase_match` row per
`(citation_key, cited_title|cited_abstract)` surface. It is always
`heuristic_advisory` / `HEURISTIC-INDICATOR`; its closed context remains
`layer: HEURISTIC-ADVISORY` and `evaluation_status: UNMEASURED`. A detected
row means only **phrase-list match requiring review**. A checked zero-match row
means only that no phrase-list match was observed on that exact checked
surface; absence is not a clean certification. It never establishes AI,
author, paper-mill, misconduct, or other origin, and it carries no precision,
recall, false-positive, false-negative, coverage, or publisher-acceptance
claim.
The local producer consumes only an explicitly named canonical snapshot and
detached manifest. The manifest binds the exact raw UTF-8 snapshot bytes by
SHA-256 and declares either `user_supplied` or `synthetic_fixture` supply;
ARS ships no native PPS parser/importer, fetch path, or redistributed PPS list
content. Matching reads no model, external API, human/model judge, system
clock, file time, or network time. Required timestamps are explicit inputs.
Each invoked corpus check keeps title and abstract coverage separate. A valid
title is checked independently. A missing abstract remains an explicit
`not_checked` / `unresolved` row with reason `ABSTRACT_MISSING`, while a present
whitespace-only abstract uses `ABSTRACT_EMPTY`; neither state can be
folded into the title result or omitted. Manual corpus entries have no
exemption. Every row binds the exact local surface bytes plus snapshot and
manifest hashes. The producer returns a new passport copy and never rewrites a
source title, abstract, citation, or input passport in place.
All v1.2 rows compose in the existing single `Bibliographic Integrity
Advisories` section. `display.marker_token` is `null`,
`terminal_policy.eligible` is `false`, and neither the finalizer nor formatter
may mint a marker, gate, terminal promotion, suggested replacement, or
automatic rewrite from the row. The separate own-draft carrier is
`tortured-phrase-advisory/1.0`; it has the same
`HEURISTIC-ADVISORY` / `UNMEASURED` boundary and is not a corpus authority.
## Retraction authority cutover (v1.1 / #651)
The v1.1 `retraction_status` row is authoritative for retraction status. It
@@ -94,3 +136,7 @@ resolver-specific record with `check_status: degraded` and
absence/unknown; it is never synthesized as a clean result. Existing
`provenance_summary.md` text is an output projection and is never parsed back
as evidence.
The v1.2 profile is not a migration of the older tortured-phrase scaffold.
Existing v1.0 heuristic rows remain readable with their original meaning, but
cannot be treated as completed title/abstract checks.
+36 -2
View File
@@ -50,8 +50,9 @@ Schemas for Material Passport input ports.
- `passport/literature_corpus_entry.schema.json` (v3.6.4) — Schema 9 `literature_corpus[]`
entries produced by user-written adapters.
- `passport/bibliographic_integrity_signal.schema.json` (#678/#651) — v1.0
additive signal carrier plus v1.1 authoritative retraction-status rows,
- `passport/bibliographic_integrity_signal.schema.json` (#678/#651/#660) — v1.0
additive signal carrier, v1.1 authoritative retraction-status rows, and the
v1.2 title/abstract tortured-phrase advisory profile,
including resolver disagreement/reinstatement, judgment context, freshness,
and opt-in finalizer policy eligibility. The separate
`retraction_status_cache_v1` namespace and pure resolver live in
@@ -429,6 +430,39 @@ Protocol:
`shared/references/authority_content_coverage_advisory_protocol.md`. Spec:
`docs/design/2026-08-09-681-authority-content-coverage-advisory-spec.md`.
## Tortured-phrase screening contracts (#660)
The #660 family is local, hash-bound, and advisory-only:
- `audit/tortured_phrase_snapshot.schema.json` defines the closed canonical
`literal` / `all` / `any` / `near` AST and rule-level `exclude_if` grammar;
- `audit/tortured_phrase_snapshot_manifest.schema.json` binds the exact raw
snapshot bytes, source/version/as-of metadata, preprocessing disclosure,
zero unsupported rules, and rights; and
- `audit/tortured_phrase_advisory.schema.json` defines the replay-bound
own-draft `HEURISTIC-ADVISORY` / `UNMEASURED` transcript.
`scripts/tortured_phrase_screening.py` accepts only explicitly named local
snapshot/manifest and input paths. A snapshot is either user supplied or a
repository-authored synthetic fixture. ARS includes no native PPS importer,
fetch option, or redistributed PPS list content, and this path invokes no
model, external API, human/model judge, ambient clock, file time, or network
time. Snapshot SHA-256 covers the exact UTF-8 file bytes; required timestamps
are explicit arguments.
The same runtime can return a new passport copy carrying
`bibliographic-integrity-signal/1.2` rows for each cited title and abstract
surface. A missing abstract is an explicit `not_checked` / `unresolved`
`ABSTRACT_MISSING` row; a present whitespace-only abstract uses
`ABSTRACT_EMPTY`. Consumers remain read-only and the enricher refuses
in-place output. A detected row is only a phrase-list match requiring review;
a zero match is not a clean certificate. No result establishes origin,
paper-mill production, contextual validity, or accuracy, and no result creates
a marker, terminal gate, replacement text, or automatic rewrite. All corpus
rows render in the one canonical `Bibliographic Integrity Advisories` section.
Spec: `docs/design/2026-08-10-660-tortured-phrase-screening-spec.md`.
## Audit artifact contracts (v3.6.7 Step 6)
The `audit/` directory carries the three wrapper-emitted artifact schemas that pair
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,208 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://github.com/Imbad0202/academic-research-skills/shared/contracts/audit/tortured_phrase_snapshot.schema.json",
"title": "Tortured-phrase canonical snapshot",
"description": "Closed canonical-AST snapshot for deterministic tortured-phrase list matching. The grammar is positive-only: exclusion contexts are represented solely by rule.exclude_if[]. There is no NOT, regex, wildcard, executable expression, or implicit operator. Cross-rule identifier uniqueness, AST depth, word budgets, and exact snapshot-byte replay are enforced by the runtime.",
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"snapshot_id",
"grammar_profile",
"normalizer_profile",
"rules"
],
"properties": {
"schema_version": {
"const": "tortured-phrase-snapshot/1.0"
},
"snapshot_id": {
"$ref": "#/$defs/identifier"
},
"grammar_profile": {
"const": "ars-tortured-phrase-canonical-ast/1.0"
},
"normalizer_profile": {
"const": "ars-nfkc-casefold-token/1.0"
},
"rules": {
"type": "array",
"minItems": 1,
"maxItems": 512,
"uniqueItems": true,
"items": {
"$ref": "#/$defs/rule"
}
}
},
"$defs": {
"identifier": {
"type": "string",
"minLength": 1,
"maxLength": 128,
"pattern": "^[a-z0-9][a-z0-9._-]{0,127}$"
},
"literal": {
"type": "string",
"minLength": 1,
"maxLength": 256,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$",
"description": "A bounded literal token sequence. Metacharacters have no special meaning: regex and wildcard interpretation are forbidden. The runtime additionally caps a literal at eight normalized tokens."
},
"expression": {
"oneOf": [
{
"$ref": "#/$defs/literal_expression"
},
{
"$ref": "#/$defs/all_expression"
},
{
"$ref": "#/$defs/any_expression"
},
{
"$ref": "#/$defs/near_expression"
}
]
},
"literal_expression": {
"type": "object",
"additionalProperties": false,
"required": [
"op",
"value"
],
"properties": {
"op": {
"const": "literal"
},
"value": {
"$ref": "#/$defs/literal"
}
}
},
"all_expression": {
"type": "object",
"additionalProperties": false,
"required": [
"op",
"terms",
"max_span_tokens"
],
"properties": {
"op": {
"const": "all"
},
"terms": {
"type": "array",
"minItems": 2,
"maxItems": 8,
"uniqueItems": true,
"items": {
"$ref": "#/$defs/expression"
}
},
"max_span_tokens": {
"type": "integer",
"minimum": 1,
"maximum": 64
}
}
},
"any_expression": {
"type": "object",
"additionalProperties": false,
"required": [
"op",
"alternatives"
],
"properties": {
"op": {
"const": "any"
},
"alternatives": {
"type": "array",
"minItems": 2,
"maxItems": 8,
"uniqueItems": true,
"items": {
"$ref": "#/$defs/expression"
}
}
}
},
"near_expression": {
"type": "object",
"additionalProperties": false,
"required": [
"op",
"left",
"right",
"max_gap_tokens",
"ordered"
],
"properties": {
"op": {
"const": "near"
},
"left": {
"$ref": "#/$defs/expression"
},
"right": {
"$ref": "#/$defs/expression"
},
"max_gap_tokens": {
"type": "integer",
"minimum": 0,
"maximum": 64
},
"ordered": {
"type": "boolean"
}
}
},
"exclusion": {
"type": "object",
"additionalProperties": false,
"required": [
"expression",
"within_tokens"
],
"properties": {
"expression": {
"$ref": "#/$defs/expression"
},
"within_tokens": {
"type": "integer",
"minimum": 0,
"maximum": 64
}
}
},
"rule": {
"type": "object",
"additionalProperties": false,
"required": [
"rule_id",
"expression"
],
"properties": {
"rule_id": {
"$ref": "#/$defs/identifier"
},
"expression": {
"$ref": "#/$defs/expression"
},
"exclude_if": {
"type": "array",
"minItems": 1,
"maxItems": 8,
"uniqueItems": true,
"items": {
"$ref": "#/$defs/exclusion"
}
}
}
}
}
}
@@ -0,0 +1,335 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://github.com/Imbad0202/academic-research-skills/shared/contracts/audit/tortured_phrase_snapshot_manifest.schema.json",
"title": "Tortured-phrase snapshot manifest",
"description": "Detached provenance, preprocessing, and rights manifest for one exact canonical tortured-phrase snapshot. snapshot_sha256 covers the exact UTF-8 snapshot file bytes; it is not a hash of a reserialized JSON value. The manifest never authorizes fetching or redistribution by itself.",
"type": "object",
"additionalProperties": false,
"required": [
"schema_version",
"snapshot_id",
"source",
"supply_mode",
"snapshot_schema_version",
"snapshot_sha256",
"grammar_profile",
"normalizer_profile",
"preprocessor",
"unsupported_rule_count",
"rule_count",
"rights"
],
"properties": {
"schema_version": {
"const": "tortured-phrase-snapshot-manifest/1.0"
},
"snapshot_id": {
"$ref": "#/$defs/identifier"
},
"source": {
"$ref": "#/$defs/source"
},
"supply_mode": {
"enum": [
"user_supplied",
"synthetic_fixture"
]
},
"snapshot_schema_version": {
"const": "tortured-phrase-snapshot/1.0"
},
"snapshot_sha256": {
"$ref": "#/$defs/sha256"
},
"grammar_profile": {
"const": "ars-tortured-phrase-canonical-ast/1.0"
},
"normalizer_profile": {
"const": "ars-nfkc-casefold-token/1.0"
},
"preprocessor": {
"$ref": "#/$defs/preprocessor"
},
"unsupported_rule_count": {
"type": "integer",
"minimum": 0,
"maximum": 512
},
"rule_count": {
"type": "integer",
"minimum": 1,
"maximum": 512
},
"rights": {
"$ref": "#/$defs/rights"
}
},
"allOf": [
{
"if": {
"properties": {
"supply_mode": {
"const": "synthetic_fixture"
}
},
"required": [
"supply_mode"
]
},
"then": {
"properties": {
"rights": {
"properties": {
"basis": {
"const": "synthetic_fixture"
},
"redistribution_status": {
"const": "permitted"
},
"reference": {
"type": "null"
},
"user_declaration": {
"type": "null"
}
}
}
}
}
},
{
"if": {
"properties": {
"rights": {
"properties": {
"basis": {
"const": "user_declared_authorized"
}
},
"required": [
"basis"
]
}
},
"required": [
"rights"
]
},
"then": {
"properties": {
"rights": {
"properties": {
"user_declaration": {
"$ref": "#/$defs/nonempty_reference"
}
}
}
}
}
},
{
"if": {
"properties": {
"rights": {
"properties": {
"basis": {
"const": "written_permission"
}
},
"required": [
"basis"
]
}
},
"required": [
"rights"
]
},
"then": {
"properties": {
"rights": {
"properties": {
"reference": {
"$ref": "#/$defs/nonempty_reference"
}
}
}
}
}
},
{
"if": {
"properties": {
"rights": {
"properties": {
"basis": {
"const": "unresolved"
}
},
"required": [
"basis"
]
}
},
"required": [
"rights"
]
},
"then": {
"properties": {
"rights": {
"properties": {
"redistribution_status": {
"const": "unresolved"
}
}
}
}
}
}
],
"$defs": {
"identifier": {
"type": "string",
"minLength": 1,
"maxLength": 128,
"pattern": "^[a-z0-9][a-z0-9._-]{0,127}$"
},
"sha256": {
"type": "string",
"pattern": "^[0-9a-f]{64}$"
},
"bounded_line": {
"type": "string",
"minLength": 1,
"maxLength": 500,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$"
},
"nonempty_reference": {
"$ref": "#/$defs/bounded_line"
},
"source": {
"type": "object",
"additionalProperties": false,
"required": [
"name",
"version",
"as_of",
"locator"
],
"properties": {
"name": {
"type": "string",
"minLength": 1,
"maxLength": 200,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$"
},
"version": {
"type": "string",
"minLength": 1,
"maxLength": 100,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$"
},
"as_of": {
"type": "string",
"format": "date"
},
"locator": {
"$ref": "#/$defs/bounded_line"
}
}
},
"preprocessor": {
"type": "object",
"additionalProperties": false,
"required": [
"name",
"version",
"native_grammar",
"reduction_notes"
],
"properties": {
"name": {
"type": "string",
"minLength": 1,
"maxLength": 100,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$"
},
"version": {
"type": "string",
"minLength": 1,
"maxLength": 100,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$"
},
"native_grammar": {
"oneOf": [
{
"type": "string",
"minLength": 1,
"maxLength": 200,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$"
},
{
"type": "null"
}
]
},
"reduction_notes": {
"type": "array",
"maxItems": 32,
"uniqueItems": true,
"items": {
"$ref": "#/$defs/bounded_line"
}
}
}
},
"rights": {
"type": "object",
"additionalProperties": false,
"required": [
"basis",
"redistribution_status",
"reference",
"user_declaration"
],
"properties": {
"basis": {
"enum": [
"synthetic_fixture",
"user_declared_authorized",
"written_permission",
"unresolved"
]
},
"redistribution_status": {
"enum": [
"permitted",
"not_permitted",
"unresolved"
]
},
"reference": {
"oneOf": [
{
"$ref": "#/$defs/nonempty_reference"
},
{
"type": "null"
}
]
},
"user_declaration": {
"oneOf": [
{
"$ref": "#/$defs/nonempty_reference"
},
{
"type": "null"
}
]
}
}
}
}
}
@@ -23,7 +23,8 @@
"schema_version": {
"enum": [
"bibliographic-integrity-signal/1.0",
"bibliographic-integrity-signal/1.1"
"bibliographic-integrity-signal/1.1",
"bibliographic-integrity-signal/1.2"
]
},
"signal_id": {
@@ -272,7 +273,465 @@
},
"marker_token": {
"type": "null",
"description": "Always null in v1.0 and v1.1. Existing CONTAMINATED-* projection remains owned by the legacy contamination carrier during migration; new signal types never mint a second marker advisory token."
"description": "Always null in v1.0, v1.1, and v1.2. Existing CONTAMINATED-* projection remains owned by the legacy contamination carrier during migration; new signal types never mint a second marker advisory token."
}
}
},
"tortured_phrase_context": {
"type": "object",
"additionalProperties": false,
"required": [
"layer",
"evaluation_status",
"surface",
"surface_binding",
"snapshot",
"reason_code",
"counts",
"matches",
"boundary"
],
"properties": {
"layer": {"const": "HEURISTIC-ADVISORY"},
"evaluation_status": {"const": "UNMEASURED"},
"surface": {"enum": ["cited_title", "cited_abstract"]},
"surface_binding": {
"type": "object",
"additionalProperties": false,
"required": ["content_sha256", "content_utf8_bytes"],
"properties": {
"content_sha256": {
"type": ["string", "null"],
"pattern": "^[0-9a-f]{64}$"
},
"content_utf8_bytes": {
"type": ["integer", "null"],
"minimum": 0,
"maximum": 8388608
}
},
"oneOf": [
{
"properties": {
"content_sha256": {"type": "string"},
"content_utf8_bytes": {"type": "integer", "minimum": 1}
}
},
{
"properties": {
"content_sha256": {"type": "null"},
"content_utf8_bytes": {"type": "null"}
}
}
]
},
"snapshot": {
"type": "object",
"additionalProperties": false,
"required": [
"status",
"reason_code",
"snapshot_sha256",
"manifest_sha256",
"snapshot_id",
"source",
"supply_mode",
"snapshot_schema_version",
"grammar_profile",
"normalizer_profile",
"unicode_data_version",
"rule_count",
"unsupported_rule_count",
"rights"
],
"properties": {
"status": {"enum": ["loaded", "not_checked", "degraded"]},
"reason_code": {
"enum": [
"CHECK_COMPLETED",
"SNAPSHOT_NOT_PROVIDED",
"SNAPSHOT_BYTES_INVALID",
"SNAPSHOT_MANIFEST_INVALID",
"SNAPSHOT_HASH_MISMATCH",
"SNAPSHOT_PROFILE_UNSUPPORTED",
"SNAPSHOT_RULES_UNSUPPORTED",
"SNAPSHOT_RESOURCE_LIMIT"
]
},
"snapshot_sha256": {
"type": ["string", "null"],
"pattern": "^[0-9a-f]{64}$"
},
"manifest_sha256": {
"type": ["string", "null"],
"pattern": "^[0-9a-f]{64}$"
},
"snapshot_id": {
"type": ["string", "null"],
"pattern": "^[a-z0-9][a-z0-9._-]{0,127}$"
},
"source": {
"oneOf": [
{"type": "null"},
{
"type": "object",
"additionalProperties": false,
"required": ["name", "version", "as_of", "locator"],
"properties": {
"name": {
"type": "string",
"minLength": 1,
"maxLength": 200,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$"
},
"version": {
"type": "string",
"minLength": 1,
"maxLength": 100,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$"
},
"as_of": {"type": "string", "format": "date"},
"locator": {
"type": "string",
"minLength": 1,
"maxLength": 500,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$"
}
}
}
]
},
"supply_mode": {
"type": ["string", "null"],
"enum": ["user_supplied", "synthetic_fixture", null]
},
"snapshot_schema_version": {
"type": ["string", "null"],
"enum": ["tortured-phrase-snapshot/1.0", null]
},
"grammar_profile": {
"type": ["string", "null"],
"enum": ["ars-tortured-phrase-canonical-ast/1.0", null]
},
"normalizer_profile": {
"type": ["string", "null"],
"enum": ["ars-nfkc-casefold-token/1.0", null]
},
"unicode_data_version": {
"type": "string",
"pattern": "^[0-9]+\\.[0-9]+\\.[0-9]+$"
},
"rule_count": {
"type": ["integer", "null"],
"minimum": 1,
"maximum": 512
},
"unsupported_rule_count": {
"type": ["integer", "null"],
"minimum": 0
},
"rights": {
"oneOf": [
{"type": "null"},
{
"type": "object",
"additionalProperties": false,
"required": [
"basis",
"redistribution_status",
"reference",
"user_declaration"
],
"properties": {
"basis": {
"enum": [
"synthetic_fixture",
"user_declared_authorized",
"written_permission",
"unresolved"
]
},
"redistribution_status": {
"enum": ["permitted", "not_permitted", "unresolved"]
},
"reference": {
"oneOf": [
{
"type": "string",
"minLength": 1,
"maxLength": 500,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$"
},
{"type": "null"}
]
},
"user_declaration": {
"oneOf": [
{
"type": "string",
"minLength": 1,
"maxLength": 500,
"pattern": "^[^\\u0000-\\u001F\\u007F]+$"
},
{"type": "null"}
]
}
}
}
]
}
},
"allOf": [
{
"if": {"properties": {"status": {"const": "loaded"}}},
"then": {
"properties": {
"reason_code": {"const": "CHECK_COMPLETED"},
"snapshot_sha256": {"type": "string"},
"manifest_sha256": {"type": "string"},
"snapshot_id": {"type": "string"},
"source": {"type": "object"},
"supply_mode": {"type": "string"},
"snapshot_schema_version": {"type": "string"},
"grammar_profile": {"type": "string"},
"normalizer_profile": {"type": "string"},
"rule_count": {"type": "integer"},
"unsupported_rule_count": {"const": 0},
"rights": {"type": "object"}
}
},
"else": {
"properties": {
"snapshot_id": {"type": "null"},
"source": {"type": "null"},
"supply_mode": {"type": "null"},
"snapshot_schema_version": {"type": "null"},
"grammar_profile": {"type": "null"},
"normalizer_profile": {"type": "null"},
"rule_count": {"type": "null"},
"unsupported_rule_count": {"type": "null"},
"rights": {"type": "null"}
}
}
},
{
"if": {"properties": {"status": {"const": "not_checked"}}},
"then": {
"properties": {
"reason_code": {"const": "SNAPSHOT_NOT_PROVIDED"},
"snapshot_sha256": {"type": "null"},
"manifest_sha256": {"type": "null"}
}
}
},
{
"if": {"properties": {"status": {"const": "degraded"}}},
"then": {
"properties": {
"reason_code": {
"enum": [
"SNAPSHOT_BYTES_INVALID",
"SNAPSHOT_MANIFEST_INVALID",
"SNAPSHOT_HASH_MISMATCH",
"SNAPSHOT_PROFILE_UNSUPPORTED",
"SNAPSHOT_RULES_UNSUPPORTED",
"SNAPSHOT_RESOURCE_LIMIT"
]
}
}
}
},
{
"if": {"properties": {"supply_mode": {"const": "synthetic_fixture"}}},
"then": {
"properties": {
"rights": {
"properties": {
"basis": {"const": "synthetic_fixture"},
"redistribution_status": {"const": "permitted"},
"reference": {"type": "null"},
"user_declaration": {"type": "null"}
}
}
}
}
},
{
"if": {
"properties": {
"rights": {
"properties": {"basis": {"const": "user_declared_authorized"}},
"required": ["basis"]
}
},
"required": ["rights"]
},
"then": {
"properties": {
"rights": {"properties": {"user_declaration": {"type": "string"}}}
}
}
},
{
"if": {
"properties": {
"rights": {
"properties": {"basis": {"const": "written_permission"}},
"required": ["basis"]
}
},
"required": ["rights"]
},
"then": {
"properties": {
"rights": {"properties": {"reference": {"type": "string"}}}
}
}
},
{
"if": {
"properties": {
"rights": {
"properties": {"basis": {"const": "unresolved"}},
"required": ["basis"]
}
},
"required": ["rights"]
},
"then": {
"properties": {
"rights": {
"properties": {"redistribution_status": {"const": "unresolved"}}
}
}
}
}
]
},
"reason_code": {
"enum": [
"CHECK_COMPLETED",
"ABSTRACT_MISSING",
"ABSTRACT_EMPTY",
"SNAPSHOT_NOT_PROVIDED",
"SNAPSHOT_BYTES_INVALID",
"SNAPSHOT_MANIFEST_INVALID",
"SNAPSHOT_HASH_MISMATCH",
"SNAPSHOT_PROFILE_UNSUPPORTED",
"SNAPSHOT_RULES_UNSUPPORTED",
"SNAPSHOT_RESOURCE_LIMIT",
"MATCH_RESOURCE_LIMIT"
]
},
"counts": {
"type": "object",
"additionalProperties": false,
"required": [
"rules_evaluated",
"matched_rule_count",
"rule_match_count",
"unique_instance_count",
"segments_total",
"unknown_segments",
"matches_by_context"
],
"properties": {
"rules_evaluated": {"type": "integer", "minimum": 0, "maximum": 512},
"matched_rule_count": {"type": "integer", "minimum": 0, "maximum": 512},
"rule_match_count": {"type": "integer", "minimum": 0, "maximum": 4096},
"unique_instance_count": {"type": "integer", "minimum": 0, "maximum": 4096},
"segments_total": {"type": "integer", "minimum": 0, "maximum": 4096},
"unknown_segments": {"type": "integer", "minimum": 0, "maximum": 4096},
"matches_by_context": {
"type": "object",
"additionalProperties": false,
"required": [
"author_prose",
"quote",
"cited_title",
"reference_entry",
"code_or_verbatim",
"unknown",
"cited_abstract"
],
"properties": {
"author_prose": {"type": "integer", "minimum": 0},
"quote": {"type": "integer", "minimum": 0},
"cited_title": {"type": "integer", "minimum": 0},
"reference_entry": {"type": "integer", "minimum": 0},
"code_or_verbatim": {"type": "integer", "minimum": 0},
"unknown": {"type": "integer", "minimum": 0},
"cited_abstract": {"type": "integer", "minimum": 0}
}
}
}
},
"matches": {
"type": "array",
"maxItems": 4096,
"items": {
"type": "object",
"additionalProperties": false,
"required": [
"match_id",
"pattern_id",
"pattern_sha256",
"segment_id",
"context",
"disposition",
"source_span",
"matched_text",
"matched_text_sha256"
],
"properties": {
"match_id": {"type": "string", "pattern": "^tpm-[0-9a-f]{24}$"},
"pattern_id": {"type": "string", "pattern": "^[a-z0-9][a-z0-9._-]{0,127}$"},
"pattern_sha256": {"type": "string", "pattern": "^[0-9a-f]{64}$"},
"segment_id": {"type": "string", "pattern": "^SEG-[0-9]{6}$"},
"context": {"enum": ["cited_title", "cited_abstract"]},
"disposition": {
"enum": [
"preserve_verbatim_review_context",
"review_cited_source_no_automatic_rewrite"
]
},
"source_span": {
"type": "object",
"additionalProperties": false,
"required": ["codepoint_start", "codepoint_end", "utf8_start", "utf8_end"],
"properties": {
"codepoint_start": {"type": "integer", "minimum": 0},
"codepoint_end": {"type": "integer", "minimum": 1},
"utf8_start": {"type": "integer", "minimum": 0},
"utf8_end": {"type": "integer", "minimum": 1}
}
},
"matched_text": {"type": "string", "minLength": 1, "maxLength": 1000},
"matched_text_sha256": {"type": "string", "pattern": "^[0-9a-f]{64}$"}
}
}
},
"boundary": {
"type": "object",
"additionalProperties": false,
"required": [
"list_match_only",
"origin_inference",
"contextual_judgment",
"automatic_rewrite",
"absence_is_clean_certificate",
"native_pps_compatibility",
"sharing_scope"
],
"properties": {
"list_match_only": {"const": true},
"origin_inference": {"const": "not_performed"},
"contextual_judgment": {"const": "not_performed"},
"automatic_rewrite": {"const": false},
"absence_is_clean_certificate": {"const": false},
"native_pps_compatibility": {"const": "not_claimed"},
"sharing_scope": {"const": "local_only"}
}
}
}
},
@@ -387,6 +846,517 @@
}
}
},
{
"description": "A tortured-phrase list match is mechanically reproducible but remains an advisory risk marker. It can never be relabeled as a factual origin finding or promoted into terminal policy.",
"if": {
"properties": {
"signal_type": {"const": "tortured_phrase_match"}
},
"required": ["signal_type"]
},
"then": {
"properties": {
"epistemic_class": {"const": "heuristic_advisory"},
"epistemic_label": {"const": "HEURISTIC-INDICATOR"},
"evidence": {
"items": {
"properties": {
"evidence_type": {
"enum": ["phrase_match", "list_record", "degradation_record"]
}
}
}
},
"terminal_policy": {
"properties": {
"eligible": {"const": false},
"owner": {"const": "none"},
"policy_key": {"type": "null"},
"current_effect": {"const": "advisory_only"}
}
},
"display": {
"properties": {
"marker_token": {"type": "null"}
}
},
"retraction_context": false
}
}
},
{
"if": {
"properties": {
"schema_version": {"const": "bibliographic-integrity-signal/1.2"}
},
"required": ["schema_version"]
},
"then": {
"required": ["tortured_phrase_context"],
"properties": {
"signal_type": {"const": "tortured_phrase_match"},
"epistemic_class": {"const": "heuristic_advisory"},
"epistemic_label": {"const": "HEURISTIC-INDICATOR"},
"evidence": {"minItems": 1, "maxItems": 1},
"provenance": {
"properties": {
"checked_at": {
"pattern": "^[0-9]{4}-[0-9]{2}-[0-9]{2}[Tt](?:[01][0-9]|2[0-3]):[0-5][0-9]:[0-5][0-9](?:\\.[0-9]{1,6})?(?:[Zz]|[+-](?:[01][0-9]|2[0-3]):[0-5][0-9])$"
},
"recorded_at": {
"pattern": "^[0-9]{4}-[0-9]{2}-[0-9]{2}[Tt](?:[01][0-9]|2[0-3]):[0-5][0-9]:[0-5][0-9](?:\\.[0-9]{1,6})?(?:[Zz]|[+-](?:[01][0-9]|2[0-3]):[0-5][0-9])$"
},
"stale_after": {"type": "null"},
"freshness": {"const": "unknown"}
}
},
"subject": {
"properties": {
"source_pointer": {"type": "string", "minLength": 1},
"affected_claims": {"maxItems": 0}
}
},
"display": {
"properties": {
"summary_label": {"const": "Phrase-list screening advisory"},
"marker_token": {"type": "null"}
}
},
"retraction_context": false
}
},
"else": {
"properties": {
"tortured_phrase_context": false
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {"surface": {"const": "cited_title"}},
"required": ["surface"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"tortured_phrase_context": {
"properties": {
"surface_binding": {
"properties": {
"content_sha256": {"type": "string"},
"content_utf8_bytes": {"type": "integer", "minimum": 1}
}
},
"matches": {
"items": {
"properties": {
"segment_id": {"const": "SEG-000001"},
"context": {"const": "cited_title"},
"disposition": {"const": "preserve_verbatim_review_context"}
}
}
},
"counts": {
"properties": {
"matches_by_context": {
"properties": {
"author_prose": {"const": 0},
"quote": {"const": 0},
"reference_entry": {"const": 0},
"code_or_verbatim": {"const": 0},
"unknown": {"const": 0},
"cited_abstract": {"const": 0}
}
}
}
}
}
}
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {"reason_code": {"const": "SNAPSHOT_NOT_PROVIDED"}},
"required": ["reason_code"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"tortured_phrase_context": {
"properties": {
"snapshot": {
"properties": {
"status": {"const": "not_checked"},
"reason_code": {"const": "SNAPSHOT_NOT_PROVIDED"}
}
}
}
}
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {
"reason_code": {
"enum": [
"SNAPSHOT_BYTES_INVALID",
"SNAPSHOT_MANIFEST_INVALID",
"SNAPSHOT_HASH_MISMATCH",
"SNAPSHOT_PROFILE_UNSUPPORTED",
"SNAPSHOT_RULES_UNSUPPORTED",
"SNAPSHOT_RESOURCE_LIMIT"
]
}
},
"required": ["reason_code"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"tortured_phrase_context": {
"properties": {
"snapshot": {
"properties": {"status": {"const": "degraded"}}
}
},
"oneOf": [
{
"properties": {
"reason_code": {"const": "SNAPSHOT_BYTES_INVALID"},
"snapshot": {"properties": {"reason_code": {"const": "SNAPSHOT_BYTES_INVALID"}}}
}
},
{
"properties": {
"reason_code": {"const": "SNAPSHOT_MANIFEST_INVALID"},
"snapshot": {"properties": {"reason_code": {"const": "SNAPSHOT_MANIFEST_INVALID"}}}
}
},
{
"properties": {
"reason_code": {"const": "SNAPSHOT_HASH_MISMATCH"},
"snapshot": {"properties": {"reason_code": {"const": "SNAPSHOT_HASH_MISMATCH"}}}
}
},
{
"properties": {
"reason_code": {"const": "SNAPSHOT_PROFILE_UNSUPPORTED"},
"snapshot": {"properties": {"reason_code": {"const": "SNAPSHOT_PROFILE_UNSUPPORTED"}}}
}
},
{
"properties": {
"reason_code": {"const": "SNAPSHOT_RULES_UNSUPPORTED"},
"snapshot": {"properties": {"reason_code": {"const": "SNAPSHOT_RULES_UNSUPPORTED"}}}
}
},
{
"properties": {
"reason_code": {"const": "SNAPSHOT_RESOURCE_LIMIT"},
"snapshot": {"properties": {"reason_code": {"const": "SNAPSHOT_RESOURCE_LIMIT"}}}
}
}
]
}
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {"reason_code": {"const": "MATCH_RESOURCE_LIMIT"}},
"required": ["reason_code"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"tortured_phrase_context": {
"properties": {
"snapshot": {
"properties": {
"status": {"const": "loaded"},
"reason_code": {"const": "CHECK_COMPLETED"}
}
}
}
}
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {"surface": {"const": "cited_abstract"}},
"required": ["surface"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"tortured_phrase_context": {
"properties": {
"matches": {
"items": {
"properties": {
"segment_id": {"const": "SEG-000001"},
"context": {"const": "cited_abstract"},
"disposition": {"const": "review_cited_source_no_automatic_rewrite"}
}
}
},
"counts": {
"properties": {
"matches_by_context": {
"properties": {
"author_prose": {"const": 0},
"quote": {"const": 0},
"cited_title": {"const": 0},
"reference_entry": {"const": 0},
"code_or_verbatim": {"const": 0},
"unknown": {"const": 0}
}
}
}
}
}
}
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {"reason_code": {"const": "CHECK_COMPLETED"}},
"required": ["reason_code"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"tortured_phrase_context": {
"properties": {
"surface_binding": {
"properties": {
"content_sha256": {"type": "string"},
"content_utf8_bytes": {"type": "integer"}
}
},
"snapshot": {
"properties": {
"status": {"const": "loaded"},
"reason_code": {"const": "CHECK_COMPLETED"},
"rule_count": {"type": "integer", "minimum": 1}
}
},
"counts": {
"properties": {
"rules_evaluated": {"minimum": 1},
"segments_total": {"const": 1},
"unknown_segments": {"const": 0}
}
}
}
}
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {
"surface_binding": {
"properties": {"content_sha256": {"type": "null"}},
"required": ["content_sha256"]
}
},
"required": ["surface_binding"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"tortured_phrase_context": {
"properties": {
"surface": {"const": "cited_abstract"},
"surface_binding": {
"properties": {"content_utf8_bytes": {"type": "null"}}
},
"reason_code": {"enum": ["ABSTRACT_MISSING", "ABSTRACT_EMPTY"]}
}
}
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {
"reason_code": {"not": {"const": "CHECK_COMPLETED"}}
},
"required": ["reason_code"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"tortured_phrase_context": {
"properties": {
"counts": {
"properties": {
"rules_evaluated": {"const": 0},
"matched_rule_count": {"const": 0},
"rule_match_count": {"const": 0},
"unique_instance_count": {"const": 0},
"segments_total": {"const": 0},
"unknown_segments": {"const": 0},
"matches_by_context": {
"properties": {
"author_prose": {"const": 0},
"quote": {"const": 0},
"cited_title": {"const": 0},
"reference_entry": {"const": 0},
"code_or_verbatim": {"const": 0},
"unknown": {"const": 0},
"cited_abstract": {"const": 0}
}
}
}
},
"matches": {"maxItems": 0}
}
}
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {
"reason_code": {"const": "CHECK_COMPLETED"},
"counts": {
"properties": {
"rule_match_count": {"minimum": 1}
},
"required": ["rule_match_count"]
}
},
"required": ["reason_code", "counts"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"check_status": {"const": "checked"},
"finding": {"const": "detected"}
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {
"reason_code": {"const": "CHECK_COMPLETED"},
"counts": {
"properties": {
"rule_match_count": {"const": 0}
},
"required": ["rule_match_count"]
}
},
"required": ["reason_code", "counts"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"check_status": {"const": "checked"},
"finding": {"const": "not_detected"}
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {
"reason_code": {
"enum": [
"ABSTRACT_MISSING",
"ABSTRACT_EMPTY",
"SNAPSHOT_NOT_PROVIDED"
]
}
},
"required": ["reason_code"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"check_status": {"const": "not_checked"},
"finding": {"const": "unresolved"}
}
}
},
{
"if": {
"properties": {
"tortured_phrase_context": {
"properties": {
"reason_code": {
"enum": [
"SNAPSHOT_BYTES_INVALID",
"SNAPSHOT_MANIFEST_INVALID",
"SNAPSHOT_HASH_MISMATCH",
"SNAPSHOT_PROFILE_UNSUPPORTED",
"SNAPSHOT_RULES_UNSUPPORTED",
"SNAPSHOT_RESOURCE_LIMIT",
"MATCH_RESOURCE_LIMIT"
]
}
},
"required": ["reason_code"]
}
},
"required": ["tortured_phrase_context"]
},
"then": {
"properties": {
"check_status": {"const": "degraded"},
"finding": {"const": "unresolved"}
}
}
},
{
"if": {
"properties": {
+27
View File
@@ -704,6 +704,33 @@ Consumer integration ships in v3.6.5: `bibliography_agent` (deep-research, Phase
See [`academic-pipeline/references/adapters/overview.md`](../academic-pipeline/references/adapters/overview.md) for the adapter contract.
### Tortured-Phrase Advisory Extension (#660)
An entry's existing optional `bibliographic_integrity_signals[]` carrier may
hold `bibliographic-integrity-signal/1.2` tortured-phrase rows. When the local
check is invoked, there is one current row for `cited_title` and one for
`cited_abstract`; the surfaces never share a rolled-up status. A missing
abstract is retained as `not_checked` / `unresolved` with
`ABSTRACT_MISSING`; a present whitespace-only abstract uses `ABSTRACT_EMPTY`.
A checked zero-match row reports only no observed match on
the exact hash-bound surface and is not a clean certificate.
The producer consumes an explicitly named, exact-byte-SHA-256-bound
user-supplied or synthetic snapshot/manifest pair. It has no native PPS
importer or fetch path, redistributes no PPS list content, and invokes no
model, API, judge, or ambient clock. It writes a new passport copy rather than
changing the source passport, title, abstract, or citation in place. Phase 1
corpus consumers remain read-only and do not use this heuristic to include,
exclude, rank, rewrite, or label a source's origin.
Rows remain `HEURISTIC-INDICATOR` with a closed
`HEURISTIC-ADVISORY` / `UNMEASURED` context. They render only in the single
`Bibliographic Integrity Advisories` section, never as a reference marker,
terminal policy, gate, replacement, or rewrite. The separate own-draft
`tortured-phrase-advisory/1.0` artifact is not a Schema 9 field and carries no
paper-mill, AI/author-origin, cleanliness, contextual-validity, or accuracy
claim. Authority: [`shared/bibliographic_integrity_signals.md`](bibliographic_integrity_signals.md).
### Audit Artifact Ledger (v3.6.7)
Schema 9 gains an optional append-only `audit_artifact[]` ledger recording cross-model audit runs that gate the three v3.6.7 downstream agents (`synthesis_agent`, `research_architect_agent` survey-designer mode, `report_compiler_agent` abstract-only mode). Each entry conforms to [`shared/contracts/passport/audit_artifact_entry.schema.json`](contracts/passport/audit_artifact_entry.schema.json).