diff --git a/docs/adr/0013-unified-gesture-plans.md b/docs/adr/0013-unified-gesture-plans.md index d9ab8c563..fd1d5c458 100644 --- a/docs/adr/0013-unified-gesture-plans.md +++ b/docs/adr/0013-unified-gesture-plans.md @@ -114,7 +114,10 @@ Platform adapters consume the canonical plan: roughly 16 ms samples using smoothstep `s(t) = 3t² - 2t³`. A linear segment has constant movement velocity through its endpoint unless a destination hold follows it; smoothstep has zero slope at both endpoints and a peak velocity 1.5 times its average. Identical endpoints and total durations - can consequently produce different recognizer and deceleration outcomes. + can consequently produce different recognizer and deceleration outcomes. Neither a destination + hold nor an analytically zero endpoint slope proves a controlled release by itself: XCTest event + sampling and app recognizer thresholds can still leave observable post-lift motion, so live + evidence must measure the resulting content offset after pointer-up. Live iOS characterization in [issue #1586](https://github.com/callstack/agent-device/issues/1586) confirmed that distinction: the schedules crossed the same fling-recognizer thresholds in the @@ -192,8 +195,9 @@ resolution and recording contracts are tested at that composition seam. plans remain cadence-bounded because their synchronized geometry is part of their contract. - Unit tests cover canonical plan shape, Android lowering, helper/provider payloads, and WebDriver action construction. They cannot prove timing or event delivery inside the private XCTest bridge, - so iOS timing changes require live simulator evidence that observes the requested content change - and records the runner start/end uptime delta alongside the requested duration. + so iOS timing changes require live simulator evidence from an independent app-observed + postcondition. Runner start/end uptime deltas can supplement that evidence, but request echoes or + internal timing alone do not prove that the app recognized the intended gesture. ## Alternatives Considered diff --git a/docs/agents/testing.md b/docs/agents/testing.md index 2eb5915f7..e0d562f3a 100644 --- a/docs/agents/testing.md +++ b/docs/agents/testing.md @@ -87,6 +87,16 @@ run the smallest owning test, record the failing test count, then restore it. Ap test relocation and structural gates: plant a type error or violation and verify the intended gate discovers and names it. +A callback-based canary must observe the subject's semantic success state, not merely lifecycle +completion. For example, React Native Gesture Handler's +[`onFinalize`](https://docs.swmansion.com/react-native-gesture-handler/docs/fundamentals/callbacks-events/) +also runs when recognition fails or is interrupted; use an activation-dependent callback or assert +the callback's success state before publishing a passing result. + +A device replay is automatic regression coverage only when an automatic PR or scheduled lane selects +and executes it. Name the owning lane and confirm the scenario ran on the exact PR head; placing a +replay in a manual or otherwise unselected tier is test material, not automatic evidence. + For structured classifiers, pair the positive case with the closest negative case. When an error message can be identical with and without a typed reason, the negative test must prove the message alone cannot activate retry, fallback, or recovery.