docs(sync) + chore: final verification, docs, and stale-bundle fix (Phase 6)

Full-matrix e2e harness (scripts/sync-matrix-e2e.ts, 39/39): fresh-device
bootstrap, >500-op offline catch-up, concurrent writes, all four mutation
types in both arrival orders, epoch reset with native-corpus re-push, and
kill-switch degradation — plus a two-real-daemon run proving the
session-start context-inject pull end-to-end.

docs/public/cloud-sync.mdx rewritten for the two-lane architecture
(durable HTTP lanes, advisory WebSocket, poll-mode fallback, mutation
convergence, settings reference) with availability framed honestly —
no live hub URLs. Canary header comment scoped to what it measures.

plugin/scripts/worker-service.cjs regenerated — the Phase 5 commit had
shipped it stale (missing pollModeOnly/onSyncModeHint/X-Sync-Mode).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Alex Newman
2026-07-19 00:33:51 -07:00
parent 528f133ad1
commit fe9caa5c88
4 changed files with 704 additions and 193 deletions
+111 -64
View File
@@ -1,82 +1,130 @@
---
title: "Cloud Sync (cmem.ai Pro)"
description: "Worker-native sync of your memory database to your cmem.ai account — no daemon, the worker syncs on write."
description: "Two-lane multi-device sync of your memory database — durable HTTP push/pull through a per-user sync hub, with an optional WebSocket speed layer."
---
# Cloud Sync (cmem.ai Pro)
Cloud sync backs up your local memory database to your [cmem.ai](https://cmem.ai)
account. It is built into the worker: **there is no daemon — the worker syncs on
write**. The same process that records every observation also uploads it, so
there is no second process to install, monitor, or restart.
Cloud sync replicates your local memory database across your devices through a
**per-user sync hub** — an ordered log of everything your devices write. It is
built into the worker: **there is no daemon — the worker syncs on write**. The
same process that records every observation also uploads it, and pulls the
other devices' writes back down.
<Note>
**The database is the queue.** Every synced table carries a `synced_at` column
(`NULL` = not in the cloud yet). After each write, the worker nudges a debounced
flusher that drains `WHERE synced_at IS NULL` in batches to cmem.ai and stamps
rows on success. That one mechanism **is** live sync, backfill, offline
catch-up, and retry — a never-synced install simply has everything `NULL`, and
anything that fails to upload stays `NULL` until a later flush picks it up.
**Two lanes, one source of truth.** Everything durable travels over plain
HTTP: pushes append to the hub's ordered log, pulls page through it with a
cursor. An optional WebSocket rides alongside as a pure *speed layer* — it
only makes the next pull happen sooner. The socket can be dropped, disabled,
or wrong with **zero data loss**; the HTTP cursor is always the truth.
</Note>
## What syncs
Three kinds of rows are uploaded to your cmem.ai account:
Three kinds of rows are replicated between your devices:
- **Observations** — the compressed memories claude-mem generates from your sessions.
- **Session summaries** — the per-session overview records.
- **User prompts** — the prompts you typed (clamped to 200&nbsp;KB each).
- **User prompts** — the prompts you typed (clamped to 200&nbsp;KB per field).
Alongside the rows, three kinds of **mutations** travel through the same log
so later edits converge everywhere: session **title** changes, **prompt→session
repairs** (a prompt captured before its session registered is re-linked on
every device), and **project remaps** (worktree adoption and cwd-based moves
retarget the affected rows on every device).
<Warning>
**Privacy:** cloud sync uploads your observation narratives and your full prompt
text to your cmem.ai account. Don't enable it if that content must stay on your
machine.
**Privacy:** cloud sync uploads your observation narratives and your full
prompt text to the sync hub under your cmem.ai account. Don't enable it if
that content must stay on your machine.
</Warning>
## Quick start
## How the push lane works
Run the `/cloud-sync` skill in Claude Code. It checks sync status and, on first
run, walks you through setup:
1. Grab your **sync token** and **user id** from **cmem.ai → Connect** and paste
them when asked. The skill writes them into `~/.claude-mem/settings.json`
(mode `0600`) without ever echoing the token.
2. The skill restarts the worker, which immediately starts draining unsynced
rows, and polls status until the pending counts fall.
If you previously used the **legacy standalone sync client**, the skill migrates
it automatically: your token is read from the old `.cloud-sync.env`, your device
identity carries over, and **nothing re-uploads** — rows the old client already
pushed are stamped as synced during migration. See
[Migrating from the standalone client](#migrating-from-the-standalone-client).
## How the flusher behaves
**The database is the queue.** Every synced table carries a `synced_at` column
(`NULL` = not in the log yet). After each write, the worker nudges a debounced
flusher that drains `WHERE synced_at IS NULL` and stamps rows on success. That
one mechanism **is** live sync, backfill, offline catch-up, and retry — a
never-synced install simply has everything `NULL`, and anything that fails to
upload stays `NULL` until a later flush picks it up.
- **Debounced:** write bursts coalesce; the flusher runs ~1.5&nbsp;s after the
last write, and only one flush runs at a time.
- **Batched:** rows drain in batches of up to 200 per kind, packed into request
bodies of at most 2&nbsp;MB.
last write (250&nbsp;ms while the speed layer is connected), and only one
flush runs at a time.
- **Batched:** ops drain in requests of up to 500 ops / 2&nbsp;MB.
- **Timeboxed:** every request has a 30&nbsp;s timeout — a dead network can
never hang the worker.
- **Retrying:** a failed upload leaves rows `NULL` and retries on the next write
plus a capped exponential backoff (30&nbsp;s → 10&nbsp;min). The write path is
never blocked — sync failures cost nothing but delay.
- **Startup drain:** the worker kicks one flush at startup, which doubles as the
initial backfill of your existing database.
- **Retrying:** a failed upload leaves rows `NULL` and retries on the next
write plus a capped exponential backoff (30&nbsp;s → 10&nbsp;min). The write
path is never blocked — sync failures cost nothing but delay.
- **Idempotent:** the hub dedupes on origin identity, so a retried push can
never create duplicates.
## How the pull lane works
Each device keeps a **cursor** into the hub's log and pulls everything after
it — in pages, so a device that was offline for a week (or is brand new and
starts from zero) catches up the same way an active one stays current:
- **Session start:** the worker pulls immediately before injecting context
(bounded to 1.5&nbsp;s, so a dead network can't stall your session), which
means memory written on another machine is there the moment you start
working.
- **Steady state:** a short poll every 30&nbsp;s while a session is active,
every 5&nbsp;min when idle, suspended entirely after an hour of no activity.
Every push response also reports the log head, so an active device learns
about remote writes without waiting out the timer.
- **Applied like native writes:** pulled rows keep their original timestamps,
carry their origin device, index into full-text and vector search through
the same paths native writes use, and are stamped so they can never be
echoed back up.
## The speed layer (optional WebSocket)
When enabled (the default), the worker also holds one WebSocket to the hub and
receives new ops the moment another device pushes them — multi-device
convergence in well under a second instead of a poll interval. The socket is
strictly **advisory**: any anomaly (a gap, a parse error, a dropped
connection) simply closes it and runs one HTTP pull, which is always correct.
Set `CLAUDE_MEM_CLOUD_SYNC_WS` to `false` to turn it off — sync stays fully
correct at poll latency.
The hub can also ask every client to fall back to polling by stamping
responses with an `X-Sync-Mode: poll` header (an operational brake on the
server side). Clients close their sockets and keep syncing over HTTP —
degraded means *slower*, never *stopped* — and resume the socket automatically
when the header disappears.
## Requirements and setup
Cloud sync is active when **all three** of the token, the user id, and the hub
URL are non-empty — there is no separate enable flag; blank any one of them to
turn sync off.
1. A **cmem.ai Pro sync token and user id** (from **cmem.ai → Connect**).
2. A **sync hub URL** in `CLAUDE_MEM_CLOUD_SYNC_HUB_URL`. The hub is a
Cloudflare Workers service (source in `workers/sync-hub/` with its own
deploy runbook); use the hub URL your cmem.ai account documentation
provides for it.
Run the `/cloud-sync` skill in Claude Code to check status and walk through
setup. It writes the settings into `~/.claude-mem/settings.json` (mode
`0600`) without ever echoing the token, restarts the worker, and polls status
until the pending counts fall. If you previously used the legacy standalone
sync client, the skill retires it and carries your device identity over — see
[Migrating from the standalone client](#migrating-from-the-standalone-client).
## Settings
Cloud sync is configured in `~/.claude-mem/settings.json`. It is **active when
both the token and the user id are non-empty** — there is no separate enable
flag; blank the token to turn sync off.
| Key | Default | Meaning |
|-----|---------|---------|
| `CLAUDE_MEM_CLOUD_SYNC_TOKEN` | `''` | Your cmem.ai sync token (from **cmem.ai → Connect**). Sent as `Authorization: Bearer`. |
| `CLAUDE_MEM_CLOUD_SYNC_USER_ID` | `''` | Your cmem.ai user id (from **cmem.ai → Connect**). |
| `CLAUDE_MEM_CLOUD_SYNC_URL` | `https://cmem.ai/api/pro/sync` | Sync endpoint base URL. Only change this if you're pointing at a non-production server. |
| `CLAUDE_MEM_CLOUD_SYNC_DEVICE_ID` | `''` | Stable identity of this machine in the cloud. Resolved at first start — adopted from the legacy client's state file if present, otherwise a fresh UUID — then persisted back here. Don't edit it: the server keys rows on it, and a changed id forks every cloud row into a duplicate. |
| `CLAUDE_MEM_CLOUD_SYNC_DEVICE_NAME` | machine hostname | Human-readable label for this device in the cmem.ai Devices panel. |
| `CLAUDE_MEM_CLOUD_SYNC_USER_ID` | `''` | Your cmem.ai user id (from **cmem.ai → Connect**). All your devices share it — it names your hub log. |
| `CLAUDE_MEM_CLOUD_SYNC_HUB_URL` | `''` | Sync hub base URL. **Empty = sync off.** |
| `CLAUDE_MEM_CLOUD_SYNC_WS` | `'true'` | Advisory WebSocket speed layer. `'false'` = HTTP polling only; sync stays fully correct at poll latency. |
| `CLAUDE_MEM_CLOUD_SYNC_DEVICE_ID` | `''` | Stable identity of this machine. Resolved at first start — adopted from the legacy client's state file if present, otherwise a fresh UUID — then persisted back here. Don't edit it: every row is attributed to its origin device, and a changed id forks every previously synced row into a duplicate. |
| `CLAUDE_MEM_CLOUD_SYNC_DEVICE_NAME` | machine hostname | Human-readable label for this device. |
| `CLAUDE_MEM_CLOUD_SYNC_URL` | (legacy) | The pre-hub per-kind endpoint base. Unused since the hub cutover; kept only so existing settings files round-trip. |
## Status endpoint
@@ -95,14 +143,14 @@ Configured:
{
"configured": true,
"deviceId": "2f6b1c9e-7d41-4c1a-9b0e-3d5f8a2c6e10",
"pending": { "observations": 0, "summaries": 0, "prompts": 2 },
"pending": { "observations": 0, "summaries": 0, "prompts": 2, "mutations": 0 },
"lastFlushAt": 1783981042731,
"lastError": null
}
```
- `pending` — rows per kind still waiting to upload (`synced_at IS NULL`).
Counts near 0 mean the cloud copy is current.
- `pending` — rows (and queued mutation ops) still waiting to upload. Counts
near 0 mean the hub's log has everything this device wrote.
- `lastFlushAt` — epoch ms of the last successful flush, `null` before the first.
- `lastError` — message from the most recent failed flush, `null` when healthy.
It never contains the token.
@@ -117,23 +165,22 @@ Not configured:
Before sync moved into the worker, cloud sync ran as a separate standalone
client (`cloud-sync.mjs` plus a `.cloud-sync.env` credentials file and a
`cloud-sync.pid` file). The worker supersedes it entirely — if the old client
kept running it would double-upload — so the `/cloud-sync` skill retires it
during setup:
`cloud-sync.pid` file). The worker supersedes it entirely, so the
`/cloud-sync` skill retires it during setup:
- **Stopped:** if `cloud-sync.pid` points at a live daemon, it is killed.
- **Retired:** `cloud-sync.mjs`, `.cloud-sync.env`, and `cloud-sync.pid` are
renamed with a `.retired` suffix (archived, not deleted). Your token is copied
from `.cloud-sync.env` into `settings.json` first, so setup needs no re-pasting.
- **Preserved:** `~/.claude-mem/cloud-sync-state.json` is deliberately **left in
place**. The worker reads two things from it: its upload cursors (used once,
by a database migration, to stamp everything the old client already pushed —
this is why migration re-uploads nothing) and its `deviceId` (adopted into
`CLAUDE_MEM_CLOUD_SYNC_DEVICE_ID` so the cloud sees the same device before and
after). Renaming or deleting that file would mint a new device identity and
fork every previously synced row into a duplicate.
- **Preserved:** `~/.claude-mem/cloud-sync-state.json` is deliberately **left
in place** — the worker adopts its `deviceId` into
`CLAUDE_MEM_CLOUD_SYNC_DEVICE_ID` so this machine keeps one identity across
the migration. Renaming or deleting that file before the first worker start
would mint a new device identity and fork every previously synced row into a
duplicate.
<Note>
The standalone client remains available from cmem.ai as an out-of-repo utility
for non-plugin users. The plugin no longer depends on it.
On first contact with the hub, every device re-pushes its own corpus once
(the hub's dedupe makes this safe) — that is how the shared log is built, and
it doubles as the initial backfill of your existing database.
</Note>
File diff suppressed because one or more lines are too long
+462
View File
@@ -0,0 +1,462 @@
#!/usr/bin/env bun
// Full-matrix sync e2e (plan Phase 6 task 1) — real SessionStore + CloudSync +
// SyncApply + SyncClient stacks against a real `wrangler dev` sync hub. The
// script starts (and restarts) its own hub on a private port with a private
// --persist-to dir, so it never collides with a dev hub on 8787 and an epoch
// reset is just "wipe the persist dir and restart".
//
// Matrix scenarios (plan: "fresh device bootstrap (since=0), week-offline
// catch-up, concurrent two-device writes, all four mutation types, epoch
// reset, kill-switch degradation"):
// 1. Row replication A→B: preserved created_at_epoch, origin attribution,
// echo-guard stamp, FTS row, prompt linking, and the set_title PARK path
// (in the hub log set_title precedes the session's row ops, so a device
// that has never seen the session parks the title, then claims it when
// the stub session materializes — all inside one applied batch).
// 2. set_prompt_session repair: a prompt pushed before its memory id
// registers is re-linked on replicas by the repair mutation.
// 3. remap_project, BOTH shapes, emitted own-connection (the
// WorktreeAdoption/ProcessManager shape — pure SQL + sync_outbox):
// {project, merged_into_project_is_null} → {merged_into_project} and
// {memory_session_id} → {project} (which also retargets sdk_sessions).
// 4. set_title DIRECT path (the other arrival order): a mutation pushed by
// a third device targeting a session that ALREADY exists on A and B
// applies as a direct UPDATE on both.
// 5. Concurrent two-device writes: A and B write + flush simultaneously;
// both converge to the identical corpus — no echo, no loss, no dupes.
// 6. Week-offline catch-up: B stops pulling while A accumulates >500 ops
// (more than one /changes page — hub cap 500), then B converges through
// paginated pulls in one cycle.
// 7. Fresh device bootstrap: a brand-new device C (empty DB, cursor 0)
// pulls since=0 and converges on the full corpus, titles included.
// 8. Epoch reset: hub storage wiped → fresh DO mints a new epoch → clients
// detect the mismatch, reset cursors, requeue their NATIVE corpora for
// re-push (the Phase 3 fix), and all devices reconverge without dupes.
//
// Kill-switch degradation (matrix scenario 9) lives in its own harness —
// scripts/sync-kill-switch-e2e.ts — because tripping the switch requires
// `wrangler kv --local` against the hub's DEFAULT persist dir. Run both:
// bun scripts/sync-matrix-e2e.ts
// (cd workers/sync-hub && bunx wrangler dev --var KILL_SWITCH_CACHE_MS:0 &)
// bun scripts/sync-kill-switch-e2e.ts
//
// Env knobs: MATRIX_PORT (default 8794), MATRIX_KEEP_TMP=1 to keep temp dirs.
import { mkdtempSync, rmSync } from 'fs';
import { tmpdir } from 'os';
import { join, resolve } from 'path';
import { SessionStore } from '../src/services/sqlite/SessionStore.js';
import { SessionSearch } from '../src/services/sqlite/SessionSearch.js';
import { openConfiguredSqliteDatabase } from '../src/services/sqlite/connection.js';
import { CloudSync } from '../src/services/sync/CloudSync.js';
import { SyncApply } from '../src/services/sync/SyncApply.js';
import { SyncClient } from '../src/services/sync/SyncClient.js';
import { emitRemapProject } from '../src/services/sync/remap-outbox.js';
const HUB_DIR = resolve(import.meta.dir, '../workers/sync-hub');
const PORT = Number(process.env.MATRIX_PORT ?? 8794);
const HUB_URL = `http://127.0.0.1:${PORT}`;
const USER_ID = `matrix-user-${Date.now().toString(36)}`;
const TOKEN = 'matrix-token';
/** >500 forces at least two /changes pages (hub page cap is 500). */
const OFFLINE_OPS = 520;
const sleep = (ms: number) => new Promise(r => setTimeout(r, ms));
let failures = 0;
function check(name: string, cond: boolean, detail?: unknown): void {
if (cond) {
console.log(` PASS ${name}`);
} else {
failures++;
console.log(` FAIL ${name}${detail !== undefined ? `${JSON.stringify(detail)}` : ''}`);
}
}
// ---------------------------------------------------------------------------
// Hub lifecycle (private port + private persist dir)
// ---------------------------------------------------------------------------
let hubProc: ReturnType<typeof Bun.spawn> | null = null;
let persistDir = '';
function authHeaders(deviceId: string): Record<string, string> {
return {
'Authorization': `Bearer ${TOKEN}`,
'X-User-Id': USER_ID,
'X-Device-Id': deviceId,
};
}
async function startHub(): Promise<void> {
hubProc = Bun.spawn(
['bunx', 'wrangler', 'dev', '--port', String(PORT), '--persist-to', persistDir],
{ cwd: HUB_DIR, stdout: 'ignore', stderr: 'ignore', env: { ...process.env } },
);
for (let i = 0; i < 120; i++) {
try {
const res = await fetch(`${HUB_URL}/v1/sync/status`, {
headers: authHeaders('probe'),
signal: AbortSignal.timeout(1000),
});
if (res.ok) return;
} catch { /* not up yet */ }
await sleep(500);
}
throw new Error('wrangler dev did not become ready');
}
async function stopHub(): Promise<void> {
if (!hubProc) return;
hubProc.kill(); // SIGTERM — a SIGKILLed workerd can wedge wrangler state
await hubProc.exited.catch(() => {});
hubProc = null;
for (let i = 0; i < 40; i++) {
try {
await fetch(`${HUB_URL}/v1/sync/status`, { headers: authHeaders('probe'), signal: AbortSignal.timeout(500) });
} catch {
return; // port refuses connections — actually down
}
await sleep(250);
}
}
interface HubStatus { epoch: string; head_seq: number; op_count: number; device_count: number }
async function hubStatus(): Promise<HubStatus> {
const res = await fetch(`${HUB_URL}/v1/sync/status`, { headers: authHeaders('probe') });
if (!res.ok) throw new Error(`hub status ${res.status}`);
return await res.json() as HubStatus;
}
/** Raw push as an arbitrary device (the synthetic third device of scenario 4). */
async function rawPush(deviceId: string, ops: Array<{ kind: string; origin_id: string; rev: number; body: string }>): Promise<void> {
const res = await fetch(`${HUB_URL}/v1/sync/ops`, {
method: 'POST',
headers: { ...authHeaders(deviceId), 'Content-Type': 'application/json' },
body: JSON.stringify({ ops }),
signal: AbortSignal.timeout(10_000),
});
if (!res.ok) throw new Error(`raw push ${res.status}: ${(await res.text().catch(() => '')).slice(0, 200)}`);
}
// ---------------------------------------------------------------------------
// Device harness — the real client stack, deterministic (no started loops)
// ---------------------------------------------------------------------------
interface Device {
name: string;
dir: string;
store: SessionStore;
search: SessionSearch;
cloudSync: CloudSync;
apply: SyncApply;
client: SyncClient;
}
function makeDevice(name: string): Device {
const dir = mkdtempSync(join(tmpdir(), `claude-mem-matrix-${name}-`));
const dbPath = join(dir, 'claude-mem.db');
const store = new SessionStore(dbPath, {
cloudSyncStatePath: join(dir, 'no-legacy.json'),
cloudSyncHubUrl: HUB_URL,
});
const search = new SessionSearch(store.db); // FTS tables + triggers
const cloudSync = new CloudSync(store.db, {
CLAUDE_MEM_CLOUD_SYNC_TOKEN: TOKEN,
CLAUDE_MEM_CLOUD_SYNC_USER_ID: USER_ID,
CLAUDE_MEM_CLOUD_SYNC_HUB_URL: HUB_URL,
CLAUDE_MEM_CLOUD_SYNC_DEVICE_ID: `device-${name}`,
CLAUDE_MEM_CLOUD_SYNC_DEVICE_NAME: `matrix-${name}`,
}, {
settingsPath: join(dir, 'settings.json'),
legacyStatePath: join(dir, 'no-legacy.json'),
debounceMs: 50,
backoffInitialMs: 100,
backoffMaxMs: 1000,
requestTimeoutMs: 15_000,
});
const apply = new SyncApply(store.db, { deviceId: `device-${name}` });
const client = new SyncClient(apply, {
hubUrl: HUB_URL,
token: TOKEN,
userId: USER_ID,
deviceId: `device-${name}`,
minPullGapMs: 0,
requestTimeoutMs: 15_000,
wsEnabled: false, // deterministic HTTP-lane matrix; the WS lane has its own e2e
});
return { name, dir, store, search, cloudSync, apply, client };
}
function destroyDevice(dev: Device): void {
dev.cloudSync.stop();
dev.client.stop();
dev.store.db.close();
if (process.env.MATRIX_KEEP_TMP !== '1') {
rmSync(dev.dir, { recursive: true, force: true });
}
}
function q<T>(dev: Device, sql: string, ...params: unknown[]): T | undefined {
return dev.store.db.prepare(sql).get(...params as never[]) as T | undefined;
}
function count(dev: Device, sql: string, ...params: unknown[]): number {
return (q<{ n: number }>(dev, sql, ...params) ?? { n: -1 }).n;
}
function pendingTotal(dev: Device): number {
const p = dev.cloudSync.status().pending;
return p.observations + p.summaries + p.prompts + p.mutations;
}
function corpusCounts(dev: Device): { obs: number; sums: number; prompts: number } {
return {
obs: count(dev, 'SELECT COUNT(*) AS n FROM observations'),
sums: count(dev, 'SELECT COUNT(*) AS n FROM session_summaries'),
prompts: count(dev, 'SELECT COUNT(*) AS n FROM user_prompts'),
};
}
function makeObs(title: string, narrative: string): Parameters<SessionStore['storeObservation']>[2] {
return {
type: 'discovery',
title,
subtitle: null,
facts: [],
narrative,
concepts: [],
files_read: [],
files_modified: [],
};
}
// ---------------------------------------------------------------------------
async function main(): Promise<void> {
persistDir = mkdtempSync(join(tmpdir(), 'claude-mem-matrix-hub-'));
console.log(`Sync matrix e2e — hub ${HUB_URL}, user ${USER_ID}`);
console.log('[hub] starting wrangler dev...');
await startHub();
console.log('[hub] ready');
const A = makeDevice('a');
const B = makeDevice('b');
let C: Device | null = null;
try {
// === Scenario 1: row replication + set_title park path =================
console.log('\nScenario 1: row replication A -> B (epoch, origin, FTS, parked title)');
const sessA = A.store.createSDKSession('sess-matrix-1', 'proj-matrix', 'hello matrix', 'Matrix Title vA', 'claude');
A.store.updateMemorySessionId(sessA, 'mem-matrix-1');
A.store.saveUserPrompt('sess-matrix-1', 1, 'find the flux capacitor', sessA);
const obsA = A.store.storeObservation('mem-matrix-1', 'proj-matrix', makeObs('Observation from A', 'the zeppelin narrative body'), 1, 7);
A.store.storeSummary('mem-matrix-1', 'proj-matrix', {
request: 'summarize matrix', investigated: 'inv', learned: 'lrn',
completed: 'done', next_steps: 'next', notes: null,
}, 1);
await A.cloudSync.flush();
check('A drain empty after flush', pendingTotal(A) === 0, A.cloudSync.status());
await B.client.pullOnce({ timeoutMs: 15_000 });
const obsB = q<{ title: string; created_at_epoch: number; origin_device_id: string; synced_at: number }>(
B, `SELECT title, created_at_epoch, origin_device_id, synced_at FROM observations WHERE origin_local_id = ?`, String(obsA.id));
check('observation replicated to B', obsB?.title === 'Observation from A');
check('created_at_epoch preserved', obsB?.created_at_epoch === obsA.createdAtEpoch, { got: obsB?.created_at_epoch, want: obsA.createdAtEpoch });
check('origin attribution recorded', obsB?.origin_device_id === 'device-a');
check('applied row pre-stamped synced_at (echo guard)', obsB?.synced_at != null);
check('FTS row present on B', count(B, `SELECT COUNT(*) AS n FROM observations_fts WHERE observations_fts MATCH 'zeppelin'`) >= 1);
check('summary replicated to B', count(B, `SELECT COUNT(*) AS n FROM session_summaries WHERE origin_device_id = 'device-a'`) === 1);
const promptB = q<{ session_db_id: number | null }>(B, `SELECT session_db_id FROM user_prompts WHERE prompt_text = 'find the flux capacitor'`);
check('prompt replicated + linked on B', promptB?.session_db_id != null);
const titleB = q<{ custom_title: string | null }>(B, `SELECT custom_title FROM sdk_sessions WHERE memory_session_id = 'mem-matrix-1'`);
check('set_title converged via park->claim (log order: title before rows)', titleB?.custom_title === 'Matrix Title vA', titleB);
// === Scenario 2: set_prompt_session repair =============================
console.log('\nScenario 2: set_prompt_session repair');
const sess2 = A.store.createSDKSession('sess-matrix-2', 'proj-matrix', 'second session');
A.store.saveUserPrompt('sess-matrix-2', 1, 'early unlinked prompt', sess2);
await A.cloudSync.flush(); // prompt travels with NULL join fields
A.store.updateMemorySessionId(sess2, 'mem-matrix-2'); // emits the repair op
await A.cloudSync.flush();
await B.client.pullOnce({ timeoutMs: 15_000 });
const repaired = q<{ session_db_id: number | null }>(B, `SELECT session_db_id FROM user_prompts WHERE prompt_text = 'early unlinked prompt'`);
const repairedSession = repaired?.session_db_id != null
? q<{ memory_session_id: string | null }>(B, `SELECT memory_session_id FROM sdk_sessions WHERE id = ?`, repaired.session_db_id)
: undefined;
check('repair re-linked the replica prompt to mem-matrix-2', repairedSession?.memory_session_id === 'mem-matrix-2', repaired);
// === Scenario 3: remap_project, both shapes, own-connection ============
console.log('\nScenario 3: remap_project (worktree shape + cwd shape)');
// Worktree-adoption shape: separate connection to A's DB, inside a tx.
let ownConn = openConfiguredSqliteDatabase(join(A.dir, 'claude-mem.db'));
try {
ownConn.transaction(() => {
emitRemapProject(ownConn, { project: 'proj-matrix', merged_into_project_is_null: true }, { merged_into_project: 'proj-parent' });
})();
} finally {
ownConn.close();
}
await A.cloudSync.flush(); // the startup-drain pickup of the outbox
await B.client.pullOnce({ timeoutMs: 15_000 });
check('worktree-shape remap converged on B (merged_into_project)',
count(B, `SELECT COUNT(*) AS n FROM observations WHERE merged_into_project = 'proj-parent'`) >= 1);
// Cwd-remap shape: {memory_session_id} -> {project}; also retargets the
// owning sdk_sessions row on apply.
ownConn = openConfiguredSqliteDatabase(join(A.dir, 'claude-mem.db'));
try {
ownConn.transaction(() => {
emitRemapProject(ownConn, { memory_session_id: 'mem-matrix-1' }, { project: 'proj-moved' });
})();
} finally {
ownConn.close();
}
await A.cloudSync.flush();
await B.client.pullOnce({ timeoutMs: 15_000 });
check('cwd-shape remap converged on B (observations.project)',
count(B, `SELECT COUNT(*) AS n FROM observations WHERE memory_session_id = 'mem-matrix-1' AND project = 'proj-moved'`) >= 1);
const remappedSession = q<{ project: string }>(B, `SELECT project FROM sdk_sessions WHERE memory_session_id = 'mem-matrix-1'`);
check('cwd-shape remap retargeted the session row on B', remappedSession?.project === 'proj-moved', remappedSession);
// === Scenario 4: set_title direct path (the other order) ===============
console.log('\nScenario 4: set_title direct path (session already exists on both)');
await rawPush('device-x', [{
kind: 'mutation',
origin_id: crypto.randomUUID(),
rev: 1,
body: JSON.stringify({
op: 'set_title',
target: { memory_session_id: 'mem-matrix-1' },
fields: { custom_title: 'Retitled by X' },
}),
}]);
await A.client.pullOnce({ timeoutMs: 15_000 });
await B.client.pullOnce({ timeoutMs: 15_000 });
check('direct set_title applied on A',
q<{ custom_title: string }>(A, `SELECT custom_title FROM sdk_sessions WHERE memory_session_id = 'mem-matrix-1'`)?.custom_title === 'Retitled by X');
check('direct set_title applied on B',
q<{ custom_title: string }>(B, `SELECT custom_title FROM sdk_sessions WHERE memory_session_id = 'mem-matrix-1'`)?.custom_title === 'Retitled by X');
// === Scenario 5: concurrent two-device writes ==========================
console.log('\nScenario 5: concurrent two-device writes (30 + 30, overlapping flushes)');
const sessCA = A.store.createSDKSession('sess-conc-a', 'proj-conc', 'conc a');
A.store.updateMemorySessionId(sessCA, 'mem-conc-a');
const sessCB = B.store.createSDKSession('sess-conc-b', 'proj-conc', 'conc b');
B.store.updateMemorySessionId(sessCB, 'mem-conc-b');
await Promise.all([A.cloudSync.flush(), B.cloudSync.flush()]); // session bootstrap ops
const writer = async (dev: Device, memId: string, prefix: string): Promise<void> => {
for (let i = 0; i < 30; i++) {
dev.store.storeObservation(memId, 'proj-conc', makeObs(`${prefix}-${i}`, `${prefix} narrative ${i}`), 1, 0);
if (i % 5 === 4) {
await Promise.all([dev.cloudSync.flush(), sleep(10)]); // overlap pushes
}
}
await dev.cloudSync.flush();
};
await Promise.all([writer(A, 'mem-conc-a', 'conc-a'), writer(B, 'mem-conc-b', 'conc-b')]);
await A.client.pullOnce({ timeoutMs: 15_000 });
await B.client.pullOnce({ timeoutMs: 15_000 });
const concCountA = count(A, `SELECT COUNT(*) AS n FROM observations WHERE project = 'proj-conc'`);
const concCountB = count(B, `SELECT COUNT(*) AS n FROM observations WHERE project = 'proj-conc'`);
check('A holds all 60 concurrent rows exactly once', concCountA === 60, { concCountA });
check('B holds all 60 concurrent rows exactly once', concCountB === 60, { concCountB });
const distinctA = count(A, `SELECT COUNT(DISTINCT title) AS n FROM observations WHERE project = 'proj-conc'`);
const distinctB = count(B, `SELECT COUNT(DISTINCT title) AS n FROM observations WHERE project = 'proj-conc'`);
check('no loss, no dupes (60 distinct titles on both)', distinctA === 60 && distinctB === 60, { distinctA, distinctB });
const opsBeforeEcho = (await hubStatus()).op_count;
await Promise.all([A.cloudSync.flush(), B.cloudSync.flush()]);
const opsAfterEcho = (await hubStatus()).op_count;
check('no echo: hub op_count stable after both re-flush', opsBeforeEcho === opsAfterEcho, { opsBeforeEcho, opsAfterEcho });
// === Scenario 6: week-offline catch-up (pagination) ====================
console.log(`\nScenario 6: week-offline catch-up (B dark while A writes ${OFFLINE_OPS} ops)`);
const sessBulk = A.store.createSDKSession('sess-bulk', 'proj-bulk', 'bulk');
A.store.updateMemorySessionId(sessBulk, 'mem-bulk');
for (let i = 0; i < OFFLINE_OPS; i++) {
A.store.storeObservation('mem-bulk', 'proj-bulk', makeObs(`bulk-${i}`, `offline backlog item ${i}`), 1, 0);
}
await A.cloudSync.flush();
check('A drained the backlog', pendingTotal(A) === 0, A.cloudSync.status());
const cursorBefore = B.apply.getCursor();
const headBefore = (await hubStatus()).head_seq;
check(`B is >500 ops behind (multiple /changes pages)`, headBefore - cursorBefore > 500, { headBefore, cursorBefore });
await B.client.pullOnce({ timeoutMs: 60_000 });
check('B converged on the full backlog', count(B, `SELECT COUNT(*) AS n FROM observations WHERE project = 'proj-bulk'`) === OFFLINE_OPS);
check('B cursor reached hub head', B.apply.getCursor() === headBefore, { cursor: B.apply.getCursor(), headBefore });
// === Scenario 7: fresh device bootstrap (since=0) ======================
console.log('\nScenario 7: fresh device bootstrap (empty DB, since=0)');
C = makeDevice('c');
check('C starts at cursor 0', C.apply.getCursor() === 0);
await C.client.pullOnce({ timeoutMs: 60_000 });
const bCounts = corpusCounts(B);
const cCounts = corpusCounts(C);
check('C converged on the full corpus (matches B)',
cCounts.obs === bCounts.obs && cCounts.sums === bCounts.sums && cCounts.prompts === bCounts.prompts,
{ cCounts, bCounts });
check('C sees the final title (full-log replay: park, claim, then overwrite)',
q<{ custom_title: string }>(C, `SELECT custom_title FROM sdk_sessions WHERE memory_session_id = 'mem-matrix-1'`)?.custom_title === 'Retitled by X');
check('C sees the remap state',
count(C, `SELECT COUNT(*) AS n FROM observations WHERE merged_into_project = 'proj-parent'`) >= 1
&& q<{ project: string }>(C, `SELECT project FROM sdk_sessions WHERE memory_session_id = 'mem-matrix-1'`)?.project === 'proj-moved');
check('C adopted the hub epoch', C.apply.getEpoch() === (await hubStatus()).epoch);
check('C pushed nothing (pure replica)', pendingTotal(C) === 0, C.cloudSync.status());
// === Scenario 8: epoch reset ==========================================
console.log('\nScenario 8: epoch reset (hub storage wiped -> reconverge, no dupes)');
const epoch1 = (await hubStatus()).epoch;
const before = { a: corpusCounts(A), b: corpusCounts(B), c: corpusCounts(C) };
const nativeA = count(A, 'SELECT COUNT(*) AS n FROM observations WHERE origin_device_id IS NULL')
+ count(A, 'SELECT COUNT(*) AS n FROM session_summaries WHERE origin_device_id IS NULL')
+ count(A, 'SELECT COUNT(*) AS n FROM user_prompts WHERE origin_device_id IS NULL');
await stopHub();
rmSync(persistDir, { recursive: true, force: true });
console.log('[hub] state wiped; restarting...');
await startHub();
const epoch2 = (await hubStatus()).epoch;
check('fresh DO storage minted a new epoch', epoch2 !== epoch1 && epoch2.length > 0, { epoch1, epoch2 });
await A.client.pullOnce({ timeoutMs: 15_000 });
check('A adopted the new epoch', A.apply.getEpoch() === epoch2);
const requeuedA = pendingTotal(A);
check('A requeued its native corpus for re-push (the Phase 3 fix)', requeuedA === nativeA && requeuedA > 0, { requeuedA, nativeA });
await A.cloudSync.flush();
await B.client.pullOnce({ timeoutMs: 60_000 });
check('B adopted the new epoch', B.apply.getEpoch() === epoch2);
check('B requeued its native corpus', pendingTotal(B) > 0, B.cloudSync.status());
await B.cloudSync.flush();
await C.client.pullOnce({ timeoutMs: 60_000 });
check('C requeued nothing (no native rows)', pendingTotal(C) === 0, C.cloudSync.status());
// Settle: everyone pulls everyone's re-pushed corpus.
await A.client.pullOnce({ timeoutMs: 60_000 });
await B.client.pullOnce({ timeoutMs: 60_000 });
await C.client.pullOnce({ timeoutMs: 60_000 });
const after = { a: corpusCounts(A), b: corpusCounts(B), c: corpusCounts(C) };
check('A reconverged with no loss and no dupes', JSON.stringify(after.a) === JSON.stringify(before.a), { before: before.a, after: after.a });
check('B reconverged with no loss and no dupes', JSON.stringify(after.b) === JSON.stringify(before.b), { before: before.b, after: after.b });
check('C reconverged with no loss and no dupes', JSON.stringify(after.c) === JSON.stringify(before.c), { before: before.c, after: after.c });
check('all drains empty after reconvergence',
pendingTotal(A) === 0 && pendingTotal(B) === 0 && pendingTotal(C) === 0,
{ a: A.cloudSync.status().pending, b: B.cloudSync.status().pending, c: C.cloudSync.status().pending });
check('rebuilt hub carries the re-pushed log', (await hubStatus()).head_seq > 0);
} finally {
destroyDevice(A);
destroyDevice(B);
if (C) destroyDevice(C);
await stopHub();
if (process.env.MATRIX_KEEP_TMP !== '1') {
rmSync(persistDir, { recursive: true, force: true });
}
}
console.log(failures === 0 ? '\nMATRIX RESULT: ALL CHECKS PASSED' : `\nMATRIX RESULT: ${failures} CHECK(S) FAILED`);
process.exit(failures === 0 ? 0 : 1);
}
main().catch(async (err) => {
console.error('Matrix harness error:', err);
await stopHub();
process.exit(1);
});
+5 -3
View File
@@ -7,9 +7,11 @@
*
* WHY IT EXISTS: the canary user's DO has a KNOWN, CONSTANT workload, so
* its duration metric is a known constant — which is what makes the
* watchdog's hibernation-defeat detector (duration GB-s ≈ 0) sensitive. A
* hibernation regression shows up in the canary's DO within hours, not on
* the invoice. It runs 24/7 on any box (launchd/systemd/cron examples in
* watchdog's hibernation-defeat detector (duration GB-s ≈ 0) sensitive.
* Scope: the canary exercises the HTTP lanes only (it holds no WebSocket),
* so its DO catches HTTP-path regressions; WS-lane hibernation defeats are
* covered by the watchdog's fleet-wide duration alerting, not the canary.
* It runs 24/7 on any box (launchd/systemd/cron examples in
* ../DEPLOY.md) and prints one structured JSON line per event to stdout —
* redirect to a file for history.
*