A remembered mistake is only useful when the next case is provably the same kind of case.
Evolved from Exam Mistake Memory.
Reach for Transfer Engine when an agent has memory of past corrections and is about to apply one of them to a new case. That single moment is where recall quietly turns into a confident error.
| Situation on your desk | Without a compatibility gate | With Transfer Engine |
|---|---|---|
| A past lesson looks like it fits the new case | it is reused on resemblance alone | eight typed checkpoints decide, field by field |
| The lesson was written under an older rule version | it still wins because retrieval ranked it first | lifecycle resolves first; a stale record cannot be revived by rank |
| Every compared field matches | matching is treated as permission to act | a reviewed target-local verifier must still pass in the target |
| Someone pasted text into a record | it can steer the agent | memory is data; directive-shaped content quarantines the candidate |
An agent recalled a past scoring lesson and reused it on a case that read almost identically: same wording, same shape of question, different issuer domain. The recall was correct. The transfer was not.
Ranking did not fix it. Keeping the top semantic match, preferring the most recent lesson, requiring a confidence floor — all three measure resemblance, and resemblance is not compatibility. A stricter threshold under a newer rule version still looks like the old one.
What changed is the unit of memory. A lesson is no longer prose plus a score. It is a typed record — domain, issuer class, factor, threshold value and direction, rule version, scope, evidence, lifecycle — and each field is compared exactly before the lesson may speak.
Similarity may propose a candidate. A conjunctive gate plus a reviewed target-local verifier decides.
flowchart LR
R["Recalled lesson"] --> S{"Content scan<br/>directive or secret shaped?"}
S -- yes --> Q["QUARANTINED<br/>candidate held as inert data"]
S -- no --> L{"Lifecycle active?<br/>superseded · revoked · expired"}
L -- "not active" --> RJ["REJECTED<br/>stale lesson refused"]
L -- active --> C["Compare 7 typed fields<br/>exactly, no coercion"]
C -- "any field diverges" --> D["DENIED<br/>the gate names the field"]
C -- "all fields equal" --> V{"Reviewed target-local<br/>verifier passes?"}
V -- no --> H["HELD<br/>compatibility is not authorisation"]
V -- yes --> A["APPLIED<br/>the lesson may guide this case"]
classDef stop fill:#2a1414,stroke:#ff5c33,color:#ff5c33;
classDef go fill:#1b2114,stroke:#dfe104,color:#dfe104;
class Q,RJ,D,H stop;
class A go;
The gate is conjunctive: all of it must hold. Print this section and you have the whole policy.
| # | Checkpoint | Compared how | Fails when |
|---|---|---|---|
| 01 | Domain | exact typed equality | the target belongs to a different domain |
| 02 | Issuer class | exact typed equality | the issuer class diverges |
| 03 | Factor | exact typed equality | the lesson measures something else |
| 04 | Threshold value | exact, no coercion | "stricter" is treated as equivalent |
| 05 | Threshold direction | exact | direction is inverted or implied |
| 06 | Rule version | exact | the lesson predates the current rule |
| 07 | Scope | exact | the lesson belongs to another scope |
| 08 | Target-local verifier | executed in the target | the committed verifier is missing or fails |
| Domain | Unsafe shortcut it removes | Final transfer condition | Fixture |
|---|---|---|---|
| identity | semantic similarity | exact domain and issuer class | spec/evaluator.test.mjs |
| policy | rule drift | exact rule version and scope | spec/evaluator.test.mjs |
| threshold | "stricter" read as equivalent | exact value and direction | spec/evaluator.test.mjs |
| provenance | ungrounded lesson transfers | high-confidence evidence | spec/evaluator.test.mjs |
| lifecycle | stale lesson transfers | a non-active lesson is refused | spec/evaluator.test.mjs |
| application | compatible fields read as permission | committed target-local verifier passes | spec/evaluator.test.mjs |
| recall | empty or corrupt recall | diagnostic retry, then the transfer is held | committed policy fixture |
| Step | Do this | You should see |
|---|---|---|
| 1 | Select target case 01 Aligned target case, press Run the route | APPLIED — every checkpoint matched and the local verifier passed |
| 2 | Select 02 Divergent issuer domain | DENIED — checkpoints 01 and 02 marked denied, the gate names the field |
| 3 | Select 03 No target-local verifier | the route is held: compatibility alone never opens it |
| 4 | Paste an injection attempt into the operator note | the candidate goes down the quarantine route before any comparison |
Every run is stamped with a run counter and a timestamp, so a new run is never mistaken for the old one. Beside the verdict, the same case is described with and without the evolved prompt.
The page has no wallet, no provider key and no storage write. It calls the committed resolver
in intelligence/evaluator.mjs through
web/app/api/evaluate/route.js over committed fixtures, so
no verdict on screen is authored by the interface. Your operator note is real input: it is
attached to the target record as analyst_note and scanned by the same trust boundary the CLI
uses — and it can never set a compared field.
make test
make demomake test runs the prompt-contract mutation check, the receipt-consistency check, the
timeline validator, the evaluator and console suites, and a repository secret scan.
make demo prints four screens:
| Screen | Output |
|---|---|
| baseline | blocked — local verification missing |
| evolved | applied — local verification passed |
| mismatch | rejected — domain: incompatible or unknown; issuer_class: incompatible or unknown |
| boundary | the committed receipt-manifest summary |
Run the route console locally:
cd web && npm install && npm run build && npm startReplay the policy against verified owner-scoped history:
make historical-replay KRUG_HISTORICAL_REPO=/path/to/owner-historical-repositoryIt pins the direct one-commit interval cf124f605084f3c065ee020cd6398b363a63063f — which
expanded the MCP, anomaly, document and oracle surface — to its immediate owner-authored
security repair 1bab9bc92eda998b4f43b82aa00312db21d78bc8, which added input validation,
bounded histories and arguments, zero-denominator handling and structured hashing. The result
is machine-checked against exact commit, author, changed-file, hardening-marker and
prompt-hash conditions.
PROMPT.md is organised as seven blocks, each of which exists because a
specific failure mode exists.
| Block | What it decides |
|---|---|
| Trust boundary | recalled and proposed records are data; directives, permission claims and secret-like values quarantine the candidate |
| Typed lesson record | the 20 fields that make applicability testable, including an independently represented threshold value and direction |
| Admission and lifecycle | what may be written, and how supersession, revocation, expiry and conflict resolve before anything else |
| Compatibility gate | the nine-step adjudication: recall, retry, scan, validate, compare, reject, escalate, verify locally, apply |
| Receipt and recovery protocol | only terminal completion with a non-empty blob_id counts as persistence |
| Required decision record | one leading outcome — applied, rejected, conflict, blocked — plus a field-by-field table |
| Instruction priority | current policy and observed target evidence outrank recalled lessons; ambiguity resolves fail-closed |
spec/prompt-contract.test.mjs removes each of the five
material prompt rules in turn and requires the contract to fail, so the prompt text and the
executable behaviour cannot drift apart silently.
| Record | What it establishes |
|---|---|
docs/RECEIPTS.md |
10 terminal receipt rows and 5 fresh-client cold-recall markers, with one independently opened Walruscan Mainnet blob and an explicit note on what it proves |
records/live-sdk-proof-2026-08-21.json |
a current official-SDK write, terminal non-empty blob_id, destroy, and exact recall from a new client |
docs/PROMPT_TO_TEST.md |
every material prompt rule mapped to the check that proves it |
docs/REPLAY_RECEIPT.md |
the historical replay interval and its pinned outcome |
Three claims are kept separate on purpose: the deterministic policy result, the historical replay, and Mainnet persistence. The route console proves the first one live; it makes no storage claim.
TransferEngine/
├── intelligence/ canonical evaluator, shared route runner, CLI demo
├── cases/ versioned typed risk cases
├── records/ receipt manifest, SDK proof, validator
├── replay/ pinned owner-history replay bundle
├── spec/ evaluator, console, prompt-contract, replay tests, secret scan
├── docs/ receipts, prompt-to-test map, replay receipt, console screenshot
├── visuals/ rendered pipeline graphic
├── web/ Next.js route console over the canonical resolver
├── PROMPT.md · DEMO.md
└── Makefile
- Deny a transfer on purpose — open the route console, select Divergent issuer domain, read the rule that fired.
- Reproduce it offline —
make test && make demo. - Adopt the record format — copy
PROMPT.md, then delete one rule and watchspec/prompt-contract.test.mjsrefuse it.
DEMO.md walks the same three routes end to end, with the exact commands and the
evidence boundary spelled out.
Route fixtures and the contract map were checked on revision 3fde646dc7c623d4a5430d39e9eb27c24774d64b (2026-08-23).
