Conversation
dirvine
left a comment
There was a problem hiding this comment.
APPROVE — reviewed 8f5af8d.
No blocking findings. Traced final-state preservation through PUT, remembered-state/file-loss handling, replication and final-state witness verification. The latest deltas correctly return the remembered identifier when rejecting a rival and require the fetched witness to back the entire claimed state, not just its identifier. Their regressions are included in the passing pointer tests.
Verification:
- Exact head:
cargo test --locked --all-features --lib pointer— 95 passed, zero failures. - Prior head e813b9c:
cargo test --locked --all-features --test pointer_convergence— 17 passed. The subsequent reviewed delta only tightens witness backing and adds its regression; do not treat the earlier convergence run as an exact-head run. - Companion client transfer E2E at its e813b9c node pin: two tests passed using local nodes and Anvil (handover and refusal of a second transfer after one node retained a final state).
Limits: ADR-0018 deliberately promises per-node final-state retention, not globally irrevocable ownership consensus. Churn, lost copies and mixed-version deployment remain important. Some current-head CI checks are still running/queued; this is not an all-green-CI or release sign-off.
Review scope: the three companion pointer-transfer PRs were read together. Independent GLM-5.2 review found no blockers on the earlier reviewed heads; I checked subsequent deltas directly and reran the relevant tests. Its cautions about mixed-version deployment and the lack of global ownership consensus are valid, documented limitations rather than demonstrated regressions. Codex CLI could not review because its credentials were revoked; it is not counted as a completed review. Additional source-review seats have not returned, so this is not a claim of full-panel consensus.
…llow A storage audit binds, in round 1, a nonced root over each pointer record a node holds, and round 2 must serve the bytes that reproduce it. The node kept only the record the last update replaced, so two paid updates to a pointer between the rounds made an honest holder fail round 2 with DigestMismatch, a confirmed failure that feeds the trust penalty. An owner who is also one of the holder's auditors knows exactly when its round 1 has been answered, so this was a cheap way to penalise a chosen honest neighbour. The round-1 session now keeps the root reported for each pointer leaf, the pointer store keeps every record an update replaces rather than only the last one, and round 2 serves the one record, held or replaced, whose root matches. Replaced records are kept for ten minutes, longer than a round 1 over the largest subtree an auditor waits for plus the session its round 2 must arrive within. Roots are capped at 65,536 across all live sessions, the oldest sessions giving theirs up first. When a node cannot serve the record round 1 read, because it aged out, was evicted, or the session kept no root and the pointer has been updated since, round 2 is rejected as Transient instead of guessing: no trust penalty, only the credit of that audit. A node that holds nothing at all for the pointer is still reported absent, as before. The auditor, the wire format and the subtree-audit protocol id are unchanged. ADR-0017 records the change and amends one point of ADR-0016, which is left as written.
Review of the first version found three ways a node could still lose the record an audit was owed. - A flood of pointer-heavy round-1 sessions made older sessions give up their roots to make room, so an honest holder's round 2 went transient. A round 1 whose roots do not fit now withholds its proof, exactly as a round 1 refused for capacity already does, and no live session gives its roots up. - Keeping a replaced record could evict another one before an update that then failed. Eviction now waits for the rename to succeed, and a failed rename takes back the record it kept. - A round 2 could read the new record before the one it replaced was kept. The replaced record is now kept before the rename, under the same lock. ADR-0017 now states what a transient round 2 costs (the auditor forgets the holder's standing for the whole pinned commitment, with no trust penalty) and that the ten-minute retention is sized for the default configuration.
ADR-0017 is now also claimed by an open PR that renumbered its own ADR, and ADR-0018 by another. 0019 is the next number neither main nor any open ADR PR uses. The ADR's text is unchanged; every reference to it in comments and tests follows, and the ADR index lists it.
… old APIs A round 1 whose pointer bindings would not fit the budget withheld its proof silently. Silence reads as a peer that did not answer, which costs the honest holder trust at the transport, so it now answers Transient, the auditor's timeout lane, with no trust penalty. ant-node 0.21.0-rc.1 already carries the pointer store and the slice handler, so this change keeps their signatures: PointerStore::superseded again returns the record last replaced, beside a new superseded_all, and handle_subtree_slice_challenge_with_pointers keeps its arguments, beside a new handle_subtree_slice_challenge_with_pointer_bindings that the engine uses. The additions are the only API change.
A round 1 refused over the pointer bindings budget, and a pointer record round 1 bound that is no longer kept, are answered Transient. The protocol docs still said Transient only follows failed read retries and that any rejection of a recent pinned commitment is a confirmed failure.
…ters as before handle_subtree_slice_challenge_with_pointers was kept for callers of its earlier signature, but it passed empty round-1 bindings, so every pointer it served came back as a transient failure, even one that never changed. It now serves a committed pointer as it did before ADR-0019: the record held and the newest one an update replaced. The engine keeps using the entry point that takes round 1's roots. A test calls the earlier entry point with an unchanged pointer and requires the auditor to pass it.
…, and those expire unaided A node that no longer held a pointer but still kept, in memory, a record an update had replaced was not reported absent: it answered Transient, or passed outright when the kept record was the one round 1 read. Replaced records prove what round 1 read, not that the pointer is still held, so a node without the pointer now reports it absent, a confirmed failure, as before ADR-0019. Replaced records past their ten minutes were dropped only when another update came, so after a burst of updates up to 11 MB could stay resident indefinitely. Each prune pass now drops them. The earlier entry point's test now also covers one update between the rounds: it serves the record held and the one it replaced, newest first, and passes.
The earlier slice-challenge entry point served the record an update replaced when the pointer itself was gone, so a node that had lost a pointer after one update could still pass an audit through it. It now reports a pointer it no longer holds as absent, as the engine's entry point does, and still serves the record held and the one it replaced otherwise. A test covers it through that entry point.
- Repair counts its quorum over the whole configured close group, so a node that sees few members of it cannot adopt an owner-signed state nobody paid for on one peer's word, and it does not repair while it is still bootstrapping. - Fetch and state requests are admitted fairly before a task exists, at most 128 outstanding and 16 per peer, as chunk fetches are. A state request larger than an honest one is dropped before admission. - This node keeps at most 8 pointer requests outstanding at any one peer, across repair, possession checks and pruning, so it never exceeds the allowance a peer gives it and is never dropped as a flood. - A peer is remembered as speaking pointers only while it is in the routing table, and that set is capped. - A fetched record's signature is verified off the async executor, as ADR-0016 says every pointer signature check is. - A commit for a record this node lost takes the lost state or a newer one, never an older one, so a write verified before the loss cannot roll the node back.
A possession check fetched the record from a peer, and a record that did not verify was penalised by the fetch and then again by the check as a record not held. The fetch now says whether it penalised an invalid record, and the check does not add a second penalty for it.
…swers A possession check looked up the close group once and then asked each peer in turn. Asking waits on that peer's other requests and on the peers before it, so a peer could leave the group before it was asked and still be penalised for not holding the record. Membership and capability are now checked again before a penalty. A record this node could not check, because the verification task failed here, was treated as no record and charged to the peer. It is now its own outcome and charges nobody. The judgement of one answer is a pure function with a unit test, which fails if an invalid record, already charged by the fetch, is charged again as missing.
…nterval Every neighbour-sync request started a hint push before admission, so a peer sending small sync requests, even refused ones, made this node scan every pointer it holds, with a routing lookup per record, as often as it liked. Hints now answer only an admitted request, at most once per peer per shortest sync interval, which an honest peer never syncs faster than, and at most eight answers run at once. A request that finds them busy gets no hints this time; the next sync round sends them anyway. The traffic summary counted the six pointer messages but never logged them. It now does, in a fourth summary line.
…nd never refuse a new one Hint answers ran before the neighbour-sync worker's freshness check, so a request later shed as stale had already cost a scan of every pointer, and any sender could be answered. A full map of answered peers also refused every new peer, an honest bootstrapping one included, for a whole sync interval. Answers now run only for a request that is admitted and still fresh, only for a routing-table peer, and a full map forgets the peer answered longest ago instead of refusing anyone.
…obody for it, and stop serving at shutdown This node's own requests to one peer waited for a permit with no limit and past shutdown, outside the request's timeout, so repeated paid updates to one close peer could pile up possession checks waiting forever and hold shutdown open. A request now waits for its permit no longer than its own timeout and not past shutdown, and one never sent is its own outcome: a possession check judges nobody on it. Serve tasks waiting for a serve permit ignored shutdown, so requests admitted just before it kept reading disk and answering after cancellation. The wait now ends at shutdown, and the work behind it.
The waits added for shutdown picked either branch at random when a permit was free and shutdown had already begun, so a request could still be sent, or a stored record read and served, after cancellation. Shutdown is now checked first and again once a permit is held. A test asks repeatedly with a free permit after shutdown and requires nothing be sent each time.
A pointer state at u64::MAX is now final (ant-protocol): nothing replaces it, not even another final state whose target sorts first. That is what lets an owner hand a pointer's address over for good, by signing one last state that points at the new owner's pointer. The store, the admission gate, fresh offers, repair and hints all compare with the protocol's replaces(), so they keep whichever final state they took first with no code of their own. The merge rule alone leaves one gap: a node holding no final state takes any, so a former owner could still finalize again on a node that joined the group after the transfer, or lost its copy. So before a node takes a final state it does not hold, from a client or a fresh offer, it asks its close group which state each holds. A peer claiming a different final state is asked for the record, and if it verifies -- only the owner could have signed it -- the write is refused as stale, naming the state the group proved. A claim alone refuses nothing, so one peer cannot block a transfer. The look runs after payment is verified, only asks peers that have sent a pointer message, and is bounded at four seconds inside the client's store timeout; silence proves nothing. Two different final states can still exist if the owner races them to different nodes. Each node keeps its first; the client decides by the close group's majority. So the possession check no longer penalises a member holding the other side of such a fork -- it holds what the merge rule told it to -- and repair, between two final states a wide group backs at quorum, adopts the larger side rather than whichever answered first. with_pointers now takes the service and wires both directions, so the node and the e2e harness cannot attach one and forget the other. ADR-0018 records the decision and amends ADR-0016's merge rule.
…nges ADR-0016 was edited in place: a backlink, the terminal-counter conclusion, the merge rule and the ownership consequence. It is restored exactly, and ADR-0018 now lists the three statements of ADR-0016 it overrides at the final counter. ADR-0018 also claimed more than the rule gives. The former owner keeps the earlier record and the key, so it can sign a second final state at any time, not only in a race, and any node the first has not reached will take it. What holds is that no node gives up a final state it holds. The claim that flooding final states pays for every round trip was false while a verified payment is cached; the ADR now describes the proven-conflict memory and the bounded, per-address looks that make it true, and the restore of a node's own lost final state without a look.
… to replay, and safe to restore The finality look took claims as they arrived but fetched each claimed rival in turn. A peer that claimed a rival and then stalled its fetch held the look until its four-second budget ran out, and a look that runs out finds nothing, so one peer could let a second final state in. Each peer's claim and fetch now run as one pipeline, all at once, and the first proof wins. A verified payment is cached, so replaying one paid final state that lost cost the node a signature check and a round of questions every time. A proven conflict is now remembered, up to 16,384 addresses, and a replayed loser is refused before its signature is checked. Looks for one address wait their turn, at most 64 run at once, and one that cannot start within two seconds answers Busy: a PUT gets a retryable error and a fresh offer is dropped, neither taking nor refusing the state for good. A node that lost the file of a final state it held was made to look before restoring it, so a peer on the other side of a fork could keep it from restoring its own copy. The store now reports what it remembers, and restoring a remembered final state asks nobody. ReplicationEngine::with_pointers keeps the signature it has on main; with_pointer_service wires the service both ways. ant-protocol is repinned to the head of its companion change, which only corrects documentation.
The companion change added documentation only: where a final state is permanent, and that conflicts below the final counter are ordered.
…ne look among queued replays A proven final state was stored one per address, so a later look that proved the other side of a fork replaced the first proof, and the state the first look had disproved could then be taken if the group went quiet. Up to two proven states are now kept per address, and a second never replaces the first; two different ones refuse every final state there. Looks for one address took turns but did not share answers, so replays queued behind a clear look each asked the group again once it finished. A clear look now answers the same state for ten seconds without asking, on the runtime's clock, so a burst of replays costs one round as ADR-0018 says. The bounds live in their own type so they are tested without a network: eight queued checks make one look, and a proof answers every replay of the loser.
…s over the address's last byte A look that found no conflict was reused for the same final state for ten seconds, even when the write it cleared then failed: a rival that landed meanwhile could be let in by a retry that skipped the look. The answer is now forgotten as soon as that write does not land, on the PUT path and the fresh-offer path, and it is reused for two seconds, as long as a replay queued behind the look can have waited for its turn. Looks took turns by the address's first byte, but the addresses one node holds share their leading bits, so nearly all its looks shared one turn and a single slow look made the rest answer Busy. They take turns by the last byte now. ant-protocol is pinned at the head of its companion change, which corrected two more documentation claims.
…ween final states adopts neither A fresh offer whose commit came back stale, because a rival had landed first, was reported as stored, so the clear look that preceded it outlived a write that never happened. A commit that loses now counts as not stored, and the look is forgotten. Repair picked the larger side between two quorum-backed final states, but on a tie, possible in a group of eight with a quorum of four, the state that answered first won, and two repairing nodes could adopt opposite sides. A tie now adopts neither until the group settles. ADR-0018 no longer claims the look keeps every final PUT inside the client's ten-second timeout: it adds at most six seconds after payment is verified, and a slow payment check can still outlast the timeout.
…lent peers, and take only the proof claimed Repair adopted the larger of two quorum-backed final states without counting peers that had not answered, so in a group of eight one node seeing four against three with one silent peer adopted the four, while another seeing the reverse adopted the other side, for good. A final state is now adopted only if it stays strictly larger with every silent peer counted for its rival, seen or not yet seen. A peer that claimed one rival final state and served another had the served record accepted as proof. The record must now be the state the peer claimed.
…r another address as no answer Repair grouped state summaries by identifier alone and kept the first one's counter and target, so one dishonest peer naming a final state's identifier with a lower counter absorbed the honest votes for it as a non-final state and slipped past the guard for final states; the fetch then took the real final record, matching the identifier only. Votes now count together only for the same whole state, and the fetched record must equal the state the quorum backed. A summary about some other address was documented as no answer but counted as a peer that answered holding nothing, which could hide a silent peer that might tie two final states. It is now left out entirely; an explicit nothing-held answer still counts.
…and test the whole-state backing A rival final state refused by a node that had lost the file of its own final state was answered Stale naming zeros, since only a state on disk was named. It now names the state the node remembers, so the sender can see what won. The check that a fetched record is exactly the state a quorum backed is its own predicate with a test that fails if it goes back to comparing identifiers only. ADR-0018 now says that a node running with no replication, a devnet or one whose replication engine failed to start, takes a final state on the merge rule alone, as a node whose group cannot answer does.
…imed The look compared only the identifier of the record a peer served with the state it had claimed, so a claim pairing a real rival's identifier with another target was taken as backed by that rival. The record must now be the whole state claimed, as repair already requires, and a test fails if the comparison goes back to identifiers.
Rust 1.99's clippy adds the pedantic assert_is_empty lint, which main now answers for its own tests: a bare assert! on is_empty() prints nothing useful when it fails. This test, new on this branch, had the last such assertion, so CI's clippy job failed once the branch sat on that main. Bind the value and print it, as main does; the condition is unchanged.
Both change the pointer possession check. #240 judges a peer's answer in judge_possession, after which the peer is charged only if it still owes the record. The final state here adds one outcome to that judgement: a peer serving a different final state from the one offered holds the other side of an owner's fork, as the merge rule told it to, so it is named loudly and not charged.
8f5af8d to
99eee09
Compare
Linear issue
Closes V2-1354
Risk tier
This changes which pointer state a node keeps at the final counter, and adds a close-group round trip before a node takes a final state.
What
Pointer ownership transfer by final redirection (ADR-0018). The owner key never changes. The owner signs one last state at
counter == u64::MAX, pointing at a pointer the new owner holds the key to. Readers are redirected there by every node that holds that state, the address stays the same, and no node that holds it ever gives it up. WithAutonomi/ant-protocol#40 makes a final state final: nothing replaces it, not even another final state whose target sorts first. This PR pins that rev and does the node's half.replaces. They keep whichever final state they took first, with no code of their own.Stale, naming the state the group proved.ReplicationEngine::with_pointerskeeps its signature;with_pointer_service(&PointerService)wires both directions, fresh writes and the finality witness, and is what the node and the e2e harness use.Compatibility
PointerStateRequestandPointerFetchRequest.ReplicationEngine::with_pointer_service,pointer::FinalStateWitness,pointer::FinalityCheck,PointerService::attach_final_state_witnessandPointerStore::remembered. Behaviour: a node refuses a second final state withStalewhere it previously took one whose target sorted first.Semver impact
Test evidence
mainat56b5770(0.21.0). Main had dropped its ant-protocol patch for the published 3.1.0, so the pin commits re-add[patch.crates-io], each at the rebased feat(pointer): make a final state final, so ownership can be handed over ant-protocol#40 commit it pinned before; the lock differs from main only in that entry. The fourteen reviewed commits are otherwise unchanged. Two merges then bring in fix(pointer): serve the record round 1 bound, however many updates follow #238 and fix(pointer): harden pointer replication against thin views and floods #240. The only code overlap is the possession check: fix(pointer): harden pointer replication against thin views and floods #240 moved the judgement intojudge_possession, and the fork exemption here is now one outcome of it,Possession::Forked, still logged at warn and never penalised, with assertions added to fix(pointer): harden pointer replication against thin views and floods #240's unit test. A last commit pins test: add comprehensive merkle payment verification tests #40's head.99eee09, locally on macOS with Rust 1.99.0:cargo test --lib --features test-utils1,270 passed;cargo test --lib --no-default-features1,225 passed;pointer_convergence17; thee2epointer replication, subtree audit, replication and fresh-offer tests, 59, passed over QUIC;webrtc_direct_devnet3, and 6 with 1 ignored undertest-utils;poc_commitment_audit_attacks19,poc_audit_handler_live16,poc_bootstrap_stall3,poc_shutdown_lmdb_drain1; a default-featureant-nodecarries no failpoint;cargo clippy --all-targets --all-features -- -D warnings,cargo fmt --check,cargo docwith-D warningsandscripts/adr-governance.pyagainstmainall clean. The full CI matrix runs on this head.ant-devnetwith Anvil) and six releasedant-node0.21.0 processes joined to it, so close groups held both (the transferred pointer sat on three 0.21.0 nodes and four of these). Withantfrom feat(pointer): hand a pointer over for good, and read forks by majority ant-client#210: create, update and read; a paid transfer reportedfinalwith 7 of 7 holding it; controller and resolve follow it; a later update is refused before paying. With the releasedant0.3.9: create, update and read its own pointer, and read and resolve the transferred one through its final record. This is honest use only: a former owner's second final state still wins on 0.21.0 nodes until they upgrade (Compatibility, Mixed fleet).8f5af8d:cargo test --lib --features test-utils: 1,250 passed. New tests, each of which fails when its fix is reverted (checked by reverting it):cargo test --features test-utils --test pointer_convergence: 17 passed. With one final state among the records every delivery order converges on it, exhaustively, and with two the first delivered is kept; a stored transfer is not displaced by a final state whose target sorts first, nor by any lower counter.cargo test --features test-utils --test e2e pointer_replicationover real QUIC: 15 passed. A transfer written to one node reaches the group and a paid final state whose target sorts first is refused by every node; a node that missed the transfer refuses a different one because a peer serves the one it holds, and still takes the group's own; the possession check does not penalise the other side of a fork and still penalises a member holding nothing.cargo test --lib --features test-utils1,250 passed;e2e113 passed, 3 ignored as onmain;migration_reclaims_disk2,migration_crash_safety5,migration_shared_volume5,storage_scale2,webrtc_direct_devnet3,poc_commitment_audit_attacks19,poc_audit_handler_live16,poc_bootstrap_stall3,poc_shutdown_lmdb_drain1, andpointer_convergence17, all passing.cargo clippy --all-targets --all-features -- -D warnings,cargo fmt --all -- --check,cargo docwith--deny=warningsandscripts/adr-governance.py: all clean.New dependency
None.
ADR
https://github.com/WithAutonomi/ant-node/blob/feat/pointer-ownership-transfer/docs/adr/ADR-0018-pointer-transfer-by-final-redirection.md:
docs/adr/ADR-0018-pointer-transfer-by-final-redirection.md, added by this PR (Proposed). It amends ADR-0016's merge rule at the final counter.Mitigation / rollback
Re-pin ant-protocol to the previous rev and revert this branch. No stored record changes shape. The only coordinated part is that nodes and clients should agree on the final-counter rule, and a mixed group degrades to reads by majority, not to lost data.