Description
OlmMachine has no startup verification that the server's stored device keys match the local account, no scoped recovery for a single-device identity-key desync, and its OTK upload path can wedge permanently when the server already holds one-time keys under the same key IDs. We hit all three at once during an incident on 2026-09-23 (Synapse homeserver, mautrix 0.21.1), and each one turned a recoverable state into a full re-registration.
Defect 1 — share_keys() silently clobbers server-side device keys when the local store is fresh but the device_id exists
Reproduction chain (verified from gateway logs + Synapse responses):
- Device
X exists on the server with uploaded device keys (normal operation).
- A second client instance starts with the same device_id but a fresh/empty crypto store.
OlmMachine.load() creates a new OlmAccount (new identity keys), account.shared = False.
- First
share_keys() call (via handle_otk_count or explicit startup call) builds device_keys because not self.account.shared (machine.py:313-317) and uploads them. Synapse's POST /keys/upload replaces the stored device keys for device X with the new ones.
- The original client (which still holds the real Olm account for X) now sees
server has different identity keys for device X and permanently refuses E2EE.
There is no server-side guard against this: Synapse only checks the signature, and the new keys are correctly self-signed. From the server's perspective the impostor's upload is indistinguishable from a legitimate re-upload. But OlmMachine could detect it: fetch the stored device keys for our own device_id first (query_keys with {mxid: []}), and if the server already has an ed25519 key that differs from account.identity_keys["ed25519"], refuse the silent clobber and surface a distinct error instead of uploading.
How the desync arises (our case): a credential leak started a duplicate client under the same user/device_id with an empty store. The same state also occurs when a crypto store is deleted/corrupted and the client restarts while the old device record survives server-side, or when a profile is copied to a second machine with the same device_id (store regenerated, device reused). None of these are exotic: "delete crypto.db and restart" is standard advice.
Defect 2 — one M_UNKNOWN: One time key … already exists in a batch wedges key sharing permanently
After the clobber, our fresh store's OTK ID counter started at AAAAAQ while the server still held the impostor's 55 one-time keys whose IDs overlapped that counter range. Every share_keys() then failed atomically:
mautrix.errors.request.MUnknown: One time key signed_curve25519:AAAAAQ already exists.
Old key: {"key":"wltvheOsyhxg…","signatures":{"@user:hs":{"ed25519:DEVICE_X":"…"}}};
new key: {'key': '45QmiSXhAQTs…', 'signatures': {'@user:hs': {'ed25519:DEVICE_X': '…'}}}
Code-level root cause in _share_keys (machine.py:302-332):
get_one_time_keys() regenerates the same key IDs from the account counter (account.py:96-98), so every retry re-emits the same colliding IDs.
resp = await self.client.upload_keys(...) raises → the function exits before self.account.shared = True, mark_keys_as_published() and put_account() (machine.py:327-331). Because mark_keys_as_published() never runs, libolm keeps returning the same unpublished keys forever.
- There is no handling for the
already exists case: no skip of colliding key IDs, no counter advance, no fallback to uploading keys one by one. The counter never advances past the server's existing IDs, so the batch can never succeed. In our case the loop ran 86 failures over 48 minutes until we deleted the device server-side and re-registered.
This is the Python-side twin of matrix-org/matrix-rust-sdk#6520 and the Synapse-side atomicity of matrix-org/synapse#7365. A transport-layer workaround is being shipped elsewhere (openclaw/openclaw#74529 rewrites the 400 to a synthetic 200 with empty counts, which makes the SDK mint fresh IDs on the next tick).
Suggested fixes (any one unwedges):
- Catch
M_UNKNOWN containing already exists in _share_keys, mark_keys_as_published() to burn the colliding batch, put_account(), and retry once with freshly generated IDs.
- Or upload keys individually (or in chunks) so one collision fails the whole batch; Synapse rejects duplicates per-key.
Defect 3 — no client API to delete a device / purge its server-side keys, and no recovery API for a single-device desync
The client API surface (mautrix/client/api/) has create_device_msc4190, logout, logout_all — but no delete_devices. We had to drop to raw DELETE /_matrix/client/v3/devices/{id} with ad-hoc UIA (m.login.password auth_data), which the client API does not support either. matrix-nio carries delete_devices with UIA support as precedent (AsyncClient.delete_devices). When a device's server-side record is poisoned (stale OTKs under wrong keys, wedged key counts), "delete the device record and start clean" is the recovery path, and the SDK should own it.
Additionally, the crypto store already has a per-account delete() (crypto/store/asyncpg/store.py:88-92) that clears crypto_account + Olm sessions + outbound Megolm sessions for one account_id — but nothing exposes a guided "reset this device's E2EE state" flow (verify server keys → if mismatch, either restore the local account if we hold the original, or delete server device + local account + re-upload) with the UIA dance handled. Today every embedder hand-rolls this, and the failure modes are non-obvious (deleting the server device revokes all its access tokens; claiming a stale OTK whose private half is gone makes inbound sessions undecryptable).
Environment
- mautrix 0.21.1 (gaps confirmed present on current master by source inspection of
machine.py / account.py / client API, 2026-09-26)
- Synapse homeserver, Matrix spec v1.12
- python-olm 3.x bindings
Suggestion
A share_keys() pre-flight check against query_keys for our own device + a delete_devices client API with UIA would have turned a multi-hour outage with full re-registration into a clean, automatic recovery. Happy to work the PR(s) if the approach sounds reasonable — the pre-flight check and the already-exists unwedge are both small, and delete_devices can follow matrix-nio's implementation closely.
Description
OlmMachinehas no startup verification that the server's stored device keys match the local account, no scoped recovery for a single-device identity-key desync, and its OTK upload path can wedge permanently when the server already holds one-time keys under the same key IDs. We hit all three at once during an incident on 2026-09-23 (Synapse homeserver, mautrix 0.21.1), and each one turned a recoverable state into a full re-registration.Defect 1 —
share_keys()silently clobbers server-side device keys when the local store is fresh but the device_id existsReproduction chain (verified from gateway logs + Synapse responses):
Xexists on the server with uploaded device keys (normal operation).OlmMachine.load()creates a newOlmAccount(new identity keys),account.shared = False.share_keys()call (viahandle_otk_countor explicit startup call) buildsdevice_keysbecausenot self.account.shared(machine.py:313-317) and uploads them. Synapse'sPOST /keys/uploadreplaces the stored device keys for device X with the new ones.server has different identity keys for device Xand permanently refuses E2EE.There is no server-side guard against this: Synapse only checks the signature, and the new keys are correctly self-signed. From the server's perspective the impostor's upload is indistinguishable from a legitimate re-upload. But
OlmMachinecould detect it: fetch the stored device keys for our own device_id first (query_keyswith{mxid: []}), and if the server already has an ed25519 key that differs fromaccount.identity_keys["ed25519"], refuse the silent clobber and surface a distinct error instead of uploading.How the desync arises (our case): a credential leak started a duplicate client under the same user/device_id with an empty store. The same state also occurs when a crypto store is deleted/corrupted and the client restarts while the old device record survives server-side, or when a profile is copied to a second machine with the same device_id (store regenerated, device reused). None of these are exotic: "delete crypto.db and restart" is standard advice.
Defect 2 — one
M_UNKNOWN: One time key … already existsin a batch wedges key sharing permanentlyAfter the clobber, our fresh store's OTK ID counter started at
AAAAAQwhile the server still held the impostor's 55 one-time keys whose IDs overlapped that counter range. Everyshare_keys()then failed atomically:Code-level root cause in
_share_keys(machine.py:302-332):get_one_time_keys()regenerates the same key IDs from the account counter (account.py:96-98), so every retry re-emits the same colliding IDs.resp = await self.client.upload_keys(...)raises → the function exits beforeself.account.shared = True,mark_keys_as_published()andput_account()(machine.py:327-331). Becausemark_keys_as_published()never runs, libolm keeps returning the same unpublished keys forever.already existscase: no skip of colliding key IDs, no counter advance, no fallback to uploading keys one by one. The counter never advances past the server's existing IDs, so the batch can never succeed. In our case the loop ran 86 failures over 48 minutes until we deleted the device server-side and re-registered.This is the Python-side twin of matrix-org/matrix-rust-sdk#6520 and the Synapse-side atomicity of matrix-org/synapse#7365. A transport-layer workaround is being shipped elsewhere (openclaw/openclaw#74529 rewrites the 400 to a synthetic 200 with empty counts, which makes the SDK mint fresh IDs on the next tick).
Suggested fixes (any one unwedges):
M_UNKNOWNcontainingalready existsin_share_keys,mark_keys_as_published()to burn the colliding batch,put_account(), and retry once with freshly generated IDs.Defect 3 — no client API to delete a device / purge its server-side keys, and no recovery API for a single-device desync
The client API surface (
mautrix/client/api/) hascreate_device_msc4190,logout,logout_all— but nodelete_devices. We had to drop to rawDELETE /_matrix/client/v3/devices/{id}with ad-hoc UIA (m.login.passwordauth_data), which the client API does not support either. matrix-nio carriesdelete_deviceswith UIA support as precedent (AsyncClient.delete_devices). When a device's server-side record is poisoned (stale OTKs under wrong keys, wedged key counts), "delete the device record and start clean" is the recovery path, and the SDK should own it.Additionally, the crypto store already has a per-account
delete()(crypto/store/asyncpg/store.py:88-92) that clearscrypto_account+ Olm sessions + outbound Megolm sessions for one account_id — but nothing exposes a guided "reset this device's E2EE state" flow (verify server keys → if mismatch, either restore the local account if we hold the original, or delete server device + local account + re-upload) with the UIA dance handled. Today every embedder hand-rolls this, and the failure modes are non-obvious (deleting the server device revokes all its access tokens; claiming a stale OTK whose private half is gone makes inbound sessions undecryptable).Environment
machine.py/account.py/ client API, 2026-09-26)Suggestion
A
share_keys()pre-flight check againstquery_keysfor our own device + adelete_devicesclient API with UIA would have turned a multi-hour outage with full re-registration into a clean, automatic recovery. Happy to work the PR(s) if the approach sounds reasonable — the pre-flight check and the already-exists unwedge are both small, anddelete_devicescan follow matrix-nio's implementation closely.