combine utf-16 surrogate pairs in cupsJSONImportString - #169
tanjiroK-coder wants to merge 1 commit into
Conversation
michaelrsweet
left a comment
There was a problem hiding this comment.
I don't really think a specific unit test for this is necessary.
Will look at the rest but I'm inclined to refactor the code a bit first.
|
Also, I don't even think that surrogates are technically valid here - they are a UTF-16 encoding side-effect using a block of reserved Unicode code points while |
8c82f81 to
e2d60b9
Compare
|
Dropped the On the surrogates: RFC 8259 §7 defines |
|
I hate JSON. Why they would even bother supporting surrogates when the industry has standardized on using UTF-8 (4 bytes UTF-8 vs. 12 bytes OK, so first I don't like the current change for a bunch of reasons; ultimately I want to simplify/refactor things, so thanks for what you've done so far but in this case I want to write my own fix and have you verify it, if you'll be so kind... WRT goals, I'll want to support valid surrogate pairs but error out on invalid hanging/second surrogates and things like escaped BOMs. Valid characters get converted back into UTF-8 with minimal escaping on output/export. |
|
Sure, happy to verify. Push it wherever suits and I'll run a valid pair (U+1D11E should come out as No argument on erroring out rather than substituting U+FFFD, strict is easier to reason about. Close this one out whenever you like. |
cupsJSONImportStringdecodes each\uXXXXescape on its own and never combines a UTF-16 surrogate pair, so"\uD834\uDD1E"(U+1D11E) decodes to two 3-byte CESU-8 sequences (ed a0 b4 ed b4 9e) instead of the 4-byte UTF-8f0 9d 84 9e, and a lone surrogate leaves a rawed a0 b4in the value. The JSON here is untrusted:cupsJSONImportURL/cupsOAuthGetTokensfeed it from an OAuth/OIDC endpoint andcupsJWTImportStringruns it over token contents, so a hostile server can drop invalid UTF-8 into strings that later get compared, re-encoded or logged. I spotted it reading the\ubranch after noticingdnssd.calready folds surrogate pairs correctly. The decoder now pairs a high surrogate (D800-DBFF) with the following low surrogate (DC00-DFFF) into one code point emitted as 4-byte UTF-8, and substitutes U+FFFD for an unpaired half rather than emitting a bare surrogate. The pre-scan already reserves five bytes per\uXXXX, so a combined pair uses four of the ten reserved bytes and the allocation is unchanged.testjsongets two cases that fail on the current code.Assisted-by: Claude Code:claude-opus-4-8