Conversation
… hosts Running the suite on s390x (QEMU, Debian sid) gave 110 failures on main. Most came from one asymmetry: V3 data types without an explicit byte order defaulted to little-endian, while NumPy dtype names mean host order. On a big-endian host `dtype="float64"` and `dtype=np.float64` produced different data types, and a reopened array changed dtype. - V3 data types default to the host byte order (`HasEndianness`), because V3 metadata carries none. Stored chunks stay little-endian by default. - `ShardingCodec` defaults its inner and index `bytes` codecs to little-endian, so sharded output no longer depends on the host. - `scale_offset` computes in native byte order and restores the input's byte order; `>f8` and `>u8` arrays failed on every host. - Base64 fill values of structured data types in V3 metadata are little-endian bytes in both directions. - Tests no longer assume a little-endian host: explicit `<` dtypes and `endian="little"` where the expectation is little, and a big-endian case for migration and pipeline parity. Refs zarr-developers#3438 Assisted-by: ClaudeCode:claude-opus-5-5
Assisted-by: ClaudeCode:claude-opus-5-5
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #4435 +/- ##
==========================================
+ Coverage 94.38% 94.43% +0.04%
==========================================
Files 93 93
Lines 13205 13219 +14
==========================================
+ Hits 12464 12483 +19
+ Misses 741 736 -5
🚀 New features to boost your workflow:
|
A weekly and path-filtered job that runs the data type, codec and metadata tests on an emulated big-endian host, in a Debian sid image that takes numpy and numcodecs from Debian packages so nothing heavy compiles under emulation. Assisted-by: ClaudeCode:claude-opus-5-5
Documentation build overview
7 files changed ·
|
d-v-b
marked this pull request as ready for review
September 27, 2026 11:40
Contributor
Author
|
cc @avalentino |
`BytesCodec()` defaulted to `sys.byteorder`, so the same code wrote big-endian chunks on s390x and little-endian chunks elsewhere. The bytes codec originally defaulted to "little"; the host default came in as a side effect of replacing attrs with dataclasses (zarr-developers#1660), and was later pinned by a test that only restated it. `"little"` also matches `default_serializer_v3`. Chunks written on big-endian hosts with a bare `BytesCodec()` are now little-endian. Chunks written before this change remain readable, because the codec records its byte order in the metadata. Refs zarr-developers#3438 Assisted-by: ClaudeCode:claude-opus-5-5
This reverts commit 8dd33e3. Changing the default of `BytesCodec()` changes what existing code writes on big-endian hosts, so it needs a deprecation cycle and will land in its own PR. Assisted-by: ClaudeCode:claude-opus-5-5
Contributor
Author
|
Internal bugfixes / platform compatibility, self merging |
Contributor
Author
|
I requested review when I thought we were changing the bytes codec default here, with that change factored out I don't think review is needed but it is of course welcome |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 AI text below 🤖
Refs #3438. The agent ran the test suite on a big-endian host: s390x emulated with QEMU, Debian sid, and Debian's numpy 2.4.6 and numcodecs 0.17. On
mainit got 110 failures across 30 test functions. With this branch, 11493 tests pass and 3 fail. Those 3 aretest_examples, which callsuv, and the image does not have it.Real bugs
dtype="float64"gave<f8anddtype=np.float64gave>f8. Every reopened array came back as<f8, so its dtype changed between create and open, and its arrays were not in native order. This one bug caused most of the failures.HasEndiannessnow defaults tosys.byteorder, because V3 metadata carries no byte order.BytesCodec(endian="little").scale_offsetfailed on arrays whose byte order is not the host's, on every host. NumPy arithmetic returns native results, which tripped the dtype-preservation check. Also,>u8 != np.uint64sent>u8data down the int64-widening path. The codec now computes in native order and restores the input's byte order.ShardingCodecdefaulted to a bareBytesCodec()for its inner and index codecs. That codec uses the host byte order, so a sharded array written with default settings had different bytes on s390x than on x86. Both defaults are nowendian="little".Tests that assumed a little-endian host
The failing tests and a static audit turned up these patterns:
'little'in reprs.tobytes()compared against little-endian stored bytesBytesCodec()/Int32()used as expected valuesThe tests now use explicit
</>dtypes and explicitendian=. Two tests gained a case in the other byte order:test_migrate_endianhas a>u2case.bytes-little-endiancase, so one case is non-native on every host.The
infodoctests useendianness=....The regression tests for bugs 2 and 3 fail on little-endian hosts without the fix, so the normal CI covers them.
test_v3_data_types_parse_in_host_byte_orderonly fails on a big-endian host.Not in this PR: the
BytesCodec()defaultA bare
BytesCodec()defaults toendian=sys.byteorder, so code that builds one itself still writes big-endian chunks on s390x. That output is valid, but it is not identical to what x86 writes.History: the bytes codec in the original zarrita import defaulted to
"little". The host-order default came in with #1660, and neither that PR's description nor its comments discuss the switch.What happens here: changing the default back changes what existing code writes on big-endian hosts, so it needs a deprecation cycle. An earlier commit in this branch changed the default; a later commit reverts it. The deprecation is proposed in #4439.
CI job
.github/workflows/big-endian.ymlruns the byte-order-sensitive tests on an emulated s390x host. It usesdocker/setup-qemu-actionplus buildx, and the image is cached withtype=gha.ci/big-endian/Dockerfile, a Debian sid image that gets numpy and numcodecs from Debian packages, so nothing heavy compiles under emulation. This is also the environment Debian tests zarr in.tests/test_dtype,test_dtype_registry.py,tests/test_codecs,tests/test_metadata,test_array.py,test_v2.py,test_info.py,test_pipeline_parity.pyandtest_migrate_v3.py.workflow_dispatch.🤖 Generated with Claude Code