Skip to content

feat(zarr-metadata)!: a shard is judged against the chunk it is handed, and its pipelines are read (layer 4d) - #369

Closed
d-v-b wants to merge 3 commits into
layer/4c-pipelinefrom
layer/4d-sharding
Closed

d-v-b wants to merge 3 commits into
layer/4c-pipelinefrom
layer/4d-sharding

Conversation

@d-v-b

@d-v-b d-v-b commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

🤖 AI text below 🤖

Upstream: zarr-developers#4443, a draft on main now that layer 3 (zarr-developers#4436) has merged. This fork PR stays as the stacked review copy.

Layer 4d of the stack, on top of #368 (4c, the codec pipeline): the sharding codec is judged against the shard it is handed, and its two pipelines are read as the array's is. This is the last of layer 4's four pieces: 4a fill values, 4b grid against shape, 4c the codec pipeline, 4d sharding.

document["shape"] = [16, 16]
document["chunk_grid"] = {"name": "regular", "configuration": {"chunk_shape": [8, 6]}}
document["codecs"] = [{"name": "sharding_indexed", "configuration": {
    "chunk_shape": [4, 4],
    "codecs": [{"name": "bytes", "configuration": {"endian": "little"}}],
    "index_codecs": ["bytes"],
}}]
validate_array_metadata_v3(document)
# codecs.0.configuration.chunk_shape.1: expected an inner chunk length that divides each length the shards take along this axis, [6], got 4
# codecs.0.configuration.index_codecs.0.configuration.endian: expected an endian, since each 'uint64' value takes several bytes

The shard. The sharding codec's chunk_shape has an inner chunk length for each axis of the shard it is handed, and each length divides every length the shards take along that axis. The shard is the chunk the codec is handed, after any codec before it. A transposed shard is divided along its transposed axes, as decided for layer 4. zarr-python reads divisibility from the untransposed grid instead.

The pipelines it holds. A codec definition gains one more hook, pipelines(configuration, nested, chunk). For each member of its configuration that holds a pipeline, it gives the chunk that pipeline's first codec is handed:

  • the inner codecs are handed the inner chunks, of chunk_shape and of the shard's data type;
  • the index codecs are handed the shard index: uint64, with the chunks per shard along each axis and a final axis of 2.

Each pipeline is read as the array's is: its order, then each codec against the chunk it is handed, with problems where they sit. A Stage keeps the inner stages as inner. An index bytes codec without an endian is now caught, and so is an inner transpose of the wrong rank. This also holds behind a codec nothing in scope claims, since the index is uint64 whatever the shard.

Like transition, pipelines is asked whatever the chunk rules found, and gives only what holds. A shard of another number of axes than its chunk_shape hands both pipelines lengths nothing is known of, since which of the two is wrong is not known. So it draws one problem, not three.

A problem inside is the inner field's own. This is the first commit. resolve used to read a field as invalid when anything inside it had a problem. So a shard holding a gzip with a level out of range, or a codec with must_understand: false, was left unread, and nothing else about it was judged. A document's fields are not read that way: a bad codec at one index leaves the others read. Now the fields a configuration holds are read the same way:

  • a problem of one is its own, reported where it sits, with its own resolution in nested;
  • the field holding it is read when its own configuration and rules are sound;
  • its rules are still asked only when each field it holds is named, since a rule may read one by name.

Every hook already treats an unread nested field as unknown, so this adds no false problems. It changes what resolve gives for a container (layer 2, merged, unreleased); two tests encoded the old answer. Without it, the correctness review found problems missing from 553 of 3,000 generated documents with shards, although the verdicts were right. The same codec also drew different problems at the top level and inside a shard.

Also here. The codec functions no codec of a kind is asked are now one declared table, _UNASKED, instead of one if each. A bytes -> bytes codec is handed bytes, and only an array -> array codec hands on a chunk.

Breaking. A document whose shard does not divide, or whose shard pipelines do not fit what they are handed, now has a problem where it sits, where the package accepted it before. A field holding a field with a problem now resolves as read, where it resolved as invalid.

Choices worth a look

  • The index's data type is a hand-built reading of the package's uint64 definition. Hooks get no scope, and the spec fixes the index type whatever the scope holds.
  • Pipelines are keyed by the member of the configuration that holds them. No codec holds a pipeline deeper than that.
  • Handed a chunk nothing is known of, a shard still hands on the inner chunks it declares, and an index of len(chunk_shape) + 1 axes. That is its own declaration, as cast_value names its data type.
  • A shard's pipelines are judged only as a document's codecs are read. resolve or canonicalize of a shard field alone does not see that codecs: [] has no array -> bytes codec. That rule needs no chunk, so it could move into the shard's rules, with the walk skipping it for held pipelines. Left for a follow-up.
  • A pipelines naming a member that holds data type fields nothing in scope claims cannot be told from one holding codecs nothing claims. It reads them as codecs nothing claims. Only a custom definition can do this.

Reviews. Two, with separate lenses.

The correctness review extended the 4c spec reference with the sharding rules:

  • 120,000 generated documents, with every inner and index stage compared, gave 0 disagreements against a reference that mirrors the walk.
  • All 17 planted bugs were caught, two only by the stage comparison, since no rule reads the index lengths.
  • 33,000 fuzzed documents raised nothing.
  • A strict reference found the problems hidden inside shards, which the first commit fixes. It also found a rank mismatch that drew three problems, now one.
  • It found a divisibility message that listed every length: on a rectilinear axis of 20,000 edges it was 134,548 characters, and it now names the smallest length not divided.

The design review's changes are in:

  • pipelines are keyed by member, not by path, so an untested path walk is gone;
  • a member holding no codecs is a TypeError;
  • the rank mismatch is fixed;
  • the tests are split per case, and the nested shard is folded into the main table.

Where zarr-python differs. The spec is the authority. The first two are filed: the transpose case on zarr-developers#2050, and the shard inside a shard as zarr-developers#4437.

  • zarr-python judges divisibility on the untransposed grid. It accepts [transpose(1,0), sharding_indexed chunk_shape [4,3]] on chunks of (4, 6), and a write and read returns different data. It refuses (6, 4), which its runtime handles.
  • It never checks a shard held inside a shard. It accepts a shard [4,4] holding a shard [3,4], and a round trip returns different data.
  • It accepts index_codecs of [crc32c], of [bytes, bytes], of bytes without an endian, or with a transpose of the wrong rank, and fails at the first write.

The stack: 0 and 1 merged upstream (zarr-developers#4420, zarr-developers#4421, zarr-developers#4422, zarr-developers#4432), 2 merged with 1b as zarr-developers#4434, 3 is #355 (upstream zarr-developers#4436), 4a is #366, 4b is #367, 4c is #368, 4d is this PR, and it completes layer 4.

🤖 Generated with Claude Code

…s own

`resolve` read a field as invalid when anything inside it had a problem:
a codec a shard holds with a gzip `level` out of range, or a
`must_understand` of false on one, left the shard unread, so nothing
else about it was judged. A document's fields are not read that way: a
problem with one codec leaves the others read. Now the fields a
configuration holds are read the same way. A problem of one is its own,
reported where it sits with its own resolution in `nested`, and the field
holding it is read when its own configuration and rules are sound. Its
rules are still asked only when each field it holds is named, since a
rule may read one by name.

Assisted-by: ClaudeCode:claude-opus-5-5
…d, and its pipelines are read

The sharding codec's `chunk_shape` has an inner chunk length for each
axis of the shard it is handed, dividing each length the shards take
along it. The shard is the chunk the codec is handed, so a transposed
shard is divided along its transposed axes.

A codec definition says what the pipelines it holds are handed,
`pipelines`, by the member of its configuration that holds each: the
shard's inner codecs are handed its inner chunks, of its data type, and
its index codecs the shard index, of `uint64` with an axis more than the
shard. Each is read as the array's pipeline is, its problems where it
sits, and a `Stage` keeps its stages as `inner`. Like the transition,
`pipelines` gives only what holds whatever the chunk rules found: a
shard of another number of axes than its `chunk_shape` hands both
pipelines lengths nothing is known of. The functions no codec of a kind
is asked are one declared table.

Assisted-by: ClaudeCode:claude-opus-5-5
@d-v-b

d-v-b commented Sep 28, 2026

Copy link
Copy Markdown
Owner Author

🤖 AI text below 🤖

Merged upstream: all of layer 4 landed as zarr-developers#4443. This fork review copy is closed.

@d-v-b d-v-b closed this Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant