Skip to content

feat(zarr-metadata)!: the codecs are read as a pipeline, each against the chunk it is handed - #4442

Closed
d-v-b wants to merge 14 commits into
zarr-developers:mainfrom
d-v-b:feat/zarr-metadata-codec-pipeline
Closed

d-v-b wants to merge 14 commits into
zarr-developers:mainfrom
d-v-b:feat/zarr-metadata-codec-pipeline

Conversation

@d-v-b

@d-v-b d-v-b commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

🤖 AI text below 🤖

The third of four pieces of layer 4, depending on #4441 (4b, grid against shape): a v3 array's codecs are read as a pipeline. The order is judged first, and then each codec is judged against the chunk it is handed. The first seven commits are #4440's and #4441's; this PR's are the last seven.

document["data_type"] = "int16"
document["codecs"] = [{"name": "transpose", "configuration": {"order": [0]}}, "bytes", "crc32c"]
validate_array_metadata_v3(document)
# codecs.0.configuration.order: expected 2 entries, one per axis of the chunk the codec is handed, got 1
# codecs.1.configuration.endian: expected an endian, since each 'int16' value takes several bytes

What a codec is handed. The array hands its first codec a Chunk:

  • the lengths the grid's chunks take along each of the array's dimensions, from the grid's new chunk_lengths, with None where nothing says them;
  • the array's data type field.

A codec definition gains two hooks. They are the spec's pair: a codec computes what it gives from what it is handed, and fails on what it cannot take.

  • chunk_rules(configuration, nested, chunk): what the spec disallows in the codec when it is handed that chunk, located in its configuration.
  • transition: for an array -> array codec, the chunk it hands on. It is asked whatever the chunk rules found, so it gives only what holds either way.

read_pipeline(codecs, chunk) judges the order, then walks the codecs. It gives each codec's Stage with the chunk it is handed, which is what zarr-python needs to prepare its codecs. chunk_grid_lengths(grid, shape) replaces 4b's private chunk_grid_problems: it gives a grid's lengths along with its problems over the shape.

A chunk has an extent for each dimension of the array, so its rank comes from shape even under a grid nothing in scope claims.

Nothing is guessed. A codec the scope did not read might do anything: nothing in scope claims it, or it has a problem of its own. The codec after it is handed a chunk nothing is known of, and so is the codec after one that says nothing of what it hands on. A codec after that hands on only what it says of its own accord: cast_value still names its data type.

The rules.

  • The order: array -> array codecs, then one array -> bytes codec, then bytes -> bytes codecs. This also refuses an empty list, so the model's separate "at least one codec" rule is gone. A codec nothing in scope claims is left out of the order check, since it might be the array -> bytes codec.
  • transpose: its order has one entry per axis of its chunk, and it hands on the axes permuted (B[i] = A[order[i]]).
  • bytes: it needs an endian when handed numbers of several bytes, and it cannot take values that vary in size.
  • A struct's field types are of fixed size.
  • cast_value:
    • it casts from and to data types that model real numbers;
    • it wraps only to an integral type;
    • each scalar_map scalar is a fill value of the data type on its side of the cast;
    • it hands on its target.
  • scale_offset: it is handed numbers, its offset and scale are fill values of that data type, and it hands on what it is handed.

What a data type knows. storage says how a data type's values are stored: in single bytes, in several bytes at a time, or each in as many as it needs.

  • bool, int8, uint8 and r8 are single bytes.
  • The other core types and the numpy time types hold numbers of several bytes.
  • string and bytes vary.
  • A struct's storage is its fields'.

storage_of(data_type) asks for it.

Rules see what is inside. A definition's rules now receive the fields its configuration holds, as the scope read them. This is the same second argument 4a's fill_value_rules and 4b's shape_rules take. It is how struct refuses a variable-length field type, and how cast_value judges its target. It changes the signature of every rules function from layer 2 (#4434, merged, unreleased). A struct-only hook was the alternative, and the general one was chosen.

Also here. A field whose configuration is not an object used to be read as claimed by nothing. It now keeps the definition its name names, so the order check sees its kind.

Breaking. A document with one of the problems above, which the package accepted, now has a problem where it sits. Four tests built float or int32 arrays with a bare bytes codec, and now give it an endian. create_default says that an overridden multi-byte data type needs its codecs too.

Choices worth a look

  • A codec with a problem of its own hands on a chunk nothing is known of, so a problem downstream of it shows once the first is fixed. This is the "unknown effect: nothing downstream is judged" decision. The alternative, running transitions on configurations their own rules refused, would weaken what a transition can assume.
  • Where the specs are silent or disagree, nothing is judged:
    • Raw bits wider than 8 bits: the spec does not say whether a byte order applies, so their storage is unknown.
    • The numpy time types say they take "any codec that supports arrays of signed 64-bit integers", while neither cast_value nor scale_offset lists them. Neither codec refuses them.
    • Complex numbers under scale_offset: their arithmetic is well defined but the list does not name them.
  • Scalars follow the core fill-value encoding, which writes a float's infinity "Infinity". The cast_value README's own example writes "+Infinity", so that example is refused here, while zarr-python accepts it. The example reads as a spec defect, filed as cast_value README example writes "+Infinity", which the Zarr V3 fill value encoding spells "Infinity" zarr-extensions#76.
  • A codec whose chunk rules fail is still asked what it hands on. scale_offset handed strings hands the bytes codec strings, which it cannot take either, and both problems are reported.
  • A problem about a whole codec, such as bytes handed a string chunk, is located at the codec's configuration, as every rule's empty location is, even for a codec written as a bare name.
  • zarr-python re-checks every codec against the array's own data type. With cast-value-rs installed, it refuses [cast_value to float32, scale_offset offset 0.5, cast_value to int8, bytes] on an int8 array: "scale_offset offset value 0.5 is not representable in dtype int8". This PR reads each codec against the chunk it is actually handed, and accepts that document.
  • create_default's default bytes codec still has no endian. Giving it one would change the default document a test pins.
  • Out of scope: value arithmetic, such as whether a fill value survives a cast_value round trip. Sharding's inner pipelines are 4d.

Reviews. Two, with separate lenses.

The correctness review found no wrong verdict and no crash:

  • A Hypothesis differential against a reference written from the spec text agreed on 100,000 generated documents. Its reach was measured per rule. A stage-by-stage comparison of what each codec is handed agreed on 43,000 more.
  • The differential caught 14 of 17 planted bugs. The three it missed change no verdict: one was caught by the stage-by-stage check, and the other two are storage choices that both mean "no endian needed".
  • 45,000 mutated or arbitrary documents raised nothing.

Its findings are in:

  • the false claim about what follows an unread codec;
  • "+Infinity" in two docstrings;
  • wrap reported twice for one fact;
  • raw bits in the bytes docstring;
  • the fragment's gaps;
  • unchecked results from a grid's chunk_lengths;
  • the malformed-configuration fix above.

The design review's cuts and changes are in:

  • the one-field Pipeline wrapper, _was_read, the public stored_as and the model's duplicate empty-codecs rule are gone;
  • transition always returns a Chunk;
  • Chunk checks what it holds, so a definition's fault is reported as the definition's;
  • the rank comes from the shape;
  • storage functions are plain module-level functions, so they pickle by reference;
  • the tests count each codec's own problems and are filed by rule.

Known and left: the partials in 4a's fill-value rules make a pickled definition compare unequal to the original. A struct nested about 250 deep raises RecursionError, as it does on main.

The stack. Layers 0 (#4420, #4421, #4422), 1 (#4432), 2 (#4434) and 3 (#4436) have merged. 4a is #4440, 4b is #4441; this is 4c; 4d (#4443) follows.

🤖 Generated with Claude Code

Each data type's definition declares the JSON shape of its fill value and
the rules for one of that shape, and the v3 array validators judge a
document's fill value against the data type the scope read. A field that
is read keeps the fields it read inside, by location, so a struct judges
each field's fill value by that field's own type.

Assisted-by: ClaudeCode:claude-opus-5-5
…resolve kept

The simplest spelling of a field that read spells each field it holds
from its reading in `Resolved.nested`, rather than reading every subtree
again at each level. A nested field of a field without problems has none,
so the branch for one without a spelling could not be taken.

Assisted-by: ClaudeCode:claude-opus-5-5
…see past unknown keys

- The JSON check walks one frame per level of nesting, as `refine_json`
  does, so a fill value checked against a JSON-typed shape reads as deep
  as `refine_json` reads: a struct fill value 600 levels deep raised
  `RecursionError`.
- The integer, float and complex fill value rules are partial applications
  of module-level functions, so a scope's definitions pickle again.
- A key a fill value's shape does not declare is reported and left out, and
  the fill value rules still judge the rest, as `judge` does for a
  configuration.
- A struct whose reading holds no reading of a field's type leaves that
  field's fill value unjudged.
- `fill_value_problems` takes the reading of any data type, whatever its
  configuration's type.
- `create_default` says that an overridden `data_type` needs its own
  `fill_value`.

Assisted-by: ClaudeCode:claude-opus-5-5
A chunk grid's definition says which arrays it fits, and the v3 array
validators judge a document's grid against its shape once both are read:
a regular grid has a chunk length for each dimension, and 0 only for a
dimension of length 0; a rectilinear grid has chunk lengths for each
dimension that cover it.

Assisted-by: ClaudeCode:claude-opus-5-5
…ength of 1

The grid `create_default` derives from an overridden `shape` had a chunk
length of 0 for a dimension of length 0. The regular grid asks for chunk
lengths greater than zero, and `zarr` refuses to open such a grid, so the
derived grid now has a chunk length of 1 there. It still fits the shape.

Assisted-by: ClaudeCode:claude-opus-5-5
Assisted-by: ClaudeCode:claude-opus-5-5
…ration holds

`rules` is handed the configuration and the fields it holds as the scope
read them (`Nested`), as `fill_value_rules` and `shape_rules` are, so a
rule can judge a configuration by what is inside it: a struct's field
types, a cast's target. `resolve` reads the nested fields before it asks
the rules, and still reports the rules' problems first; `judge` reads in
no scope, and hands them nothing read. Every definition's rules take the
new argument.

Assisted-by: ClaudeCode:claude-opus-5-5
… the chunk it is handed

A v3 array's codecs are read in order, and each codec handed an array is
judged against the chunk it is handed. The array hands its first codec a
`Chunk`: the lengths its grid's chunks take along each of the array's
dimensions, which the grid's new `chunk_lengths` says, and its data
type. A codec definition says what the spec disallows in it handed a
chunk, `chunk_rules`, and an array -> array codec says what it hands on,
`transition`, whatever its chunk rules found. `transpose` has both: its
`order` has one entry per axis of its chunk, and it hands on the axes
permuted.

`read_pipeline(codecs, chunk)` judges the order first (array -> array
codecs, one array -> bytes codec, bytes -> bytes codecs), which refuses
an empty list too, and gives each codec's `Stage` with the chunk it is
handed. Nothing is guessed: the codec after one the scope did not read,
or after one that says nothing of what it hands on, is handed a chunk
nothing is known of. `chunk_grid_lengths` replaces the private
`chunk_grid_problems`, giving a grid field's lengths with its problems,
and a `Chunk` checks what it holds, so a definition's fault is reported
as the definition's.

Assisted-by: ClaudeCode:claude-opus-5-5
A data type definition says how its values are stored, `storage`: in
single bytes, in several bytes at a time, or each in as many as it
needs. `bool`, `int8`, `uint8` and `r8` are made of single bytes; the
other core types, and the numpy time types, hold numbers of several
bytes; `string` and `bytes` vary. Of raw bits wider than a byte the spec
does not say whether a byte order applies, so their storage is unknown.
A struct's is its fields', packed together. `storage_of` asks it of a
data type field a scope read.

Two rules follow. The `bytes` codec takes an `endian` when it is handed
numbers of several bytes, and is not handed values that vary in size. A
struct's field types are of fixed size.

Assisted-by: ClaudeCode:claude-opus-5-5
…the data types they meet

`cast_value` casts to a data type that models real numbers, wraps only to
an integral one, and is handed one that models real numbers; each scalar
of its `scalar_map` is a fill value of the data type on its side of the
cast. It hands on its data type, whatever it is handed. `scale_offset`
is handed a data type with arithmetic, and its `offset` and `scale` are
fill values of it; it hands on what it is handed.

Each codec names the data types it takes "defined in this repository",
so one it does not name is refused only where the spec's own words refuse
it: truth values, raw bits, text, byte strings and records are no
numbers, and complex numbers model no real number. The numpy time types
say they take any codec of 64-bit integers, and scale_offset does not
say whether complex numbers are its, so those are left be. A scalar is
judged by the core fill value encoding, which writes a float's infinity
`"Infinity"`, where the cast_value codec's own example writes
`"+Infinity"`.

Assisted-by: ClaudeCode:claude-opus-5-5
…till claimed by its name

`resolve` read a field whose configuration was not an object as claimed
by nothing, although its name names a definition, so the codec order was
judged without it: `[{"name": "gzip", "configuration": 5}, "bytes"]` drew
no problem for the gzip before the bytes codec. The field keeps the
definition its name names, unread.

Assisted-by: ClaudeCode:claude-opus-5-5
@d-v-b
d-v-b force-pushed the feat/zarr-metadata-codec-pipeline branch from da062bd to fb7064a Compare September 28, 2026 07:56
@github-actions github-actions Bot added needs release notes Automatically applied to PRs which haven't added release notes zarr-metadata Specific to the zarr-metadata sub-package labels Sep 28, 2026
@read-the-docs-community

Copy link
Copy Markdown

Documentation build overview

📚 zarr-metadata | 🛠️ Build #34798178 | 📁 Comparing fb7064a against latest (0401a7f)

  🔍 Preview build  

7 files changed · ± 7 modified

± Modified

@d-v-b

d-v-b commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

🤖 AI text below 🤖

This landed with #4443, which carried its commits and merged all of layer 4. Nothing is left to merge here. #4444 renames this PR's changelog fragment to #4443.

@d-v-b d-v-b closed this Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs release notes Automatically applied to PRs which haven't added release notes zarr-metadata Specific to the zarr-metadata sub-package

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant