Skip to content

rational M3: the single stages, every ratio of the vocabulary - #58

Draft
tap wants to merge 1 commit into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6
Draft

tap wants to merge 1 commit into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6

Conversation

@tap

@tap tap commented Oct 2, 2026

Copy link
Copy Markdown
Owner

What this changes

Milestone M3 of rational/PLAN.md: the single stages ship for every ratio of the vocabulary — ↑2, ↓2, ↑3, ↓3, ↑6, ↓6, ↑8, ↓8 and the mixed 3/2, 2/3, 4/3, 3/4, 8/3, 3/8 — in all four sample formats.

  • stage.h — basic_stage<S, R>: one Nyquist stage at ratio L/M with bridge's call shapes (process, pull, exact outputs_for / frames_needed, flush, reset) and the latency as an exact rational, (N − 1) / (2M) output frames (R7). Two machines, chosen by the ratio at compile time:
    • an interpolator or a mixed ratio runs the L-phase schedule machine, each phase row stored tap-reversed and trimmed to its nonzero span, so an interpolator's centre phase is the single tap 1 (a copy) and a mixed ratio's zero-padded last phase costs its real taps only — skipping an exact-zero product changes no bit in any format;
    • a decimator runs the M-branch commutator, its nonzero branches summed in branch order through one accumulator (tap::dsp::accumulate_row, fir_kernels: accumulate_row, the dot without its finalize DspTap#51) and finalized once, the centre branch trimmed to its one tap: MACs per output are the design's nonzero count, 2m(M − 1) + 1 (R4). The half-band decimator dots 23 of its 43 taps.
  • Table gain: going up the L phases are the design's branches bit for bit; a decimator's table is h / M; a mixed ratio going down has each L-phase normalized to DC gain 1 (bridge's argument). Fixed-point rows are quantized with row-sum preservation (per phase; the decimator's whole filter as one row), so DC gain is exactly 1 at every output in every format. Channels planar inside, interleaved at the API. The stage models tap::dsp::sync_stage (size_t accounting, so on the 32-bit targets too).
  • converter.h — basic_converter<S, R> = basic_stage<S, R> with the converter<R> / converter_q15<R> / converter_q31<R> aliases; rational.h is now the umbrella over all five headers (the family's header-count pin moves to 5).
  • The scipy leg — tools/reference/make_reference_vectors.py and the committed tests/reference/reference_vectors.h: 14 ratios at economy and the by-2 pair at transparent, 480 inputs each, scipy.signal.upfirdn in float64 over the stage's table (the designer and the gain rule mirrored in numpy), cast to float32.
  • Tests (62 rational entries, from 25): the impulse reproduces the table bit for bit from every input phase of every ratio in every format (the trait's own mac + finalize as the oracle); the scipy vectors sample for sample; accounting exact from every position of every superblock; chunking bit-identical in every format; pull against process bit-exact and the pull draws exactly frames_needed; dry-then-resume; reset; flush equals zero padding bit for bit; the latency by impulse; DC exact in every format; the MAC counts; two channels independent; DspTap's decimators as a second golden at by 2 / 3 / 6; output hashes printed per ratio and format (not pinned, bridge's reasoning).
  • Submodule pin → DspTap's accumulate_row (fir_kernels: accumulate_row, the dot without its finalize DspTap#51). Draft until Ignore the migration's rename commits in git blame #51 merges: the pin names the DspTap branch commit and will be re-pointed at the identical tree on main.
  • rational/PLAN.md v0.6 records M3 done against its acceptance criteria; the family PLAN.md, both READMEs and both CLAUDE files follow.

Why

PLAN.md section 6, M3. The plan sequenced the half- and third-band stages first and the rest after; since the two machines are generic over the ratio, the whole vocabulary landed at once with one battery. The symmetry-halved table named in the plan's layout is deferred to the codegen levers after M6's baselines, as bridge did (its M7 lever 3): it changes storage, not bits, and the ratchet should measure it.

Verification

  • Clang 18 -Werror and GCC -Werror (all three TAP_SR_*_WERROR): the whole family battery, 235 / 235.
  • The scipy vectors: float within 6e-8 of the float64 reference for every ratio (the 3e-5 tolerance is bridge's); double within the reference's own float32 cast.
  • scripts/tidy.sh over the new TUs: clean.
  • Cortex-M33 under QEMU: the one-shot battery passes in 70 s with M2's filter (no new exclusions).
  • M55, Hexagon, MSVC and macOS: CI.

MACs per output at economy, measured by the stages' own counters: ↑2 11.5, ↓2 23, ↑3 15, ↓3 45, ↑6 20.2, ↓6 121, ↑8 21.1, ↓8 169, 3/2 15, 2/3 32.5, 4/3 16.75, 3/4 29, 8/3 21.1, 3/8 63.7.

Notes for the reviewer

  • Submodule pin moved — to fir_kernels: accumulate_row, the dot without its finalize DspTap#51's commit; re-pointed at its rebase merge on main before this leaves draft.
  • Contract, new: the stages' alignment, gain rule and accumulation order (branch order for decimators, tap order within a row) are now what the committed vectors and the output hashes pin; a later structural lever that changes float bits is a documented numeric change.
  • The generated reference header is committed as clang-format leaves it (five values per line), as bridge's is; the generator writes eight per line and the hook reflows.

🤖 Generated with Claude Code

https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA


Generated by Claude Code

stage.h: basic_stage<S, R>, one Nyquist stage at ratio L/M with bridge's
call shapes (process, pull, exact outputs_for / frames_needed, flush,
reset) and the latency as an exact rational, (N - 1) / (2 M) output
frames (R7). Two machines, chosen by the ratio at compile time:

- an interpolator or a mixed ratio runs the L-phase schedule machine
  (phase (n M) mod L, advance floor((n + 1) M / L) - floor(n M / L)), each
  phase row stored tap-reversed and trimmed to its nonzero span, so an
  interpolator's centre phase is the single tap 1 (a copy) and a mixed
  ratio's zero-padded last phase costs its real taps only; skipping an
  exact-zero product changes no bit in any format;
- a decimator runs the M-branch commutator (input n to sub-line n mod M,
  output k as x[k M] arrives), its nonzero branches summed in branch order
  through ONE accumulator (tap::dsp::accumulate_row, tap/DspTap#51) and
  finalized once, the centre branch trimmed to its one tap: MACs per
  output are the design's nonzero count, 2 m (M - 1) + 1 (R4) — the
  half-band decimator dots 23 of its 43 taps.

Table gain: going up the L phases are the design's branches bit for bit;
a decimator's table is h / M, a mixed ratio going down has each L-phase
normalized to DC gain 1 (bridge's argument); fixed-point rows are
quantized with row-sum preservation (per phase; the decimator's whole
filter as one row), so DC gain is exactly 1 at every output in every
format. Channels planar inside, interleaved at the API. The stage models
tap::dsp::sync_stage (size_t accounting, so on the 32-bit targets too).
converter.h: basic_converter<S, R> = basic_stage<S, R> with the
converter<R> / converter_q15<R> / converter_q31<R> aliases; rational.h
is the umbrella over all five headers (the family's header count pin 5).

Tests (62 rational entries): the impulse reproduces the table bit for
bit from every input phase of every ratio in every format (the trait's
own mac + finalize as the oracle); the committed scipy upfirdn vectors
(tools/reference/make_reference_vectors.py: 14 ratios at economy, the
by-2 pair at transparent, 480 inputs) sample for sample — float within
6e-8 (the tolerance is bridge's 3e-5), double within the reference's
float32 cast; accounting exact from every position of every superblock;
chunking bit-identical in every format; pull against process bit-exact
and the pull draws exactly frames_needed; dry-then-resume; reset; flush
equals zero padding bit for bit; the latency by impulse (an impulse at
the input phase that lands the centre tap on an output); DC exact; the
MAC counts; two channels independent; DspTap's decimators as a second
golden at by 2 / 3 / 6 (each engine within its own documented ripple of
the analytic sine); output hashes printed per ratio and format.
bare_metal_main.cpp's filter and the Hexagon row are unchanged: the
battery runs on the M33 leg in 70 s.

rational/PLAN.md v0.6 records M3 done against its acceptance criteria
(the whole vocabulary at once, the two machines being generic; the
symmetry-halved table deferred to the codegen levers after M6, as
bridge's M7 lever 3); the family plan, READMEs and CLAUDE files follow.
The submodule pin names tap/DspTap#51's branch commit until it merges.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
@tap
tap force-pushed the claude/sample-rate-expansion-strategies-ezqzu6 branch from a7cf96b to 814287b Compare October 2, 2026 01:26

tap commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

CI note: the red Cortex-M33 cross async (QEMU) check is on the superseded head 1ca7762 (the M2 tree, merged as #57), not on this PR's current head 814287b. Its tap_sr_async_tests_emulated one-shot hit ctest's 1800 s timeout at 1802.83 s with no test failure: the async battery's usual ~25 min under emulation, pushed past the limit while four runs of this branch shared the runner pool (the concurrency group deliberately does not cancel this branch's runs). Async is untouched by this PR.

The current head's run is queued behind that one; if the same timeout recurs there I will treat it as this PR's to address. No re-run of the stale head: its tree is already on main.


Generated by Claude Code

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants