Skip to content

Align the icount marker's format string and re-record the Hexagon baselines - #52

Merged
tap merged 1 commit into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6
Oct 1, 2026
Merged

tap merged 1 commit into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6

Conversation

@tap

@tap tap commented Oct 1, 2026

Copy link
Copy Markdown
Owner

What this changes

The first code follow-up PLAN.md step 5 lists after the migration (async/PLAN.md §4 item 6): the three icount workload mains print their *_ICOUNT_DONE line through an alignas(64) constexpr format string, and the Hexagon baselines of both engines are re-recorded. The engine README tables are regenerated from the baselines, the book's quoted snippet is updated, and the family plan, async/PLAN.md, bridge/PLAN.md §7 and async/docs/PERFORMANCE.md record the measurement. Harness only: no shipped code changes.

Why

The migration's renames changed the length of exception-message literals, .rodata shifted, and each workload's format string landed at an alignment where static musl's memcpy (which printf uses for the format's literal runs) takes a shorter path. Every Hexagon count fell by 47–87 instructions with every shipped function counting identically (the migration's per-function proof, PLAN.md §7), and the committed baselines were left high by those amounts, inside the ±3 % gate. A marker whose cost depends on where the linker put an unrelated literal is noise the exact (--exact) comparisons cannot tolerate.

Verification

All measured here, on toolchains that match the committed baselines to the instruction (the unchanged tree counts +0 on every target: Hexagon kernels, M33 and M55 spot checks).

  • The dependence, reproduced. A sweep injecting 0–13 bytes of .rodata ahead of the workload's literals (-include of a used static const char[]) moves the unaligned tree's counts with period 4: up_q15_se +7 / −107 / −204 / −47 for shifts 0–3, then repeating at 4–7; pipeline_q15 −88 / +34 / 0 / −87 likewise.
  • The fix, invariant. The same sweep on the aligned string gives one count at every shift on both workloads (up_q15_se 27,424,077; pipeline_q15 204,476,476).
  • Arm legs unchanged. icount.py --exact against the committed baselines passes on M33 and M55 for all 34 workloads of both engines (+0 each).
  • Hexagon re-recorded (icount.py --update): async pipelines −95, kernels +7/+7/+38; bridge −211 (Q15/Q31), −242 (float). The CI ratchet on this PR is the first run of the new baselines on the runner image.
  • Host clang -Wall -Wextra -Wconversion -Wformat=2 -Werror builds the two workload TUs and the comparison TU (cmp_main.cpp, engine 0); pre-commit clean; mdbook build clean; update_icount_docs.py produces exactly the committed README tables.

Notes for the reviewer

  • A rejected variant, recorded as a finding. The first version also gave stdout a 64-byte-aligned static buffer via setvbuf. On Hexagon it was equally invariant, but on both Arm legs the two extra statics in main flipped GCC's decision to inline run<float>() (per-function attribution: main 29.2 M → 21, run<float> 0 → 28.95 M) and moved bridge's float counts by −0.3…−0.6 % on M55 and +0.0005…+0.014 % on M33 with no change to the measured code. The format-only fix leaves the Arm legs at +0. The finding (a harness edit that changes main's size can move Arm baselines through the workload function's inlining alone; pinning run<S>() out of line would re-record every baseline) is in async/docs/PERFORMANCE.md as deferred.
  • The async kernels moved up (+7/+7/+38) while the pipelines moved down (−95): the kernels' old string happened to sit at a cheaper alignment than the pipelines'. The aligned string now costs the same in every workload's image, which is the point.
  • The guest markers SRT_ICOUNT_DONE / RATIO_ICOUNT_DONE are unchanged (CLAUDE.md).
  • The remaining half of async/PLAN.md §4 item 6, bringing bench/icount/ under clang-tidy, is left for its own PR.

🤖 Generated with Claude Code

https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA


Generated by Claude Code

…elines

The family migration's renames changed the length of exception-message
literals, .rodata shifted, and each workload's *_ICOUNT_DONE format string
landed at an alignment where static musl's memcpy (printf copies the
format's literal runs with it) takes a shorter path: every Hexagon count
fell by 47-87 instructions with every shipped function identical, and the
committed baselines were left high by those amounts (PLAN.md section 7).

A sweep of one-byte .rodata shifts (0-13 bytes injected ahead of the
workload's literals with -include) reproduces the dependence with period 4
on the unaligned string and shows a 64-byte-aligned format string invariant
at every shift, so the three workload mains now print through an
alignas(64) constexpr format string. The variant that also gave stdout an
aligned static buffer was rejected: the extra statics in main flipped GCC's
inlining of run<float>() on the Arm legs and moved bridge's float counts by
0.3-0.6 % with no change to the measured code (recorded in
async/docs/PERFORMANCE.md as a finding).

Measured locally on toolchains that match the committed baselines to the
instruction (the unchanged tree counts +0 on every target): M33 and M55
exact (+0) on all 34 workloads; Hexagon async pipelines -95, kernels
+7/+7/+38, bridge -211 (Q15/Q31) and -242 (float). Hexagon baselines
re-recorded with icount.py --update; the engine README tables regenerated;
the book's quoted snippet and the plans' records updated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
@tap
tap merged commit db8e203 into main Oct 1, 2026
24 checks passed
@tap
tap deleted the claude/sample-rate-expansion-strategies-ezqzu6 branch October 1, 2026 11:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants