From a67928fe8a701dff9804f0919f68e38cafd8d294 Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 19:17:32 +0000 Subject: [PATCH 01/19] Make OCaml rewatch benchmarks reproducible across build scenarios Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 88 +++++++++++++++++++++---- rewatch-ocaml/bench/filesystem_audit.sh | 2 + rewatch-ocaml/bench/performance_gate.sh | 77 +++++++++++++++------- rewatch/tests/add-belt-dependencies.mjs | 2 +- 4 files changed, 133 insertions(+), 36 deletions(-) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index f28b755e93..ab486d6888 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -43,6 +43,66 @@ byte-identical edited JavaScript. The watcher held 12 file descriptors and four tasks; RSS rose from 26,680 to 27,688 KiB. That small fixture dirties only one module per edit and does not establish watch-time scaling. +## Native Linux testrepo checkpoint + +At compiler revision `49951ac49f78be57f8f7ae347bf577a2839f90ca`, five +interleaved runs on a 12-CPU Linux ARM64 host with OCaml 5.5.1 measured the +472-module testrepo fixture. Both implementations used the same local compiler +and runtime. The fixture includes installed, lockfile-pinned dependencies; +the benchmark copies them into isolated roots and applies the canonical +test-suite Belt-dependency correction to those copies. No company-project +source was used. + +| scenario | Rust median wall | OCaml median wall | Rust median peak tree RSS | OCaml median peak tree RSS | +| --- | ---: | ---: | ---: | ---: | +| Clean, 8 OCaml workers | 3,625 ms | 1,511 ms | 264,644 KiB | 358,256 KiB | +| Unchanged after clean, 8 workers | 142 ms | 71 ms | 29,156 KiB | 25,764 KiB | +| One source edit, 8 workers | 163 ms | 72 ms | 45,120 KiB | 25,520 KiB | +| Clean, 6 OCaml workers | 3,724 ms | 1,673 ms | 252,272 KiB | 303,536 KiB | +| Clean, 4 OCaml workers | 3,642 ms | 2,020 ms | 250,644 KiB | 219,364 KiB | + +Eight workers gave a 2.40x clean-build wall-time gain and a 1.35x peak-RSS +ratio relative to Rust in the complete clean/unchanged/edit run. The existing +125% clean-memory gate therefore failed at eight workers. Six workers passed +the complete gate, at a 2.23x clean-build gain and 1.20x peak-RSS ratio. Four +workers also passed, with a 1.80x clean-build gain. These are separate runs, +so compare ratios within a row rather than treating small cross-run differences +as an effect of worker count. Peak tree RSS is sampled every 20 ms and can +miss short-lived child peaks. + +All three worker counts matched Rust's clean, unchanged, and edit compiler +work. The clean build made 1,031 logical compiler requests in each +implementation: 512 parse, seven namespace, and 512 compile requests. The +unchanged build made four requests because the warning in the testrepo's +`ModuleA` deliberately invalidates its AST for diagnostic replay. Every +comparison matched the complete post-build file set and the selected stable +artifact bytes. + +In a separate seven-edit retained-watch run at the default eight workers, +Rust's median edit-to-hook latency was 128 ms and OCaml's was 82 ms. Both made +seven parser and seven compiler requests, produced identical edited JavaScript, +and held stable file descriptor, task, and RSS counts. Watch-mode samples use +a small one-module fixture, so they do not establish scaling on a large +dependency graph. + +One interactive clean build reported Rust parse/compile times of 1.35/2.09 s +and OCaml times of 0.16/1.24 s. These single-run phase timings are diagnostic, +not medians. The remaining OCaml compile phase alone exceeds the roughly +0.73 s total required for a 5x gain over the measured Rust clean median. +Further work toward 5x therefore needs substantial compilation or scheduling +improvement; eliminating the already short parse phase cannot reach it alone. + +The filesystem audit counted 6,604 Rust versus 9,056 OCaml project-local +metadata calls on clean builds, with nearly equal open counts (17,636 and +17,633). The high-count missing CMI lookups, including 313 opens of the +fixture's `Pervasives.cmi` path, were identical in both implementations. The +extra metadata checks merit investigation on slower filesystems, but the +shared CMI lookup pattern does not identify an OCaml-specific optimization. +An exploratory `OCAMLRUNPARAM=o=70` run reduced the eight-worker median peak +tree RSS to 330,232 KiB with a 1,505 ms clean median; it still missed the +125% memory gate in that interleaved comparison. The default GC setting is +unchanged pending broader workload evidence. + ## AST I/O checkpoint Temporary counters on the same host and eight-domain fixture measured 917 @@ -62,27 +122,33 @@ behavior, then measure RSS, work, and artifact parity on a larger project. ## Rust comparison gates -Build both release executables, then run the Linux clean-build, work, resource, +Build the local runtime, lockfile-pinned testrepo dependencies, and both +release executables. Then run the Linux clean, unchanged, edit, work, resource, and artifact comparison: ```sh +yarn --cwd rewatch/testrepo install --immutable +opam exec -- make lib cargo build --manifest-path rewatch/Cargo.toml --release opam exec -- dune build --profile release rewatch-ocaml/rescript_ocaml.exe +export RESCRIPT_BSC_EXE="$PWD/_build/default/compiler/bsc/rescript_compiler_main.exe" +export RESCRIPT_RUNTIME="$PWD/packages/@rescript/runtime" rewatch-ocaml/bench/performance_gate.sh \ rewatch/target/release/rescript \ _build/default/rewatch-ocaml/rescript_ocaml.exe 5 ``` -The harness isolates dependency trees, interleaves builds, traces Rust `bsc` -requests and OCaml's logical compiler-request log, compares clean, unchanged, -and one-edit work, and compares complete file sets and stable artifact bytes at -the same absolute path. `KEEP_REWATCH_BENCHMARK_WORKDIR=1` retains raw outputs. -The fixture currently does work on an unchanged build. In a one-run smoke check, -both implementations performed four compiler requests there, and their clean -and single-edit request counts matched too. The source of those unexpected -unchanged requests is unresolved. The same check found byte differences in -some `rescript-bun` CMI files and an AST despite equal generated file sets; -investigate these before using the artifact gate as an acceptance result. +The harness isolates dependency trees, interleaves timed clean, unchanged, +and single-edit builds, and samples process-tree RSS and task counts. It then +traces Rust `bsc` requests and OCaml's logical compiler-request log, compares +work in all three scenarios, and compares complete file sets and stable +artifact bytes at the same absolute path. The timed source edits add unique +comments to `packages/watch-warnings/src/B.res` in both isolated fixtures. +The unchanged workload replays the fixture's local `ModuleA` warning, so four +compiler requests there are expected. `KEEP_REWATCH_BENCHMARK_WORKDIR=1` +retains raw outputs and `results.csv`. The 125% wall-time and memory limits +currently apply to the clean scenario; the other scenarios are measured and +checked for equivalent work and artifacts. ## Filesystem-work audit diff --git a/rewatch-ocaml/bench/filesystem_audit.sh b/rewatch-ocaml/bench/filesystem_audit.sh index 19a1ffc92b..e896dd6176 100755 --- a/rewatch-ocaml/bench/filesystem_audit.sh +++ b/rewatch-ocaml/bench/filesystem_audit.sh @@ -44,6 +44,8 @@ prepare_fixture() { cp -a --reflink=auto "$dependency_tree" "$destination/$relative_tree" done < <(find "$repo_root/rewatch/testrepo" -type d -name node_modules \ -prune -print) + node "$repo_root/rewatch/tests/add-belt-dependencies.mjs" \ + "$destination/rewatch/testrepo" } if [[ -z ${RESCRIPT_BSC_EXE:-} || -z ${RESCRIPT_RUNTIME:-} ]]; then diff --git a/rewatch-ocaml/bench/performance_gate.sh b/rewatch-ocaml/bench/performance_gate.sh index b9aa48e942..e6cde0158a 100755 --- a/rewatch-ocaml/bench/performance_gate.sh +++ b/rewatch-ocaml/bench/performance_gate.sh @@ -62,6 +62,8 @@ prepare_fixture() { cp -a --reflink=auto "$dependency_tree" "$destination/$relative_tree" done < <(find "$repo_root/rewatch/testrepo" -type d -name node_modules \ -prune -print) + node "$repo_root/rewatch/tests/add-belt-dependencies.mjs" \ + "$destination/rewatch/testrepo" } rust_root="$work_root/rust" @@ -77,7 +79,7 @@ fi export RESCRIPT_BSC_EXE RESCRIPT_RUNTIME results="$work_root/results.csv" -echo "implementation,iteration,wall_ms,peak_tree_rss_kib,peak_tree_tasks" \ +echo "scenario,implementation,iteration,wall_ms,peak_tree_rss_kib,peak_tree_tasks" \ >"$results" tree_resources() { @@ -105,9 +107,16 @@ clean_and_build() { } measure() { - local implementation=$1 executable=$2 fixture=$3 iteration=$4 - local output="$work_root/${implementation}-${iteration}" - "$executable" clean "$fixture" >/dev/null 2>&1 + local scenario=$1 implementation=$2 executable=$3 fixture=$4 iteration=$5 + local output="$work_root/${scenario}-${implementation}-${iteration}" + case "$scenario" in + clean) "$executable" clean "$fixture" >/dev/null 2>&1 ;; + unchanged) ;; + edit) + printf '\n// timed single edit %d\n' "$iteration" \ + >>"$fixture/packages/watch-warnings/src/B.res" ;; + *) echo "Unknown benchmark scenario: $scenario" >&2; exit 2 ;; + esac local start_ns root_pid peak_rss=0 peak_tasks=0 rss tasks end_ns wall_ms start_ns=$(date +%s%N) "$executable" build "$fixture" >"$output" 2>"$output.stderr" & @@ -125,15 +134,17 @@ measure() { wait "$root_pid" end_ns=$(date +%s%N) wall_ms=$(((end_ns - start_ns) / 1000000)) - echo "$implementation,$iteration,$wall_ms,$peak_rss,$peak_tasks" >>"$results" - printf '%-5s run %d: %6d ms %8d KiB %4d tasks\n' \ - "$implementation" "$iteration" "$wall_ms" "$peak_rss" "$peak_tasks" + echo "$scenario,$implementation,$iteration,$wall_ms,$peak_rss,$peak_tasks" \ + >>"$results" + printf '%-9s %-5s run %d: %6d ms %8d KiB %4d tasks\n' \ + "$scenario" "$implementation" "$iteration" "$wall_ms" "$peak_rss" \ + "$peak_tasks" } median_column() { - local implementation=$1 column=$2 middle=$((runs / 2 + 1)) - awk -F, -v implementation="$implementation" \ - '$1 == implementation { print $'"$column"' }' "$results" \ + local scenario=$1 implementation=$2 column=$3 middle=$((runs / 2 + 1)) + awk -F, -v scenario="$scenario" -v implementation="$implementation" \ + '$1 == scenario && $2 == implementation { print $'"$column"' }' "$results" \ | sort -n | sed -n "${middle}p" } @@ -142,32 +153,50 @@ echo "commit: $(git -C "$repo_root" rev-parse HEAD)" echo "host: $(uname -a)" echo "cpus: $(getconf _NPROCESSORS_ONLN 2>/dev/null || echo unknown)" echo "runs: $runs (interleaved after one warm-up each)" -echo "threshold: ${threshold_percent}% of Rust median wall and RSS" +echo "clean threshold: ${threshold_percent}% of Rust median wall and RSS" clean_and_build "$rust_executable" "$rust_fixture" "$work_root/rust-warmup" clean_and_build "$ocaml_executable" "$ocaml_fixture" "$work_root/ocaml-warmup" for ((iteration = 1; iteration <= runs; iteration++)); do if ((iteration % 2 == 1)); then - measure rust "$rust_executable" "$rust_fixture" "$iteration" - measure ocaml "$ocaml_executable" "$ocaml_fixture" "$iteration" + measure clean rust "$rust_executable" "$rust_fixture" "$iteration" + measure clean ocaml "$ocaml_executable" "$ocaml_fixture" "$iteration" else - measure ocaml "$ocaml_executable" "$ocaml_fixture" "$iteration" - measure rust "$rust_executable" "$rust_fixture" "$iteration" + measure clean ocaml "$ocaml_executable" "$ocaml_fixture" "$iteration" + measure clean rust "$rust_executable" "$rust_fixture" "$iteration" fi done -rust_wall=$(median_column rust 3) -ocaml_wall=$(median_column ocaml 3) -rust_rss=$(median_column rust 4) -ocaml_rss=$(median_column ocaml 4) -rust_tasks=$(median_column rust 5) -ocaml_tasks=$(median_column ocaml 5) -printf 'median Rust: %6d ms %8d KiB %4d peak tasks\n' \ +rust_wall=$(median_column clean rust 4) +ocaml_wall=$(median_column clean ocaml 4) +rust_rss=$(median_column clean rust 5) +ocaml_rss=$(median_column clean ocaml 5) +rust_tasks=$(median_column clean rust 6) +ocaml_tasks=$(median_column clean ocaml 6) +printf 'clean median Rust: %6d ms %8d KiB %4d peak tasks\n' \ "$rust_wall" "$rust_rss" "$rust_tasks" -printf 'median OCaml: %6d ms %8d KiB %4d peak tasks\n' \ +printf 'clean median OCaml: %6d ms %8d KiB %4d peak tasks\n' \ "$ocaml_wall" "$ocaml_rss" "$ocaml_tasks" +for scenario in unchanged edit; do + for ((iteration = 1; iteration <= runs; iteration++)); do + if ((iteration % 2 == 1)); then + measure "$scenario" rust "$rust_executable" "$rust_fixture" "$iteration" + measure "$scenario" ocaml "$ocaml_executable" "$ocaml_fixture" "$iteration" + else + measure "$scenario" ocaml "$ocaml_executable" "$ocaml_fixture" "$iteration" + measure "$scenario" rust "$rust_executable" "$rust_fixture" "$iteration" + fi + done + printf '%s median Rust: %6d ms %8d KiB peak tree RSS\n' \ + "$scenario" "$(median_column "$scenario" rust 4)" \ + "$(median_column "$scenario" rust 5)" + printf '%s median OCaml: %6d ms %8d KiB peak tree RSS\n' \ + "$scenario" "$(median_column "$scenario" ocaml 4)" \ + "$(median_column "$scenario" ocaml 5)" +done + trace_and_classify() { local implementation=$1 scenario=$2 executable=$3 fixture=$4 manifest=$5 local clean_first=$6 @@ -387,5 +416,5 @@ fi if ((runs < 5)); then echo "PASS: correctness smoke checks passed; performance gate not evaluated." else - echo "PASS: timing, memory, compiler-work, and artifact-equivalence gates passed." + echo "PASS: clean timing and memory, compiler-work, and artifact-equivalence gates passed." fi diff --git a/rewatch/tests/add-belt-dependencies.mjs b/rewatch/tests/add-belt-dependencies.mjs index a50502a34a..418d2c7d1e 100644 --- a/rewatch/tests/add-belt-dependencies.mjs +++ b/rewatch/tests/add-belt-dependencies.mjs @@ -3,7 +3,7 @@ import * as fs from "node:fs/promises"; import * as path from "node:path"; -const testrepo = path.join(import.meta.dirname, "..", "testrepo"); +const testrepo = process.argv[2] ?? path.join(import.meta.dirname, "..", "testrepo"); const configs = [ path.join(testrepo, "node_modules", "rescript-nodejs", "rescript.json"), path.join( From e415568fd13f573ab4fd30e2c24832a70d705565 Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 19:27:40 +0000 Subject: [PATCH 02/19] Add opt-in compiler request timing trace Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 32 +++++++ .../bench/analyze_compiler_timing.js | 93 +++++++++++++++++++ rewatch-ocaml/compiler_process.ml | 54 +++++++---- rewatch-ocaml/tests/run.sh | 8 +- 4 files changed, 168 insertions(+), 19 deletions(-) create mode 100644 rewatch-ocaml/bench/analyze_compiler_timing.js diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index ab486d6888..5f4f632fde 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -103,6 +103,38 @@ tree RSS to 330,232 KiB with a 1,505 ms clean median; it still missed the 125% memory gate in that interleaved comparison. The default GC setting is unchanged pending broader workload evidence. +An opt-in per-request timing trace resolves the OCaml compile phase further. +On one eight-worker clean build of the same fixture, 512 parse requests +spanned 95 ms and 512 implementation/interface requests spanned 1,172 ms. +The compiler workers were active for virtually the entire compile span, with +7.39 of eight workers active on average and all eight active at peak. The compile +requests summed to 8,666 ms of worker time; the 95th percentile request took +48 ms. `DOMAPI.ast`, `Net.ast`, and `Http.ast` were among the slowest requests. +This is a single diagnostic run with logging enabled, not a benchmark median. +It indicates that the clean-build limit is largely compiler work rather than +idle scheduler time on this fixture. At eight workers, the summed work alone +has a 1.08 s lower bound without faster individual requests. + +To collect another trace, set `REWATCH_COMPILER_TIMING_LOG` to an absolute +path for an OCaml build and analyze the resulting tab-separated file: + +```sh +export RESCRIPT_BSC_EXE="$PWD/_build/default/compiler/bsc/rescript_compiler_main.exe" +export RESCRIPT_RUNTIME="$PWD/packages/@rescript/runtime" +_build/default/rewatch-ocaml/rescript_ocaml.exe clean rewatch/testrepo +rm -f /tmp/rewatch-compiler-timing.tsv +REWATCH_COMPILER_TIMING_LOG=/tmp/rewatch-compiler-timing.tsv \ + _build/default/rewatch-ocaml/rescript_ocaml.exe build rewatch/testrepo +node rewatch-ocaml/bench/analyze_compiler_timing.js \ + /tmp/rewatch-compiler-timing.tsv +``` + +Each row records phase, working directory, input, start time, and end time. +Request time includes any PPX command that the compiler invokes. +The analyzer reports elapsed phase span, summed compiler time, average and +peak active requests, idle time inside each phase, and the longest compile +requests. Remove an old trace before a new run; the compiler appends rows. + ## AST I/O checkpoint Temporary counters on the same host and eight-domain fixture measured 917 diff --git a/rewatch-ocaml/bench/analyze_compiler_timing.js b/rewatch-ocaml/bench/analyze_compiler_timing.js new file mode 100644 index 0000000000..2184444e04 --- /dev/null +++ b/rewatch-ocaml/bench/analyze_compiler_timing.js @@ -0,0 +1,93 @@ +#!/usr/bin/env node + +import fs from "node:fs"; + +if (process.argv.length !== 3) { + console.error("Usage: analyze_compiler_timing.js TIMING_LOG"); + process.exit(2); +} + +const rows = fs + .readFileSync(process.argv[2], "utf8") + .trim() + .split("\n") + .filter(Boolean) + .map((line, index) => { + const [phase, cwd, input, startText, endText, ...extra] = line.split("\t"); + const start = Number(startText); + const end = Number(endText); + if ( + extra.length > 0 || + !["parse", "namespace", "interface", "implementation"].includes(phase) || + !cwd || + !input || + !Number.isFinite(start) || + !Number.isFinite(end) || + end <= start + ) { + throw new Error(`Invalid timing row ${index + 1}`); + } + return { phase, cwd, input, start, end }; + }); + +function summarize(name, requests) { + if (requests.length === 0) return; + const durations = requests + .map(({ start, end }) => (end - start) * 1000) + .sort((a, b) => a - b); + const events = requests + .flatMap(({ start, end }) => [ + { time: start, delta: 1 }, + { time: end, delta: -1 }, + ]) + .sort((a, b) => a.time - b.time || a.delta - b.delta); + let active = 0; + let peak = 0; + let zeroActive = 0; + let previous = events[0].time; + for (const { time, delta } of events) { + if (active === 0) zeroActive += time - previous; + active += delta; + peak = Math.max(peak, active); + previous = time; + } + const span = (events.at(-1).time - events[0].time) * 1000; + const summed = durations.reduce((total, duration) => total + duration, 0); + console.log( + [ + name, + requests.length, + span.toFixed(1), + summed.toFixed(1), + (summed / span).toFixed(2), + peak, + (zeroActive * 1000).toFixed(1), + durations[Math.ceil(0.95 * durations.length) - 1].toFixed(2), + ].join(","), + ); +} + +console.log( + "phase,requests,span_ms,summed_job_ms,mean_active,peak_active,zero_active_ms,p95_job_ms", +); +summarize( + "parse", + rows.filter(({ phase }) => phase === "parse"), +); +summarize( + "namespace", + rows.filter(({ phase }) => phase === "namespace"), +); +summarize( + "compile", + rows.filter(({ phase }) => phase === "interface" || phase === "implementation"), +); + +console.log("longest compile jobs:"); +rows + .filter(({ phase }) => phase === "interface" || phase === "implementation") + .sort((a, b) => (b.end - b.start) - (a.end - a.start)) + .slice(0, 5) + .forEach(({ phase, cwd, input, start, end }) => { + console.log(`${((end - start) * 1000).toFixed(1)} ms ${phase} ${cwd}/${input}`); + }); diff --git a/rewatch-ocaml/compiler_process.ml b/rewatch-ocaml/compiler_process.ml index 9d447949cf..84eaa669b8 100644 --- a/rewatch-ocaml/compiler_process.ml +++ b/rewatch-ocaml/compiler_process.ml @@ -24,17 +24,34 @@ let compiler_phase args = in (phase, input) +let append_log path line = + let channel = open_out_gen [Open_creat; Open_append; Open_text] 0o644 path in + Fun.protect + ~finally:(fun () -> close_out_noerr channel) + (fun () -> output_string channel line) + let log_compiler_request (job : Process.job) = match Sys.getenv_opt "REWATCH_COMPILER_CALL_LOG" with | None -> () | Some path -> let phase, input = compiler_phase job.args in - let channel = - open_out_gen [Open_creat; Open_append; Open_text] 0o644 path - in + append_log path (Printf.sprintf "%s\t%s\t%s\n" phase job.cwd input) + +let compiler_timing_log = Sys.getenv_opt "REWATCH_COMPILER_TIMING_LOG" + +let time_compiler_request (job : Process.job) run = + match compiler_timing_log with + | None -> run () + | Some path -> + let started = Unix.gettimeofday () in Fun.protect - ~finally:(fun () -> close_out_noerr channel) - (fun () -> Printf.fprintf channel "%s\t%s\t%s\n" phase job.cwd input) + ~finally:(fun () -> + let finished = Unix.gettimeofday () in + let phase, input = compiler_phase job.args in + append_log path + (Printf.sprintf "%s\t%s\t%s\t%.9f\t%.9f\n" phase job.cwd input started + finished)) + run let exit_code = function | Unix.WEXITED code -> code @@ -51,19 +68,20 @@ let run_in_process ?poll (job : Process.job) = } | input :: reversed_argv -> let result = - Rescript_compiler_driver.run_request ~cwd:job.cwd - ~argv:(List.rev reversed_argv) ~input - ~run_external: - (Some - (fun command -> - let command = Platform.shell_command command in - (* Signal handlers are process-wide; domain workers launch PPXs - without replacing the scheduler domain's handlers. *) - let result = - Process.run ?poll ~defer_signals:false ~cwd:job.cwd - command.program command.args - in - (exit_code result.status, result.stdout, result.stderr))) + time_compiler_request job (fun () -> + Rescript_compiler_driver.run_request ~cwd:job.cwd + ~argv:(List.rev reversed_argv) ~input + ~run_external: + (Some + (fun command -> + let command = Platform.shell_command command in + (* Signal handlers are process-wide; domain workers launch + PPXs without replacing the scheduler domain's handlers. *) + let result = + Process.run ?poll ~defer_signals:false ~cwd:job.cwd + command.program command.args + in + (exit_code result.status, result.stdout, result.stderr)))) in { Process.status = Unix.WEXITED result.exit_code; diff --git a/rewatch-ocaml/tests/run.sh b/rewatch-ocaml/tests/run.sh index 0a3b596ffb..a433c0aa18 100644 --- a/rewatch-ocaml/tests/run.sh +++ b/rewatch-ocaml/tests/run.sh @@ -869,7 +869,13 @@ rm -f "$basic/src/A.mjs" mkdir -p "$basic/lib/bs/other" touch "$basic/lib/bs/other/Authored.js" -"$port" build --after-build 'test -f src/A.mjs' "$basic" +REWATCH_COMPILER_TIMING_LOG=$(native_path "$work/compiler-timing.tsv") \ + "$port" build --after-build 'test -f src/A.mjs' "$basic" +test -s "$work/compiler-timing.tsv" +node "$root/rewatch-ocaml/bench/analyze_compiler_timing.js" \ + "$work/compiler-timing.tsv" >"$work/compiler-timing.summary" +grep -Eq '^parse,[1-9][0-9]*,' "$work/compiler-timing.summary" +grep -Eq '^compile,[1-9][0-9]*,' "$work/compiler-timing.summary" test -f "$basic/lib/bs/build.ninja" test -f "$basic/src/A.mjs" test -f "$basic/src/Authored.js" From 19cd9bf6d81f2b696db91fd090731c79a6c65ca2 Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 19:28:10 +0000 Subject: [PATCH 03/19] Record retained watch filesystem audit Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index 5f4f632fde..140f4a4150 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -229,6 +229,12 @@ tables, process attribution, and command output. As with the short-lived audit, project-local repeated paths and compiler work are the useful comparison; raw runtime-wide syscall totals are diagnostic rather than an acceptance limit. +On the small `basic` fixture, one retained-watch edit produced 19 Rust versus +26 OCaml project-local metadata calls and 43 versus 42 opens. The most +repeated source and compiler-artifact opens were similar in both versions. +This one-edit trace does not indicate a large OCaml-specific filesystem cost +on the watch path; it says little about larger dependency graphs. + The retained-watch performance and resource gate exercises several ordinary edits through the same long-lived watcher: From d49ad430bf84a5b05ddefeeea4be2dad062a3efd Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 19:31:25 +0000 Subject: [PATCH 04/19] Document compiler-core timing bottlenecks Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 22 ++++++++++++++++++++++ 1 file changed, 22 insertions(+) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index 140f4a4150..c424cbddec 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -135,6 +135,28 @@ The analyzer reports elapsed phase span, summed compiler time, average and peak active requests, idle time inside each phase, and the longest compile requests. Remove an old trace before a new run; the compiler appends rows. +A separate temporary compiler-core trace split 479 implementation requests +from one instrumented clean build. These are summed concurrent-worker times, +not elapsed build time: + +| compiler-core phase | summed worker time | +| --- | ---: | +| Initial environment setup | 1,837 ms | +| Type checking, including CMI/CMT work | 6,736 ms | +| Lambda translation | 30 ms | +| Lambda compilation | 135 ms | +| JavaScript emission | 46 ms | + +The remaining 33 interface requests and outer request setup are outside this +split. The instrumented build's 512 compile requests spanned 1,253 ms, so do +not compare that span directly with the uninstrumented benchmark median. +Environment setup and type checking account for nearly all measured +implementation work. Reusing a prepared environment across requests would +have to preserve the compiler's fresh per-request type and identifier state; +it is an architectural change with correctness and memory risks. Faster +JavaScript emission alone has little headroom on this fixture. The temporary +compiler-core instrumentation was removed after the measurement. + ## AST I/O checkpoint Temporary counters on the same host and eight-domain fixture measured 917 From 95499eda7e68283f7c6f198628f4cce6018f6dec Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 19:55:19 +0000 Subject: [PATCH 05/19] Align benchmark compiler profiles and refresh results Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 94 +++++++++++++++---------- rewatch-ocaml/bench/performance_gate.sh | 1 + 2 files changed, 59 insertions(+), 36 deletions(-) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index c424cbddec..c8485d0719 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -5,7 +5,9 @@ The OCaml rewatch now uses parallel in-process compiler domains by default. the worker count is `min(8, max(1, available CPUs - 1))`. Rust rewatch is the build-level reference for compiler work, generated file sets, and stable artifact bytes. Keep both implementations on the same compiler and runtime -build when comparing them. +build and the same Dune profile when comparing them. The profile affects +serialized AST and CMI bytes for at least one testrepo dependency, even when +compiler request arguments are identical. ## Domain baseline @@ -45,30 +47,41 @@ only one module per edit and does not establish watch-time scaling. ## Native Linux testrepo checkpoint -At compiler revision `49951ac49f78be57f8f7ae347bf577a2839f90ca`, five +At revision `7688129cd0bd6d8b66fec797327db3455398a4a1`, five interleaved runs on a 12-CPU Linux ARM64 host with OCaml 5.5.1 measured the -472-module testrepo fixture. Both implementations used the same local compiler -and runtime. The fixture includes installed, lockfile-pinned dependencies; -the benchmark copies them into isolated roots and applies the canonical -test-suite Belt-dependency correction to those copies. No company-project -source was used. +472-module testrepo fixture. Standalone `bsc` and the embedded compiler were +built with the same Dune `release` profile; Rust Rewatch was a Cargo release +build. Both implementations used the same local runtime. The fixture includes +installed, lockfile-pinned dependencies; the benchmark copies them into +isolated roots and applies the canonical test-suite Belt-dependency correction +to those copies. No company-project source was used. Earlier measurements at +`49951ac49f78be57f8f7ae347bf577a2839f90ca` did not pin both compiler +executables to the same Dune profile and are superseded by this checkpoint. | scenario | Rust median wall | OCaml median wall | Rust median peak tree RSS | OCaml median peak tree RSS | | --- | ---: | ---: | ---: | ---: | -| Clean, 8 OCaml workers | 3,625 ms | 1,511 ms | 264,644 KiB | 358,256 KiB | -| Unchanged after clean, 8 workers | 142 ms | 71 ms | 29,156 KiB | 25,764 KiB | -| One source edit, 8 workers | 163 ms | 72 ms | 45,120 KiB | 25,520 KiB | -| Clean, 6 OCaml workers | 3,724 ms | 1,673 ms | 252,272 KiB | 303,536 KiB | -| Clean, 4 OCaml workers | 3,642 ms | 2,020 ms | 250,644 KiB | 219,364 KiB | - -Eight workers gave a 2.40x clean-build wall-time gain and a 1.35x peak-RSS -ratio relative to Rust in the complete clean/unchanged/edit run. The existing -125% clean-memory gate therefore failed at eight workers. Six workers passed -the complete gate, at a 2.23x clean-build gain and 1.20x peak-RSS ratio. Four -workers also passed, with a 1.80x clean-build gain. These are separate runs, -so compare ratios within a row rather than treating small cross-run differences -as an effect of worker count. Peak tree RSS is sampled every 20 ms and can -miss short-lived child peaks. +| Clean, 8 OCaml workers | 3,585 ms | 1,413 ms | 252,464 KiB | 368,032 KiB | +| Unchanged after clean, 8 workers | 142 ms | 71 ms | 30,056 KiB | 26,072 KiB | +| One source edit, 8 workers | 140 ms | 72 ms | 28,648 KiB | 26,156 KiB | +| Clean, 7 OCaml workers | 3,712 ms | 1,527 ms | 260,280 KiB | 334,112 KiB | +| Clean, 6 OCaml workers | 3,617 ms | 1,574 ms | 275,592 KiB | 289,640 KiB | + +Eight workers gave a 2.54x clean-build wall-time gain and a 1.46x sampled +peak-tree-RSS ratio relative to Rust. Seven workers gave a 2.43x gain but +still failed the existing 125% clean-memory gate in its comparison. Six +workers passed the complete gate, at a 2.30x gain and a 1.05x sampled RSS +ratio. These are separate runs, so compare ratios within a row. Peak tree RSS +is sampled every 20 ms and can miss short-lived compiler-child peaks; Rust's +sampled clean peak varied substantially across runs. The memory gate gives a +directional constraint, especially near its threshold, rather than a precise +cross-architecture memory ratio. + +In one clean build, GNU `/usr/bin/time -v` reported 49,644 KiB maximum RSS +for Rust and 385,944 KiB for OCaml. Rust launches many compiler children, +whereas OCaml compiles mostly in-process. GNU time's maximum RSS does not sum +the concurrent process tree, so those two values are not a total-build memory +comparison. Keep the sampled tree figure alongside any per-executable GNU +time reading. All three worker counts matched Rust's clean, unchanged, and edit compiler work. The clean build made 1,031 logical compiler requests in each @@ -79,29 +92,26 @@ comparison matched the complete post-build file set and the selected stable artifact bytes. In a separate seven-edit retained-watch run at the default eight workers, -Rust's median edit-to-hook latency was 128 ms and OCaml's was 82 ms. Both made +Rust's median edit-to-hook latency was 142 ms and OCaml's was 83 ms. Both made seven parser and seven compiler requests, produced identical edited JavaScript, and held stable file descriptor, task, and RSS counts. Watch-mode samples use a small one-module fixture, so they do not establish scaling on a large dependency graph. -One interactive clean build reported Rust parse/compile times of 1.35/2.09 s -and OCaml times of 0.16/1.24 s. These single-run phase timings are diagnostic, -not medians. The remaining OCaml compile phase alone exceeds the roughly -0.73 s total required for a 5x gain over the measured Rust clean median. -Further work toward 5x therefore needs substantial compilation or scheduling -improvement; eliminating the already short parse phase cannot reach it alone. +The later per-request timing trace measured an OCaml compile span of 1.17 s. +A 5x clean gain over the aligned Rust median would require the entire build +to finish in about 0.72 s. The compiler work alone exceeds that budget on +this fixture; faster parsing or orchestration alone cannot reach it. -The filesystem audit counted 6,604 Rust versus 9,056 OCaml project-local +The aligned filesystem audit counted 6,604 Rust versus 9,042 OCaml project-local metadata calls on clean builds, with nearly equal open counts (17,636 and 17,633). The high-count missing CMI lookups, including 313 opens of the fixture's `Pervasives.cmi` path, were identical in both implementations. The extra metadata checks merit investigation on slower filesystems, but the shared CMI lookup pattern does not identify an OCaml-specific optimization. -An exploratory `OCAMLRUNPARAM=o=70` run reduced the eight-worker median peak -tree RSS to 330,232 KiB with a 1,505 ms clean median; it still missed the -125% memory gate in that interleaved comparison. The default GC setting is -unchanged pending broader workload evidence. +Exploratory lower GC space-overhead settings reduced OCaml RSS in individual +runs, but did not establish a validated advantage over six default-GC +workers. The default GC setting is unchanged. An opt-in per-request timing trace resolves the OCaml compile phase further. On one eight-worker clean build of the same fixture, 512 parse requests @@ -157,6 +167,16 @@ it is an architectural change with correctness and memory risks. Faster JavaScript emission alone has little headroom on this fixture. The temporary compiler-core instrumentation was removed after the measurement. +A second temporary trace split `Typemod.type_implementation_more` on the same +fixture. Among 479 implementation requests, `type_structure` used 5,253 ms +of summed worker time, inclusion and delayed checks 820 ms, CMI saving 267 ms, +and CMT saving 443 ms. Signature simplification and reset took under 3 ms +combined. This is a separate instrumented run, so its totals differ slightly +from the compiler-core split above. The type-structure work is the main part +of the remaining compiler cost; eliminating artifact writes alone has a +limited bound. The temporary type-checker instrumentation was removed and +the normal release binary rebuilt afterward. + ## AST I/O checkpoint Temporary counters on the same host and eight-domain fixture measured 917 @@ -184,7 +204,9 @@ and artifact comparison: yarn --cwd rewatch/testrepo install --immutable opam exec -- make lib cargo build --manifest-path rewatch/Cargo.toml --release -opam exec -- dune build --profile release rewatch-ocaml/rescript_ocaml.exe +opam exec -- dune build --profile release \ + compiler/bsc/rescript_compiler_main.exe \ + rewatch-ocaml/rescript_ocaml.exe export RESCRIPT_BSC_EXE="$PWD/_build/default/compiler/bsc/rescript_compiler_main.exe" export RESCRIPT_RUNTIME="$PWD/packages/@rescript/runtime" rewatch-ocaml/bench/performance_gate.sh \ @@ -280,8 +302,8 @@ as well as resource growth that a one-event syscall trace cannot show. The default median-latency limit is 150% of Rust because individual watch events include operating-system notification and 50 ms polling intervals; override it with `REWATCH_WATCH_PERFORMANCE_THRESHOLD_PERCENT` only for investigation. -The build gate's lower-noise 125% clean/incremental threshold remains the -authoritative general performance criterion. Set +The build gate's 125% clean-build wall-time limit remains the general latency +criterion; unchanged and single-edit builds are measured separately. Set `KEEP_REWATCH_WATCH_PERFORMANCE=1` to retain output, compiler-call logs, latencies, and fixtures. This gate requires Linux `/proc`, GNU-compatible millisecond `date`, and `setsid`. diff --git a/rewatch-ocaml/bench/performance_gate.sh b/rewatch-ocaml/bench/performance_gate.sh index e6cde0158a..dc435fc7a5 100755 --- a/rewatch-ocaml/bench/performance_gate.sh +++ b/rewatch-ocaml/bench/performance_gate.sh @@ -407,6 +407,7 @@ if ((file_set_equivalence == 0)); then fi if ((artifact_equivalence == 0)); then echo "FAIL: Rust and OCaml generated different artifacts." >&2 + echo "Check that standalone bsc and OCaml rewatch use the same Dune profile." >&2 failed=1 fi From 87f58d4b60ae024036e8f935efcdd08608ae0dbe Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 20:02:14 +0000 Subject: [PATCH 06/19] Record matched before-and-after OCaml benchmarks Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 49 +++++++++++++++++++++++++++++++++++ 1 file changed, 49 insertions(+) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index c8485d0719..40b1186b0f 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -45,6 +45,55 @@ byte-identical edited JavaScript. The watcher held 12 file descriptors and four tasks; RSS rose from 26,680 to 27,688 KiB. That small fixture dirties only one module per edit and does not establish watch-time scaling. +## Direct before-and-after comparison + +Revision `95467b2bf` is the parent of the parallel in-process compiler change. +It still launches standalone `bsc` requests. On the same 12-CPU Linux ARM64 +host, five interleaved testrepo runs compared its release executable with +revision `7494d76ece1f33e8a808c26206194b7a20158766`. Both used the +current lockfile-pinned fixture, release-profile `bsc`, and local runtime. +Compiler source files did not change between these revisions: + +| scenario | before median wall | after median wall | before peak tree RSS | after peak tree RSS | +| --- | ---: | ---: | ---: | ---: | +| Clean | 4,045 ms | 1,420 ms | 302,984 KiB | 365,916 KiB | +| Unchanged | 187 ms | 71 ms | 42,336 KiB | 26,080 KiB | +| One source edit | 185 ms | 70 ms | 53,052 KiB | 26,152 KiB | + +The clean build became 2.85x faster with 1.21x sampled peak tree RSS. Both +revisions made the same 1,031 clean compiler requests and the same unchanged +and edit requests. Complete post-build file sets and selected stable artifact +bytes matched. The before/after gate passed its 125% clean RSS limit in this +comparison. The first executable is labeled `Rust` by the reusable benchmark +script below, but it is the older OCaml build system; its external compiler +calls are observed by the same `bsc` proxy. This comparison measures the +revision range, including the in-process compiler and parallel scheduling, +not an isolated compiler micro-optimization. + +In a separate seven-edit retained-watch comparison, median latency fell from +161 to 83 ms. Both revisions made seven parser and seven compiler requests +and produced identical edited JavaScript. The older watcher's RSS rose from +10,960 to 11,736 KiB and the current watcher's from 26,800 to 27,932 KiB; +both held stable file descriptor and task counts. + +To reproduce the before/after gate after preparing the dependencies and +release compiler/runtime as described below: + +```sh +git worktree add --detach /tmp/rewatch-before-954 95467b2bf +(cd /tmp/rewatch-before-954 && \ + opam exec -- dune build --profile release rewatch-ocaml/rescript_ocaml.exe) +export RESCRIPT_BSC_EXE="$PWD/_build/default/compiler/bsc/rescript_compiler_main.exe" +export RESCRIPT_RUNTIME="$PWD/packages/@rescript/runtime" +REWATCH_COMPILER_DOMAINS=8 rewatch-ocaml/bench/performance_gate.sh \ + /tmp/rewatch-before-954/_build/default/rewatch-ocaml/rescript_ocaml.exe \ + _build/default/rewatch-ocaml/rescript_ocaml.exe 5 +REWATCH_WATCH_COMPILER_DOMAINS=8 \ + rewatch-ocaml/bench/watch_performance_gate.sh \ + /tmp/rewatch-before-954/_build/default/rewatch-ocaml/rescript_ocaml.exe \ + _build/default/rewatch-ocaml/rescript_ocaml.exe 7 +``` + ## Native Linux testrepo checkpoint At revision `7688129cd0bd6d8b66fec797327db3455398a4a1`, five From 1830b9ecd3658ab0b056afb5dd2b2d0bf3e4fa1a Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 20:12:05 +0000 Subject: [PATCH 07/19] Document CMI lookup cost and reuse bound Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 13 +++++++++++++ 1 file changed, 13 insertions(+) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index 40b1186b0f..22108df08f 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -226,6 +226,19 @@ of the remaining compiler cost; eliminating artifact writes alone has a limited bound. The temporary type-checker instrumentation was removed and the normal release binary rebuilt afterward. +A third temporary trace timed CMI loading in `Bs_cmi_load` on one eight-worker +clean build. The 512 compile requests made 2,952 successful persistent-module +lookups, taking 2,711 ms summed across workers, including file selection and +failed candidate opens. They decoded 2,992 CMIs totaling 194 MB of repeated +input, which took 1,511 ms summed across workers; the extra 40 reads came +through other CMI call sites. These measurements include tracing overhead and +are not wall-time savings. Even eliminating all 2,711 ms of lookup work would +have an ideal eight-worker bound of about 0.34 s, short of closing the 5x gap. +A raw-byte cache would still pay most decoding cost, while reusing decoded +type graphs across fresh compiler requests would need safe copying and CMI +invalidation to preserve dependency correctness. The temporary trace code +was removed and both release compiler executables rebuilt afterward. + ## AST I/O checkpoint Temporary counters on the same host and eight-domain fixture measured 917 From 64de04177f43b447dc1eda16eaa2fd066089b3c3 Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 20:17:17 +0000 Subject: [PATCH 08/19] Document repeated signature opening cost Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index 22108df08f..059c1b34ce 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -239,6 +239,21 @@ type graphs across fresh compiler requests would need safe copying and CMI invalidation to preserve dependency correctness. The temporary trace code was removed and both release compiler executables rebuilt afterward. +A further temporary single-worker trace separated module lookup from opening +the resolved signature. Across 1,897 opens, lookup used about 360 ms of +process CPU time and signature opening about 2,063 ms. The 137 opens of +`WebAPI.DOMAPI` alone used about 230 ms for lookup and 1,687 ms for signature +opening. Its CMI is roughly 987 kB, and expanding its large signature is +repeated in fresh compiler requests. A separate structure-item trace counted +429 source `open` items using about 1,784 ms of process CPU time, including +those `DOMAPI` opens; these two traces are separate runs and their times are +not additive. The one-worker timings are diagnostic, include tracing overhead +and some build-system CPU activity, and cannot be read as eight-worker wall +time savings. Even an ideal eight-way division of all `DOMAPI` opening work +would save only about 0.21 s. Reusing expanded components would need to keep +the mutable type graphs isolated and invalidate them when a CMI changes. The +temporary instrumentation was removed after these measurements. + ## AST I/O checkpoint Temporary counters on the same host and eight-domain fixture measured 917 From b2f9f94e1f43d2fc5037cfffad8f9fe082ed6987 Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 20:50:55 +0000 Subject: [PATCH 09/19] Build opened signature label tables in bulk Signed-off-by: Codex --- compiler/ml/env.ml | 20 +++++++++++++++++- rewatch-ocaml/bench/performance_gate.sh | 19 ++++++++++++++--- rewatch-ocaml/bench/watch_performance_gate.sh | 21 +++++++++++++------ 3 files changed, 50 insertions(+), 10 deletions(-) diff --git a/compiler/ml/env.ml b/compiler/ml/env.ml index 48af05a901..c5e072aa71 100644 --- a/compiler/ml/env.ml +++ b/compiler/ml/env.ml @@ -1514,6 +1514,8 @@ and components_of_module_maker (env, sub, path, mty) = let pl, sub = prefix_idents path sub sg in let env = ref env in let pos = ref 0 in + let labels_by_name = Hashtbl.create 127 in + let label_names_rev = ref [] in List.iter2 (fun item path -> match item with @@ -1540,7 +1542,13 @@ and components_of_module_maker (env, sub, path, mty) = constructors; List.iter (fun descr -> - c.comp_labels <- add_to_tbl descr.lbl_name descr c.comp_labels) + let name = descr.lbl_name in + match Hashtbl.find labels_by_name name with + | _, previous -> + Hashtbl.replace labels_by_name name (name, descr :: previous) + | exception Not_found -> + Hashtbl.add labels_by_name name (name, [descr]); + label_names_rev := name :: !label_names_rev) labels; env := store_type_infos id decl !env | Sig_typext (id, ext, _) -> @@ -1568,6 +1576,16 @@ and components_of_module_maker (env, sub, path, mty) = Tbl.add (Ident.name id) (decl', nopos) c.comp_modtypes; env := store_modtype id decl !env) sg pl; + (* Large signatures often repeat label names. Keep first appearance order + to preserve Tbl's shape and the latest key and declarations to preserve + its contents. *) + c.comp_labels <- + List.fold_left + (fun table name -> + let latest_name, declarations = Hashtbl.find labels_by_name name in + Tbl.add latest_name declarations table) + Tbl.empty + (List.rev !label_names_rev); Some (Structure_comps c) | Mty_functor (param, _ty_arg, ty_res) -> Some diff --git a/rewatch-ocaml/bench/performance_gate.sh b/rewatch-ocaml/bench/performance_gate.sh index dc435fc7a5..207fabc25d 100755 --- a/rewatch-ocaml/bench/performance_gate.sh +++ b/rewatch-ocaml/bench/performance_gate.sh @@ -16,6 +16,11 @@ if [[ ! -x "$rust_executable" || ! -x "$ocaml_executable" ]]; then echo "Both rewatch executables must exist and be executable." >&2 exit 2 fi +if [[ ${REWATCH_FIRST_EMBEDDED:-0} == 1 && + $(dirname "$rust_executable") != $(dirname "$ocaml_executable") ]]; then + echo "Place both embedded executables in the same directory to avoid startup-path bias." >&2 + exit 2 +fi if [[ ! "$runs" =~ ^[1-9][0-9]*$ || $((runs % 2)) -eq 0 ]]; then echo "RUNS must be a positive odd integer so the median is unambiguous." >&2 exit 2 @@ -154,6 +159,9 @@ echo "host: $(uname -a)" echo "cpus: $(getconf _NPROCESSORS_ONLN 2>/dev/null || echo unknown)" echo "runs: $runs (interleaved after one warm-up each)" echo "clean threshold: ${threshold_percent}% of Rust median wall and RSS" +if [[ ${REWATCH_FIRST_EMBEDDED:-0} == 1 ]]; then + echo "first executable uses embedded compiler tracing; 'Rust' labels mean baseline" +fi clean_and_build "$rust_executable" "$rust_fixture" "$work_root/rust-warmup" clean_and_build "$ocaml_executable" "$ocaml_fixture" "$work_root/ocaml-warmup" @@ -200,14 +208,19 @@ done trace_and_classify() { local implementation=$1 scenario=$2 executable=$3 fixture=$4 manifest=$5 local clean_first=$6 + local embedded=0 + if [[ $implementation == ocaml || ${REWATCH_FIRST_EMBEDDED:-0} == 1 ]]; then + embedded=1 + fi local trace_prefix="$work_root/${implementation}-${scenario}.execve" local call_log="$work_root/${implementation}-${scenario}.compiler" if [[ "$clean_first" == 1 ]]; then "$executable" clean "$fixture" >/dev/null 2>&1 fi - if [[ $implementation == ocaml ]]; then + if ((embedded)); then # Embedded requests have no compiler execve. Record them at their shared # logical boundary while tracing PPXs as external processes. + : >"$call_log" strace -f -ff -qq -s 4096 -e trace=execve,chdir -o "$trace_prefix" \ env REWATCH_COMPILER_CALL_LOG="$call_log" \ "$executable" build "$fixture" \ @@ -224,7 +237,7 @@ trace_and_classify() { local trace_file exec_line argv cwd_line cwd phase input identity : >"$manifest.unsorted" for trace_file in "${trace_files[@]}"; do - if [[ $implementation == ocaml ]]; then + if ((embedded)); then exec_line=$(grep -m1 -E 'execve\("[^"]*sury-ppx' "$trace_file" || true) else exec_line=$(grep -m1 -F "execve(\"$RESCRIPT_BSC_EXE\"" "$trace_file" \ @@ -257,7 +270,7 @@ trace_and_classify() { | sed "s#$implementation_root##g" >>"$manifest.unsorted" done local invocations parse namespace compile interface ppx - if [[ $implementation == ocaml ]]; then + if ((embedded)); then while IFS=$'\t' read -r phase cwd input; do if [[ $phase != parse && $phase != namespace ]]; then phase=compile; fi printf '%s\t%s\t"%s"\n' "$cwd" "$phase" "$input" \ diff --git a/rewatch-ocaml/bench/watch_performance_gate.sh b/rewatch-ocaml/bench/watch_performance_gate.sh index 58dbf7737b..48a0ff8514 100755 --- a/rewatch-ocaml/bench/watch_performance_gate.sh +++ b/rewatch-ocaml/bench/watch_performance_gate.sh @@ -140,7 +140,7 @@ start_watcher() { cp -R "$repo_root/rewatch-ocaml/tests/basic" "$fixture" : >"$work_root/$implementation.bsc" : >"$work_root/$implementation.compiler" - if [[ $implementation == ocaml ]]; then + if [[ $implementation == ocaml || ${REWATCH_FIRST_EMBEDDED:-0} == 1 ]]; then local -a domain_count_env=() if [[ ${REWATCH_WATCH_COMPILER_DOMAINS+x} ]]; then domain_count_env=(REWATCH_COMPILER_DOMAINS="$REWATCH_WATCH_COMPILER_DOMAINS") @@ -245,17 +245,26 @@ median() { rust_median=$(median "$work_root/rust.latencies") ocaml_median=$(median "$work_root/ocaml.latencies") -rust_parse_count=$(grep -cF -- '-bs-ast' "$work_root/rust.bsc" || true) -rust_total_count=$(wc -l <"$work_root/rust.bsc") +if [[ ${REWATCH_FIRST_EMBEDDED:-0} == 1 ]]; then + rust_parse_count=$(grep -c '^parse' "$work_root/rust.compiler" || true) + rust_compile_count=$(grep -c '^implementation' \ + "$work_root/rust.compiler" || true) + rust_total_count=$(wc -l <"$work_root/rust.compiler") +else + rust_parse_count=$(grep -cF -- '-bs-ast' "$work_root/rust.bsc" || true) + rust_total_count=$(wc -l <"$work_root/rust.bsc") + rust_compile_count=$((rust_total_count - rust_parse_count)) +fi ocaml_parse_count=$(grep -c '^parse' "$work_root/ocaml.compiler" || true) ocaml_compile_count=$(grep -c '^implementation' \ "$work_root/ocaml.compiler" || true) ocaml_total_count=$(wc -l <"$work_root/ocaml.compiler") -if ((rust_parse_count != runs || rust_total_count != runs * 2 || +if ((rust_parse_count != runs || rust_compile_count != runs || + rust_total_count != runs * 2 || ocaml_parse_count != runs || ocaml_compile_count != runs || ocaml_total_count != runs * 2)); then - printf 'retained work mismatch: Rust %d parser / %d total; embedded OCaml %d parser / %d compiler / %d total; expected %d / %d.\n' \ - "$rust_parse_count" "$rust_total_count" "$ocaml_parse_count" \ + printf 'retained work mismatch: first executable %d parser / %d compiler / %d total; embedded OCaml %d parser / %d compiler / %d total; expected %d / %d.\n' \ + "$rust_parse_count" "$rust_compile_count" "$rust_total_count" "$ocaml_parse_count" \ "$ocaml_compile_count" "$ocaml_total_count" "$runs" "$((runs * 2))" >&2 exit 1 fi From cacbf09a48a504323b3c0f72bbb91f97b8a9c9f7 Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 20:57:40 +0000 Subject: [PATCH 10/19] Time builds independently of resource sampling Signed-off-by: Codex --- rewatch-ocaml/bench/performance_gate.sh | 31 ++++++++++++++++--------- 1 file changed, 20 insertions(+), 11 deletions(-) diff --git a/rewatch-ocaml/bench/performance_gate.sh b/rewatch-ocaml/bench/performance_gate.sh index 207fabc25d..9cca5c9e89 100755 --- a/rewatch-ocaml/bench/performance_gate.sh +++ b/rewatch-ocaml/bench/performance_gate.sh @@ -122,22 +122,31 @@ measure() { >>"$fixture/packages/watch-warnings/src/B.res" ;; *) echo "Unknown benchmark scenario: $scenario" >&2; exit 2 ;; esac - local start_ns root_pid peak_rss=0 peak_tasks=0 rss tasks end_ns wall_ms + local start_ns root_pid sampler_pid peak_rss peak_tasks end_ns wall_ms + local resource_file="$output.resources" start_ns=$(date +%s%N) "$executable" build "$fixture" >"$output" 2>"$output.stderr" & root_pid=$! - while kill -0 "$root_pid" 2>/dev/null; do - read -r rss tasks < <(tree_resources "$root_pid") - if ((rss > peak_rss)); then - peak_rss=$rss - fi - if ((tasks > peak_tasks)); then - peak_tasks=$tasks - fi - sleep 0.02 - done + ( + peak_rss=0 + peak_tasks=0 + while kill -0 "$root_pid" 2>/dev/null; do + read -r rss tasks < <(tree_resources "$root_pid") + if ((rss > peak_rss)); then + peak_rss=$rss + fi + if ((tasks > peak_tasks)); then + peak_tasks=$tasks + fi + sleep 0.02 + done + printf '%d %d\n' "$peak_rss" "$peak_tasks" >"$resource_file" + ) & + sampler_pid=$! wait "$root_pid" end_ns=$(date +%s%N) + wait "$sampler_pid" + read -r peak_rss peak_tasks <"$resource_file" wall_ms=$(((end_ns - start_ns) / 1000000)) echo "$scenario,$implementation,$iteration,$wall_ms,$peak_rss,$peak_tasks" \ >>"$results" From 515f4f59aa4c1fe3d950ac9e02c208e1d6fe2e51 Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 20:58:37 +0000 Subject: [PATCH 11/19] Record label-table benchmark and reproduction steps Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 69 +++++++++++++++++++++++++++++++++++ 1 file changed, 69 insertions(+) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index 059c1b34ce..a749697245 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -254,6 +254,75 @@ would save only about 0.21 s. Reusing expanded components would need to keep the mutable type graphs isolated and invalidate them when a CMI changes. The temporary instrumentation was removed after these measurements. +## Bulk label table checkpoint + +Revision `56164c19e3b0cc751301e4344cc0e4ecff46df20` builds the opened +signature's record-label table once per distinct label name. The previous +revision was `8d1fa55ed0f88bfdebef17bbde797109dfa1e52f`. A temporary +single-worker trace of the testrepo's `WebAPI.DOMAPI` signature found 6,133 +label entries under 991 names in each of 137 expansions. Insertion into the +persistent table took about 461 ms summed across those expansions. The new +builder retains the first-seen insertion order, latest key, and per-name +declaration order; the temporary trace code was removed. + +On the same 12-CPU Linux ARM64 host, five interleaved eight-worker testrepo +runs used release-profile executables placed in the same directory, the same +standalone `bsc` and runtime, and one warm-up per executable. The benchmark +gate, updated at `719e4bd231bf59b21029dfff7584267550fb1799`, times process +completion separately from its 20 ms process-tree resource sampler: + +| scenario | before median wall | after median wall | before peak tree RSS | after peak tree RSS | +| --- | ---: | ---: | ---: | ---: | +| Clean | 1,424 ms | 1,414 ms | 373,612 KiB | 365,252 KiB | +| Unchanged | 45 ms | 44 ms | 26,892 KiB | 26,540 KiB | +| One source edit | 45 ms | 46 ms | 26,980 KiB | 26,628 KiB | + +The five-run clean difference is small beside run-to-run variation. A separate +ten-pair, high-resolution interleaved clean comparison without the resource +sampler measured 1,457 ms before and 1,430 ms after (1.9% faster). Twenty +warmed unchanged builds on isolated fixtures measured 45.10 and 45.01 ms; +there was no measurable incremental gain. These direct timings used +`process.hrtime.bigint()` around each child build, after cleaning before each +clean sample. The resource figures above are sampled peaks, not exact maximum +RSS, and do not establish a memory reduction. + +Both versions made the same 1,031 clean, four unchanged, and six edit compiler +requests. The complete post-build file sets and stable artifact bytes matched. +The clean time and memory gate passed. A seven-edit retained-watch comparison +measured 84 ms before and 82 ms after, with seven parse and seven compile +requests each, identical edited JavaScript, and stable watcher resources. +The small watch fixture cannot establish a latency gain. `make test`, +`make test-rewatch`, the OCaml Rewatch integration script, and the Rewatch +unit tests passed with the new compiler. No company-project performance is +inferred from these repository measurements. + +To reproduce the before/after gates after installing the dependencies shown +below, build both revisions with the Dune `release` profile and put their +executables in the same directory. For two embedded compiler executables, +`REWATCH_FIRST_EMBEDDED=1` makes the first argument use the logical request +trace; the harness still labels that first executable `Rust` in its output: + +```sh +git worktree add --detach /tmp/rewatch-before-bulk 8d1fa55ed +(cd /tmp/rewatch-before-bulk && opam exec -- dune build --profile release \ + compiler/bsc/rescript_compiler_main.exe rewatch-ocaml/rescript_ocaml.exe) +opam exec -- dune build --profile release \ + compiler/bsc/rescript_compiler_main.exe rewatch-ocaml/rescript_ocaml.exe +mkdir -p /tmp/rewatch-bulk-binaries +cp /tmp/rewatch-before-bulk/_build/default/rewatch-ocaml/rescript_ocaml.exe \ + /tmp/rewatch-bulk-binaries/before +cp _build/default/rewatch-ocaml/rescript_ocaml.exe \ + /tmp/rewatch-bulk-binaries/after +export RESCRIPT_BSC_EXE=/tmp/rewatch-before-bulk/_build/default/compiler/bsc/rescript_compiler_main.exe +export RESCRIPT_RUNTIME="$PWD/packages/@rescript/runtime" +REWATCH_FIRST_EMBEDDED=1 REWATCH_COMPILER_DOMAINS=8 \ + rewatch-ocaml/bench/performance_gate.sh \ + /tmp/rewatch-bulk-binaries/before /tmp/rewatch-bulk-binaries/after 5 +REWATCH_FIRST_EMBEDDED=1 REWATCH_WATCH_COMPILER_DOMAINS=8 \ + rewatch-ocaml/bench/watch_performance_gate.sh \ + /tmp/rewatch-bulk-binaries/before /tmp/rewatch-bulk-binaries/after 7 +``` + ## AST I/O checkpoint Temporary counters on the same host and eight-domain fixture measured 917 From 0af49e41a22d790707ab52e2186c10989e593269 Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 20:59:35 +0000 Subject: [PATCH 12/19] Record unexplained single-worker outliers Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index a749697245..082a137ef2 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -286,6 +286,12 @@ there was no measurable incremental gain. These direct timings used clean sample. The resource figures above are sampled peaks, not exact maximum RSS, and do not establish a memory reduction. +One exploratory five-pair single-worker comparison had isolated long clean +builds in both versions (15 s before and 52 s after). Five later traced clean +builds per version did not reproduce those outliers and made the expected 512 +compile requests each. The cause is unknown, so these single-worker samples +are not evidence for or against a stable tail-latency change. + Both versions made the same 1,031 clean, four unchanged, and six edit compiler requests. The complete post-build file sets and stable artifact bytes matched. The clean time and memory gate passed. A seven-edit retained-watch comparison From 02b518e3faaea02a0189891bfb2708baa68434a4 Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 21:07:14 +0000 Subject: [PATCH 13/19] Record compiler filesystem effect on Rewatch baselines Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 46 +++++++++++++++++++++++++ rewatch-ocaml/bench/performance_gate.sh | 5 +++ 2 files changed, 51 insertions(+) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index 082a137ef2..efc1c80947 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -329,6 +329,52 @@ REWATCH_FIRST_EMBEDDED=1 REWATCH_WATCH_COMPILER_DOMAINS=8 \ /tmp/rewatch-bulk-binaries/before /tmp/rewatch-bulk-binaries/after 7 ``` +## Current Rust comparison and compiler placement + +At revision `f8c60996fbb50a1f3678723f4910549e9fe3994d`, five interleaved +clean, unchanged, and edit runs used the Cargo release Rust executable, the +Dune release standalone `bsc` from `_build/default`, the Dune release OCaml +executable, and the same local runtime. This used the gate's independent wall +timer and 20 ms process-tree RSS sampler: + +| scenario | Rust median wall | OCaml median wall | Rust sampled peak tree RSS | OCaml sampled peak tree RSS | +| --- | ---: | ---: | ---: | ---: | +| Clean, 8 OCaml workers | 3,377 ms | 1,368 ms | 272,316 KiB | 361,520 KiB | +| Unchanged, 8 workers | 125 ms | 57 ms | 28,976 KiB | 26,320 KiB | +| One source edit, 8 workers | 126 ms | 56 ms | 38,784 KiB | 26,400 KiB | +| Clean, 6 OCaml workers | 3,626 ms | 1,572 ms | 263,240 KiB | 303,368 KiB | +| Unchanged, 6 workers | 131 ms | 57 ms | 29,880 KiB | 26,276 KiB | +| One source edit, 6 workers | 132 ms | 58 ms | 32,536 KiB | 26,268 KiB | + +Eight workers gave a 2.47x clean wall-time gain but exceeded the gate's 125% +sampled-memory limit. Six workers gave a 2.31x gain and passed that limit. +Both counts matched the 1,031 clean, four unchanged, and six edit logical +compiler requests, complete post-build file sets, and stable artifact bytes. +These are separate runs; compare ratios within the same worker-count group. +The memory sampler can miss brief child-process peaks, so treat its ratios as +directional. This refreshes the earlier native Linux checkpoint above; no +company-project performance is implied. + +The standalone compiler's filesystem location materially affects the Rust +baseline on this host. The workspace's Dune build directory is on `virtiofs`; +`/tmp` is on `overlay`. A copied release `bsc` had the same SHA-256 digest +(`cfefc4fe91cd78b7906fde1f546026d971f29775557654269d7a34415d15dfa9`) +as the Dune-path executable. Thirty interleaved `bsc -version` launches +measured about 5 ms median from `/tmp` versus 15 ms from the Dune path. This +is consistent with executable loading from the different mounts; it does not +show a compiler-code difference. + +With that byte-identical compiler on `/tmp`, a separate five-run eight-worker +comparison measured 1,483 ms Rust versus 1,358 ms OCaml clean wall time, with +sampled peak tree RSS of 336,308 versus 353,276 KiB. Unchanged medians were +84 versus 59 ms and single-edit medians 83 versus 59 ms. Work counts, complete +file sets, and stable artifact bytes matched, and the clean time and memory +gate passed. The clean OCaml gain was only 1.09x under this placement, versus +2.47x with the compiler on the workspace mount. Compare each pair only within +its run. The gate now prints executable and compiler hashes, runtime path, and +worker settings, so a benchmark can be reproduced with its actual compiler +storage layout. Neither layout predicts the closed-source company project. + ## AST I/O checkpoint Temporary counters on the same host and eight-domain fixture measured 917 diff --git a/rewatch-ocaml/bench/performance_gate.sh b/rewatch-ocaml/bench/performance_gate.sh index 9cca5c9e89..7110100bd5 100755 --- a/rewatch-ocaml/bench/performance_gate.sh +++ b/rewatch-ocaml/bench/performance_gate.sh @@ -166,6 +166,11 @@ echo "Rewatch clean-build performance gate" echo "commit: $(git -C "$repo_root" rev-parse HEAD)" echo "host: $(uname -a)" echo "cpus: $(getconf _NPROCESSORS_ONLN 2>/dev/null || echo unknown)" +echo "runtime: $RESCRIPT_RUNTIME" +echo "compiler domains: ${REWATCH_COMPILER_DOMAINS:-default}" +echo "Rust Rayon threads: ${RAYON_NUM_THREADS:-default}" +echo "executable and compiler SHA-256:" +sha256sum "$rust_executable" "$ocaml_executable" "$RESCRIPT_BSC_EXE" echo "runs: $runs (interleaved after one warm-up each)" echo "clean threshold: ${threshold_percent}% of Rust median wall and RSS" if [[ ${REWATCH_FIRST_EMBEDDED:-0} == 1 ]]; then From 9f5079a104654d5d8cb9a897e577b7e046f204ca Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 21:10:34 +0000 Subject: [PATCH 14/19] Restore source baseline between edit benchmark runs Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 8 ++++++-- rewatch-ocaml/bench/performance_gate.sh | 12 +++++++++++- 2 files changed, 17 insertions(+), 3 deletions(-) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index efc1c80947..0121e0a8eb 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -416,8 +416,12 @@ The harness isolates dependency trees, interleaves timed clean, unchanged, and single-edit builds, and samples process-tree RSS and task counts. It then traces Rust `bsc` requests and OCaml's logical compiler-request log, compares work in all three scenarios, and compares complete file sets and stable -artifact bytes at the same absolute path. The timed source edits add unique -comments to `packages/watch-warnings/src/B.res` in both isolated fixtures. +artifact bytes at the same absolute path. Each timed source edit adds a +comment to `packages/watch-warnings/src/B.res` in an isolated fixture. The +harness then restores the original source and completes an untimed build, so +every edit sample starts from the same compiled baseline. +The incremental figures recorded above used the earlier cumulative-comment +procedure; the clean-build figures are unaffected by this harness change. The unchanged workload replays the fixture's local `ModuleA` warning, so four compiler requests there are expected. `KEEP_REWATCH_BENCHMARK_WORKDIR=1` retains raw outputs and `results.csv`. The 125% wall-time and memory limits diff --git a/rewatch-ocaml/bench/performance_gate.sh b/rewatch-ocaml/bench/performance_gate.sh index 7110100bd5..13c42017c5 100755 --- a/rewatch-ocaml/bench/performance_gate.sh +++ b/rewatch-ocaml/bench/performance_gate.sh @@ -86,6 +86,10 @@ export RESCRIPT_BSC_EXE RESCRIPT_RUNTIME results="$work_root/results.csv" echo "scenario,implementation,iteration,wall_ms,peak_tree_rss_kib,peak_tree_tasks" \ >"$results" +edit_source_relative=packages/watch-warnings/src/B.res +edit_baseline="$work_root/B.res.baseline" +cp "$rust_fixture/$edit_source_relative" "$edit_baseline" +cmp "$edit_baseline" "$ocaml_fixture/$edit_source_relative" tree_resources() { local root_pid=$1 @@ -119,7 +123,7 @@ measure() { unchanged) ;; edit) printf '\n// timed single edit %d\n' "$iteration" \ - >>"$fixture/packages/watch-warnings/src/B.res" ;; + >>"$fixture/$edit_source_relative" ;; *) echo "Unknown benchmark scenario: $scenario" >&2; exit 2 ;; esac local start_ns root_pid sampler_pid peak_rss peak_tasks end_ns wall_ms @@ -153,6 +157,12 @@ measure() { printf '%-9s %-5s run %d: %6d ms %8d KiB %4d tasks\n' \ "$scenario" "$implementation" "$iteration" "$wall_ms" "$peak_rss" \ "$peak_tasks" + if [[ $scenario == edit ]]; then + # Return to the same compiled baseline before the next timed edit. + cp "$edit_baseline" "$fixture/$edit_source_relative" + "$executable" build "$fixture" >"$output.restore" \ + 2>"$output.restore.stderr" + fi } median_column() { From d54439798f8830e892ee0e7f4276a75e326de426 Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 21:17:55 +0000 Subject: [PATCH 15/19] Measure retained watches without compiler proxy overhead Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 88 +++++++++++++++++-- rewatch-ocaml/bench/watch_performance_gate.sh | 74 ++++++++++++++-- 2 files changed, 151 insertions(+), 11 deletions(-) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index 0121e0a8eb..bdd60aa81a 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -75,6 +75,64 @@ In a separate seven-edit retained-watch comparison, median latency fell from and produced identical edited JavaScript. The older watcher's RSS rose from 10,960 to 11,736 KiB and the current watcher's from 26,800 to 27,932 KiB; both held stable file descriptor and task counts. +That older watch gate timed the external-compiler side through its counting +proxy, which added process launches to its latency samples. + +A filesystem-controlled follow-up at revision +`9fa158aee436b0804ae7f6d0bb5d72144e038053` put both release Rewatch +executables and the byte-identical release `bsc` on `/tmp` (`overlay`) instead +of launching `bsc` from the workspace's `virtiofs` mount. The older executable +was rebuilt from `95467b2bfab9d6edd8fd00a8897a5308794a211d`; both used +the same current `bsc` and runtime. Five interleaved runs after one warm-up +each measured: + +| scenario | before median wall | after median wall | before peak tree RSS | after peak tree RSS | +| --- | ---: | ---: | ---: | ---: | +| Clean | 1,717 ms | 1,374 ms | 365,952 KiB | 361,164 KiB | +| Unchanged | 145 ms | 45 ms | 29,480 KiB | 25,848 KiB | +| One source edit | 148 ms | 47 ms | 52,456 KiB | 25,220 KiB | + +The clean gain was 1.25x under this compiler placement, compared with 2.85x +when `bsc` launched from the slower workspace mount. This supports compiler +launch location as a large part of the earlier measured gain on this host. +The unchanged and single-edit medians still fell by about 3x. Each timed edit +started from the same restored source and compiled baseline. The gate passed +its clean time and memory limits; both versions made the same 1,031 clean, +four unchanged, and six edit compiler requests, and produced identical +complete file sets and stable artifact bytes. The first executable is the +older OCaml Rewatch, despite the gate's `Rust` label. + +The revised retained-watch gate times both executables with the real `bsc` +and replays the external side through a counting proxy only after timing. In +a separate seven-edit run with the same `/tmp` compiler, the older OCaml +watcher measured 143 ms median versus 83 ms for current OCaml. Both made seven +parse and seven compile requests and generated identical edited JavaScript; +file descriptors, task counts, and retained RSS stayed within the gate's +limits. Each edit changed the generated JavaScript, proving that the timed +watchers rebuilt the edited module. These medians replace the proxy-influenced +161 versus 83 ms comparison above for latency purposes. + +After creating the older worktree and building both release executables as +shown below, reproduce this placement with: + +```sh +mkdir -p /tmp/rewatch-before-after-fast +cp /tmp/rewatch-before-954/_build/default/rewatch-ocaml/rescript_ocaml.exe \ + /tmp/rewatch-before-after-fast/before +cp _build/default/rewatch-ocaml/rescript_ocaml.exe \ + /tmp/rewatch-before-after-fast/after +cp _build/default/compiler/bsc/rescript_compiler_main.exe \ + /tmp/rewatch-before-after-fast/bsc +export RESCRIPT_BSC_EXE=/tmp/rewatch-before-after-fast/bsc +export RESCRIPT_RUNTIME="$PWD/packages/@rescript/runtime" +REWATCH_COMPILER_DOMAINS=8 rewatch-ocaml/bench/performance_gate.sh \ + /tmp/rewatch-before-after-fast/before \ + /tmp/rewatch-before-after-fast/after 5 +REWATCH_WATCH_COMPILER_DOMAINS=8 \ + rewatch-ocaml/bench/watch_performance_gate.sh \ + /tmp/rewatch-before-after-fast/before \ + /tmp/rewatch-before-after-fast/after 7 +``` To reproduce the before/after gate after preparing the dependencies and release compiler/runtime as described below: @@ -375,6 +433,21 @@ its run. The gate now prints executable and compiler hashes, runtime path, and worker settings, so a benchmark can be reproduced with its actual compiler storage layout. Neither layout predicts the closed-source company project. +With the fast-placement Rust median as the reference, a 5x clean-build gain +would require about 297 ms total. The separate instrumented OCaml compile +span was 1,172 ms on this fixture, before accounting for the rest of the +build. That trace has overhead and is not a same-run lower bound, but it shows +why scheduling and startup changes alone are unlikely to reach the target; +compiler work would need a several-fold reduction as well. + +The earlier 142 versus 83 ms Rust/OCaml retained-watch comparison timed Rust +through the counting compiler proxy. With the revised gate and the same real +`bsc` on `/tmp` for both implementations, seven retained edits measured 92 +ms Rust versus 81 ms OCaml median. Each made seven parse and seven compile +requests; edited JavaScript changed and matched, and file descriptor, task, +and RSS growth stayed within the gate's limits. This small watch fixture measures +single-module edits only. + ## AST I/O checkpoint Temporary counters on the same host and eight-domain fixture measured 917 @@ -420,6 +493,7 @@ artifact bytes at the same absolute path. Each timed source edit adds a comment to `packages/watch-warnings/src/B.res` in an isolated fixture. The harness then restores the original source and completes an untimed build, so every edit sample starts from the same compiled baseline. + The incremental figures recorded above used the earlier cumulative-comment procedure; the clean-build figures are unaffected by this harness change. The unchanged workload replays the fixture's local `ModuleA` warning, so four @@ -494,11 +568,15 @@ rewatch-ocaml/bench/watch_performance_gate.sh \ Set `REWATCH_WATCH_COMPILER_DOMAINS` to measure a specific worker count; otherwise the compiler uses its CPU-based heuristic. -It warms both implementations, interleaves an odd number of timed edits, -requires byte-identical generated JavaScript and equal logical parser/compiler -work counts, and samples file descriptors, tasks, and RSS after every build. -Rust work is observed through the counting `bsc` proxy; embedded OCaml work is -observed at the shared logical compiler-request boundary. +It warms both implementations, interleaves an odd number of timed edits with +the real `bsc` path, requires byte-identical generated JavaScript and equal +logical parser/compiler work counts, and samples file descriptors, tasks, and +RSS after every build. External compiler work is counted in a separate, +untimed replay through the `bsc` proxy; embedded OCaml work is observed at the +shared logical compiler-request boundary during the timed run. Every edit +changes `B.res`'s generated JavaScript; the gate verifies each timed result +changed and compares both implementations and the replay. The proxy therefore +adds no launch overhead to the measured edits. This catches retained-state implementations that appear fast by skipping work, as well as resource growth that a one-event syscall trace cannot show. The default median-latency limit is 150% of Rust because individual watch events diff --git a/rewatch-ocaml/bench/watch_performance_gate.sh b/rewatch-ocaml/bench/watch_performance_gate.sh index 48a0ff8514..8234ef888d 100755 --- a/rewatch-ocaml/bench/watch_performance_gate.sh +++ b/rewatch-ocaml/bench/watch_performance_gate.sh @@ -16,7 +16,8 @@ if ((runs < 5 || runs % 2 == 0)); then echo "RUNS must be an odd number of at least five." >&2 exit 2 fi -for command in awk cat cmp cp date find grep mktemp node sed seq setsid sleep sort tail wc; do +for command in awk cat cmp cp date find git grep mktemp node sed seq setsid \ + sha256sum sleep sort tail wc; do command -v "$command" >/dev/null || { echo "Missing required command: $command" >&2 exit 2 @@ -38,6 +39,14 @@ real_bsc=$RESCRIPT_BSC_EXE runtime=$RESCRIPT_RUNTIME counting_bsc="$repo_root/_build/default/tests/rewatch_ounit_tests/rewatch_bsc_test_proxy.exe" +echo "Rewatch retained-watch performance gate" +echo "commit: $(git -C "$repo_root" rev-parse HEAD)" +echo "runtime: $runtime" +echo "compiler domains: ${REWATCH_WATCH_COMPILER_DOMAINS:-default}" +echo "executable and compiler SHA-256:" +sha256sum "$rust_executable" "$ocaml_executable" "$real_bsc" +echo "runs: $runs (interleaved after one warm-up edit each)" + work_root=$(mktemp -d "${TMPDIR:-/tmp}/rewatch-watch-performance.XXXXXX") declare -A pids=() terminate_group() { @@ -135,7 +144,7 @@ resource_value() { } start_watcher() { - local implementation=$1 executable=$2 fixture + local implementation=$1 executable=$2 count_external=${3:-0} fixture fixture="$work_root/$implementation" cp -R "$repo_root/rewatch-ocaml/tests/basic" "$fixture" : >"$work_root/$implementation.bsc" @@ -154,7 +163,7 @@ start_watcher() { "$executable" watch --after-build "node $marker_script" "$fixture" \ >"$work_root/$implementation.stdout" \ 2>"$work_root/$implementation.stderr" & - else + elif [[ $count_external == 1 ]]; then setsid env \ RESCRIPT_BSC_EXE="$counting_bsc" \ REWATCH_BSC_PROXY_MODE=counting \ @@ -165,6 +174,14 @@ start_watcher() { "$executable" watch --after-build "node $marker_script" "$fixture" \ >"$work_root/$implementation.stdout" \ 2>"$work_root/$implementation.stderr" & + else + setsid env \ + RESCRIPT_BSC_EXE="$real_bsc" \ + RESCRIPT_RUNTIME="$runtime" \ + REWATCH_WATCH_MARKER="$work_root/$implementation.marker" \ + "$executable" watch --after-build "node $marker_script" "$fixture" \ + >"$work_root/$implementation.stdout" \ + 2>"$work_root/$implementation.stderr" & fi pids[$implementation]=$! wait_for_lines "$work_root/$implementation.marker" 1 @@ -187,6 +204,8 @@ for implementation in rust ocaml; do wait_for_idle "${pids[$implementation]}" : >"$work_root/$implementation.bsc" : >"$work_root/$implementation.compiler" + cp "$work_root/$implementation/src/B.mjs" \ + "$work_root/$implementation.prior.mjs" done declare -A baseline_fd baseline_tasks baseline_rss max_fd max_tasks max_rss @@ -205,7 +224,8 @@ measure_edit() { local implementation=$1 round=$2 expected=$((round + 2)) local started finished latency pid value started=$(date +%s%3N) - printf 'let answer = A.value + 1\n// retained edit %d\n' "$round" \ + printf 'let answer = A.value + %d\n// retained edit %d\n' \ + "$((round + 1))" "$round" \ >"$work_root/$implementation/src/B.res" wait_for_lines "$work_root/$implementation.marker" "$expected" finished=$(tail -n 1 "$work_root/$implementation.marker") @@ -214,6 +234,13 @@ measure_edit() { wait_for_text_count "$work_root/$implementation.stdout" \ "Finished incremental compilation" "$((round + 1))" wait_for_idle "${pids[$implementation]}" + if cmp -s "$work_root/$implementation.prior.mjs" \ + "$work_root/$implementation/src/B.mjs"; then + echo "$implementation output did not change after retained edit $round." >&2 + exit 1 + fi + cp "$work_root/$implementation/src/B.mjs" \ + "$work_root/$implementation.prior.mjs" pid=${pids[$implementation]} for kind in fd tasks rss; do value=$(resource_value "$pid" "$kind") @@ -237,8 +264,43 @@ for round in $(seq 1 "$runs"); do echo "Generated output differs after retained edit $round." >&2 exit 1 fi + cp "$work_root/rust/src/B.mjs" "$work_root/rust-round-$round.mjs" done +if [[ ${REWATCH_FIRST_EMBEDDED:-0} != 1 ]]; then + # Count external compiler requests in an untimed replay. A proxy in the + # timed watcher would add a process launch to every parse and compile. + start_watcher rust_work "$rust_executable" 1 + printf 'let answer = A.value + 1\n// warm retained edit\n' \ + >"$work_root/rust_work/src/B.res" + wait_for_lines "$work_root/rust_work.marker" 2 + wait_for_text_count "$work_root/rust_work.stdout" \ + "Finished incremental compilation" 1 + wait_for_idle "${pids[rust_work]}" + : >"$work_root/rust_work.bsc" + cp "$work_root/rust_work/src/B.mjs" "$work_root/rust_work.prior.mjs" + for round in $(seq 1 "$runs"); do + printf 'let answer = A.value + %d\n// retained edit %d\n' \ + "$((round + 1))" "$round" \ + >"$work_root/rust_work/src/B.res" + wait_for_lines "$work_root/rust_work.marker" "$((round + 2))" + wait_for_text_count "$work_root/rust_work.stdout" \ + "Finished incremental compilation" "$((round + 1))" + wait_for_idle "${pids[rust_work]}" + if cmp -s "$work_root/rust_work.prior.mjs" \ + "$work_root/rust_work/src/B.mjs"; then + echo "Replay output did not change after retained edit $round." >&2 + exit 1 + fi + cp "$work_root/rust_work/src/B.mjs" "$work_root/rust_work.prior.mjs" + cmp "$work_root/rust_work/src/B.mjs" \ + "$work_root/rust-round-$round.mjs" + done + rm -f "$work_root/rust_work/lib/watch.lock" + wait "${pids[rust_work]}" + unset 'pids[rust_work]' +fi + median() { sort -n "$1" | sed -n "$((runs / 2 + 1))p" } @@ -251,8 +313,8 @@ if [[ ${REWATCH_FIRST_EMBEDDED:-0} == 1 ]]; then "$work_root/rust.compiler" || true) rust_total_count=$(wc -l <"$work_root/rust.compiler") else - rust_parse_count=$(grep -cF -- '-bs-ast' "$work_root/rust.bsc" || true) - rust_total_count=$(wc -l <"$work_root/rust.bsc") + rust_parse_count=$(grep -cF -- '-bs-ast' "$work_root/rust_work.bsc" || true) + rust_total_count=$(wc -l <"$work_root/rust_work.bsc") rust_compile_count=$((rust_total_count - rust_parse_count)) fi ocaml_parse_count=$(grep -c '^parse' "$work_root/ocaml.compiler" || true) From cfbb43f4c8e8b2a3520ae985d69d4797e45fc88b Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 21:18:49 +0000 Subject: [PATCH 16/19] Record restored-edit fast compiler comparison Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 9 +++++++++ 1 file changed, 9 insertions(+) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index bdd60aa81a..1115027a22 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -433,6 +433,15 @@ its run. The gate now prints executable and compiler hashes, runtime path, and worker settings, so a benchmark can be reproduced with its actual compiler storage layout. Neither layout predicts the closed-source company project. +After the gate began restoring the source between timed edit samples, a +five-run repeat at `9fa158aee436b0804ae7f6d0bb5d72144e038053` with the +same `/tmp` compiler measured 1,482 ms Rust versus 1,398 ms OCaml clean wall +time, and 339,716 versus 359,656 KiB sampled peak tree RSS. Unchanged +medians were 80 versus 59 ms; edit medians were 77 versus 57 ms. Equal +compiler work, complete file sets, and stable artifact bytes passed the gate. +The clean gain in this repeat was 1.06x. The older fast-placement results +above used cumulative comment edits; compare medians only within each run. + With the fast-placement Rust median as the reference, a 5x clean-build gain would require about 297 ms total. The separate instrumented OCaml compile span was 1,172 ms on this fixture, before accounting for the rest of the From 68e372d7361d28aff9e704719a43be40be4efc3c Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 21:24:54 +0000 Subject: [PATCH 17/19] Record parser environment allocation probe Signed-off-by: Codex --- rewatch-ocaml/bench/README.md | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index 1115027a22..4eb79babe7 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -274,6 +274,16 @@ it is an architectural change with correctness and memory risks. Faster JavaScript emission alone has little headroom on this fixture. The temporary compiler-core instrumentation was removed after the measurement. +An exploratory change deferred construction of `Env.initial_safe_string` for +`-bs-ast` requests while keeping it eager for type-checking requests. Fifteen +interleaved clean testrepo builds of matched release executables on `/tmp` +measured 1,374 ms before versus 1,372 ms after; unchanged and edit medians +were 46 ms in both versions. Sampled clean peak tree RSS was 364,392 versus +368,052 KiB. Compiler requests, complete file sets, and stable artifact bytes +matched. The extra lazy-state handling had no useful measured gain, so it was +reverted. This probe does not split the cost of opening the implicit modules +from the rest of initial-environment setup. + A second temporary trace split `Typemod.type_implementation_more` on the same fixture. Among 479 implementation requests, `type_structure` used 5,253 ms of summed worker time, inclusion and delayed checks 820 ms, CMI saving 267 ms, From f43bfddd294bd5b1c57a0a53d4e90c27836063bc Mon Sep 17 00:00:00 2001 From: Codex Date: Thu, 24 Sep 2026 21:30:07 +0000 Subject: [PATCH 18/19] Document in-process OCaml rewatch in changelog Signed-off-by: Codex --- CHANGELOG.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 311d5eee7f..35f5a71ba7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,6 +18,8 @@ #### :rocket: New Feature +- Use the OCaml rewatch build system with parallel in-process compiler workers as the default `rescript` executable; keep the Rust implementation available as `rescript-rust`. https://github.com/rescript-lang/rescript/pull/8653 + #### :bug: Bug fix - Make rewatch compile independent modules after an unrelated failure and recompile blocked dependents when a changed interface survives a failed implementation, including across full watcher rebuilds. https://github.com/rescript-lang/rescript/pull/8667 From 80eb95963a8d960e6f36a93784f3b65c77ecaf8a Mon Sep 17 00:00:00 2001 From: Florian Hammerschmidt Date: Fri, 25 Sep 2026 10:37:43 +0200 Subject: [PATCH 19/19] Clarify OCaml rewatch compiler timing breakdowns Signed-off-by: Florian Hammerschmidt --- rewatch-ocaml/bench/README.md | 79 +++++++++++++++++++++-------------- 1 file changed, 48 insertions(+), 31 deletions(-) diff --git a/rewatch-ocaml/bench/README.md b/rewatch-ocaml/bench/README.md index 4eb79babe7..e00d655782 100644 --- a/rewatch-ocaml/bench/README.md +++ b/rewatch-ocaml/bench/README.md @@ -252,27 +252,67 @@ The analyzer reports elapsed phase span, summed compiler time, average and peak active requests, idle time inside each phase, and the longest compile requests. Remove an old trace before a new run; the compiler appends rows. -A separate temporary compiler-core trace split 479 implementation requests -from one instrumented clean build. These are summed concurrent-worker times, -not elapsed build time: +A temporary compiler-core trace split 479 implementation requests from one +instrumented clean build. These are summed concurrent-worker times, not elapsed +build time: | compiler-core phase | summed worker time | | --- | ---: | | Initial environment setup | 1,837 ms | -| Type checking, including CMI/CMT work | 6,736 ms | +| Implementation typing and persistence (split below) | 6,736 ms | | Lambda translation | 30 ms | | Lambda compilation | 135 ms | | JavaScript emission | 46 ms | The remaining 33 interface requests and outer request setup are outside this -split. The instrumented build's 512 compile requests spanned 1,253 ms, so do -not compare that span directly with the uninstrumented benchmark median. -Environment setup and type checking account for nearly all measured +pipeline split. The instrumented build's 512 compile requests spanned 1,253 +ms, so do not compare that span directly with the uninstrumented benchmark +median. + +A second instrumented clean build split the implementation typing and +persistence phase for the same 479 implementation requests: + +| work inside `Typemod.type_implementation_more` | summed worker time | +| --- | ---: | +| Structure typing, including imported CMI and signature work (`type_structure`) | 5,253 ms | +| Interface inclusion and delayed checks | 820 ms | +| CMI saving | 267 ms | +| CMT saving | 443 ms | +| Signature simplification and reset | under 3 ms | + +This second run totals about 6,783 ms, rather than the first run's 6,736 ms; +the rows are a breakdown of the same phase on a different run, not an exact +subtraction from 6,736 ms. `type_structure` includes import loading and +signature expansion, so its 5,253 ms is not a measurement of source typing +alone. Eliminating CMI/CMT saves alone has a limited bound. + +A third instrumented clean build measured imported CMI work across **all 512 +compile requests**, including the 33 interface requests excluded from the +tables above: + +| imported CMI operation | summed worker time | +| --- | ---: | +| 2,952 successful persistent-module lookups, including path search and failed candidate opens | 2,711 ms | +| 2,992 CMI decodes, including 40 through other CMI call sites | 1,511 ms | + +CMI decoding overlaps the lookup row, and CMI loading overlaps structure +typing and possibly other phases. These rows came from a separate run and +cover more requests, so 2,711 / 6,736 (about 40%) is only a numerical ratio, +not the measured lookup share of the typing phase. The remaining source typing +time was not isolated. A disjoint split requires nested timers in the same +run. The CMI trace decoded 194 MB of repeated input; its measurements include +tracing overhead and are not wall-time savings. Even eliminating all 2,711 ms +of lookup work would have an ideal eight-worker bound of about 0.34 s, short +of closing the 5x gap. A raw-byte cache would still pay most decoding cost, +while reusing decoded type graphs across fresh compiler requests would need +safe copying and CMI invalidation to preserve dependency correctness. + +Environment setup and implementation typing account for nearly all measured implementation work. Reusing a prepared environment across requests would have to preserve the compiler's fresh per-request type and identifier state; it is an architectural change with correctness and memory risks. Faster JavaScript emission alone has little headroom on this fixture. The temporary -compiler-core instrumentation was removed after the measurement. +instrumentation was removed after each measurement. An exploratory change deferred construction of `Env.initial_safe_string` for `-bs-ast` requests while keeping it eager for type-checking requests. Fifteen @@ -284,29 +324,6 @@ matched. The extra lazy-state handling had no useful measured gain, so it was reverted. This probe does not split the cost of opening the implicit modules from the rest of initial-environment setup. -A second temporary trace split `Typemod.type_implementation_more` on the same -fixture. Among 479 implementation requests, `type_structure` used 5,253 ms -of summed worker time, inclusion and delayed checks 820 ms, CMI saving 267 ms, -and CMT saving 443 ms. Signature simplification and reset took under 3 ms -combined. This is a separate instrumented run, so its totals differ slightly -from the compiler-core split above. The type-structure work is the main part -of the remaining compiler cost; eliminating artifact writes alone has a -limited bound. The temporary type-checker instrumentation was removed and -the normal release binary rebuilt afterward. - -A third temporary trace timed CMI loading in `Bs_cmi_load` on one eight-worker -clean build. The 512 compile requests made 2,952 successful persistent-module -lookups, taking 2,711 ms summed across workers, including file selection and -failed candidate opens. They decoded 2,992 CMIs totaling 194 MB of repeated -input, which took 1,511 ms summed across workers; the extra 40 reads came -through other CMI call sites. These measurements include tracing overhead and -are not wall-time savings. Even eliminating all 2,711 ms of lookup work would -have an ideal eight-worker bound of about 0.34 s, short of closing the 5x gap. -A raw-byte cache would still pay most decoding cost, while reusing decoded -type graphs across fresh compiler requests would need safe copying and CMI -invalidation to preserve dependency correctness. The temporary trace code -was removed and both release compiler executables rebuilt afterward. - A further temporary single-worker trace separated module lookup from opening the resolved signature. Across 1,897 opens, lookup used about 360 ms of process CPU time and signature opening about 2,063 ms. The 137 opens of