Skip to content

fix(sim): report failed, canceled and incomplete runs instead of a false pass - #58

Open
scott-lowe-vapi wants to merge 1 commit into
ci/test-workflowfrom
fix/sim-false-green
Open

scott-lowe-vapi wants to merge 1 commit into
ci/test-workflowfrom
fix/sim-false-green

Conversation

@scott-lowe-vapi

Copy link
Copy Markdown
Contributor

Value

V.A.L.U.E. tier: project — PR 2 of 10 for inline simulation PR checks (TEST-141). Stacked on #57.

  • Problem: npm run sim reported every run as passed. It read a results field the simulation-run API doesn't return and counted status === "pass" (items are passed/failed), so it always summarised 0/0 and exited 0, failing runs included.
  • Who it affects: anyone gating on npm run sim, locally or in CI. The PR check (later in this stack) reuses this verdict, so it has to be right.
  • What changes:
    • A strict verdict (src/sim-result.ts).
    • Item fetching that handles both response shapes and late results.
    • The run link printed, plus each failing judge with expected vs extracted values.
    • --timeout and Ctrl-C both cancel the run.
    • Exit codes: 0 passed, 1 failed, 2 usage, 3 incomplete.
    • A config-free client that never retries run creation on a 5xx, because the run may already be queued.

Evidence of value

Same stub API, one passed and one failed item:

Result
main's sim.ts {"pass":0,"fail":0} → exits 0 (false green)
This branch failed — 1 of 1 simulations failed → exits 1

Live, against a test org (chat transport):

Suite Exit Output
Designed to fail (judge: "open 24 hours?") 1 ✗ … open-24h (expected = true, got false) — run
Designed to pass (judge: "open 8–5 on Fridays?") 0 passed — 1 of 1 simulations passed — run

The temporary resources were deleted afterwards.

Testing plan

  • tests/sim-result.test.ts is a verdict table covering:
    • the old false-green shape (no results);
    • 0 items, a short item list, and a count mismatch;
    • a failed item, with the failing judge listed;
    • canceled items;
    • all required evaluations skipped, and an optional skip alongside a scored required judge;
    • missing itemCounts, and a run that hasn't ended.
  • tests/sim-run.test.ts runs runSimulation against a local HTTP stub:
    • pass, and fail using the bare-array item shape;
    • late item results;
    • timeout cancels the run, and an interrupt cancels it with the 400 "already ended" swallowed;
    • no retry of a 502 on create;
    • --no-watch;
    • pagination with overlapping pages deduped.
  • npm run build and npm test pass (377 tests).
  • Not tested: a live run that's still running when it's canceled (cancel was only exercised against the stub), and voice transport (the live runs used chat). Default behaviour change: unknown CLI arguments are now an error (exit 2) instead of being silently ignored.

Refs TEST-141

🤖 Generated with Claude Code

…lse pass

`npm run sim` read a `results` field the simulation-run API doesn't
return and counted `status === "pass"` (items are `passed`/`failed`), so
every watched run summarised as 0/0 and exited 0, including failing
runs.

- src/sim-result.ts: strict, pure verdict. Passed only when the run
  ended, every expected item exists and passed, and each had a required
  evaluation that was actually scored; otherwise failed or incomplete
  with a reason.
- src/sim.ts: read run items (paginated or bare array, deduped by id),
  wait for late item results, print the run link from the create
  response and the failing judges, cancel the run on --timeout (default
  20 min) or Ctrl-C.
- src/vapi-client.ts: config-free client; run creation is never retried
  on a 5xx, since the run may already be queued.
- src/sim-cmd.ts: exported simCommandRun; exit 0 passed, 1 failed,
  2 usage, 3 incomplete; --no-watch prints the link and exits 0.
- Docs: run/item shapes and the verdict in docs/learnings/simulations.md,
  sim rows in README and AGENTS.md, improvements.md #33.

Refs TEST-141

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants