Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -902,7 +902,7 @@ npm run promote -- --pipeline <name> --from <source> --to <target> --apply # For
npm run validate -- <org> # Lint resources locally (fails fast on schema drift)
npm run audit -- <org> # Read-only drift detector: orphan YAML, state ghosts, content-identical clusters, sibling base-slugs, dashboard orphans, inline model.tools. Exit 1 on findings.
npm run audit -- <org> --type assistants # Scope audit to a single resource type
npm run sim -- <org> --suite <name> --target <name> # Run a simulation suite against an assistant/squad
npm run sim -- <org> --suite <name> --target <name> # Run a simulation suite against an assistant/squad (exit 0 pass, 1 fail, 3 incomplete; --timeout <min>)
npm run rollback -- <org> --to <ISO-timestamp> # Re-apply a snapshot taken before a push
npm run rollback -- <org> --list # List available snapshots

Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -103,7 +103,7 @@ Every command works in two modes:
| `npm run cleanup` | ✅ | `npm run cleanup -- <org> [--force --confirm <org>]` | Inspect (default) or delete orphaned remote resources. Destructive run requires `--confirm <org>`. |
| `npm run rollback` | — | `npm run rollback -- <org> --list` or `--to <ISO>` | Restore from a snapshot in `.vapi-state.<org>.snapshots/` (one is written before every push/apply). |
| `npm run call` | ✅ | `npm run call -- <org> -a <name>` or `-s <squad>` | Start an interactive WebSocket call against an assistant or squad. |
| `npm run sim` | — | `npm run sim -- <org> --suite <name> --target <name>` | Run a simulation suite (or specific simulations) against a deployed assistant/squad. |
| `npm run sim` | — | `npm run sim -- <org> --suite <name> --target <name> [--timeout <min>]` | Run a simulation suite (or specific simulations) against a deployed assistant/squad. Prints the run link; exits 0 passed, 1 failed, 3 incomplete (timeout, Ctrl-C, missing results). |
| `npm run migrate` | — | `npm run migrate` | One-time, all orgs at once: slim legacy state files to pure `name → uuid` and seed the per-developer `.vapi-state-hash/` baseline store from the old hashes. Required once after upgrading to the hash-store engine — `pull`/`push`/`apply` refuse legacy-shaped state until it runs. Idempotent. |
| `npm run build` | — | — | Type-check the codebase (`tsc --noEmit`). |
| `npm test` | — | — | Run regression tests (`node:test`). |
Expand Down
21 changes: 20 additions & 1 deletion docs/learnings/simulations.md
Original file line number Diff line number Diff line change
Expand Up @@ -471,6 +471,23 @@ This is the unified executor. A "run" is a batch — it expands into many `runIt
}
```

The **create** response (`CreateSimulationRunResponse`) also carries two fields nothing else returns:

- `url` — the dashboard link to this run. `GET /eval/simulation/run/:id` does **not** return it, so read it from the create response.
- `simulationRunItemIds` — one id per item queued (simulations × iterations), i.e. how many items the run should end with.

A run has **no `results` field**. Pass/fail lives in `itemCounts` and on the items themselves (`GET …/item`, below). Creating a run queues paid work before the response returns, so a 5xx on create may still have started it — don't blindly retry the POST.

### How `npm run sim` decides pass/fail

`src/sim-result.ts` (`simRunVerdict`) reports **passed** only when all of these hold; otherwise **failed** (any failed item) or **incomplete**:

- the run `ended` and has `itemCounts`, with `total` equal to the number of items created (`simulationRunItemIds`) and greater than 0;
- nothing is `queued`, `running` or `canceled`, and `passed === total`;
- every item was fetched, is `passed`, and has at least one **required** evaluation that wasn't skipped (the all-skipped trap from the chat-mode gotcha above).

Exit codes: 0 passed, 1 failed, 2 usage error, 3 incomplete (timeout, Ctrl-C, canceled items, results that never arrived). On timeout or Ctrl-C the run is canceled.

### List runs — `GET /eval/simulation/run`

Query params:
Expand All @@ -486,7 +503,7 @@ Query params:
### Get / Cancel run

- `GET /eval/simulation/run/:id` → `SimulationRun`
- `PATCH /eval/simulation/run/:id` → cancels the run **and** all its queued items. No body required.
- `PATCH /eval/simulation/run/:id` → cancels the run **and** all its queued items. No body required. Returns 400 `Run has already ended` for an ended run and 409 when a concurrent cancel wins; both mean "nothing left to cancel".

---

Expand All @@ -498,6 +515,8 @@ Run items are system-managed — there's no create/update API for users; they're

Query params: `limit`, `page`, `simulationId`, `runId`, `status` (`queued` | `running` | `evaluating` | `passed` | `failed` | `canceled`).

Response shape depends on the query: **with** `limit` or `page` it's paginated (`{ results, metadata: { totalItems, itemsPerPage, currentPage } }`, `limit` up to 1000); **without** them it's a bare array. Pages are ordered only by creation time and a run's items share it, so OFFSET pages can overlap — dedupe by `id`. Items can also lag the run: a run can be `ended` while an item is `passed` but its `results` aren't written yet.

### Get a run item — `GET /eval/simulation/run/:id/item/:itemId`

**Response** — `SimulationRunItem` is rich; the highlights:
Expand Down
52 changes: 52 additions & 0 deletions improvements.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,7 @@ you which stack PR closes the row.**
| 30 | Tool-linking pass could PATCH a raw assistant slug | Mid-push 400 naming the wrong resource | None | RESOLVED 2026-08-03 (#51) |
| 31 | Unresolved references handled 3 inconsistent ways, no dangling-ref check | Same authoring mistake, three different failure modes | None | Open |
| 32 | Test suite never ran in CI; 20 tests rotted after the hash store | Regression guards for #22/#23 silently stopped running | None | RESOLVED 2026-09-30 (#56) |
| 33 | `npm run sim` reported every run as passed | A failing suite exited 0 — false green | None | RESOLVED 2026-10-01 |

**Active backlog after cleanup:** `#2`, `#6`, `#8`, `#12`, `#20`, `#24–#26`, `#31`, and the open remainder of `#27` (wiring the listing-completeness verdict into push/delete/audit, and moving `cleanup.ts` onto the shared pager). Resolved entries stay in this file as historical incident notes per the maintenance directive; stale superseded backlog rows are not duplicated.

Expand Down Expand Up @@ -1721,6 +1722,57 @@ fields.

---

## 33. `npm run sim` reported every run as passed

**[RESOLVED 2026-10-01]**

**Discovered:** 2026-10-01, while designing simulation PR checks (TEST-141).

### Problem

`npm run sim` scored a run by reading `results[]` on the run and counting
`status === "pass"`. The simulation-run API has no `results` field and
item statuses are `passed` / `failed`, so every watched run summarised
as 0 pass / 0 fail and the command exited 0 — including runs whose
simulations failed.

### Current behavior (Verified, before the fix)

- `src/sim.ts` `runSimulation` computed pass/fail from `last.results`,
which `GET /eval/simulation/run/:id` never returns; pass/fail is in
`itemCounts` and on the run items (`GET /eval/simulation/run/:id/item`).
- `src/sim-cmd.ts` exited 1 only when `fail > 0`, which could never happen.
- Polling also stopped on statuses that don't exist (`failed`,
`completed`), never canceled a timed-out run, and didn't print the run
link (only the create response carries `url`).

### Risk

Anyone gating on `npm run sim` (locally or in CI) got a green result for a
failing suite.

### Current mitigation

None needed once the fix below lands.

### Possible fix (landed)

- `src/sim-result.ts` `simRunVerdict`: passed only when the run ended,
every expected item exists, passed, and had a required evaluation that was
actually scored; otherwise failed or incomplete with a reason.
- `src/sim.ts` reads items (paginated or bare-array, deduped by id),
waits for late item results, prints the run link and failing judges, and
cancels the run on `--timeout` (default 20 min) or Ctrl-C.
- `src/vapi-client.ts`: a config-free client that never retries run
creation on a 5xx (the run may already be queued).
- Exit codes: 0 passed, 1 failed, 2 usage, 3 incomplete.

### Status

**RESOLVED 2026-10-01.**

---

## Out of scope (intentionally not improvements)

- **State file is identity-only and not git-ignored.** It's intentionally
Expand Down
200 changes: 125 additions & 75 deletions src/sim-cmd.ts
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,12 @@
//
// Wraps `POST /eval/simulation/run` (the simulations API). See AGENTS.md for
// usage.
//
// Exit codes: 0 passed, 1 failed, 2 usage/config error, 3 incomplete
// (timed out, interrupted, canceled, or results that never fully arrived).

import { resolve } from "node:path";
import { fileURLToPath } from "node:url";
import {
formatSummary,
loadEnvFile,
Expand All @@ -12,27 +17,28 @@ import {
runSimulation,
} from "./sim.ts";

function printUsage(): void {
console.error(
[
"Usage:",
" npm run sim -- <org> --suite <suite-name> --target <assistant-or-squad-name>",
" npm run sim -- <org> --simulations <name1>,<name2> --target <assistant-name>",
"",
"Options:",
" --suite <name> Run an entire simulation suite by local resource name",
" --simulations <list> Run one or more simulations by comma-separated local names",
" --target <name> Local assistant or squad name (resolves to UUID via state)",
" --transport voice|chat Transport (default: voice; chat is faster/cheaper)",
" --iterations N Override default iteration count",
" --watch Tail status until completion (default: on)",
"",
"Examples:",
" npm run sim -- my-org --suite booking-tests --target intake-agent",
" npm run sim -- my-org --simulations happy-path,edge-case --target main-agent --transport chat",
].join("\n"),
);
}
const USAGE = [
"Usage:",
" npm run sim -- <org> --suite <suite-name> --target <assistant-or-squad-name>",
" npm run sim -- <org> --simulations <name1>,<name2> --target <assistant-name>",
"",
"Options:",
" --suite <name> Run an entire simulation suite by local resource name",
" --simulations <list> Run one or more simulations by comma-separated local names",
" --target <name> Local assistant or squad name (resolves to UUID via state)",
" --transport voice|chat Transport (default: voice; chat is faster/cheaper)",
" --iterations N Override default iteration count",
" --timeout <minutes> Give up and cancel the run after this long (default: 20)",
" --no-watch Start the run, print its link, and exit 0 without a verdict",
"",
"Exit codes: 0 passed, 1 failed, 2 usage error, 3 incomplete (timeout, interrupt, missing results)",
"",
"Examples:",
" npm run sim -- my-org --suite booking-tests --target intake-agent",
" npm run sim -- my-org --simulations happy-path,edge-case --target main-agent --transport chat",
].join("\n");

class UsageError extends Error {}

interface ParsedArgs {
env: string;
Expand All @@ -42,71 +48,83 @@ interface ParsedArgs {
squad?: string;
transport?: "voice" | "chat";
iterations?: number;
timeoutMinutes?: number;
watch: boolean;
help: boolean;
}

function parseArgs(): ParsedArgs {
const args = process.argv.slice(2);
function argsParse(args: string[]): ParsedArgs {
const env = args[0];
if (!env) {
printUsage();
process.exit(1);
if (!env || env === "--help" || env === "-h") {
return { env: env ?? "", watch: true, help: true };
}
const SLUG_RE = /^[a-z0-9]([a-z0-9-]*[a-z0-9])?$/;
if (!SLUG_RE.test(env)) {
console.error(`❌ Invalid org name: ${env}`);
process.exit(1);
}

const parsed: ParsedArgs = { env, watch: true };
if (!SLUG_RE.test(env)) throw new UsageError(`Invalid org name: ${env}`);

const parsed: ParsedArgs = { env, watch: true, help: false };
for (let i = 1; i < args.length; i++) {
const arg = args[i];
if (arg === "--suite") parsed.suite = args[++i];
else if (arg === "--simulations") parsed.simulations = args[++i];
else if (arg === "--target") {
// We don't know yet whether target is an assistant or squad — defer
// resolution to the state lookup. Try assistant first; resolveTarget()
// accepts either argument key, so we set the candidate in `assistant`
// and let `resolveTarget` fall through to `squad` if not found.
// For clarity, we accept --assistant / --squad as explicit alternatives.
// resolution to the state lookup (see simCommandRun).
parsed.assistant = args[++i];
} else if (arg === "--assistant") parsed.assistant = args[++i];
else if (arg === "--squad") parsed.squad = args[++i];
else if (arg === "--transport") {
const v = args[++i];
if (v === "voice" || v === "chat") parsed.transport = v;
else {
console.error(`❌ --transport must be "voice" or "chat" (got "${v}")`);
process.exit(1);
if (v !== "voice" && v !== "chat") {
throw new UsageError(
`--transport must be "voice" or "chat" (got "${v}")`,
);
}
parsed.transport = v;
} else if (arg === "--iterations") {
parsed.iterations = parseInt(args[++i] ?? "", 10);
parsed.iterations = Number.parseInt(args[++i] ?? "", 10);
if (Number.isNaN(parsed.iterations)) {
console.error("❌ --iterations requires a number");
process.exit(1);
throw new UsageError("--iterations requires a number");
}
} else if (arg === "--timeout") {
parsed.timeoutMinutes = Number(args[++i]);
if (
!Number.isFinite(parsed.timeoutMinutes) ||
parsed.timeoutMinutes <= 0
) {
throw new UsageError("--timeout requires a positive number of minutes");
}
} else if (arg === "--no-watch") parsed.watch = false;
else if (arg === "--watch") parsed.watch = true;
else if (arg === "--help" || arg === "-h") {
printUsage();
process.exit(0);
}
else if (arg === "--help" || arg === "-h") parsed.help = true;
else throw new UsageError(`Unknown argument: ${arg}`);
}

return parsed;
}

async function main(): Promise<void> {
const args = parseArgs();
const cfg = loadEnvFile(args.env);
const state = loadStateFile(args.env);
export async function simCommandRun(
args = process.argv.slice(2),
): Promise<number> {
let parsed: ParsedArgs;
try {
parsed = argsParse(args);
} catch (error) {
if (!(error instanceof UsageError)) throw error;
console.error(`❌ ${error.message}\n\n${USAGE}`);
return 2;
}
if (parsed.help) {
console.log(USAGE);
return parsed.env ? 0 : 2;
}

const cfg = loadEnvFile(parsed.env);
const state = loadStateFile(parsed.env);

// Disambiguate --target: if the bare value matches a squad name in state
// and not an assistant, treat it as a squad. Explicit --assistant / --squad
// override the heuristic.
let assistant = args.assistant;
let squad = args.squad;
let assistant = parsed.assistant;
let squad = parsed.squad;
if (assistant && !squad) {
const isSquad =
typeof state.squads[assistant] !== "undefined" &&
Expand All @@ -120,39 +138,71 @@ async function main(): Promise<void> {
console.log(
"═══════════════════════════════════════════════════════════════",
);
console.log(`🧪 Vapi GitOps Sim Runner — Environment: ${args.env}`);
console.log(`🧪 Vapi GitOps Sim Runner — Environment: ${parsed.env}`);
console.log(` API: ${cfg.baseUrl}`);
console.log(
"═══════════════════════════════════════════════════════════════\n",
);

const selection = resolveSelection(state, {
suite: args.suite,
simulations: args.simulations,
suite: parsed.suite,
simulations: parsed.simulations,
});
const target = resolveTarget(state, { assistant, squad });

const summary = await runSimulation(cfg, selection, target, {
watch: args.watch,
iterations: args.iterations,
transport: args.transport,
});
// Ctrl-C / SIGTERM cancel the run instead of leaving it running unwatched.
const controller = new AbortController();
const stop = () => controller.abort();
process.on("SIGINT", stop);
process.on("SIGTERM", stop);
let summary;
try {
summary = await runSimulation(cfg, selection, target, {
watch: parsed.watch,
iterations: parsed.iterations,
transport: parsed.transport,
timeoutMs:
parsed.timeoutMinutes === undefined
? undefined
: parsed.timeoutMinutes * 60_000,
signal: controller.signal,
});
} finally {
process.off("SIGINT", stop);
process.off("SIGTERM", stop);
}

console.log(`\n${formatSummary(summary)}\n`);

if (summary.fail > 0) {
console.error(
`❌ Simulation run failed (${summary.fail} fail / ${summary.pass} pass)`,
);
process.exit(1);
if (!summary.verdict) {
console.log("▶️ Run started (not watched).");
return 0;
}
if (summary.verdict.status === "passed") {
console.log("✅ Simulation run passed.");
return 0;
}
console.log("✅ Simulation run passed.");
if (summary.verdict.status === "failed") {
console.error(`❌ Simulation run failed: ${summary.verdict.reason}`);
return 1;
}
console.error(`⚠️ Simulation run incomplete: ${summary.verdict.reason}`);
return 3;
}

main().catch((error) => {
console.error(
"\n❌ Sim failed:",
error instanceof Error ? error.message : error,
const isMainModule =
process.argv[1] !== undefined &&
resolve(process.argv[1]) === fileURLToPath(import.meta.url);
if (isMainModule) {
simCommandRun().then(
(code) => {
process.exitCode = code;
},
(error: unknown) => {
console.error(
"\n❌ Sim failed:",
error instanceof Error ? error.message : error,
);
process.exitCode = 2;
},
);
process.exit(1);
});
}
Loading
Loading