Skip to content

feat(check): build inline simulation-run payloads and add a dry-run CLI - #61

Open
scott-lowe-vapi wants to merge 1 commit into
feat/check-configfrom
feat/check-inline-payload
Open

scott-lowe-vapi wants to merge 1 commit into
feat/check-configfrom
feat/check-inline-payload

Conversation

@scott-lowe-vapi

Copy link
Copy Markdown
Contributor

Value

V.A.L.U.E. tier: project — PR 5 of 10 for inline simulation PR checks (TEST-141); this PR adds an offline dry run, and nothing is sent.

  • Problem: to test a PR's changes without deploying them, the check has to turn the branch's files into one inline POST /eval/simulation/run body: the target with every tool, handoff and structured output, plus every scenario, judge and personality. It has to assemble that body the way push and the runtime would, or the check tests something other than what ships.
  • Who it affects: gitops users, who can now see exactly what a check would send (npm run check -- core --dry-run --print-payload) before any minutes are spent. Every later PR (mock policy, live runs, workflow, promotion gate) builds on this payload.
  • What changes:
    • src/check-payload.ts (with check-payload-assistant.ts and check-payload-refs.ts) builds the body from orgResourcesRead and the org states. It's pure, makes no network calls, and collects every problem so one dry run reports them all.
      • Tools:
        • runtime tool order: model.tools, then toolIds, then toolRefs. This is what callAssistantsGet does, and the parity run's likely endCall difference came from getting it wrong;
        • ## comments stripped, and UUID references resolved through state;
        • a toolRefs pin wins over a duplicate toolIds entry, with a warning that the version pin is ignored;
        • knowledgeBase tools are kept by run-org UUID, since the API refuses them inline.
      • Squads: members are inlined (from a file, by UUID, or inline), and handoffs to members switch from ID to member name.
      • Assistants:
        • hook do[].toolId becomes an inline tool;
        • artifactPlan.structuredOutputIds, plus structured outputs that link the assistant through their own assistant_ids, are inlined;
        • linkage and server fields are stripped.
      • Simulations:
        • suites and simulations are deduplicated into entries with unique names of at most 80 characters;
        • judges' structuredOutputIds are inlined;
        • local personalities are inlined, and stock ones pass by ID.
      • Credentials are bound by name to the run org's UUIDs using promotion's binding policy (bind / omit).
      • Fails the build, naming the field:
        • a missing file;
        • a handoff leaving the target, or an assistant target that hands off;
        • legacy assistantDestinations by ID;
        • members without a name, or with the same name;
        • two inlined tools with the same type and name, which stored would keep but inline would silently drop;
        • tools referenced by ID inside overrides (under strict mocks);
        • over chat: audio judges, scenario hooks, or no required text judge;
        • any reference still a name after the build (the backstop for docs: document orphan-YAML gate + --allow-new-files in README and AGENTS #31's silent drops);
        • a body over 4.5 MB.
      • Warns when text mentions handoff_to_… but a handoff is auto-named, because generated names differ inline and stored.
    • src/check-cmd.ts: npm run check -- <check>|--all --dry-run [--print-payload [dir]].
      • Offline: no API key needed, and it honours VAPI_GITOPS_ROOT.
      • Exits 0 when every payload builds, and 2 on a usage, config or build error.
      • Without --dry-run it exits 2; live runs land in PR 7.
    • package.json gets a check script, and the README and AGENTS.md command tables get npm run check rows.
    • New fixture tests/fixtures/check-parity/: the TEST-141 parity squad written as gitops files.
  • Deliberately not here: the fail-closed mock policy (dead servers, default error mocks, the tool-type allowlist), which is the next PR. Until then the payload carries tools as written, and there's no live path.

Evidence of value

The builder reproduces the payload that scored 15/15 in the parity run.

  • How: the dental squad from the 2026-10-01 parity experiment was written as gitops files (.md assistants with toolIds, a handoff tool by assistantId, judges by structuredOutputId). npm run check -- core --dry-run --print-payload was run on it, and the output was diffed against the inline body the experiment sent, rebuilt from parity.mjs.
  • Result: the only differences are these.
Path Experiment sent Built from files Why
members[*].assistant.model.tools order [lookup_patient, handoff, endCall], [check_availability, endCall] [endCall, lookup_patient, handoff], [endCall, check_availability] Intended: the runtime order of the stored arm (model.tools then toolIds). Same tools, byte-for-byte, order aside
squad.name, personality.name inline, caller Bright Smile Dental, Dental caller Fixture names
iterations, transport 5, (API default) 1, vapi.webchat vapi-checks.yml defaults
  • Every prompt, judge, tool mock, tool definition and handoff destination is identical. The handoff by assistantId: scheduler came out as assistantName: "Scheduler", exactly what the experiment sent.

Tests: npm test goes from 404 to 430 passing (26 new), and npm run build is clean.

Testing plan

  • tests/check-payload.test.ts (21 tests, temp-dir fixtures):
    • the parity fixture end to end: tool order, member-name handoffs, .md prompt, inline judges, entries, transport;
    • .md body as the only system message;
    • missing toolIds, and toolIds by UUID with server fields stripped;
    • the toolRefs pin and warning;
    • knowledgeBase kept by UUID, and failing with no UUID;
    • duplicate tool names;
    • hook tools;
    • structured outputs from artifactPlan and from assistant_ids (by slug and by UUID);
    • an assistant target that hands off;
    • squad members by file, UUID and inline;
    • every squad problem reported at once;
    • strict vs off overrides;
    • credentials bound, omitted and missing;
    • a leftover reference, with free-form parameters ignored;
    • the chat rules, and voice allowing them;
    • stock and local personalities, name truncation and dedupe;
    • missing suites, simulations, scenarios, personalities and targets;
    • the auto-handoff warning;
    • the 4.5 MB limit.
  • tests/check-cmd.test.ts (5 tests):
    • a dry run with --print-payload;
    • --all with one broken check exiting 2;
    • no config exiting 0;
    • usage, selection, config and live-mode errors exiting 2;
    • a missing state file as a warning.
  • Not tested:
    • Any live run: nothing is sent until PR 7, so API acceptance of a built body is shown only for the parity fixture (the experiment's 201).
    • Real customer squads: override merge semantics and integration tool shapes are untested.
    • Phone-number bindings beyond the unit-tested promotion helper.
    • EU base URLs.

Stacked on #60.

Refs TEST-141

🤖 Generated with Claude Code

`npm run check -- <check>|--all --dry-run` reads vapi-checks.yml and,
for each check target, builds the inline POST /eval/simulation/run body
from the files on disk: the assistant or squad with its tools, handoffs
and structured outputs, plus the scenarios, judges and personalities the
check's suites and simulations name. Nothing is looked up on the
platform and no key is needed; --print-payload writes the JSON.

The builder follows push and the runtime so the check tests what push
deploys:
- tools in runtime order (model.tools, then toolIds, then toolRefs);
  knowledgeBase tools stay in toolIds by run-org UUID;
- `##` comments stripped; UUID references resolved through state;
- squad members inlined, handoffs to members by name;
- hook toolIds, artifactPlan structuredOutputIds and linked structured
  outputs inlined; judges' structuredOutputId inlined;
- credentials bound by name to the run org (promotion's bindings).

Anything it can't place fails the build with the field named: missing
files, handoffs leaving the target, legacy assistantDestinations by ID,
tools by ID inside overrides (strict mocks), duplicate tool names,
audio judges/hooks/no required judge over chat, leftover non-UUID
references, and payloads over 4.5 MB.

The tool mock policy and live runs follow in later changes; without
--dry-run the command exits 2.

Refs TEST-141

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants