feat(check): build inline simulation-run payloads and add a dry-run CLI - #61
Open
scott-lowe-vapi wants to merge 1 commit into
Open
scott-lowe-vapi wants to merge 1 commit into
scott-lowe-vapi wants to merge 1 commit into
Conversation
`npm run check -- <check>|--all --dry-run` reads vapi-checks.yml and, for each check target, builds the inline POST /eval/simulation/run body from the files on disk: the assistant or squad with its tools, handoffs and structured outputs, plus the scenarios, judges and personalities the check's suites and simulations name. Nothing is looked up on the platform and no key is needed; --print-payload writes the JSON. The builder follows push and the runtime so the check tests what push deploys: - tools in runtime order (model.tools, then toolIds, then toolRefs); knowledgeBase tools stay in toolIds by run-org UUID; - `##` comments stripped; UUID references resolved through state; - squad members inlined, handoffs to members by name; - hook toolIds, artifactPlan structuredOutputIds and linked structured outputs inlined; judges' structuredOutputId inlined; - credentials bound by name to the run org (promotion's bindings). Anything it can't place fails the build with the field named: missing files, handoffs leaving the target, legacy assistantDestinations by ID, tools by ID inside overrides (strict mocks), duplicate tool names, audio judges/hooks/no required judge over chat, leftover non-UUID references, and payloads over 4.5 MB. The tool mock policy and live runs follow in later changes; without --dry-run the command exits 2. Refs TEST-141 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 1, 2026
Contributor
Author
|
Warning This pull request is not mergeable via GitHub because a downstack PR is open. Once all requirements are satisfied, merge this PR as a stack on Graphite.
This stack of pull requests is managed by Graphite. Learn more about stacking. |
This was referenced Oct 1, 2026
scott-lowe-vapi
marked this pull request as ready for review
October 1, 2026 23:52
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Value
V.A.L.U.E. tier: project — PR 5 of 10 for inline simulation PR checks (TEST-141); this PR adds an offline dry run, and nothing is sent.
POST /eval/simulation/runbody: the target with every tool, handoff and structured output, plus every scenario, judge and personality. It has to assemble that body the way push and the runtime would, or the check tests something other than what ships.npm run check -- core --dry-run --print-payload) before any minutes are spent. Every later PR (mock policy, live runs, workflow, promotion gate) builds on this payload.src/check-payload.ts(withcheck-payload-assistant.tsandcheck-payload-refs.ts) builds the body fromorgResourcesReadand the org states. It's pure, makes no network calls, and collects every problem so one dry run reports them all.model.tools, thentoolIds, thentoolRefs. This is whatcallAssistantsGetdoes, and the parity run's likelyendCalldifference came from getting it wrong;##comments stripped, and UUID references resolved through state;toolRefspin wins over a duplicatetoolIdsentry, with a warning that the version pin is ignored;knowledgeBasetools are kept by run-org UUID, since the API refuses them inline.do[].toolIdbecomes an inline tool;artifactPlan.structuredOutputIds, plus structured outputs that link the assistant through their ownassistant_ids, are inlined;structuredOutputIds are inlined;bind/omit).assistantDestinationsby ID;handoff_to_…but a handoff is auto-named, because generated names differ inline and stored.src/check-cmd.ts:npm run check -- <check>|--all --dry-run [--print-payload [dir]].VAPI_GITOPS_ROOT.--dry-runit exits 2; live runs land in PR 7.package.jsongets acheckscript, and the README and AGENTS.md command tables getnpm run checkrows.tests/fixtures/check-parity/: the TEST-141 parity squad written as gitops files.Evidence of value
The builder reproduces the payload that scored 15/15 in the parity run.
.mdassistants withtoolIds, a handoff tool byassistantId, judges bystructuredOutputId).npm run check -- core --dry-run --print-payloadwas run on it, and the output was diffed against the inline body the experiment sent, rebuilt fromparity.mjs.members[*].assistant.model.toolsorder[lookup_patient, handoff, endCall],[check_availability, endCall][endCall, lookup_patient, handoff],[endCall, check_availability]model.toolsthentoolIds). Same tools, byte-for-byte, order asidesquad.name,personality.nameinline,callerBright Smile Dental,Dental calleriterations,transportvapi.webchatvapi-checks.ymldefaultsassistantId: schedulercame out asassistantName: "Scheduler", exactly what the experiment sent.Tests:
npm testgoes from 404 to 430 passing (26 new), andnpm run buildis clean.Testing plan
tests/check-payload.test.ts(21 tests, temp-dir fixtures):.mdprompt, inline judges, entries, transport;.mdbody as the only system message;toolIds, andtoolIdsby UUID with server fields stripped;toolRefspin and warning;knowledgeBasekept by UUID, and failing with no UUID;artifactPlanand fromassistant_ids(by slug and by UUID);parametersignored;tests/check-cmd.test.ts(5 tests):--print-payload;--allwith one broken check exiting 2;Stacked on #60.
Refs TEST-141
🤖 Generated with Claude Code