feat(check): run checks live and report Vapi Evals statuses - #63
Merged
Merged
Conversation
This was referenced Oct 1, 2026
Contributor
Author
This was referenced Oct 1, 2026
scott-lowe-vapi
marked this pull request as ready for review
October 1, 2026 23:52
vtkovapi
approved these changes
Oct 3, 2026
Contributor
Author
Merge activity
|
scott-lowe-vapi
changed the base branch from
feat/check-fail-closed-mocks
to
graphite-base/63
October 3, 2026 06:08
`npm run check -- <check>|--all` now runs each built target payload in the check's run org: one POST /eval/simulation/run per target, at most three at once, judged by the same strict verdict as `npm run sim` (the create/poll/cancel/hydrate loop moves into simRunExecute, which both share). - Budget and cancellation: each run's deadline is the earlier of its timeoutMinutes and --budget-minutes; no run starts with under five minutes left; the deadline, SIGINT and SIGTERM cancel in-flight runs. - Keys: VAPI_CHECK_TOKENS, then .env.<runOrg>, then VAPI_PRIVATE_API_KEY when every selected check runs in one org. --refresh-bindings runs the same read-only bindings pull promotion uses. - Selection: --changed-since <ref> runs only checks affected by changes since the merge base (the org's files and state, the run org's, the config, the engine, and the check's own paths). - Statuses (when GITHUB_TOKEN, GITHUB_REPOSITORY and HEAD_SHA are set): `Vapi Evals / <check> / <target>` goes pending with the run's link, then success/failure/error; --all runs post the aggregate `Vapi Evals` from a finally block, so config errors and exceptions report too. - Report: a markdown summary (also appended to $GITHUB_STEP_SUMMARY) with failing evaluations' expected vs extracted values and "unmocked tool called" notices from transcripts, plus --json. No PR comments. - Runs send User-Agent vapi-gitops-check/<version>. Exit codes: 0 passed, 1 failed, 2 config or build error, 3 incomplete. Refs TEST-141 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
scott-lowe-vapi
force-pushed
the
feat/check-live-run
branch
from
October 3, 2026 06:10
b055db7 to
cb88b94
Compare
scott-lowe-vapi
added a commit
that referenced
this pull request
Oct 3, 2026
## Value **V.A.L.U.E. tier:** project — PR 8 of 10 for inline simulation PR checks ([TEST-141](https://linear.app/vapi/issue/TEST-141/gitops-run-simulation-suites-against-pr-changes-inline-as-ci-checks)). This PR turns `npm run check` into a PR check and documents setup end to end. - **Problem:** `npm run check` (PR 7) gives a verdict locally, but a reviewer needs it on the PR: a `Vapi Evals` status whose **Details** link opens the exact run (PAL-608's contract), run on every affected push, safe for forks and Dependabot. - **Who it affects:** gitops users, single-org and CI-org alike, who get a PR check by adding one variable and one secret; and the maintainers of forks like Hoag's and Mudflap's, who hand-maintain shell workflows for this today (PAL-608). - **What changes:** - **`.github/workflows/vapi-checks.yml`** (opt-in): - **Triggers:** `pull_request` (opened, synchronize, reopened, ready_for_review), plus `workflow_dispatch` with an optional `check` input. Never `pull_request_target`. - **Opt-in:** it runs only when `vars.VAPI_CHECKS_ENABLED == 'true'`, and does nothing without a `vapi-checks.yml`, so on the upstream template it stays dormant. - **Live vs dry run:** - live only for same-repository, non-Dependabot PRs (checking both the actor and the PR author) and for dispatch; - everything else gets a keyless dry run, and secrets are passed only to live runs. - **Checkout:** the head SHA, with `fetch-depth: 0` so `--changed-since origin/<base>` diffs from the merge base, and `persist-credentials: false`. - **Statuses:** `--all` posts the aggregate `Vapi Evals`; dispatching one named check never changes it. - **Cancellation:** concurrency is keyed per PR with `cancel-in-progress`, and the run step `exec`s node, so GitHub's cancel reaches it and it cancels the superseded runs. - **Time:** `--budget-minutes 22` under `timeout-minutes: 30`. - **Permissions:** `contents: read` and `statuses: write`. - **README "PR Checks":** the end-user guide, covering: - test files with a judge example, and the check config; - the dry run, plus the build-failure table (what fails and how to fix it); - live runs, and turning on the workflow; - what a PR shows, making `Vapi Evals` required, and fork/Dependabot unblocking; - dedicated CI orgs; - cost, and what stays real in the run org. The project tree gains `vapi-checks.example.yml` and `check-cmd.ts`. - **AGENTS.md:** the "Testing with Simulations" step 5 now points at `npm run sim` vs `npm run check`. - **`docs/learnings/simulations.md` "Inline PR Checks":** - the parity results and the two known differences (handoff names, tool order); - mock behaviour, and hook-fired tools bypassing mocks; - why servers are replaced rather than deleted; - transfers; - what still reaches real systems; - the chat-mode limits; - where run items keep tool results. - **Indexes:** the learnings index row, and `improvements.md` #34 (RESOLVED). No `docs/changelog.md` edit. ## Evidence of value - **Dormant on the template:** this PR adds the workflow, and `VapiAI/gitops` has no `VAPI_CHECKS_ENABLED` variable, so the `vapi-checks` job is **skipped** in this PR's own checks ([run 36941407785](https://github.com/VapiAI/gitops/actions/runs/36941407785): `completed / skipped`). That also shows GitHub parsed the workflow and evaluated its `if:`. Customers who haven't opted in see no change and no spend. - **The command the workflow runs** (`--all`, `--changed-since`, live and dry run, statuses, job summary, JSON) is covered end to end by PR 7's tests against a stub of the simulations and GitHub APIs. It was also run live in the owner's test org: innocuous exit 0, degraded exit 1, no resources created, every tool result a mock. - `npm test`: 474 passing (docs and workflow only; no code change). ## Testing plan - The workflow YAML parses. `actionlint` isn't available in this environment, so it isn't linted. To check by hand: `if:` expressions, the `secrets` ternaries in step `env`, and the `exec` line. - **Not yet tested — needs the owner,** because it needs a repository secret holding an API key: 1. Create a private dogfood repo from this branch, with `resources/<test-org>/` holding the parity squad from `tests/fixtures/check-parity/resources/parity/`, renamed, and that fixture's `vapi-checks.yml`. 2. Add the `VAPI_PRIVATE_API_KEY` secret and `VAPI_CHECKS_ENABLED=true`. 3. Open a degraded-prompt PR and an innocuous one. Expect red and green `Vapi Evals` statuses whose **Details** links land on the runs, plus the job summaries (screenshots to add here). 4. Open a PR from a fork. Expect a dry run and no status, because the token is read-only. 5. Push a `package.json` change as Dependabot would. Expect `Vapi Evals` = `error`. 6. Push twice quickly. Expect the first run canceled (`itemCounts.canceled > 0`). Stacked on #63. Refs TEST-141 🤖 Generated with [Claude Code](https://claude.com/claude-code)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Value
V.A.L.U.E. tier: project — PR 7 of 10 for inline simulation PR checks (TEST-141); this is the first PR that sends anything.
Vapi Evalscommit status whose Details link opens the exact run.npm run check -- corelocally against their branch's files and get a strict verdict with a run link. The PR workflow (PR 8) and the promotion gate (PR 10) call the same command.src/check-run.ts: one run per target, at most 3 at once.npm run sim's create/poll/cancel/hydrate/verdict loop, now extracted fromrunSimulationassimRunExecute, so a check is judged by the same strict rules (all-skipped judges, canceled or missing items are never a pass).timeoutMinutesand--budget-minutes, no run starts with under 5 minutes left, and the deadline, SIGINT and SIGTERM cancel in-flight runs.src/check-select.ts:--changed-since <ref>diffs from the merge base. A check is affected by:resources/**and state files;vapi-checks.ymlorpromotion.yml;src/**orpackage*.json;paths.src/check-status.ts: GitHub commit statuses with plainfetch, whenGITHUB_TOKEN,GITHUB_REPOSITORYandHEAD_SHAare set.Vapi Evals / <check> / <target>goes pending with the run's canonicalurlas soon as the run exists, then to success, failure or error.--allruns post the aggregateVapi Evalsfrom afinally: the worst state,successif nothing is affected,errorfor a dry run of an affected PR, anderrorwhen config fails to parse or an exception escapes.src/check-report.ts: a markdown summary, also appended to$GITHUB_STEP_SUMMARY. It has one row per target with the run link, then failing evaluations with expected vs extracted values, mock notices and warnings.--jsonwrites the same data. No PR comments.src/check-cmd.tsgains live runs and--changed-since,--budget-minutes,--refresh-bindings(the read-only bindings pull promotion uses) and--json.VAPI_CHECK_TOKENS, then.env.<runOrg>, thenVAPI_PRIVATE_API_KEYwhen every selected check runs in one org.User-Agent: vapi-gitops-check/<version>, for the post-deploy PostHog measure.npm run checkrows now describe live mode.Evidence of value
The definition-of-done pair, run live in the owner's test org on the TEST-141 parity squad (2 members, function tools, a handoff, 3 scenarios; built from
tests/fixtures/check-parity/):npm run check -- corenpm run check -- coreThe degraded run's report named each failing judge with expected vs extracted:
S2 (the hours question, which never reaches the scheduler) correctly still passed.
Nothing was created in the org. Counts of assistants, tools, squads, structured outputs, personalities, scenarios, simulations, suites and credentials were identical before and after both runs:
{"/assistant":12,"/tool":8,"/squad":2,"/structured-output":8,"/eval/simulation/personality":10,"/eval/simulation/scenario":10,"/eval/simulation":10,"/eval/simulation/suite":1,"/credential":0}. This stands in fornpm run cleanup, which needs a gitops-managed org.Every tool call returned a mock. Every
tool_call_resultin both runs' transcripts was matched to its call bytoolCallId:The run items also confirmed that personalities kept the dead server and
serverMessages: [].A bug the live run caught, fixed in this PR:
metadata.scenario, default error mocks included. Scanning the whole item would have reported "unmocked tool called" for every default mock, called or not.tool_call_resultmessages inmetadata.call.messages, mapped to tool names throughtool_calls.S1 book cleaning: unmocked tool called: lookup_patient.Tests:
npm testgoes from 450 to 474 passing;npm run buildis clean.Testing plan
tests/check-run.test.ts(10 tests, stateful local stub):tests/check-cmd.test.ts(11 tests): the live--allpath against one stub serving both the simulations API and GitHub:--changed-sincewith an unaffected change runs nothing and posts green;errorand sends nothing;error.The file clears
VAPI_*andGITHUB_*first, so a developer shell with a key exported can't reach a real org.tests/check-select.test.ts: the affected-path table, and the merge-base case in a temp git repo (commits on the base after branching don't count).tests/check-status.test.ts,tests/check-report.test.ts: env parsing, POST shape and 140-character truncation, 403 only warning, state ordering, the exact markdown, and the JSON.Not tested:
--refresh-bindingsagainst a live org (it runs the same child pull promotion does);runningrun live (cancel was verified only on a queued run in the parity experiment);Stacked on #62.
Refs TEST-141
🤖 Generated with Claude Code