Skip to content

feat(check): run checks live and report Vapi Evals statuses - #63

Merged
scott-lowe-vapi merged 1 commit into
mainfrom
feat/check-live-run
Oct 3, 2026
Merged

scott-lowe-vapi merged 1 commit into
mainfrom
feat/check-live-run

Conversation

@scott-lowe-vapi

Copy link
Copy Markdown
Contributor

Value

V.A.L.U.E. tier: project — PR 7 of 10 for inline simulation PR checks (TEST-141); this is the first PR that sends anything.

  • Problem: the payload builder (PRs 5–6) produces a safe inline body, but nothing runs it or reports the result where a reviewer looks. PAL-608 asks for a stable Vapi Evals commit status whose Details link opens the exact run.
  • Who it affects: gitops users, who can now run npm run check -- core locally against their branch's files and get a strict verdict with a run link. The PR workflow (PR 8) and the promotion gate (PR 10) call the same command.
  • What changes:
    • src/check-run.ts: one run per target, at most 3 at once.
      • It reuses npm run sim's create/poll/cancel/hydrate/verdict loop, now extracted from runSimulation as simRunExecute, so a check is judged by the same strict rules (all-skipped judges, canceled or missing items are never a pass).
      • Each run's deadline is the earlier of timeoutMinutes and --budget-minutes, no run starts with under 5 minutes left, and the deadline, SIGINT and SIGTERM cancel in-flight runs.
      • A 402 is reported as billing.
      • A 5xx on create is never retried, because a duplicate paid run is worse than a retry.
      • Transcripts are scanned for default-mock answers ("unmocked tool called").
    • src/check-select.ts: --changed-since <ref> diffs from the merge base. A check is affected by:
      • its org's and run org's resources/** and state files;
      • vapi-checks.yml or promotion.yml;
      • src/** or package*.json;
      • its own paths.
    • src/check-status.ts: GitHub commit statuses with plain fetch, when GITHUB_TOKEN, GITHUB_REPOSITORY and HEAD_SHA are set.
      • Vapi Evals / <check> / <target> goes pending with the run's canonical url as soon as the run exists, then to success, failure or error.
      • --all runs post the aggregate Vapi Evals from a finally: the worst state, success if nothing is affected, error for a dry run of an affected PR, and error when config fails to parse or an exception escapes.
      • A rejected post (a fork's read-only token) only warns.
    • src/check-report.ts: a markdown summary, also appended to $GITHUB_STEP_SUMMARY. It has one row per target with the run link, then failing evaluations with expected vs extracted values, mock notices and warnings. --json writes the same data. No PR comments.
    • src/check-cmd.ts gains live runs and --changed-since, --budget-minutes, --refresh-bindings (the read-only bindings pull promotion uses) and --json.
      • Keys: VAPI_CHECK_TOKENS, then .env.<runOrg>, then VAPI_PRIVATE_API_KEY when every selected check runs in one org.
      • Exit codes: 0 passed, 1 failed, 2 config or build error, 3 incomplete.
    • Runs send User-Agent: vapi-gitops-check/<version>, for the post-deploy PostHog measure.
    • The README and AGENTS.md npm run check rows now describe live mode.

Evidence of value

The definition-of-done pair, run live in the owner's test org on the TEST-141 parity squad (2 members, function tools, a handoff, 3 scenarios; built from tests/fixtures/check-parity/):

Run Command Exit Result
Innocuous (fixture as-is) npm run check -- core 0 ✅ 3/3 passed — run f68ba89b
Degraded scheduler prompt ("always say there are no openings, never book") npm run check -- core 1 ❌ 1/3 — run 19711623

The degraded run's report named each failing judge with expected vs extracted:

Simulation Evaluation Comparator Expected Got
S3 alternative slot offered-alternative = true false
S3 alternative slot booked-wednesday = true false
S1 book cleaning booking-confirmed = true false
S1 book cleaning only-real-slots = true false

S2 (the hours question, which never reaches the scheduler) correctly still passed.

Nothing was created in the org. Counts of assistants, tools, squads, structured outputs, personalities, scenarios, simulations, suites and credentials were identical before and after both runs: {"/assistant":12,"/tool":8,"/squad":2,"/structured-output":8,"/eval/simulation/personality":10,"/eval/simulation/scenario":10,"/eval/simulation":10,"/eval/simulation/suite":1,"/credential":0}. This stands in for npm run cleanup, which needs a gitops-managed org.

Every tool call returned a mock. Every tool_call_result in both runs' transcripts was matched to its call by toolCallId:

Run Function results equal to the scenario's mock Built-in results (handoff, endCall) Anything else
Innocuous 6 4 0
Degraded 2 2 0

The run items also confirmed that personalities kept the dead server and serverMessages: [].

A bug the live run caught, fixed in this PR:

  • The trap: run items echo the scenario under metadata.scenario, default error mocks included. Scanning the whole item would have reported "unmocked tool called" for every default mock, called or not.
  • The fix: the scan now reads only tool_call_result messages in metadata.call.messages, mapped to tool names through tool_calls.
  • Checked on both runs' real items: no notices; with one result swapped for a default-mock answer, it reported S1 book cleaning: unmocked tool called: lookup_patient.
  • A regression test covers the echoed-defaults case.

Tests: npm test goes from 450 to 474 passing; npm run build is clean.

Testing plan

  • tests/check-run.test.ts (10 tests, stateful local stub):

    • pass with link and User-Agent;
    • two targets with one failing, naming the evaluation;
    • all-skipped is incomplete;
    • timeout cancels;
    • abort cancels in-flight runs and starts none;
    • the budget gate;
    • 402 as billing, and no retry on a create 502;
    • a build error sends nothing;
    • mock notices (echoed defaults ignored);
    • exactly 3 in flight with results in job order.
  • tests/check-cmd.test.ts (11 tests): the live --all path against one stub serving both the simulations API and GitHub:

    • pending, then success per target, then the aggregate with the run link, plus the job summary and JSON;
    • a failing run exits 1 with a red aggregate;
    • a named check never posts the aggregate;
    • --changed-since with an unaffected change runs nothing and posts green;
    • a dry run of an affected PR posts error and sends nothing;
    • invalid config posts error.

    The file clears VAPI_* and GITHUB_* first, so a developer shell with a key exported can't reach a real org.

  • tests/check-select.test.ts: the affected-path table, and the merge-base case in a temp git repo (commits on the base after branching don't count).

  • tests/check-status.test.ts, tests/check-report.test.ts: env parsing, POST shape and 140-character truncation, 403 only warning, state ordering, the exact markdown, and the JSON.

  • Not tested:

    • --refresh-bindings against a live org (it runs the same child pull promotion does);
    • EU base URLs;
    • canceling a running run live (cancel was verified only on a queued run in the parity experiment);
    • GitHub's real statuses API (stubbed here; PR 8's dogfood repo covers it);
    • a real customer-shaped squad, which needs the owner's permission first.

Stacked on #62.

Refs TEST-141

🤖 Generated with Claude Code

scott-lowe-vapi commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor Author

Merge activity

  • Oct 3, 5:59 AM UTC: A user started a stack merge that includes this pull request via Graphite.
  • Oct 3, 6:11 AM UTC: Graphite rebased this pull request as part of a merge.
  • Oct 3, 6:11 AM UTC: @scott-lowe-vapi merged this pull request with Graphite.

@scott-lowe-vapi
scott-lowe-vapi changed the base branch from feat/check-fail-closed-mocks to graphite-base/63 October 3, 2026 06:08
@scott-lowe-vapi
scott-lowe-vapi changed the base branch from graphite-base/63 to main October 3, 2026 06:09
`npm run check -- <check>|--all` now runs each built target payload in
the check's run org: one POST /eval/simulation/run per target, at most
three at once, judged by the same strict verdict as `npm run sim`
(the create/poll/cancel/hydrate loop moves into simRunExecute, which
both share).

- Budget and cancellation: each run's deadline is the earlier of its
  timeoutMinutes and --budget-minutes; no run starts with under five
  minutes left; the deadline, SIGINT and SIGTERM cancel in-flight runs.
- Keys: VAPI_CHECK_TOKENS, then .env.<runOrg>, then VAPI_PRIVATE_API_KEY
  when every selected check runs in one org. --refresh-bindings runs the
  same read-only bindings pull promotion uses.
- Selection: --changed-since <ref> runs only checks affected by changes
  since the merge base (the org's files and state, the run org's, the
  config, the engine, and the check's own paths).
- Statuses (when GITHUB_TOKEN, GITHUB_REPOSITORY and HEAD_SHA are set):
  `Vapi Evals / <check> / <target>` goes pending with the run's link,
  then success/failure/error; --all runs post the aggregate `Vapi Evals`
  from a finally block, so config errors and exceptions report too.
- Report: a markdown summary (also appended to $GITHUB_STEP_SUMMARY) with
  failing evaluations' expected vs extracted values and "unmocked tool
  called" notices from transcripts, plus --json. No PR comments.
- Runs send User-Agent vapi-gitops-check/<version>.

Exit codes: 0 passed, 1 failed, 2 config or build error, 3 incomplete.

Refs TEST-141

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@scott-lowe-vapi
scott-lowe-vapi merged commit e28e7d7 into main Oct 3, 2026
2 checks passed
scott-lowe-vapi added a commit that referenced this pull request Oct 3, 2026
## Value

**V.A.L.U.E. tier:** project — PR 8 of 10 for inline simulation PR checks ([TEST-141](https://linear.app/vapi/issue/TEST-141/gitops-run-simulation-suites-against-pr-changes-inline-as-ci-checks)). This PR turns `npm run check` into a PR check and documents setup end to end.

- **Problem:** `npm run check` (PR 7) gives a verdict locally, but a reviewer needs it on the PR: a `Vapi Evals` status whose **Details** link opens the exact run (PAL-608's contract), run on every affected push, safe for forks and Dependabot.
- **Who it affects:** gitops users, single-org and CI-org alike, who get a PR check by adding one variable and one secret; and the maintainers of forks like Hoag's and Mudflap's, who hand-maintain shell workflows for this today (PAL-608).
- **What changes:**
  - **`.github/workflows/vapi-checks.yml`** (opt-in):
    - **Triggers:** `pull_request` (opened, synchronize, reopened, ready_for_review), plus `workflow_dispatch` with an optional `check` input. Never `pull_request_target`.
    - **Opt-in:** it runs only when `vars.VAPI_CHECKS_ENABLED == 'true'`, and does nothing without a `vapi-checks.yml`, so on the upstream template it stays dormant.
    - **Live vs dry run:**
      - live only for same-repository, non-Dependabot PRs (checking both the actor and the PR author) and for dispatch;
      - everything else gets a keyless dry run, and secrets are passed only to live runs.
    - **Checkout:** the head SHA, with `fetch-depth: 0` so `--changed-since origin/<base>` diffs from the merge base, and `persist-credentials: false`.
    - **Statuses:** `--all` posts the aggregate `Vapi Evals`; dispatching one named check never changes it.
    - **Cancellation:** concurrency is keyed per PR with `cancel-in-progress`, and the run step `exec`s node, so GitHub's cancel reaches it and it cancels the superseded runs.
    - **Time:** `--budget-minutes 22` under `timeout-minutes: 30`.
    - **Permissions:** `contents: read` and `statuses: write`.
  - **README "PR Checks":** the end-user guide, covering:
    - test files with a judge example, and the check config;
    - the dry run, plus the build-failure table (what fails and how to fix it);
    - live runs, and turning on the workflow;
    - what a PR shows, making `Vapi Evals` required, and fork/Dependabot unblocking;
    - dedicated CI orgs;
    - cost, and what stays real in the run org.

    The project tree gains `vapi-checks.example.yml` and `check-cmd.ts`.
  - **AGENTS.md:** the "Testing with Simulations" step 5 now points at `npm run sim` vs `npm run check`.
  - **`docs/learnings/simulations.md` "Inline PR Checks":**
    - the parity results and the two known differences (handoff names, tool order);
    - mock behaviour, and hook-fired tools bypassing mocks;
    - why servers are replaced rather than deleted;
    - transfers;
    - what still reaches real systems;
    - the chat-mode limits;
    - where run items keep tool results.
  - **Indexes:** the learnings index row, and `improvements.md` #34 (RESOLVED). No `docs/changelog.md` edit.

## Evidence of value

- **Dormant on the template:** this PR adds the workflow, and `VapiAI/gitops` has no `VAPI_CHECKS_ENABLED` variable, so the `vapi-checks` job is **skipped** in this PR's own checks ([run 36941407785](https://github.com/VapiAI/gitops/actions/runs/36941407785): `completed / skipped`). That also shows GitHub parsed the workflow and evaluated its `if:`. Customers who haven't opted in see no change and no spend.
- **The command the workflow runs** (`--all`, `--changed-since`, live and dry run, statuses, job summary, JSON) is covered end to end by PR 7's tests against a stub of the simulations and GitHub APIs. It was also run live in the owner's test org: innocuous exit 0, degraded exit 1, no resources created, every tool result a mock.
- `npm test`: 474 passing (docs and workflow only; no code change).

## Testing plan

- The workflow YAML parses. `actionlint` isn't available in this environment, so it isn't linted. To check by hand: `if:` expressions, the `secrets` ternaries in step `env`, and the `exec` line.
- **Not yet tested — needs the owner,** because it needs a repository secret holding an API key:
  1. Create a private dogfood repo from this branch, with `resources/<test-org>/` holding the parity squad from `tests/fixtures/check-parity/resources/parity/`, renamed, and that fixture's `vapi-checks.yml`.
  2. Add the `VAPI_PRIVATE_API_KEY` secret and `VAPI_CHECKS_ENABLED=true`.
  3. Open a degraded-prompt PR and an innocuous one. Expect red and green `Vapi Evals` statuses whose **Details** links land on the runs, plus the job summaries (screenshots to add here).
  4. Open a PR from a fork. Expect a dry run and no status, because the token is read-only.
  5. Push a `package.json` change as Dependabot would. Expect `Vapi Evals` = `error`.
  6. Push twice quickly. Expect the first run canceled (`itemCounts.canceled > 0`).

Stacked on #63.

Refs TEST-141

🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants