Skip to content

feat(check): add the opt-in Vapi checks PR workflow and user docs - #64

Open
scott-lowe-vapi wants to merge 1 commit into
feat/check-live-runfrom
feat/vapi-checks-workflow
Open

scott-lowe-vapi wants to merge 1 commit into
feat/check-live-runfrom
feat/vapi-checks-workflow

Conversation

@scott-lowe-vapi

@scott-lowe-vapi scott-lowe-vapi commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Value

V.A.L.U.E. tier: project — PR 8 of 10 for inline simulation PR checks (TEST-141). This PR turns npm run check into a PR check and documents setup end to end.

  • Problem: npm run check (PR 7) gives a verdict locally, but a reviewer needs it on the PR: a Vapi Evals status whose Details link opens the exact run (PAL-608's contract), run on every affected push, safe for forks and Dependabot.
  • Who it affects: gitops users, single-org and CI-org alike, who get a PR check by adding one variable and one secret; and the maintainers of forks like Hoag's and Mudflap's, who hand-maintain shell workflows for this today (PAL-608).
  • What changes:
    • .github/workflows/vapi-checks.yml (opt-in):

      • Triggers: pull_request (opened, synchronize, reopened, ready_for_review), plus workflow_dispatch with an optional check input. Never pull_request_target.
      • Opt-in: it runs only when vars.VAPI_CHECKS_ENABLED == 'true', and does nothing without a vapi-checks.yml, so on the upstream template it stays dormant.
      • Live vs dry run:
        • live only for same-repository, non-Dependabot PRs (checking both the actor and the PR author) and for dispatch;
        • everything else gets a keyless dry run, and secrets are passed only to live runs.
      • Checkout: the head SHA, with fetch-depth: 0 so --changed-since origin/<base> diffs from the merge base, and persist-credentials: false.
      • Statuses: --all posts the aggregate Vapi Evals; dispatching one named check never changes it.
      • Cancellation: concurrency is keyed per PR with cancel-in-progress, and the run step execs node, so GitHub's cancel reaches it and it cancels the superseded runs.
      • Time: --budget-minutes 22 under timeout-minutes: 30.
      • Permissions: contents: read and statuses: write.
    • README "PR Checks": the end-user guide, covering:

      • test files with a judge example, and the check config;
      • the dry run, plus the build-failure table (what fails and how to fix it);
      • live runs, and turning on the workflow;
      • what a PR shows, making Vapi Evals required, and fork/Dependabot unblocking;
      • dedicated CI orgs;
      • cost, and what stays real in the run org.

      The project tree gains vapi-checks.example.yml and check-cmd.ts.

    • AGENTS.md: the "Testing with Simulations" step 5 now points at npm run sim vs npm run check.

    • docs/learnings/simulations.md "Inline PR Checks":

      • the parity results and the two known differences (handoff names, tool order);
      • mock behaviour, and hook-fired tools bypassing mocks;
      • why servers are replaced rather than deleted;
      • transfers;
      • what still reaches real systems;
      • the chat-mode limits;
      • where run items keep tool results.
    • Indexes: the learnings index row, and improvements.md refactor(push): extract reconcileStateKeyForResource — fold two ensure-fns into one generic helper #34 (RESOLVED). No docs/changelog.md edit.

Evidence of value

  • Dormant on the template: this PR adds the workflow, and VapiAI/gitops has no VAPI_CHECKS_ENABLED variable, so the vapi-checks job is skipped in this PR's own checks (run 36941407785: completed / skipped). That also shows GitHub parsed the workflow and evaluated its if:. Customers who haven't opted in see no change and no spend.
  • The command the workflow runs (--all, --changed-since, live and dry run, statuses, job summary, JSON) is covered end to end by PR 7's tests against a stub of the simulations and GitHub APIs. It was also run live in the owner's test org: innocuous exit 0, degraded exit 1, no resources created, every tool result a mock.
  • npm test: 474 passing (docs and workflow only; no code change).

Testing plan

  • The workflow YAML parses. actionlint isn't available in this environment, so it isn't linted. To check by hand: if: expressions, the secrets ternaries in step env, and the exec line.
  • Not yet tested — needs the owner, because it needs a repository secret holding an API key:
    1. Create a private dogfood repo from this branch, with resources/<test-org>/ holding the parity squad from tests/fixtures/check-parity/resources/parity/, renamed, and that fixture's vapi-checks.yml.
    2. Add the VAPI_PRIVATE_API_KEY secret and VAPI_CHECKS_ENABLED=true.
    3. Open a degraded-prompt PR and an innocuous one. Expect red and green Vapi Evals statuses whose Details links land on the runs, plus the job summaries (screenshots to add here).
    4. Open a PR from a fork. Expect a dry run and no status, because the token is read-only.
    5. Push a package.json change as Dependabot would. Expect Vapi Evals = error.
    6. Push twice quickly. Expect the first run canceled (itemCounts.canceled > 0).

Stacked on #63.

Refs TEST-141

🤖 Generated with Claude Code

.github/workflows/vapi-checks.yml runs `npm run check` on pull requests
(opened, synchronize, reopened, ready_for_review) and on manual dispatch,
only when the repository variable VAPI_CHECKS_ENABLED is 'true' and a
vapi-checks.yml exists — so on the upstream template it is dormant.

- Live runs only for same-repository, non-Dependabot PRs and dispatch;
  forks and Dependabot get a keyless dry run, and secrets are passed only
  to live runs. Never pull_request_target.
- Checks out the head SHA with full history (for --changed-since against
  origin/<base>) and persist-credentials: false.
- --all posts the aggregate `Vapi Evals` status; dispatching one named
  check doesn't. The run step execs node so GitHub's cancel reaches it,
  inside a 22-minute budget under a 30-minute job timeout; concurrency
  cancels a superseded push's runs.
- permissions: contents: read, statuses: write.

Docs: a README "PR Checks" section (setup from test files to required
status, the build-failure table, fork/Dependabot handling, CI orgs, cost
and what stays real), the AGENTS.md simulations step, a "Inline PR
Checks" section in docs/learnings/simulations.md, and improvements.md #34.

Refs TEST-141

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants