From 680a03d632c07f8d1248432f2fde68d0c61a85b4 Mon Sep 17 00:00:00 2001 From: Scott Lowe Date: Thu, 1 Oct 2026 16:06:42 -0700 Subject: [PATCH] feat(check): build inline simulation-run payloads and add a dry-run CLI `npm run check -- |--all --dry-run` reads vapi-checks.yml and, for each check target, builds the inline POST /eval/simulation/run body from the files on disk: the assistant or squad with its tools, handoffs and structured outputs, plus the scenarios, judges and personalities the check's suites and simulations name. Nothing is looked up on the platform and no key is needed; --print-payload writes the JSON. The builder follows push and the runtime so the check tests what push deploys: - tools in runtime order (model.tools, then toolIds, then toolRefs); knowledgeBase tools stay in toolIds by run-org UUID; - `##` comments stripped; UUID references resolved through state; - squad members inlined, handoffs to members by name; - hook toolIds, artifactPlan structuredOutputIds and linked structured outputs inlined; judges' structuredOutputId inlined; - credentials bound by name to the run org (promotion's bindings). Anything it can't place fails the build with the field named: missing files, handoffs leaving the target, legacy assistantDestinations by ID, tools by ID inside overrides (strict mocks), duplicate tool names, audio judges/hooks/no required judge over chat, leftover non-UUID references, and payloads over 4.5 MB. The tool mock policy and live runs follow in later changes; without --dry-run the command exits 2. Refs TEST-141 Co-Authored-By: Claude Opus 5.5 --- AGENTS.md | 2 + README.md | 1 + package.json | 1 + src/check-cmd.ts | 226 +++++++ src/check-payload-assistant.ts | 325 ++++++++++ src/check-payload-refs.ts | 127 ++++ src/check-payload.ts | 538 ++++++++++++++++ tests/check-cmd.test.ts | 132 ++++ tests/check-payload.test.ts | 572 ++++++++++++++++++ .../check-parity/.vapi-state.parity.json | 1 + .../parity/assistants/receptionist.md | 18 + .../resources/parity/assistants/scheduler.md | 15 + .../personalities/dental-caller.yml | 9 + .../scenarios/s1-book-cleaning.yml | 21 + .../scenarios/s2-hours-question.yml | 21 + .../scenarios/s3-alternative-slot.yml | 21 + .../parity/simulations/suites/core.yml | 5 + .../simulations/tests/s1-book-cleaning.yml | 3 + .../simulations/tests/s2-hours-question.yml | 3 + .../simulations/tests/s3-alternative-slot.yml | 3 + .../resources/parity/squads/dental.yml | 25 + .../structuredOutputs/booked-wednesday.yml | 4 + .../structuredOutputs/booking-confirmed.yml | 4 + .../structuredOutputs/hours-correct.yml | 4 + .../parity/structuredOutputs/no-booking.yml | 4 + .../structuredOutputs/offered-alternative.yml | 4 + .../structuredOutputs/only-real-slots.yml | 4 + .../parity/tools/check-availability.yml | 15 + .../parity/tools/handoff-to-scheduler.yml | 5 + .../resources/parity/tools/lookup-patient.yml | 13 + tests/fixtures/check-parity/vapi-checks.yml | 6 + 31 files changed, 2132 insertions(+) create mode 100644 src/check-cmd.ts create mode 100644 src/check-payload-assistant.ts create mode 100644 src/check-payload-refs.ts create mode 100644 src/check-payload.ts create mode 100644 tests/check-cmd.test.ts create mode 100644 tests/check-payload.test.ts create mode 100644 tests/fixtures/check-parity/.vapi-state.parity.json create mode 100644 tests/fixtures/check-parity/resources/parity/assistants/receptionist.md create mode 100644 tests/fixtures/check-parity/resources/parity/assistants/scheduler.md create mode 100644 tests/fixtures/check-parity/resources/parity/simulations/personalities/dental-caller.yml create mode 100644 tests/fixtures/check-parity/resources/parity/simulations/scenarios/s1-book-cleaning.yml create mode 100644 tests/fixtures/check-parity/resources/parity/simulations/scenarios/s2-hours-question.yml create mode 100644 tests/fixtures/check-parity/resources/parity/simulations/scenarios/s3-alternative-slot.yml create mode 100644 tests/fixtures/check-parity/resources/parity/simulations/suites/core.yml create mode 100644 tests/fixtures/check-parity/resources/parity/simulations/tests/s1-book-cleaning.yml create mode 100644 tests/fixtures/check-parity/resources/parity/simulations/tests/s2-hours-question.yml create mode 100644 tests/fixtures/check-parity/resources/parity/simulations/tests/s3-alternative-slot.yml create mode 100644 tests/fixtures/check-parity/resources/parity/squads/dental.yml create mode 100644 tests/fixtures/check-parity/resources/parity/structuredOutputs/booked-wednesday.yml create mode 100644 tests/fixtures/check-parity/resources/parity/structuredOutputs/booking-confirmed.yml create mode 100644 tests/fixtures/check-parity/resources/parity/structuredOutputs/hours-correct.yml create mode 100644 tests/fixtures/check-parity/resources/parity/structuredOutputs/no-booking.yml create mode 100644 tests/fixtures/check-parity/resources/parity/structuredOutputs/offered-alternative.yml create mode 100644 tests/fixtures/check-parity/resources/parity/structuredOutputs/only-real-slots.yml create mode 100644 tests/fixtures/check-parity/resources/parity/tools/check-availability.yml create mode 100644 tests/fixtures/check-parity/resources/parity/tools/handoff-to-scheduler.yml create mode 100644 tests/fixtures/check-parity/resources/parity/tools/lookup-patient.yml create mode 100644 tests/fixtures/check-parity/vapi-checks.yml diff --git a/AGENTS.md b/AGENTS.md index 05ba2b2..a147c92 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -90,6 +90,7 @@ If setup reports the org is "already set up locally", do not delete anything to | Push with new resources | `npm run push -- --allow-new-files` — bypass orphan-YAML gate. **AI agents**: do NOT auto-pass this flag; confirm with the human first (see push section below) | | Test a call | `npm run call -- -a ` or `-s ` | | Run a simulation suite | `npm run sim -- --suite --target ` | +| Build PR-check payloads offline | `npm run check -- --dry-run` (or `--all`); `--print-payload` writes the JSON | | Migrate a legacy state file | `npm run migrate` — one-shot, all orgs; required once after upgrading to the hash-store engine | --- @@ -903,6 +904,7 @@ npm run validate -- # Lint resources locally (fai npm run audit -- # Read-only drift detector: orphan YAML, state ghosts, content-identical clusters, sibling base-slugs, dashboard orphans, inline model.tools. Exit 1 on findings. npm run audit -- --type assistants # Scope audit to a single resource type npm run sim -- --suite --target # Run a simulation suite against an assistant/squad (exit 0 pass, 1 fail, 3 incomplete; --timeout ) +npm run check -- --dry-run # Build a vapi-checks.yml check's inline run payload offline (exit 0 built, 2 config/build error; --print-payload [dir]) npm run rollback -- --to # Re-apply a snapshot taken before a push npm run rollback -- --list # List available snapshots diff --git a/README.md b/README.md index 180da18..61da8da 100644 --- a/README.md +++ b/README.md @@ -104,6 +104,7 @@ Every command works in two modes: | `npm run rollback` | — | `npm run rollback -- --list` or `--to ` | Restore from a snapshot in `.vapi-state..snapshots/` (one is written before every push/apply). | | `npm run call` | ✅ | `npm run call -- -a ` or `-s ` | Start an interactive WebSocket call against an assistant or squad. | | `npm run sim` | — | `npm run sim -- --suite --target [--timeout ]` | Run a simulation suite (or specific simulations) against a deployed assistant/squad. Prints the run link; exits 0 passed, 1 failed, 3 incomplete (timeout, Ctrl-C, missing results). | +| `npm run check` | — | `npm run check -- \|--all --dry-run [--print-payload [dir]]` | Build each `vapi-checks.yml` check target's inline simulation-run payload from the files on disk, with its tools, handoffs, judges and personalities. Offline: no API key, nothing sent. `--print-payload` writes the JSON (default `tmp/check-payloads/`). Exits 0 when every payload builds, 2 on a config or build error. | | `npm run migrate` | — | `npm run migrate` | One-time, all orgs at once: slim legacy state files to pure `name → uuid` and seed the per-developer `.vapi-state-hash/` baseline store from the old hashes. Required once after upgrading to the hash-store engine — `pull`/`push`/`apply` refuse legacy-shaped state until it runs. Idempotent. | | `npm run build` | — | — | Type-check the codebase (`tsc --noEmit`). | | `npm test` | — | — | Run regression tests (`node:test`). | diff --git a/package.json b/package.json index e7869df..8c8415f 100644 --- a/package.json +++ b/package.json @@ -19,6 +19,7 @@ "validate": "tsx src/validate-cmd.ts", "audit": "tsx src/audit-cmd.ts", "sim": "tsx src/sim-cmd.ts", + "check": "tsx src/check-cmd.ts", "rollback": "tsx src/rollback-cmd.ts", "promote": "tsx src/promote-cmd.ts", "build": "tsc --noEmit", diff --git a/src/check-cmd.ts b/src/check-cmd.ts new file mode 100644 index 0000000..e9655bf --- /dev/null +++ b/src/check-cmd.ts @@ -0,0 +1,226 @@ +// CLI entry: `npm run check -- |--all --dry-run [--print-payload [dir]]` +// +// Builds each check target's inline simulation-run payload from the files +// on disk (see check-payload.ts) and reports what would run. Config-free: +// it needs no API key and makes no network calls in --dry-run. +// +// Exit codes: 0 every payload built, 2 usage, config or build error. + +import { existsSync, mkdirSync, readFileSync, writeFileSync } from "node:fs"; +import { join, relative, resolve } from "node:path"; +import { fileURLToPath } from "node:url"; +import type { CheckDefinition } from "./check-config.ts"; +import { CHECKS_CONFIG_FILE, checksConfigLoad } from "./check-config.ts"; +import type { CheckPayloadResult } from "./check-payload.ts"; +import { checkPayloadBuild } from "./check-payload.ts"; +import { promotionBindingsResolve, promotionStateParse } from "./promotion.ts"; +import { orgResourcesRead } from "./resource-parse.ts"; +import type { StateFile } from "./types.ts"; + +const USAGE = [ + "Usage:", + " npm run check -- --dry-run [--print-payload [dir]]", + " npm run check -- --all --dry-run [--print-payload [dir]]", + "", + "Options:", + ` A check name from ${CHECKS_CONFIG_FILE}`, + " --all Every check", + " --dry-run Build the payloads offline; nothing is sent", + " --print-payload [dir] Write each payload as JSON (default: tmp/check-payloads)", + "", + "Exit codes: 0 payloads built, 2 usage, config or build error", +].join("\n"); + +const DEFAULT_PAYLOAD_DIR = "tmp/check-payloads"; + +interface CheckArgs { + check?: string; + all: boolean; + dryRun: boolean; + payloadDir?: string; + help: boolean; +} + +class UsageError extends Error {} + +function argsParse(args: string[]): CheckArgs { + const parsed: CheckArgs = { all: false, dryRun: false, help: false }; + for (let index = 0; index < args.length; index++) { + const arg = args[index]!; + if (arg === "--help" || arg === "-h") parsed.help = true; + else if (arg === "--all") parsed.all = true; + else if (arg === "--dry-run") parsed.dryRun = true; + else if (arg === "--print-payload") { + const next = args[index + 1]; + parsed.payloadDir = + next && !next.startsWith("--") ? args[++index] : DEFAULT_PAYLOAD_DIR; + } else if (arg.startsWith("-")) + throw new UsageError(`Unknown option: ${arg}`); + else if (parsed.check) throw new UsageError(`Unexpected argument: ${arg}`); + else parsed.check = arg; + } + if (parsed.help) return parsed; + if (parsed.all === Boolean(parsed.check)) + throw new UsageError("Name one check, or pass --all"); + return parsed; +} + +function emptyState(): StateFile { + return promotionStateParse("{}"); +} + +function stateRead(root: string, org: string, warnings: string[]): StateFile { + const path = join(root, `.vapi-state.${org}.json`); + if (!existsSync(path)) { + warnings.push( + `no .vapi-state.${org}.json: references by UUID and credential bindings can't resolve`, + ); + return emptyState(); + } + return promotionStateParse(readFileSync(path, "utf8")); +} + +function payloadFileName(check: string, target: string): string { + return `${check}--${target.replace(/\//g, "--")}.json`; +} + +function resultPrint( + label: string, + result: CheckPayloadResult, + check: CheckDefinition, +): void { + console.log(`\n${label}`); + if (result.body) { + const count = result.body.simulations.length; + console.log( + ` ✅ ${count} simulation${count === 1 ? "" : "s"} × ${check.iterations} iteration${check.iterations === 1 ? "" : "s"} over ${check.transport}, ${(result.bytes / 1024).toFixed(1)} KB`, + ); + } else { + console.log( + ` ❌ ${result.errors.length} problem${result.errors.length === 1 ? "" : "s"}`, + ); + } + for (const warning of result.warnings) console.log(` ⚠️ ${warning}`); + for (const error of result.errors) console.log(` ❌ ${error}`); +} + +async function checkDryRun( + root: string, + check: CheckDefinition, + payloadDir: string | undefined, +): Promise { + if (!existsSync(join(root, "resources", check.org))) { + console.log(`\n${check.name}\n ❌ resources/${check.org}/ does not exist`); + return false; + } + const warnings: string[] = []; + const resources = await orgResourcesRead(root, check.org); + const sourceState = stateRead(root, check.org, warnings); + const runState = + check.runOrg === check.org + ? sourceState + : stateRead(root, check.runOrg, warnings); + const bindingsResolved = await promotionBindingsResolve( + root, + check.org, + check.runOrg, + sourceState, + ); + let ok = true; + for (const target of check.targets) { + const targetLabel = `${target.type}/${target.id}`; + const result = checkPayloadBuild({ + check, + target, + resources, + sourceState, + runState, + bindingsResolved, + }); + result.warnings.unshift(...warnings); + resultPrint(`${check.name} / ${targetLabel}`, result, check); + if (!result.body) { + ok = false; + continue; + } + if (payloadDir) { + const dir = resolve(root, payloadDir); + mkdirSync(dir, { recursive: true }); + const file = join(dir, payloadFileName(check.name, targetLabel)); + writeFileSync(file, `${JSON.stringify(result.body, null, 2)}\n`); + console.log(` 📝 ${relative(root, file)}`); + } + } + return ok; +} + +export async function checkCommandRun( + args = process.argv.slice(2), + root = resolve( + process.env.VAPI_GITOPS_ROOT ?? + fileURLToPath(new URL("..", import.meta.url)), + ), +): Promise { + let parsed: CheckArgs; + try { + parsed = argsParse(args); + } catch (error) { + if (!(error instanceof UsageError)) throw error; + console.error(`❌ ${error.message}\n\n${USAGE}`); + return 2; + } + if (parsed.help) { + console.log(USAGE); + return 0; + } + let config; + try { + config = checksConfigLoad(root); + } catch (error) { + console.error(`❌ ${CHECKS_CONFIG_FILE}: ${(error as Error).message}`); + return 2; + } + if (!config) { + console.log( + `No ${CHECKS_CONFIG_FILE} at the repository root; nothing to check.`, + ); + return 0; + } + const names = parsed.all ? Object.keys(config.checks) : [parsed.check!]; + const missing = names.filter((name) => !config.checks[name]); + if (missing.length > 0) { + console.error( + `❌ No check named ${missing.join(", ")} in ${CHECKS_CONFIG_FILE} (checks: ${Object.keys(config.checks).join(", ")})`, + ); + return 2; + } + if (!parsed.dryRun) { + console.error( + "❌ Live runs aren't available yet: pass --dry-run to build the payloads offline.", + ); + return 2; + } + let ok = true; + for (const name of names) { + if (!(await checkDryRun(root, config.checks[name]!, parsed.payloadDir))) + ok = false; + } + console.log( + ok ? "\n✅ Every payload built." : "\n❌ Some payloads could not be built.", + ); + return ok ? 0 : 2; +} + +const isMainModule = + resolve(process.argv[1] ?? "") === resolve(fileURLToPath(import.meta.url)); +if (isMainModule) { + checkCommandRun().then( + (code) => process.exit(code), + (error: unknown) => { + console.error( + `❌ Check failed: ${error instanceof Error ? error.message : String(error)}`, + ); + process.exit(2); + }, + ); +} diff --git a/src/check-payload-assistant.ts b/src/check-payload-assistant.ts new file mode 100644 index 0000000..acb7cbe --- /dev/null +++ b/src/check-payload-assistant.ts @@ -0,0 +1,325 @@ +// Inline an assistant from its file the way the runtime would assemble the +// stored one, so the check tests what push deploys: +// - tools in runtime order: existing model.tools, then toolIds in order, +// then toolRefs (callAssistantsGet appends toolIds tools after +// model.tools — order changes which tool the model reaches for); +// - knowledgeBase tools stay in toolIds by UUID (the API won't take them +// inline); +// - hook `do[].toolId` → `do[].tool`; +// - artifactPlan.structuredOutputIds, plus structured outputs that link +// the assistant from their own `assistant_ids`, → structuredOutputs; +// - handoffs to squad members by ID → by member name. + +import type { + CheckPayloadContext, + SquadMembers, +} from "./check-payload-refs.ts"; +import { + clone, + isObject, + linkedFieldsStrip, + missingRef, + refClean, + refResolve, + serverFieldsStrip, + sourceUuid, +} from "./check-payload-refs.ts"; + +export interface AssistantInlineArgs { + data: Record; + label: string; + squad?: SquadMembers; + // The assistant's own slug and UUID, for structured outputs that link it + // through their `assistant_ids`. Empty for transient assistants. + selfRefs?: string[]; +} + +function toolPrepare(data: Record): Record { + return linkedFieldsStrip(data); +} + +function toolKey(tool: Record): string | undefined { + const fn = isObject(tool.function) ? tool.function : undefined; + const name = typeof fn?.name === "string" ? fn.name : tool.name; + return typeof name === "string" ? `${String(tool.type)} ${name}` : undefined; +} + +function toolsDuplicateCheck( + ctx: CheckPayloadContext, + tools: Record[], + label: string, +): void { + const seen = new Set(); + for (const tool of tools) { + const key = toolKey(tool); + if (!key) continue; + if (seen.has(key)) + ctx.errors.push( + `${label}: two tools named "${key.split(" ").slice(1).join(" ")}" (type ${String(tool.type)}); stored, both are kept, but inline the second is dropped — rename one`, + ); + seen.add(key); + } +} + +// A toolIds / toolRefs entry: inlined, or kept as a run-org UUID for +// knowledge-base tools. +function toolRefInline( + ctx: CheckPayloadContext, + ref: string, + label: string, + inline: Record[], + keepIds: string[], +): void { + const resolved = refResolve(ctx, "tools", ref); + if (!resolved.resource) { + missingRef(ctx, label, "tools", ref); + return; + } + if (resolved.resource.data.type === "knowledgeBase") { + const uuid = ctx.runState.tools[resolved.slug]?.uuid; + if (!uuid) + ctx.errors.push( + `${label}: knowledgeBase tool "${resolved.slug}" can't be sent inline and has no UUID in the run org's state`, + ); + else keepIds.push(uuid); + return; + } + inline.push(toolPrepare(resolved.resource.data)); +} + +function modelToolsInline( + ctx: CheckPayloadContext, + model: Record, + label: string, +): void { + const inline: Record[] = Array.isArray(model.tools) + ? model.tools.map((tool) => + isObject(tool) ? toolPrepare(tool) : (tool as Record), + ) + : []; + const keepIds: string[] = []; + const toolRefs = Array.isArray(model.toolRefs) ? model.toolRefs : []; + // The toolRefs pin wins over a toolIds entry for the same tool. + const pinned = new Set( + toolRefs + .filter(isObject) + .map((ref) => + typeof ref.toolId === "string" + ? refResolve(ctx, "tools", ref.toolId).slug + : "", + ), + ); + (Array.isArray(model.toolIds) ? model.toolIds : []).forEach((ref, index) => { + if (typeof ref !== "string") { + ctx.errors.push(`${label}.model.toolIds[${index}] must be a string`); + return; + } + if (pinned.has(refResolve(ctx, "tools", ref).slug)) return; + toolRefInline( + ctx, + ref, + `${label}.model.toolIds[${index}]`, + inline, + keepIds, + ); + }); + toolRefs.forEach((ref, index) => { + const at = `${label}.model.toolRefs[${index}]`; + if (!isObject(ref) || typeof ref.toolId !== "string") { + ctx.errors.push(`${at} must have a toolId`); + return; + } + ctx.warnings.push( + `${at}: "${refClean(ref.toolId)}" is inlined from its file; the version pin is ignored`, + ); + toolRefInline(ctx, ref.toolId, at, inline, keepIds); + }); + delete model.toolRefs; + if (inline.length > 0) model.tools = inline; + else delete model.tools; + if (keepIds.length > 0) model.toolIds = keepIds; + else delete model.toolIds; + toolsDuplicateCheck(ctx, inline, `${label}.model.tools`); +} + +// Handoff destinations that name an assistant by ID become member names in +// a squad. Anything else that leaves the target is a build error. +export function handoffsResolve( + ctx: CheckPayloadContext, + tools: unknown, + label: string, + squad?: SquadMembers, +): void { + if (!Array.isArray(tools)) return; + tools.forEach((tool, toolIndex) => { + if (!isObject(tool) || tool.type !== "handoff") return; + if (!Array.isArray(tool.destinations)) return; + tool.destinations.forEach((destination, index) => { + const at = `${label}[${toolIndex}].destinations[${index}]`; + if (!isObject(destination)) return; + if (isObject(destination.assistant)) + destination.assistant = assistantInline(ctx, { + data: destination.assistant, + label: `${at}.assistant`, + squad, + }); + if (isObject(destination.assistantOverrides)) + overridesProcess( + ctx, + destination.assistantOverrides, + `${at}.assistantOverrides`, + squad, + ); + if (typeof destination.assistantId !== "string") return; + const resolved = refResolve(ctx, "assistants", destination.assistantId); + const name = + squad?.get(resolved.slug) ?? + squad?.get(refClean(destination.assistantId)); + if (name !== undefined) { + destination.assistantName = name; + delete destination.assistantId; + return; + } + ctx.errors.push( + squad + ? `${at}: hands off to "${resolved.slug}", which isn't a member of the squad; add it as a member` + : `${at}: an assistant target can't hand off to "${resolved.slug}"; make the target a squad of those assistants`, + ); + }); + }); +} + +function hooksInline( + ctx: CheckPayloadContext, + hooks: unknown, + label: string, +): void { + if (!Array.isArray(hooks)) return; + hooks.forEach((hook, hookIndex) => { + if (!isObject(hook) || !Array.isArray(hook.do)) return; + hook.do.forEach((action, index) => { + if (!isObject(action) || typeof action.toolId !== "string") return; + const at = `${label}[${hookIndex}].do[${index}].toolId`; + const resolved = refResolve(ctx, "tools", action.toolId); + if (!resolved.resource) { + missingRef(ctx, at, "tools", action.toolId); + return; + } + action.tool = toolPrepare(resolved.resource.data); + delete action.toolId; + }); + }); +} + +function structuredOutputsInline( + ctx: CheckPayloadContext, + assistant: Record, + label: string, + selfRefs: string[], +): void { + const plan = isObject(assistant.artifactPlan) ? assistant.artifactPlan : {}; + const inline: Record[] = Array.isArray( + plan.structuredOutputs, + ) + ? [...plan.structuredOutputs] + : []; + const added = new Set(); + const add = (slug: string, data: Record) => { + if (added.has(slug)) return; + added.add(slug); + inline.push(linkedFieldsStrip(data)); + }; + (Array.isArray(plan.structuredOutputIds) + ? plan.structuredOutputIds + : [] + ).forEach((ref, index) => { + if (typeof ref !== "string") return; + const resolved = refResolve(ctx, "structuredOutputs", ref); + if (!resolved.resource) + missingRef( + ctx, + `${label}.artifactPlan.structuredOutputIds[${index}]`, + "structuredOutputs", + ref, + ); + else add(resolved.slug, resolved.resource.data); + }); + if (selfRefs.length > 0) { + for (const resource of ctx.resources.values()) { + if (resource.type !== "structuredOutputs") continue; + const links = resource.data.assistant_ids ?? resource.data.assistantIds; + if (!Array.isArray(links)) continue; + const linked = links.some( + (link) => typeof link === "string" && selfRefs.includes(refClean(link)), + ); + if (linked) add(resource.id, resource.data); + } + } + if (inline.length === 0 && plan.structuredOutputIds === undefined) return; + delete plan.structuredOutputIds; + if (inline.length > 0) plan.structuredOutputs = inline; + assistant.artifactPlan = plan; +} + +// assistantOverrides, membersOverrides and scenario targetOverrides. The +// runtime merges override tools differently (it ignores +// membersOverrides.model.toolIds and rejects toolRefs in overrides), so under +// strict mocks their tools must already be inline. +export function overridesProcess( + ctx: CheckPayloadContext, + overrides: Record, + label: string, + squad?: SquadMembers, +): void { + const model = isObject(overrides.model) ? overrides.model : undefined; + if (ctx.strict) { + for (const key of ["toolIds", "toolRefs"]) { + if (model?.[key] !== undefined) + ctx.errors.push( + `${label}.model.${key}: tools inside overrides must be inline (model.tools or tools:append)`, + ); + } + if (Array.isArray(overrides.hooks)) { + overrides.hooks.forEach((hook, hookIndex) => { + if (!isObject(hook) || !Array.isArray(hook.do)) return; + hook.do.forEach((action, index) => { + if (isObject(action) && action.toolId !== undefined) + ctx.errors.push( + `${label}.hooks[${hookIndex}].do[${index}].toolId: tools inside overrides must be inline`, + ); + }); + }); + } + } + handoffsResolve(ctx, model?.tools, `${label}.model.tools`, squad); + handoffsResolve( + ctx, + overrides["tools:append"], + `${label}.tools:append`, + squad, + ); +} + +export function assistantInline( + ctx: CheckPayloadContext, + args: AssistantInlineArgs, +): Record { + const assistant = serverFieldsStrip(clone(args.data)); + const { label, squad } = args; + if (isObject(assistant.model)) { + modelToolsInline(ctx, assistant.model, label); + handoffsResolve(ctx, assistant.model.tools, `${label}.model.tools`, squad); + } + hooksInline(ctx, assistant.hooks, `${label}.hooks`); + structuredOutputsInline(ctx, assistant, label, args.selfRefs ?? []); + return assistant; +} + +// The slug and source UUID an assistant file is known by. +export function assistantSelfRefs( + ctx: CheckPayloadContext, + slug: string, +): string[] { + const uuid = sourceUuid(ctx, "assistants", slug); + return uuid ? [slug, uuid] : [slug]; +} diff --git a/src/check-payload-refs.ts b/src/check-payload-refs.ts new file mode 100644 index 0000000..38f64a6 --- /dev/null +++ b/src/check-payload-refs.ts @@ -0,0 +1,127 @@ +// Shared pieces of the inline payload builder: the build context, reference +// lookup (local file by slug, or by UUID through the source org's state), +// and field stripping. Pure; no I/O. + +import type { OrgResource } from "./resource-parse.ts"; +import { FOLDER_MAP } from "./resource-parse.ts"; +import type { ResourceType, StateFile } from "./types.ts"; + +// Squad member names keyed by every way a file can refer to the member: its +// slug and its source-org UUID. +export type SquadMembers = Map; + +export interface CheckPayloadContext { + org: string; + resources: Map; + sourceState: StateFile; + runState: StateFile; + // toolMocks: strict — tools inside overrides must already be inline. + strict: boolean; + // Set once a squad target is built, for scenario targetOverrides. + targetMembers?: SquadMembers; + errors: string[]; + warnings: string[]; +} + +export interface ResolvedRef { + slug: string; + resource?: OrgResource; +} + +const UUID_RE = + /^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$/i; + +// Server-managed fields push never sends meaningfully; inline copies drop +// them so the payload reads like a create. +const SERVER_FIELDS = [ + "id", + "orgId", + "createdAt", + "updatedAt", + "analyticsMetadata", + "isDeleted", + "isServerUrlSecretSet", + "_platformDefault", +]; + +// Linkage fields that only mean something on stored resources. +const LINK_FIELDS = [ + "assistant_ids", + "assistantIds", + "workflow_ids", + "workflowIds", +]; + +export function isUuid(value: string): boolean { + return UUID_RE.test(value); +} + +export function isObject(value: unknown): value is Record { + return !!value && typeof value === "object" && !Array.isArray(value); +} + +export function clone(value: T): T { + return structuredClone(value); +} + +// A `##` suffix is a human comment on a reference (`lookup ## CRM tool`), as +// resolver.ts treats it. +export function refClean(ref: string): string { + return ref.split("##")[0]?.trim() ?? ""; +} + +// Find the local file a reference points at: a slug directly, or a UUID +// through the source org's state. +export function refResolve( + ctx: CheckPayloadContext, + type: ResourceType, + ref: string, +): ResolvedRef { + const clean = refClean(ref); + let slug = clean; + if (isUuid(clean)) { + const entry = Object.entries(ctx.sourceState[type]).find( + ([, value]) => value.uuid === clean, + ); + if (!entry) return { slug: clean }; + slug = entry[0]; + } + return { slug, resource: ctx.resources.get(`${type}:${slug}`) }; +} + +// The UUID a slug has in the source org, so stored-side references written +// as UUIDs (squad member maps, `assistant_ids`) can be matched. +export function sourceUuid( + ctx: CheckPayloadContext, + type: ResourceType, + slug: string, +): string | undefined { + return ctx.sourceState[type][slug]?.uuid; +} + +export function serverFieldsStrip( + data: Record, +): Record { + const copy = clone(data); + for (const key of SERVER_FIELDS) delete copy[key]; + return copy; +} + +export function linkedFieldsStrip( + data: Record, +): Record { + const copy = serverFieldsStrip(data); + for (const key of LINK_FIELDS) delete copy[key]; + return copy; +} + +export function missingRef( + ctx: CheckPayloadContext, + label: string, + type: ResourceType, + ref: string, +): void { + ctx.errors.push( + `${label}: "${refClean(ref)}" has no file in resources/${ctx.org}/${FOLDER_MAP[type]}/`, + ); +} diff --git a/src/check-payload.ts b/src/check-payload.ts new file mode 100644 index 0000000..6f5c20e --- /dev/null +++ b/src/check-payload.ts @@ -0,0 +1,538 @@ +// Build the inline `POST /eval/simulation/run` body for one check target +// from the PR branch's files: the target assistant or squad with its tools, +// handoffs and structured outputs, plus every scenario, judge and +// personality the check's suites and simulations name. Nothing is looked up +// on the platform — what's in the files is what runs. +// +// Pure: inputs come from orgResourcesRead and promotionStateParse. Every +// problem is collected, so one dry run reports them all. + +import type { CheckDefinition, CheckTarget } from "./check-config.ts"; +import { + assistantInline, + assistantSelfRefs, + overridesProcess, +} from "./check-payload-assistant.ts"; +import type { + CheckPayloadContext, + SquadMembers, +} from "./check-payload-refs.ts"; +import { + clone, + isObject, + isUuid, + linkedFieldsStrip, + missingRef, + refClean, + refResolve, + serverFieldsStrip, +} from "./check-payload-refs.ts"; +import { credentialForwardMap, replaceCredentialRefs } from "./credentials.ts"; +import type { PromotionBindingsResolved } from "./promotion.ts"; +import { promotionBindingsApply } from "./promotion.ts"; +import type { OrgResource } from "./resource-parse.ts"; +import type { StateFile } from "./types.ts"; + +export interface CheckPayloadInput { + check: CheckDefinition; + target: CheckTarget; + resources: Map; + sourceState: StateFile; + runState: StateFile; + // Phone-number bindings between the org and the run org (from + // promotionBindingsResolve). Credentials resolve through the states. + bindingsResolved?: PromotionBindingsResolved; +} + +export interface CheckSimulationEntry { + type: "simulation"; + name: string; + scenario: Record; + personality?: Record; + personalityId?: string; +} + +export interface CheckRunBody { + simulations: CheckSimulationEntry[]; + target: + | { type: "assistant"; assistant: Record } + | { type: "squad"; squad: Record }; + transport: { provider: "vapi.webchat" | "vapi.websocket" }; + iterations: number; +} + +export interface CheckPayloadResult { + body?: CheckRunBody; + errors: string[]; + warnings: string[]; + bytes: number; +} + +export const MAX_PAYLOAD_BYTES = 4.5 * 1024 * 1024; +const MAX_ENTRY_NAME = 80; +const STOCK_PERSONALITY_RE = /^a0000000-0000-4000-8000-00000000000\d$/; +const AUDIO_TARGET = "messages-with-audio"; + +// Keys whose string values must be platform UUIDs once the body is built. +// A slug left in one of these means a reference the builder couldn't place; +// resolveReferences would silently drop it (improvements.md #31), so the +// check refuses instead. +const REFERENCE_KEYS = new Set([ + "toolId", + "toolIds", + "structuredOutputId", + "structuredOutputIds", + "assistantId", + "assistantIds", + "squadId", + "personalityId", + "scenarioId", + "simulationId", + "simulationIds", + "credentialId", + "credentialIds", + "phoneNumberId", +]); + +// Free-form data: a function parameter called `toolId` isn't a reference. +const FREE_FORM_KEYS = new Set([ + "parameters", + "schema", + "metadata", + "variableValues", +]); + +function squadInline( + ctx: CheckPayloadContext, + squadFile: OrgResource, +): Record | undefined { + const squad = serverFieldsStrip(squadFile.data); + const label = `squads/${squadFile.id}`; + if (!Array.isArray(squad.members) || squad.members.length === 0) { + ctx.errors.push(`${label}: a squad target needs at least one member`); + return undefined; + } + // First pass: every member's name, so handoffs by ID can become names. + const members: SquadMembers = new Map(); + const memberFiles: Array = []; + const names = new Set(); + squad.members.forEach((member, index) => { + const at = `${label}.members[${index}]`; + let file: OrgResource | undefined; + let data: Record | undefined; + if (isObject(member) && typeof member.assistantId === "string") { + const resolved = refResolve(ctx, "assistants", member.assistantId); + if (!resolved.resource) + missingRef(ctx, `${at}.assistantId`, "assistants", member.assistantId); + file = resolved.resource; + data = file?.data; + } else if (isObject(member) && isObject(member.assistant)) { + data = member.assistant; + } + memberFiles.push(file); + if (!data) { + if (!isObject(member) || member.assistantId === undefined) + ctx.errors.push(`${at}: needs an assistantId or an inline assistant`); + return; + } + const name = data.name; + if (typeof name !== "string" || name.trim() === "") { + ctx.errors.push( + `${at}: the member assistant needs a name (handoffs inside a squad target it by name)`, + ); + return; + } + if (names.has(name)) + ctx.errors.push( + `${at}: two members are named "${name}"; member names must be unique`, + ); + names.add(name); + if (file) + for (const ref of assistantSelfRefs(ctx, file.id)) members.set(ref, name); + }); + // Second pass: inline each member. + squad.members = squad.members.map((member, index) => { + if (!isObject(member)) return member; + const at = `${label}.members[${index}]`; + const file = memberFiles[index]; + const next: Record = clone(member); + const data = + file?.data ?? (isObject(member.assistant) ? member.assistant : undefined); + if (data) + next.assistant = assistantInline(ctx, { + data, + label: file ? `assistants/${file.id}` : `${at}.assistant`, + squad: members, + selfRefs: file ? assistantSelfRefs(ctx, file.id) : [], + }); + delete next.assistantId; + if (next.assistantVersion !== undefined) { + ctx.warnings.push( + `${at}: assistantVersion is ignored; the member is built from its file`, + ); + delete next.assistantVersion; + } + if (Array.isArray(next.assistantDestinations)) { + next.assistantDestinations.forEach((destination, destIndex) => { + if (isObject(destination) && destination.assistantId !== undefined) + ctx.errors.push( + `${at}.assistantDestinations[${destIndex}]: legacy assistantDestinations by ID aren't supported inline; convert it to a handoff tool`, + ); + }); + } + if (isObject(next.assistantOverrides)) + overridesProcess( + ctx, + next.assistantOverrides, + `${at}.assistantOverrides`, + members, + ); + return next; + }); + if (isObject(squad.membersOverrides)) + overridesProcess( + ctx, + squad.membersOverrides, + `${label}.membersOverrides`, + members, + ); + ctx.targetMembers = members; + return squad; +} + +function targetBuild( + ctx: CheckPayloadContext, + target: CheckTarget, +): CheckRunBody["target"] | undefined { + const file = ctx.resources.get(`${target.type}:${target.id}`); + if (!file) { + missingRef( + ctx, + `target ${target.type}/${target.id}`, + target.type, + target.id, + ); + return undefined; + } + if (target.type === "assistants") + return { + type: "assistant", + assistant: assistantInline(ctx, { + data: file.data, + label: `assistants/${file.id}`, + selfRefs: assistantSelfRefs(ctx, file.id), + }), + }; + const squad = squadInline(ctx, file); + return squad ? { type: "squad", squad } : undefined; +} + +function simulationIdsCollect( + ctx: CheckPayloadContext, + check: CheckDefinition, +): string[] { + const ids: string[] = []; + const push = (id: string) => { + if (!ids.includes(id)) ids.push(id); + }; + for (const suiteId of check.suites) { + const suite = ctx.resources.get(`simulationSuites:${suiteId}`); + if (!suite) { + missingRef( + ctx, + `checks.${check.name}.suites`, + "simulationSuites", + suiteId, + ); + continue; + } + const simulationIds = suite.data.simulationIds; + if (!Array.isArray(simulationIds) || simulationIds.length === 0) { + ctx.errors.push(`simulations/suites/${suiteId}: lists no simulationIds`); + continue; + } + for (const ref of simulationIds) { + if (typeof ref !== "string") continue; + push(refResolve(ctx, "simulations", ref).slug); + } + } + for (const id of check.simulations) push(id); + return ids; +} + +function judgesInline( + ctx: CheckPayloadContext, + scenario: Record, + label: string, +): void { + if (!Array.isArray(scenario.evaluations)) return; + scenario.evaluations.forEach((evaluation, index) => { + if ( + !isObject(evaluation) || + typeof evaluation.structuredOutputId !== "string" + ) + return; + const at = `${label}.evaluations[${index}].structuredOutputId`; + const resolved = refResolve( + ctx, + "structuredOutputs", + evaluation.structuredOutputId, + ); + if (!resolved.resource) { + missingRef(ctx, at, "structuredOutputs", evaluation.structuredOutputId); + return; + } + evaluation.structuredOutput = linkedFieldsStrip(resolved.resource.data); + delete evaluation.structuredOutputId; + }); +} + +// Chat transport has no audio and runs no scenario hooks. +function chatRulesCheck( + ctx: CheckPayloadContext, + scenario: Record, + label: string, +): void { + const evaluations = Array.isArray(scenario.evaluations) + ? scenario.evaluations + : []; + const isAudio = (evaluation: unknown) => + isObject(evaluation) && + isObject(evaluation.structuredOutput) && + evaluation.structuredOutput.target === AUDIO_TARGET; + evaluations.forEach((evaluation, index) => { + if (isAudio(evaluation)) + ctx.errors.push( + `${label}.evaluations[${index}]: a ${AUDIO_TARGET} judge needs transport: voice`, + ); + }); + const textRequired = evaluations.some( + (evaluation) => + isObject(evaluation) && + evaluation.required !== false && + !isAudio(evaluation), + ); + if (!textRequired) + ctx.errors.push( + `${label}: needs at least one required text judge (a scenario with none can't fail)`, + ); + if (Array.isArray(scenario.hooks) && scenario.hooks.length > 0) + ctx.errors.push( + `${label}: scenario hooks don't run over chat; remove them or use transport: voice`, + ); +} + +function entryName(name: string, used: Set): string { + let candidate = name.slice(0, MAX_ENTRY_NAME); + for (let n = 2; used.has(candidate); n++) { + const suffix = ` (${n})`; + candidate = name.slice(0, MAX_ENTRY_NAME - suffix.length) + suffix; + } + used.add(candidate); + return candidate; +} + +function personalityBuild( + ctx: CheckPayloadContext, + ref: unknown, + label: string, +): Pick | undefined { + if (typeof ref !== "string") { + ctx.errors.push(`${label}: needs a personalityId`); + return undefined; + } + const resolved = refResolve(ctx, "personalities", ref); + if (resolved.resource) { + const personality = serverFieldsStrip(resolved.resource.data); + if (isObject(personality.assistant)) + personality.assistant = assistantInline(ctx, { + data: personality.assistant, + label: `simulations/personalities/${resolved.slug}.assistant`, + }); + return { personality }; + } + if (STOCK_PERSONALITY_RE.test(refClean(ref))) + return { personalityId: refClean(ref) }; + missingRef(ctx, label, "personalities", ref); + return undefined; +} + +function entriesBuild( + ctx: CheckPayloadContext, + check: CheckDefinition, +): CheckSimulationEntry[] { + const entries: CheckSimulationEntry[] = []; + const used = new Set(); + for (const id of simulationIdsCollect(ctx, check)) { + const label = `simulations/tests/${id}`; + const simulation = ctx.resources.get(`simulations:${id}`); + if (!simulation) { + missingRef(ctx, `checks.${check.name}`, "simulations", id); + continue; + } + const scenarioRef = simulation.data.scenarioId; + const scenarioFile = + typeof scenarioRef === "string" + ? refResolve(ctx, "scenarios", scenarioRef).resource + : undefined; + if (!scenarioFile) { + if (typeof scenarioRef === "string") + missingRef(ctx, `${label}.scenarioId`, "scenarios", scenarioRef); + else ctx.errors.push(`${label}: needs a scenarioId`); + continue; + } + const scenarioLabel = `simulations/scenarios/${scenarioFile.id}`; + const scenario = serverFieldsStrip(scenarioFile.data); + judgesInline(ctx, scenario, scenarioLabel); + if (isObject(scenario.targetOverrides)) + overridesProcess( + ctx, + scenario.targetOverrides, + `${scenarioLabel}.targetOverrides`, + ctx.targetMembers, + ); + if (check.transport === "chat") + chatRulesCheck(ctx, scenario, scenarioLabel); + const personality = personalityBuild( + ctx, + simulation.data.personalityId, + `${label}.personalityId`, + ); + if (!personality) continue; + const name = + typeof simulation.data.name === "string" ? simulation.data.name : id; + entries.push({ + type: "simulation", + name: entryName(name, used), + scenario, + ...personality, + }); + } + if (entries.length === 0 && ctx.errors.length === 0) + ctx.errors.push(`checks.${check.name}: selects no simulations`); + return entries; +} + +function credentialsBind( + ctx: CheckPayloadContext, + input: CheckPayloadInput, + value: unknown, +): unknown { + const resolved: PromotionBindingsResolved = input.bindingsResolved ?? { + credentialReverse: new Map( + Object.entries(ctx.sourceState.credentials).map(([alias, entry]) => [ + entry.uuid, + alias, + ]), + ), + sourcePhones: new Map(), + targetPhones: new Map(), + }; + try { + const bound = promotionBindingsApply( + value, + input.check.bindings, + resolved, + ctx.runState, + ); + return replaceCredentialRefs(bound, credentialForwardMap(ctx.runState)); + } catch (error) { + ctx.errors.push( + `run org ${input.check.runOrg}: ${(error as Error).message}`, + ); + return value; + } +} + +function leftoverReferencesCheck( + ctx: CheckPayloadContext, + value: unknown, + path: string, +): void { + if (Array.isArray(value)) { + value.forEach((item, index) => + leftoverReferencesCheck(ctx, item, `${path}[${index}]`), + ); + return; + } + if (!isObject(value)) return; + for (const [key, child] of Object.entries(value)) { + if (FREE_FORM_KEYS.has(key)) continue; + const at = `${path}.${key}`; + if (REFERENCE_KEYS.has(key)) { + const values = Array.isArray(child) ? child : [child]; + for (const item of values) { + if (typeof item === "string" && !isUuid(item)) + ctx.errors.push( + `${at}: "${item}" is still a name, not a UUID; it couldn't be resolved for the run org`, + ); + } + continue; + } + leftoverReferencesCheck(ctx, child, at); + } +} + +function autoHandoffNamesWarn( + ctx: CheckPayloadContext, + body: CheckRunBody, +): void { + const serialized = JSON.stringify(body); + if (!serialized.includes("handoff_to_")) return; + let autoNamed = false; + const visit = (value: unknown) => { + if (Array.isArray(value)) return value.forEach(visit); + if (!isObject(value)) return; + const fn = isObject(value.function) ? value.function : undefined; + if (value.type === "handoff" && typeof fn?.name !== "string") + autoNamed = true; + Object.values(value).forEach(visit); + }; + visit(body.target); + if (autoNamed) + ctx.warnings.push( + 'text mentions "handoff_to_…", but a handoff tool has no explicit function.name; generated names differ inline (handoff_to_) and stored (handoff_to_), so give it an explicit function.name', + ); +} + +export function checkPayloadBuild( + input: CheckPayloadInput, +): CheckPayloadResult { + const ctx: CheckPayloadContext = { + org: input.check.org, + resources: input.resources, + sourceState: input.sourceState, + runState: input.runState, + strict: input.check.toolMocks === "strict", + errors: [], + warnings: [], + }; + const target = targetBuild(ctx, input.target); + const simulations = entriesBuild(ctx, input.check); + if (!target) return { errors: ctx.errors, warnings: ctx.warnings, bytes: 0 }; + const body: CheckRunBody = { + simulations: credentialsBind( + ctx, + input, + simulations, + ) as CheckSimulationEntry[], + target: credentialsBind(ctx, input, target) as CheckRunBody["target"], + transport: { + provider: + input.check.transport === "voice" ? "vapi.websocket" : "vapi.webchat", + }, + iterations: input.check.iterations, + }; + // A backstop for reference fields the builder doesn't handle; run only + // when nothing else failed, so a known problem isn't reported twice. + if (ctx.errors.length === 0) leftoverReferencesCheck(ctx, body, "body"); + autoHandoffNamesWarn(ctx, body); + const bytes = Buffer.byteLength(JSON.stringify(body)); + if (bytes > MAX_PAYLOAD_BYTES) + ctx.errors.push( + `payload is ${(bytes / 1024 / 1024).toFixed(1)} MB; the limit is 4.5 MB`, + ); + if (ctx.errors.length > 0) + return { errors: ctx.errors, warnings: ctx.warnings, bytes }; + return { body, errors: [], warnings: ctx.warnings, bytes }; +} diff --git a/tests/check-cmd.test.ts b/tests/check-cmd.test.ts new file mode 100644 index 0000000..98813fd --- /dev/null +++ b/tests/check-cmd.test.ts @@ -0,0 +1,132 @@ +import assert from "node:assert/strict"; +import { + cpSync, + existsSync, + mkdtempSync, + readFileSync, + rmSync, + writeFileSync, +} from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import test from "node:test"; +import { fileURLToPath } from "node:url"; +import { checkCommandRun } from "../src/check-cmd.ts"; + +const PARITY_ROOT = fileURLToPath( + new URL("./fixtures/check-parity/", import.meta.url), +); + +async function run( + args: string[], + root: string, +): Promise<{ code: number; out: string }> { + const lines: string[] = []; + const log = console.log; + const error = console.error; + console.log = (...parts: unknown[]) => lines.push(parts.join(" ")); + console.error = (...parts: unknown[]) => lines.push(parts.join(" ")); + try { + return { code: await checkCommandRun(args, root), out: lines.join("\n") }; + } finally { + console.log = log; + console.error = error; + } +} + +function parityCopy(): string { + const root = mkdtempSync(join(tmpdir(), "check-cmd-")); + cpSync(PARITY_ROOT, root, { recursive: true }); + return root; +} + +test("a dry run of the parity fixture builds the payload and writes it with --print-payload", async () => { + const root = parityCopy(); + try { + const result = await run(["core", "--dry-run", "--print-payload"], root); + const file = join(root, "tmp/check-payloads/core--squads--dental.json"); + const body = JSON.parse(readFileSync(file, "utf8")); + assert.deepEqual( + [ + result.code, + result.out.includes("✅ 3 simulations × 1 iteration over chat"), + body.simulations.length, + ], + [0, true, 3], + ); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("--all runs every check; a build error exits 2 and names the problem", async () => { + const root = parityCopy(); + try { + writeFileSync( + join(root, "vapi-checks.yml"), + "version: 1\nchecks:\n core:\n org: parity\n targets: [squads/dental]\n suites: [core]\n broken:\n org: parity\n targets: [assistants/receptionist]\n suites: [core]\n", + ); + const result = await run(["--all", "--dry-run"], root); + assert.deepEqual( + [ + result.code, + result.out.includes("core / squads/dental\n ✅"), + result.out.includes( + `❌ assistants/receptionist.model.tools[2].destinations[0]: an assistant target can't hand off to "scheduler"`, + ), + ], + [2, true, true], + ); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("no vapi-checks.yml means nothing to check, and exits 0", async () => { + const root = mkdtempSync(join(tmpdir(), "check-cmd-")); + try { + assert.deepEqual(await run(["--all", "--dry-run"], root), { + code: 0, + out: "No vapi-checks.yml at the repository root; nothing to check.", + }); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("usage, config and selection errors exit 2", async () => { + const root = parityCopy(); + try { + const codes = [ + (await run([], root)).code, + (await run(["core", "--all", "--dry-run"], root)).code, + (await run(["core", "--bogus"], root)).code, + (await run(["nope", "--dry-run"], root)).code, + // Live runs aren't available in this command yet. + (await run(["core"], root)).code, + ]; + writeFileSync(join(root, "vapi-checks.yml"), "version: 2\n"); + codes.push((await run(["core", "--dry-run"], root)).code); + assert.deepEqual(codes, [2, 2, 2, 2, 2, 2]); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); + +test("a missing state file is a warning, not an error", async () => { + const root = parityCopy(); + try { + rmSync(join(root, ".vapi-state.parity.json")); + const result = await run(["core", "--dry-run"], root); + assert.deepEqual( + [ + result.code, + result.out.includes("⚠️ no .vapi-state.parity.json"), + existsSync(join(root, "tmp")), + ], + [0, true, false], + ); + } finally { + rmSync(root, { recursive: true, force: true }); + } +}); diff --git a/tests/check-payload.test.ts b/tests/check-payload.test.ts new file mode 100644 index 0000000..6209e91 --- /dev/null +++ b/tests/check-payload.test.ts @@ -0,0 +1,572 @@ +import assert from "node:assert/strict"; +import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { dirname, join } from "node:path"; +import test from "node:test"; +import { fileURLToPath } from "node:url"; +import type { CheckDefinition, CheckTarget } from "../src/check-config.ts"; +import { checksConfigParse } from "../src/check-config.ts"; +import type { CheckPayloadResult } from "../src/check-payload.ts"; +import { checkPayloadBuild } from "../src/check-payload.ts"; +import { promotionStateParse } from "../src/promotion.ts"; +import { orgResourcesRead } from "../src/resource-parse.ts"; +import type { StateFile } from "../src/types.ts"; + +const PARITY_ROOT = fileURLToPath( + new URL("./fixtures/check-parity/", import.meta.url), +); +const STOCK_PERSONALITY = "a0000000-0000-4000-8000-000000000001"; +const UUID_A = "11111111-1111-4111-8111-111111111111"; +const UUID_B = "22222222-2222-4222-8222-222222222222"; +const UUID_CRED = "33333333-3333-4333-8333-333333333333"; + +// One passing scenario + simulation + suite, so tests only add what they test. +const BASE_TESTS: Record = { + "structuredOutputs/ok.yml": "name: ok\nschema:\n type: boolean\n", + "simulations/scenarios/s1.yml": + "name: S1\ninstructions: Ask a question.\nevaluations:\n - structuredOutputId: ok\n comparator: '='\n value: true\n required: true\n", + "simulations/tests/t1.yml": `name: T1\npersonalityId: ${STOCK_PERSONALITY}\nscenarioId: s1\n`, + "simulations/suites/core.yml": "name: Core\nsimulationIds: [t1]\n", +}; + +interface BuildArgs { + files: Record; + target?: string; + check?: string; + sourceState?: Partial; + runState?: Partial; +} + +function state(entries: Partial = {}): StateFile { + return { ...promotionStateParse("{}"), ...entries }; +} + +function checkDefinition(target: string, extra = ""): CheckDefinition { + return checksConfigParse( + `version: 1\nchecks:\n core:\n org: acme\n targets: [${target}]\n suites: [core]\n${extra}`, + ).checks.core!; +} + +async function build(args: BuildArgs): Promise { + const root = mkdtempSync(join(tmpdir(), "check-payload-")); + try { + for (const [path, content] of Object.entries({ + ...BASE_TESTS, + ...args.files, + })) { + const full = join(root, "resources", "acme", path); + mkdirSync(dirname(full), { recursive: true }); + writeFileSync(full, content); + } + const check = checkDefinition( + args.target ?? "assistants/main", + args.check ?? "", + ); + const resources = await orgResourcesRead(root, "acme"); + const sourceState = state(args.sourceState); + return checkPayloadBuild({ + check, + target: check.targets[0] as CheckTarget, + resources, + sourceState, + runState: args.runState ? state(args.runState) : sourceState, + }); + } finally { + rmSync(root, { recursive: true, force: true }); + } +} + +function assistantOf(result: CheckPayloadResult): Record { + const target = result.body?.target; + assert.ok(target?.type === "assistant", result.errors.join("\n")); + return target.assistant; +} + +const MAIN = (model: string, extra = "") => + `name: Main\nmodel:\n provider: openai\n model: gpt-4o\n${model}${extra}`; + +test("parity squad: builds from files with stored tool order, member handoffs and inline judges", async () => { + const check = checksConfigParse( + `version: 1\nchecks:\n core:\n org: parity\n targets: [squads/dental]\n suites: [core]\n`, + ).checks.core!; + const sourceState = state(); + const result = checkPayloadBuild({ + check, + target: check.targets[0]!, + resources: await orgResourcesRead(PARITY_ROOT, "parity"), + sourceState, + runState: sourceState, + }); + const squad = + result.body?.target.type === "squad" ? result.body.target.squad : {}; + const members = squad.members as Array<{ + assistant: { + name: string; + model: { + tools: Array>; + messages: Array<{ role: string; content: string }>; + }; + }; + }>; + const handoff = members[0]!.assistant.model.tools[2]!; + assert.deepEqual( + { + errors: result.errors, + warnings: result.warnings, + tools: members.map((m) => + m.assistant.model.tools.map( + (t) => (t.function as { name?: string } | undefined)?.name ?? t.type, + ), + ), + handoff: handoff.destinations, + prompt: members[1]!.assistant.model.messages[0]!.content.slice(0, 40), + entries: result + .body!.simulations.map((e) => [ + e.name, + e.personality?.name, + e.scenario.evaluations, + ]) + .map(([name, personality, evaluations]) => [ + name, + personality, + (evaluations as Array<{ structuredOutput: { name: string } }>).map( + (e) => e.structuredOutput.name, + ), + ]), + transport: result.body!.transport, + }, + { + errors: [], + warnings: [], + // Runtime order: model.tools first, then toolIds in order. + tools: [ + ["endCall", "lookup_patient", "handoff"], + ["endCall", "check_availability"], + ], + handoff: [ + { + type: "assistant", + description: "Books and changes appointments, checks availability.", + assistantName: "Scheduler", + }, + ], + prompt: "You are the scheduler for Bright Smile D", + entries: [ + [ + "S1 book cleaning", + "Dental caller", + ["booking-confirmed", "only-real-slots"], + ], + ["S2 hours question", "Dental caller", ["hours-correct", "no-booking"]], + [ + "S3 alternative slot", + "Dental caller", + ["offered-alternative", "booked-wednesday"], + ], + ], + transport: { provider: "vapi.webchat" }, + }, + ); +}); + +test("a .md assistant's body is its only system message, exactly as push loads it", async () => { + const result = await build({ + files: { + "assistants/main.md": + "---\nname: Main\nmodel:\n provider: openai\n messages:\n - role: system\n content: stale\n---\n\nThe real prompt.\n", + }, + }); + assert.deepEqual(assistantOf(result).model, { + provider: "openai", + messages: [{ role: "system", content: "The real prompt." }], + }); +}); + +test("toolIds that don't name a local file fail the build", async () => { + const result = await build({ + files: { "assistants/main.yml": MAIN(' toolIds: ["missing ## gone"]\n') }, + }); + assert.deepEqual(result.errors, [ + 'assistants/main.model.toolIds[0]: "missing" has no file in resources/acme/tools/', + ]); +}); + +test("toolIds resolve by UUID through the source state, and server fields are stripped", async () => { + const result = await build({ + files: { + "assistants/main.yml": MAIN(` toolIds: [${UUID_A}]\n`), + "tools/lookup.yml": `id: ${UUID_A}\norgId: org\nassistant_ids: [main]\ntype: function\nfunction:\n name: lookup\n`, + }, + sourceState: { tools: { lookup: { uuid: UUID_A } } }, + }); + assert.deepEqual((assistantOf(result).model as { tools: unknown }).tools, [ + { type: "function", function: { name: "lookup" } }, + ]); +}); + +test("toolRefs are inlined after toolIds with a warning, and the pin wins over a toolIds duplicate", async () => { + const result = await build({ + files: { + "assistants/main.yml": MAIN( + " toolIds: [a, b]\n toolRefs:\n - toolId: a\n version: 3\n", + ), + "tools/a.yml": "type: function\nfunction:\n name: a\n", + "tools/b.yml": "type: function\nfunction:\n name: b\n", + }, + }); + const model = assistantOf(result).model as { + tools: Array<{ function: { name: string } }>; + toolRefs?: unknown; + }; + assert.deepEqual( + [model.tools.map((t) => t.function.name), model.toolRefs, result.warnings], + [ + ["b", "a"], + undefined, + [ + 'assistants/main.model.toolRefs[0]: "a" is inlined from its file; the version pin is ignored', + ], + ], + ); +}); + +test("knowledgeBase tools stay in toolIds by their run-org UUID", async () => { + const files = { + "assistants/main.yml": MAIN(" toolIds: [kb, lookup]\n"), + "tools/kb.yml": "type: knowledgeBase\nname: docs\n", + "tools/lookup.yml": "type: function\nfunction:\n name: lookup\n", + }; + const kept = await build({ + files, + sourceState: { tools: { kb: { uuid: UUID_B } } }, + }); + const missing = await build({ files }); + assert.deepEqual( + [ + (assistantOf(kept).model as { toolIds: string[] }).toolIds, + missing.errors, + ], + [ + [UUID_B], + [ + `assistants/main.model.toolIds[0]: knowledgeBase tool "kb" can't be sent inline and has no UUID in the run org's state`, + ], + ], + ); +}); + +test("two inlined tools with the same type and name fail the build", async () => { + const result = await build({ + files: { + "assistants/main.yml": MAIN( + " tools:\n - type: function\n function:\n name: lookup\n toolIds: [lookup]\n", + ), + "tools/lookup.yml": "type: function\nfunction:\n name: lookup\n", + }, + }); + assert.deepEqual(result.errors, [ + 'assistants/main.model.tools: two tools named "lookup" (type function); stored, both are kept, but inline the second is dropped — rename one', + ]); +}); + +test("hook do[].toolId becomes an inline tool", async () => { + const result = await build({ + files: { + "assistants/main.yml": MAIN( + "", + "hooks:\n - on: call.ending\n do:\n - type: tool\n toolId: notify\n", + ), + "tools/notify.yml": "type: function\nfunction:\n name: notify\n", + }, + }); + assert.deepEqual(assistantOf(result).hooks, [ + { + on: "call.ending", + do: [ + { + type: "tool", + tool: { type: "function", function: { name: "notify" } }, + }, + ], + }, + ]); +}); + +test("structured outputs from artifactPlan and from their own assistant_ids are inlined once", async () => { + const result = await build({ + files: { + "assistants/main.yml": MAIN( + "", + "artifactPlan:\n structuredOutputIds: [summary]\n", + ), + "structuredOutputs/summary.yml": + "name: summary\nassistant_ids: [main]\nschema:\n type: string\n", + "structuredOutputs/sentiment.yml": `name: sentiment\nassistant_ids: [${UUID_A}]\nschema:\n type: string\n`, + "structuredOutputs/other.yml": + "name: other\nassistant_ids: [someone-else]\nschema:\n type: string\n", + }, + sourceState: { assistants: { main: { uuid: UUID_A } } }, + }); + assert.deepEqual(assistantOf(result).artifactPlan, { + structuredOutputs: [ + { name: "summary", schema: { type: "string" } }, + { name: "sentiment", schema: { type: "string" } }, + ], + }); +}); + +test("an assistant target that hands off to another assistant fails the build", async () => { + const result = await build({ + files: { + "assistants/main.yml": MAIN(" toolIds: [to-other]\n"), + "tools/to-other.yml": + "type: handoff\ndestinations:\n - type: assistant\n assistantId: other\n", + }, + }); + assert.deepEqual(result.errors, [ + 'assistants/main.model.tools[0].destinations[0]: an assistant target can\'t hand off to "other"; make the target a squad of those assistants', + ]); +}); + +test("squad members are inlined by file, UUID or inline assistant, and handoffs become member names", async () => { + const result = await build({ + target: "squads/team", + files: { + "squads/team.yml": `name: Team\nmembers:\n - assistantId: a\n assistantVersion: 4\n - assistantId: ${UUID_B}\n - assistant:\n name: C\n model:\n provider: openai\n tools:\n - type: handoff\n destinations:\n - type: assistant\n assistantId: a\n`, + "assistants/a.yml": + "name: A\nmodel:\n provider: openai\n toolIds: [to-b]\n", + "assistants/b.yml": "name: B\nmodel:\n provider: openai\n", + "tools/to-b.yml": `type: handoff\ndestinations:\n - type: assistant\n assistantId: ${UUID_B}\n`, + }, + sourceState: { assistants: { b: { uuid: UUID_B } } }, + }); + const squad = + result.body?.target.type === "squad" ? result.body.target.squad : {}; + const members = squad.members as Array<{ + assistant: { + name: string; + model: { tools?: Array<{ destinations: unknown }> }; + }; + assistantId?: string; + assistantVersion?: number; + }>; + assert.deepEqual( + { + errors: result.errors, + warnings: result.warnings, + names: members.map((m) => m.assistant.name), + ids: members.map((m) => m.assistantId ?? m.assistantVersion), + handoffs: [ + members[0]!.assistant.model.tools![0]!.destinations, + members[2]!.assistant.model.tools![0]!.destinations, + ], + }, + { + errors: [], + warnings: [ + "squads/team.members[0]: assistantVersion is ignored; the member is built from its file", + ], + names: ["A", "B", "C"], + ids: [undefined, undefined, undefined], + handoffs: [ + [{ type: "assistant", assistantName: "B" }], + [{ type: "assistant", assistantName: "A" }], + ], + }, + ); +}); + +test("squad problems are all reported: unnamed and duplicate members, outside handoffs, legacy destinations", async () => { + const result = await build({ + target: "squads/team", + files: { + "squads/team.yml": + "name: Team\nmembers:\n - assistantId: a\n - assistantId: a2\n - assistantId: nameless\n - assistantId: b\n assistantDestinations:\n - assistantId: a\n", + "assistants/a.yml": + "name: A\nmodel:\n provider: openai\n toolIds: [to-outside]\n", + "assistants/a2.yml": "name: A\nmodel:\n provider: openai\n", + "assistants/nameless.yml": "model:\n provider: openai\n", + "assistants/b.yml": "name: B\nmodel:\n provider: openai\n", + "assistants/outside.yml": "name: Outside\n", + "tools/to-outside.yml": + "type: handoff\ndestinations:\n - type: assistant\n assistantId: outside\n", + }, + }); + assert.deepEqual(result.errors, [ + 'squads/team.members[1]: two members are named "A"; member names must be unique', + "squads/team.members[2]: the member assistant needs a name (handoffs inside a squad target it by name)", + 'assistants/a.model.tools[0].destinations[0]: hands off to "outside", which isn\'t a member of the squad; add it as a member', + "squads/team.members[3].assistantDestinations[0]: legacy assistantDestinations by ID aren't supported inline; convert it to a handoff tool", + ]); +}); + +test("under strict mocks, tools referenced by ID inside overrides fail; with toolMocks off they pass", async () => { + const files = { + "squads/team.yml": + "name: Team\nmembers:\n - assistantId: a\nmembersOverrides:\n model:\n toolIds: [lookup]\n", + "assistants/a.yml": "name: A\nmodel:\n provider: openai\n", + "tools/lookup.yml": "type: function\nfunction:\n name: lookup\n", + }; + const strict = await build({ target: "squads/team", files }); + const off = await build({ + target: "squads/team", + files, + check: " toolMocks: off\n", + }); + assert.deepEqual( + [strict.errors, off.errors.filter((e) => !e.includes("still a name"))], + [ + [ + "squads/team.membersOverrides.model.toolIds: tools inside overrides must be inline (model.tools or tools:append)", + ], + [], + ], + ); +}); + +test("credentials bind by name to the run org's UUIDs; omit drops them; a missing bind fails", async () => { + const files = { + "assistants/main.yml": MAIN( + "", + "voice:\n provider: 11labs\n credentialId: eleven\n", + ), + }; + const bound = await build({ + files, + runState: { credentials: { eleven: { uuid: UUID_CRED } } }, + }); + const omitted = await build({ + files, + check: " bindings:\n credentials:\n default: omit\n", + }); + const missing = await build({ files }); + assert.deepEqual( + [assistantOf(bound).voice, assistantOf(omitted).voice, missing.errors], + [ + { provider: "11labs", credentialId: UUID_CRED }, + { provider: "11labs" }, + [ + 'run org acme: Credential binding "eleven" is required in the target state', + ], + ], + ); +}); + +test("a reference left as a name anywhere fails the build; free-form parameters are ignored", async () => { + const result = await build({ + files: { + "assistants/main.yml": MAIN( + " tools:\n - type: function\n function:\n name: f\n parameters:\n type: object\n properties:\n toolId: { type: string }\n", + "transferPlan:\n assistantId: somewhere\n", + ), + }, + }); + assert.deepEqual(result.errors, [ + 'body.target.assistant.transferPlan.assistantId: "somewhere" is still a name, not a UUID; it couldn\'t be resolved for the run org', + ]); +}); + +test("over chat: audio judges, scenarios without a required text judge, and scenario hooks fail", async () => { + const files = { + "assistants/main.yml": MAIN(""), + "structuredOutputs/audio.yml": + "name: audio\ntarget: messages-with-audio\nschema:\n type: boolean\n", + "simulations/scenarios/s1.yml": + "name: S1\ninstructions: Hi.\nevaluations:\n - structuredOutputId: audio\n comparator: '='\n value: true\n - structuredOutputId: ok\n comparator: '='\n value: true\n required: false\nhooks:\n - on: call.started\n do: []\n", + }; + const chat = await build({ files }); + const voice = await build({ files, check: " transport: voice\n" }); + assert.deepEqual( + [chat.errors, voice.errors, voice.body?.transport], + [ + [ + "simulations/scenarios/s1.evaluations[0]: a messages-with-audio judge needs transport: voice", + "simulations/scenarios/s1: needs at least one required text judge (a scenario with none can't fail)", + "simulations/scenarios/s1: scenario hooks don't run over chat; remove them or use transport: voice", + ], + [], + { provider: "vapi.websocket" }, + ], + ); +}); + +test("simulations: stock personalities pass by ID, local ones inline, names stay unique within 80 characters", async () => { + const long = "x".repeat(90); + const result = await build({ + files: { + "assistants/main.yml": MAIN(""), + "simulations/personalities/calm.yml": + "name: Calm\nassistant:\n model:\n provider: openai\n", + "simulations/tests/t1.yml": `name: ${long}\npersonalityId: ${STOCK_PERSONALITY}\nscenarioId: s1\n`, + "simulations/tests/t2.yml": `name: ${long}\npersonalityId: calm\nscenarioId: s1 ## same scenario\n`, + "simulations/suites/core.yml": + "name: Core\nsimulationIds: [t1, t2, t1]\n", + }, + }); + const entries = result.body!.simulations; + assert.deepEqual( + entries.map((e) => [ + e.name.length, + e.name.slice(-4), + e.personalityId, + e.personality?.name, + ]), + [ + [80, "xxxx", STOCK_PERSONALITY, undefined], + [80, " (2)", undefined, "Calm"], + ], + ); +}); + +test("missing suites, simulations, scenarios and personalities are each named", async () => { + const result = await build({ + files: { + "assistants/main.yml": MAIN(""), + "simulations/tests/t2.yml": + "name: T2\npersonalityId: nobody\nscenarioId: s1\n", + "simulations/tests/t3.yml": `name: T3\npersonalityId: ${STOCK_PERSONALITY}\nscenarioId: nowhere\n`, + "simulations/suites/core.yml": + "name: Core\nsimulationIds: [t2, t3, ghost]\n", + }, + check: " simulations: [t1]\n", + }); + assert.deepEqual(result.errors, [ + 'simulations/tests/t2.personalityId: "nobody" has no file in resources/acme/simulations/personalities/', + 'simulations/tests/t3.scenarioId: "nowhere" has no file in resources/acme/simulations/scenarios/', + 'checks.core: "ghost" has no file in resources/acme/simulations/tests/', + ]); +}); + +test("a missing target is named", async () => { + const result = await build({ files: {}, target: "squads/nope" }); + assert.deepEqual(result.errors, [ + 'target squads/nope: "nope" has no file in resources/acme/squads/', + ]); +}); + +test("auto-named handoffs warn when text mentions handoff_to_", async () => { + const result = await build({ + target: "squads/team", + files: { + "squads/team.yml": + "name: Team\nmembers:\n - assistantId: a\n - assistantId: b\n", + "assistants/a.md": + "---\nname: A\nmodel:\n provider: openai\n toolIds: [to-b]\n---\nCall handoff_to_B when booking.\n", + "assistants/b.yml": "name: B\nmodel:\n provider: openai\n", + "tools/to-b.yml": + "type: handoff\ndestinations:\n - type: assistant\n assistantId: b\n", + }, + }); + assert.deepEqual(result.warnings, [ + 'text mentions "handoff_to_…", but a handoff tool has no explicit function.name; generated names differ inline (handoff_to_) and stored (handoff_to_), so give it an explicit function.name', + ]); +}); + +test("payloads above 4.5 MB fail", async () => { + const result = await build({ + files: { + "assistants/main.md": `---\nname: Main\nmodel:\n provider: openai\n---\n${"p".repeat(5 * 1024 * 1024)}\n`, + }, + }); + assert.deepEqual(result.errors, ["payload is 5.0 MB; the limit is 4.5 MB"]); +}); diff --git a/tests/fixtures/check-parity/.vapi-state.parity.json b/tests/fixtures/check-parity/.vapi-state.parity.json new file mode 100644 index 0000000..0967ef4 --- /dev/null +++ b/tests/fixtures/check-parity/.vapi-state.parity.json @@ -0,0 +1 @@ +{} diff --git a/tests/fixtures/check-parity/resources/parity/assistants/receptionist.md b/tests/fixtures/check-parity/resources/parity/assistants/receptionist.md new file mode 100644 index 0000000..defbce8 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/assistants/receptionist.md @@ -0,0 +1,18 @@ +--- +name: Receptionist +firstMessage: Thanks for calling Bright Smile Dental, how can I help? +model: + provider: openai + model: gpt-4.1 + temperature: 0 + tools: + - type: endCall + toolIds: + - lookup-patient + - handoff-to-scheduler +--- + +You are the receptionist for Bright Smile Dental, 123 Main St. Clinic hours: Monday to Friday, 8am to 5pm; closed weekends. +First, ask for the caller's phone number and call lookup_patient with it. +If the caller wants to book or change an appointment or check availability, hand off to the Scheduler. Do not book appointments yourself. +Answer general questions (hours, address) yourself, briefly. Keep replies short. diff --git a/tests/fixtures/check-parity/resources/parity/assistants/scheduler.md b/tests/fixtures/check-parity/resources/parity/assistants/scheduler.md new file mode 100644 index 0000000..aaebf24 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/assistants/scheduler.md @@ -0,0 +1,15 @@ +--- +name: Scheduler +model: + provider: openai + model: gpt-4.1 + temperature: 0 + tools: + - type: endCall + toolIds: + - check-availability +--- + +You are the scheduler for Bright Smile Dental. Call check_availability for the date and service the caller wants, and offer only the slots it returns. +If the requested date has no slots, say so and offer the alternatives it returns. +When the caller accepts a slot, call book_appointment, then confirm the booked date and time. Keep replies short. diff --git a/tests/fixtures/check-parity/resources/parity/simulations/personalities/dental-caller.yml b/tests/fixtures/check-parity/resources/parity/simulations/personalities/dental-caller.yml new file mode 100644 index 0000000..16023cb --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/simulations/personalities/dental-caller.yml @@ -0,0 +1,9 @@ +name: Dental caller +assistant: + model: + provider: openai + model: gpt-4.1-mini + temperature: 0 + messages: + - role: system + content: You are a patient calling a dental clinic. Follow your scenario exactly, speak only as the patient, answer what you're asked briefly, and don't invent facts beyond the scenario. Let the clinic finish its steps; a transfer between clinic staff is normal. diff --git a/tests/fixtures/check-parity/resources/parity/simulations/scenarios/s1-book-cleaning.yml b/tests/fixtures/check-parity/resources/parity/simulations/scenarios/s1-book-cleaning.yml new file mode 100644 index 0000000..7b09dc3 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/simulations/scenarios/s1-book-cleaning.yml @@ -0,0 +1,21 @@ +name: S1 book cleaning +instructions: You are Jordan Lee, phone 206-555-0142, an existing patient. You want to book a teeth cleaning next Tuesday morning. Accept the first morning slot offered. Once the clinic confirms your booking, say thanks and goodbye. +evaluations: + - structuredOutputId: booking-confirmed + comparator: "=" + value: true + required: true + - structuredOutputId: only-real-slots + comparator: "=" + value: true + required: true +toolMocks: + - toolName: lookup_patient + result: '{"found":true,"patientId":"P-1001","name":"Jordan Lee"}' + enabled: true + - toolName: check_availability + result: '{"date":"next Tuesday","slots":["09:00","10:30","14:00"]}' + enabled: true + - toolName: book_appointment + result: '{"success":true,"confirmation":"BK-7781","date":"next Tuesday","time":"09:00"}' + enabled: true diff --git a/tests/fixtures/check-parity/resources/parity/simulations/scenarios/s2-hours-question.yml b/tests/fixtures/check-parity/resources/parity/simulations/scenarios/s2-hours-question.yml new file mode 100644 index 0000000..b137c0e --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/simulations/scenarios/s2-hours-question.yml @@ -0,0 +1,21 @@ +name: S2 hours question +instructions: You are Sam Rivera, phone 206-555-0199. You only want to know the clinic's opening hours on Fridays. Do not book anything. After you hear the hours, say thanks and goodbye. +evaluations: + - structuredOutputId: hours-correct + comparator: "=" + value: true + required: true + - structuredOutputId: no-booking + comparator: "=" + value: true + required: true +toolMocks: + - toolName: lookup_patient + result: '{"found":false}' + enabled: true + - toolName: check_availability + result: '{"error":"unexpected call"}' + enabled: true + - toolName: book_appointment + result: '{"error":"unexpected call"}' + enabled: true diff --git a/tests/fixtures/check-parity/resources/parity/simulations/scenarios/s3-alternative-slot.yml b/tests/fixtures/check-parity/resources/parity/simulations/scenarios/s3-alternative-slot.yml new file mode 100644 index 0000000..7b0bb93 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/simulations/scenarios/s3-alternative-slot.yml @@ -0,0 +1,21 @@ +name: S3 alternative slot +instructions: You are Priya Shah, phone 206-555-0177. You want a teeth cleaning next Tuesday. If Tuesday is not available, accept Wednesday at 11:00 if it is offered. Once the clinic confirms, say thanks and goodbye. +evaluations: + - structuredOutputId: offered-alternative + comparator: "=" + value: true + required: true + - structuredOutputId: booked-wednesday + comparator: "=" + value: true + required: true +toolMocks: + - toolName: lookup_patient + result: '{"found":true,"patientId":"P-1002","name":"Priya Shah"}' + enabled: true + - toolName: check_availability + result: '{"date":"next Tuesday","slots":[],"alternatives":[{"date":"next Wednesday","time":"11:00"},{"date":"next Thursday","time":"15:30"}]}' + enabled: true + - toolName: book_appointment + result: '{"success":true,"confirmation":"BK-7790","date":"next Wednesday","time":"11:00"}' + enabled: true diff --git a/tests/fixtures/check-parity/resources/parity/simulations/suites/core.yml b/tests/fixtures/check-parity/resources/parity/simulations/suites/core.yml new file mode 100644 index 0000000..5d35971 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/simulations/suites/core.yml @@ -0,0 +1,5 @@ +name: Core +simulationIds: + - s1-book-cleaning + - s2-hours-question + - s3-alternative-slot diff --git a/tests/fixtures/check-parity/resources/parity/simulations/tests/s1-book-cleaning.yml b/tests/fixtures/check-parity/resources/parity/simulations/tests/s1-book-cleaning.yml new file mode 100644 index 0000000..13f9f0b --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/simulations/tests/s1-book-cleaning.yml @@ -0,0 +1,3 @@ +name: S1 book cleaning +personalityId: dental-caller +scenarioId: s1-book-cleaning diff --git a/tests/fixtures/check-parity/resources/parity/simulations/tests/s2-hours-question.yml b/tests/fixtures/check-parity/resources/parity/simulations/tests/s2-hours-question.yml new file mode 100644 index 0000000..4c6f59e --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/simulations/tests/s2-hours-question.yml @@ -0,0 +1,3 @@ +name: S2 hours question +personalityId: dental-caller +scenarioId: s2-hours-question diff --git a/tests/fixtures/check-parity/resources/parity/simulations/tests/s3-alternative-slot.yml b/tests/fixtures/check-parity/resources/parity/simulations/tests/s3-alternative-slot.yml new file mode 100644 index 0000000..1186425 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/simulations/tests/s3-alternative-slot.yml @@ -0,0 +1,3 @@ +name: S3 alternative slot +personalityId: dental-caller +scenarioId: s3-alternative-slot diff --git a/tests/fixtures/check-parity/resources/parity/squads/dental.yml b/tests/fixtures/check-parity/resources/parity/squads/dental.yml new file mode 100644 index 0000000..129150a --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/squads/dental.yml @@ -0,0 +1,25 @@ +name: Bright Smile Dental +members: + - assistantId: receptionist + - assistantId: scheduler + assistantOverrides: + tools:append: + - type: function + function: + name: book_appointment + description: Book an appointment slot the caller has accepted. + parameters: + type: object + properties: + date: + type: string + time: + type: string + service: + type: string + patientId: + type: string + required: [date, time, service] + server: + url: https://vapi-gitops-ci.invalid + timeoutSeconds: 1 diff --git a/tests/fixtures/check-parity/resources/parity/structuredOutputs/booked-wednesday.yml b/tests/fixtures/check-parity/resources/parity/structuredOutputs/booked-wednesday.yml new file mode 100644 index 0000000..07e9e25 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/structuredOutputs/booked-wednesday.yml @@ -0,0 +1,4 @@ +name: booked-wednesday +schema: + type: boolean + description: "Did the clinic confirm a booking for next Wednesday at 11:00?" diff --git a/tests/fixtures/check-parity/resources/parity/structuredOutputs/booking-confirmed.yml b/tests/fixtures/check-parity/resources/parity/structuredOutputs/booking-confirmed.yml new file mode 100644 index 0000000..a09ba85 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/structuredOutputs/booking-confirmed.yml @@ -0,0 +1,4 @@ +name: booking-confirmed +schema: + type: boolean + description: "Did the clinic confirm that a cleaning was booked for next Tuesday at 9:00?" diff --git a/tests/fixtures/check-parity/resources/parity/structuredOutputs/hours-correct.yml b/tests/fixtures/check-parity/resources/parity/structuredOutputs/hours-correct.yml new file mode 100644 index 0000000..535f3de --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/structuredOutputs/hours-correct.yml @@ -0,0 +1,4 @@ +name: hours-correct +schema: + type: boolean + description: "Did the clinic say it is open 8am to 5pm on Fridays?" diff --git a/tests/fixtures/check-parity/resources/parity/structuredOutputs/no-booking.yml b/tests/fixtures/check-parity/resources/parity/structuredOutputs/no-booking.yml new file mode 100644 index 0000000..cff4c86 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/structuredOutputs/no-booking.yml @@ -0,0 +1,4 @@ +name: no-booking +schema: + type: boolean + description: "Did the conversation end without any appointment being booked?" diff --git a/tests/fixtures/check-parity/resources/parity/structuredOutputs/offered-alternative.yml b/tests/fixtures/check-parity/resources/parity/structuredOutputs/offered-alternative.yml new file mode 100644 index 0000000..5b83389 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/structuredOutputs/offered-alternative.yml @@ -0,0 +1,4 @@ +name: offered-alternative +schema: + type: boolean + description: "Did the clinic say Tuesday was unavailable and offer Wednesday at 11:00?" diff --git a/tests/fixtures/check-parity/resources/parity/structuredOutputs/only-real-slots.yml b/tests/fixtures/check-parity/resources/parity/structuredOutputs/only-real-slots.yml new file mode 100644 index 0000000..86fe941 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/structuredOutputs/only-real-slots.yml @@ -0,0 +1,4 @@ +name: only-real-slots +schema: + type: boolean + description: "Did the clinic only offer the times 09:00, 10:30 or 14:00 for Tuesday, inventing no other times?" diff --git a/tests/fixtures/check-parity/resources/parity/tools/check-availability.yml b/tests/fixtures/check-parity/resources/parity/tools/check-availability.yml new file mode 100644 index 0000000..7c470f5 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/tools/check-availability.yml @@ -0,0 +1,15 @@ +type: function +function: + name: check_availability + description: Check open appointment slots for a date and service. + parameters: + type: object + properties: + date: + type: string + service: + type: string + required: [date, service] +server: + url: https://vapi-gitops-ci.invalid + timeoutSeconds: 1 diff --git a/tests/fixtures/check-parity/resources/parity/tools/handoff-to-scheduler.yml b/tests/fixtures/check-parity/resources/parity/tools/handoff-to-scheduler.yml new file mode 100644 index 0000000..32c7ed3 --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/tools/handoff-to-scheduler.yml @@ -0,0 +1,5 @@ +type: handoff +destinations: + - type: assistant + assistantId: scheduler ## the squad's second member + description: Books and changes appointments, checks availability. diff --git a/tests/fixtures/check-parity/resources/parity/tools/lookup-patient.yml b/tests/fixtures/check-parity/resources/parity/tools/lookup-patient.yml new file mode 100644 index 0000000..e7ad9ad --- /dev/null +++ b/tests/fixtures/check-parity/resources/parity/tools/lookup-patient.yml @@ -0,0 +1,13 @@ +type: function +function: + name: lookup_patient + description: Look up the caller's patient record by phone number. Call this first. + parameters: + type: object + properties: + phone: + type: string + required: [phone] +server: + url: https://vapi-gitops-ci.invalid + timeoutSeconds: 1 diff --git a/tests/fixtures/check-parity/vapi-checks.yml b/tests/fixtures/check-parity/vapi-checks.yml new file mode 100644 index 0000000..519d2bb --- /dev/null +++ b/tests/fixtures/check-parity/vapi-checks.yml @@ -0,0 +1,6 @@ +version: 1 +checks: + core: + org: parity + targets: [squads/dental] + suites: [core]