Skip to content

Security: ExaDev/claude-code-action

docs/security.md

Security notes

The full detail behind the README's security summary: who can trigger a run, the fork pull request token downgrade, the self-modifying-workflow trap, untrusted text handling, the Headroom trust addition, interactive mode's unscoped Bash grant, and the review-widening inputs and what each one actually exposes.

Security notes

Who can trigger a run. By default only users with repository write access, because the underlying action checks write permission before doing anything. The allowed_non_write_users input widens that, and it is off by default in all three modes deliberately.

  • For triage, you will probably want allowed_non_write_users: "*", since the issues worth triaging are usually opened by people without write access. That is the setting in the triage example. It is defensible there because triage holds no code-write or pull-request scope: the worst case is a misleading label and a wrong comment.
  • For interactive, leave it unset. That mode can edit code, commit, and push a branch, so allowing users without write access to invoke it hands them write access by proxy through text the model interprets. If you want Claude responding to outside reporters, use triage.

Fork pull requests and the GITHUB_TOKEN downgrade. GitHub automatically restricts the ambient GITHUB_TOKEN to read-only for any pull_request-triggered run originating from a fork, regardless of what permissions: a workflow declares — this is a platform-level protection, not something any input here controls. That's why review mode's own example (examples/direct/claude-review.yml) uses the direct form and deliberately leaves github_token unset: doing so makes the run authenticate as the Claude Code GitHub App (claude[bot]) instead, a separately-minted credential via OIDC that isn't subject to that downgrade. Reviewing a fork pull request through the reusable workflow's default GITHUB_TOKEN would appear to run successfully but silently fail to post — worth knowing if a repository that accepts fork pull requests seems to review nothing. Separately, GitHub's own "Require approval for first-time contributors" repository setting (on by default for public repositories) gates any workflow run — including this one — from an unknown contributor's fork pull request until a maintainer approves it once in the Actions UI, on top of everything above.

A pull request cannot reliably self-test a change to its own OIDC-authenticated workflow file, and this action fails loudly when that happens rather than silently doing nothing. Any workflow using the App-identity/OIDC path (leaving github_token unset, as claude-review.yml, claude-triage.yml, and claude-interactive.yml all do here) is subject to a GitHub platform check that the calling workflow file be byte-identical to the version on the default branch — the same mechanism the "Automated upstream bumps" section below documents for dependabot.yml. A pull request that edits one of these three workflow files fails that check on its own first run against itself; the same check also fails, independent of anything the pull request itself touches, when the branch the workflow file lives on has simply fallen behind a since-changed copy of the file on the default branch. Upstream itself treats a failure either way as a soft skip (Skipping action due to workflow validation, a warning, not a crash) and exits successfully having done nothing — reasonable for a repository adopting one of these workflows for the very first time, before the file exists on its own default branch yet, but actively misleading for an already-working workflow that has quietly stopped reviewing anything with no visible signal. Resolve Claude Code result detects this specific case (a winning attempt with no conclusion output at all — the one upstream code path this can come from, since the only other cause of a missing conclusion, "no trigger found," is provably unreachable through this wrapper, which always sets a prompt) and fails the job outright instead, with an ::error:: naming the two possible causes and pointing back here — confirmed live against adpeak/adpeak-mono, where a branch that had fallen behind a same-day rewrite of its own claude-review.yml on main hit exactly this. Passing github_token: ${{ github.token }} explicitly sidesteps this entirely by skipping OIDC, but doing that permanently on one of these three files would give up the App identity's own capabilities (resolve_stale_threads and friends) for every future run, not just self-modifying ones — not a trade worth making for a narrow, rare edge case. Validate a change to one of these files with a separate, temporary workflow instead (one that does not itself modify the file under test), then remove it once confirmed; the change becomes fully self-testable again the moment it lands on the default branch. If the loud failure fires and neither cause applies — this pull request does not touch the workflow file, and the file already matches the default branch — merge or rebase the latest default branch into this branch, which resolves it in the branch-staleness case even without a workflow-file conflict of its own.

Untrusted text reaches the model. Issue bodies, comments, pull request descriptions, and the contents of a pull request's own files are attacker-influenced input on a public repository or one accepting fork pull requests. The shared prompt tells Claude to treat all of it as data rather than instructions and to report injection attempts, but a prompt is mitigation, not a guarantee. The real controls are the per-mode token scopes and tool allowlists above: assume the prompt can be subverted and check that the blast radius is acceptable if it is.

headroom_enabled adds a new third-party dependency that processes all of the above, and it's on by default. Every diff, CI log, issue body, and file read already covered by the paragraph above passes through the Headroom proxy before it reaches Anthropic — a real supply-chain trust addition, even though Headroom's own design claims compression runs entirely locally with no prompt or file content sent elsewhere. A repository that does not want that trust addition sets headroom_enabled: false. The proxy's data plane is unauthenticated by default; that is acceptable here specifically because it is bound to 127.0.0.1:<headroom_port> on the runner — directly, for the pip install method (confirmed by the proxy's own startup banner, which reports "loopback-only" for this exact bind, unlike the Docker method's own "non-loopback bind" warning about its internal 0.0.0.0 listen address, necessary there only so Docker's port-publish can reach it); via a Docker port publish that never touches 0.0.0.0 on the host side, for the docker method — so nothing beyond this job's own processes can reach it either way. This action never widens that bind. headroom_show_savings (on by default) additionally posts a comment reporting the numbers — needs no token scope beyond what the mode already holds, and its body is built entirely from numeric stats this action's own bash computes, never from anything the model or untrusted PR content produced, so it carries none of the prompt-injection surface the rest of this section is about.

One consequence is worth being explicit about: on a pull_request event the checkout includes the pull request's own changes, so a fork pull request can modify .github/claude/review.md (see below) or the repository's CLAUDE.md and thereby change the instructions used to review it. Review mode holds no write scope to repository contents — it cannot edit files, commit, or push — but that no longer bounds the damage to a misleading review by default, because four on-by-default-or-opt-in inputs widen what a subverted review (or its own follow-up call) can do within the token's pull-requests scope. include_suggestions lets it emit a fenced suggestion block an author might one-click apply (not a permission escalation — the author's Apply is the write, the bot has none — but it removes the friction of reading and re-typing a prose fix, so a malicious "fix" gets less scrutiny). fix_pr_metadata grants the whole mcp__github__update_pull_request tool, not just a title/body-scoped slice of it — the tool itself also accepts state (it can close the pull request), base, draft, maintainer_can_modify, and reviewers, so a subverted review can call it with any of those, no gh api:* needed; only the prompt restricts its own use of the tool to title and body. mark_draft_if_unready (off by default, unlike fix_pr_metadata) grants that same tool for a different, equally prompt-enforced restriction — setting only draft: true, and only when the pull request shows an explicit, concrete unreadiness signal — but a subverted review could abuse the same grant to convert an unrelated, genuinely-ready pull request to draft, which does not destroy anything but does pull it out of reviewers' default "open" queue and off the notifications a ready-for-review state normally generates, reducing scrutiny rather than merely misrepresenting a finding; this is why it defaults to false even though fix_pr_metadata does not. resolve_stale_threads adds gh api:* to the review allowlist, so a subverted review gains the whole REST/GraphQL surface within that pull-requests scope (editing the PR's title, body, or labels; posting arbitrary comments; blanking or rewriting any review's body, not only the bot's own, since nothing at the API layer restricts the call to reviews the bot itself submitted — only the prompt's own filter-by-login step does) — a second, broader route to the same title/body write, and to more besides. structured_review_summary, though opt-in rather than on by default, grants the same gh api:* surface to its own follow-up call for the same reason (reading inline review comments has no narrower gh pr view --json field) — only the prompt restricts that call to reading. include_ci_logs, never previously listed here, belongs on this list too now that it does more than inform prose: it grants tools that fetch this pull request's own workflow logs, which are attacker-influenced on a fork pull request (a test name, an assertion message, or anything else the change causes CI to print lands in the model's context), and it is the input fix_ci_failures (below) reads to decide what to change. On its own it is still read-only and low-risk; paired with fix_ci_failures it is the input path into a write. On a repository that accepts untrusted fork pull requests, set include_suggestions: false, fix_pr_metadata: false, resolve_stale_threads: false, verify_prior_findings: false, and leave structured_review_summary, fix_ci_failures, fix_diff_findings, add_regression_tests, and mark_draft_if_unready off (their defaults) until a human has reviewed the change, and do not treat a review of an untrusted pull request as a security control.

fix_ci_failures and fix_diff_findings are different in kind from everything above, and the difference is worth stating exactly. Everything above widens what a subverted review can do within the token's pull-requests scope — write a comment, retitle, close. These two add a call that can write files and commit, and a job that holds contents: write. Three things bound that, and none of them is the prompt: the review call itself is unchanged (Write, Edit, and NotebookEdit stay on its denylist, so the call that reads untrusted diff content and the call that can write a file are different calls with different allowlists); the push target is fixed in the action (one refspec, this pull request's own head branch, computed in bash and unreachable from any prompt — the model has no git push tool at all, and the whole write is discarded with the runner if the publishing step does not push it); and a fork pull request never reaches the fix call, because the job's token cannot push to a fork regardless of maintainer_can_modify, so the pass is skipped before a call is even paid for. That last point is what makes this defensible at all: the untrusted-input scenario this section is about is precisely the scenario in which the write path is mechanically unreachable. What is not bounded mechanically is the content of a fix on a same-repository pull request, and the contents: write the direct form's job must hold is broader than the one branch the action pushes to — branch protection on the default branch is what stops that becoming a direct push, exactly as it already is for interactive mode. add_regression_tests adds no new capability of its own, only more written lines alongside a fix already being made. Leave all three off (their default) unless every pull request in the repository comes from someone you would already give write access to — they are not something to enable estate-wide.

verify_prior_findings is a third kind of risk, different again from both clusters above. Every input discussed so far governs how much a subverted review can do within GitHub's own API surface — writing a comment, retitling, closing, or (for fix_ci_failures/fix_diff_findings) committing to a fixed, bounded branch. verify_prior_findings grants WebFetch instead, which is not GitHub-scoped at all: it is arbitrary network egress to any host the model chooses to fetch, mediated by nothing this action's token permissions bound. The prompt restricts it to a URL the model reasoned its own way to (a package's own registry or documentation site) and forbids following any URL that appears in the diff, the pull request description, or a comment — exactly the kind of untrusted content a subverted review could otherwise use to redirect the fetch — but, as with every other prompt-only restriction in this list, that boundary is enforced by the prompt alone, not by anything technical. It is deliberately not domain-scoped: this action's stack fragments already span enough different package registries and documentation hosts that a maintained allowlist would either lag behind real usage or need constant upkeep, so set verify_prior_findings: false (see above) rather than relying on a narrower grant that does not exist.

Interactive mode's Bash grant is unscoped, and this is a real widening, not a cost-free refactor. Unlike every other mode, interactive grants the bare Bash tool rather than an enumerated Bash(<cmd>:*) list, so it can run any shell command, not just the specific prefixes named in action.yml. This is a deliberate trade against the Claude Agent SDK's own permission model: a scoped Bash(<cmd>:*) entry only auto-approves a command that matches it exactly, and the SDK decomposes a compound command (&&, pipes, $(...), case statements) into its individual operations, requiring every one of them to match an allow rule — so a single unlisted step anywhere in an otherwise-ordinary one-liner (a plain find, wc, or echo used to survey a repository, say) blocks the whole command, with no human present in a headless run to approve it. Deny rules are evaluated before allow rules and are never overridden by a bare allow (see the Claude Agent SDK's permission docs), so action.yml's own deny list still blocks the exact command spellings it names — but those are spellings, not capabilities: gh pr merge is denied, but gh api -X PUT .../pulls/N/merge reaches the same endpoint within the same token's scope, and the same route defeats the gh pr close, gh issue close, and gh release denials too. Bash(gh api:*) is denied for exactly this reason, closing the one-liner route, though not curl, node -e, or any other way of reaching the GitHub API this token can authenticate. Treat the deny list as defence in depth against an accidental invocation, not as the actual security boundary — what genuinely bounds this run is the token itself, independently capped at contents/pull_requests/issues write (setupGitHubToken's DEFAULT_PERMISSIONS, since interactive sets no ADDITIONAL_PERMISSIONS), branch protection on the default branch, and that only a write-access user can invoke this mode at all. What genuinely is new here: interactive can now run any command within that token's reach — curl, arbitrary scripts, reading environment variables — not just the build/test/git/gh prefixes previously enumerated. That is a real widening of what "assume the prompt can be subverted and check the blast radius" needs to account for in this specific mode.

Branch protection matters. Interactive mode holds contents: write. Its prompt forbids committing to the default branch, force-pushing, and rewriting history, but branch protection on the default branch is what actually enforces it. Turn it on.

Cost and loops. Interactive mode is gated on the trigger phrase in the reusable workflow's job condition rather than inside the action. This is load-bearing: supplying a prompt puts the underlying action into agent mode, which runs unconditionally and does not check the phrase itself. Without that condition every comment on every issue would start a paid run. That same job condition also excludes any bot-authored event — confirmed directly: Anthropic's own native Claude Code Review GitHub App posts a boilerplate "Comment @claude review for a one-time review..." notice on every pull request when its automatic review is set to manual, and that boilerplate's own text satisfies a naive trigger-phrase match, starting a paid run from the bot's own comment. The review example includes a concurrency block that cancels a superseded review when the author pushes again.

There aren't any published security advisories