fix(antigravity): enforce tool lists, non-interactive commands, bounded polling - #215
Merged
Merged
Conversation
…ed polling Upgrade google-antigravity to 0.1.20 and close three Antigravity gaps that made its runs diverge from Claude Code and Codex: - allowed_tools / disallowed_tools are now mapped onto the harness builtin tools (Bash->run_command, Read->view_file, ...) and passed as CapabilitiesConfig; finish always stays on, subagents follow Task. - run_command gets CI=1, npm_config_yes, GIT_TERMINAL_PROMPT=0, DEBIAN_FRONTEND=noninteractive, PIP_NO_INPUT=1 and PAGER=cat unless the environment already sets them. - Only an orphaned run_command keeps the turn polling, and the wait is capped at 10 minutes (or 80% of turn_timeout if shorter). Orphaned edit_file / view_file / start_subagent calls no longer hold the turn. On a 370-task SkillSpec run this cut wall time from 174 to 75 minutes and poll-budget hangs from 21 (24 min each) to 3 (10 min each, all servers the model left running). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
akshaylive
requested review from
CarlesUIPath,
bai-uipath,
tmatup and
uipreliga
as code owners
October 3, 2026 06:22
bai-uipath
approved these changes
Oct 3, 2026
bai-uipath
left a comment
Collaborator
There was a problem hiding this comment.
LGTM, one fix before merge. Clean and well-evidenced; the before/after run makes the case.
Subagents named Agent
- Scope: the skills suite allows the subagent tool under Claude Code's current name
Agentin 260 of ~1,390 tasks (19%): 250 uipath-troubleshoot, 7 planner, 4 maestro-case. The new map only knowsTask, so on all 260start_subagentis turned off even though the task allows it. - Effect: uipath-troubleshoot uses subagents for its 2-4 parallel escalation probes and to delegate an approved fix to the owning skill. Without one it runs the probes serially and, for fixes, presents the diff and stops, so Antigravity scores on those tasks would move for a config reason, not a model one.
- Fix: map
Agenttostart_subagentalongsideTaskand add a test case; the rest of coder_eval already treats both names as the subagent tool.
…udit Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
bai-uipath
added a commit
that referenced
this pull request
Oct 3, 2026
…by policy The harness builds its system prompt from the toolset. With start_subagent off it drops the whole subagents section, which carries the prompt's only "you do NOT need to poll, you will be notified" guidance. Under an allowlist without Task, Gemini then polls backgrounded commands with status checks and short liveness timers: 1% of rows on 0.12.9 vs 8-10% after #215, and max-turn rows 2 vs 12 of 148 on the SkillsBench Gemini rerun. start_subagent now stays on, so the prompt matches the harness default, and when the tool lists do not allow Task the call that runs a subagent (invoke_subagent) is denied by policy. That covers the built-in research subagent and model-defined ones, which both get search_web regardless of the parent's toolset. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
google-antigravity0.1.18 → 0.1.20.allowed_tools/disallowed_tools: map Claude tool names to the harness builtins (Bash→run_command,Read→view_file,Write→create_file,Edit→edit_file,Glob→find_file,Grep→search_directory,Task→start_subagent,WebSearch→search_web,WebFetch→read_url_content) and pass them asCapabilitiesConfig.finishalways stays on; subagents are enabled only whenTaskis.run_commandgetsCI=1,npm_config_yes=true,GIT_TERMINAL_PROMPT=0,DEBIAN_FRONTEND=noninteractive,PIP_NO_INPUT=1,PAGER=cat/GIT_PAGER=cat, each only where the environment doesn't already set it.run_commandkeeps the turn polling (orphanededit_file, pagedview_fileandstart_subagentcalls used to hold it too), and the wait is capped at 10 minutes, or 80% ofturn_timeoutif that is shorter.ANTIGRAVITY.md,HARNESS_PARITY.mdand.claude/notes/agents.mdupdated to match.Why
On a 370-task SkillSpec run (
gemini-3.5-flash,-j 4×2), 21 tasks sat in "Polling for backgrounded work" for the full 24 minutes (80% of a 1800 s turn_timeout). The causes were Edit on read-only skill files, never-ending shell commands, pagedview_fileandstart_subagent. The model had already finished in every case.Results
Same task set, rerun on a locally built
coder-eval-agent:0.12.9image with this change:The 3 remaining hangs are servers the model left running (
node src/index.js,python3 server.py). All 3 tasks still graded normally and passed.Test plan
pytest tests/test_antigravity_agent.py: 111 passed🤖 Generated with Claude Code