Skip to content

fix(antigravity): enforce tool lists, non-interactive commands, bounded polling - #215

Merged
akshaylive merged 2 commits into
mainfrom
akshaya/antigravity-tools-noninteractive
Oct 3, 2026
Merged

akshaylive merged 2 commits into
mainfrom
akshaya/antigravity-tools-noninteractive

Conversation

@akshaylive

Copy link
Copy Markdown
Collaborator

Summary

  • Upgrade google-antigravity 0.1.18 → 0.1.20.
  • Enforce allowed_tools / disallowed_tools: map Claude tool names to the harness builtins (Bash→run_command, Read→view_file, Write→create_file, Edit→edit_file, Glob→find_file, Grep→search_directory, Task→start_subagent, WebSearch→search_web, WebFetch→read_url_content) and pass them as CapabilitiesConfig. finish always stays on; subagents are enabled only when Task is.
  • Non-interactive commands: run_command gets CI=1, npm_config_yes=true, GIT_TERMINAL_PROMPT=0, DEBIAN_FRONTEND=noninteractive, PIP_NO_INPUT=1, PAGER=cat / GIT_PAGER=cat, each only where the environment doesn't already set it.
  • Bounded background polling: only an orphaned run_command keeps the turn polling (orphaned edit_file, paged view_file and start_subagent calls used to hold it too), and the wait is capped at 10 minutes, or 80% of turn_timeout if that is shorter.
  • Docs: ANTIGRAVITY.md, HARNESS_PARITY.md and .claude/notes/agents.md updated to match.

Why

On a 370-task SkillSpec run (gemini-3.5-flash, -j 4 ×2), 21 tasks sat in "Polling for backgrounded work" for the full 24 minutes (80% of a 1800 s turn_timeout). The causes were Edit on read-only skill files, never-ending shell commands, paged view_file and start_subagent. The model had already finished in every case.

Results

Same task set, rerun on a locally built coder-eval-agent:0.12.9 image with this change:

Before After
Wall time 174 min 75 min
Poll-cap hangs 21 × 24 min 3 × 10 min
Errors 0 0

The 3 remaining hangs are servers the model left running (node src/index.js, python3 server.py). All 3 tasks still graded normally and passed.

Test plan

  • pytest tests/test_antigravity_agent.py: 111 passed
  • ruff, pyright clean
  • Full SkillSpec Antigravity run (370 tasks) on the rebuilt image; logs confirm the restricted builtin tool set is in effect

🤖 Generated with Claude Code

…ed polling

Upgrade google-antigravity to 0.1.20 and close three Antigravity gaps that
made its runs diverge from Claude Code and Codex:

- allowed_tools / disallowed_tools are now mapped onto the harness builtin
  tools (Bash->run_command, Read->view_file, ...) and passed as
  CapabilitiesConfig; finish always stays on, subagents follow Task.
- run_command gets CI=1, npm_config_yes, GIT_TERMINAL_PROMPT=0,
  DEBIAN_FRONTEND=noninteractive, PIP_NO_INPUT=1 and PAGER=cat unless the
  environment already sets them.
- Only an orphaned run_command keeps the turn polling, and the wait is
  capped at 10 minutes (or 80% of turn_timeout if shorter). Orphaned
  edit_file / view_file / start_subagent calls no longer hold the turn.

On a 370-task SkillSpec run this cut wall time from 174 to 75 minutes and
poll-budget hangs from 21 (24 min each) to 3 (10 min each, all servers the
model left running).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@bai-uipath bai-uipath left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, one fix before merge. Clean and well-evidenced; the before/after run makes the case.

Subagents named Agent

  • Scope: the skills suite allows the subagent tool under Claude Code's current name Agent in 260 of ~1,390 tasks (19%): 250 uipath-troubleshoot, 7 planner, 4 maestro-case. The new map only knows Task, so on all 260 start_subagent is turned off even though the task allows it.
  • Effect: uipath-troubleshoot uses subagents for its 2-4 parallel escalation probes and to delegate an approved fix to the owning skill. Without one it runs the probes serially and, for fixes, presents the diff and stops, so Antigravity scores on those tasks would move for a config reason, not a model one.
  • Fix: map Agent to start_subagent alongside Task and add a test case; the rest of coder_eval already treats both names as the subagent tool.

…udit

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@akshaylive
akshaylive merged commit 7a5a4d0 into main Oct 3, 2026
17 checks passed
@akshaylive
akshaylive deleted the akshaya/antigravity-tools-noninteractive branch October 3, 2026 06:59
bai-uipath added a commit that referenced this pull request Oct 3, 2026
…by policy

The harness builds its system prompt from the toolset. With start_subagent
off it drops the whole subagents section, which carries the prompt's only
"you do NOT need to poll, you will be notified" guidance. Under an allowlist
without Task, Gemini then polls backgrounded commands with status checks and
short liveness timers: 1% of rows on 0.12.9 vs 8-10% after #215, and
max-turn rows 2 vs 12 of 148 on the SkillsBench Gemini rerun.

start_subagent now stays on, so the prompt matches the harness default, and
when the tool lists do not allow Task the call that runs a subagent
(invoke_subagent) is denied by policy. That covers the built-in research
subagent and model-defined ones, which both get search_web regardless of the
parent's toolset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants