Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 34 additions & 1 deletion .claude/notes/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -434,7 +434,10 @@ agent had produced; bounding it only by a fixed cycle count disconnected from `t
path unreachable — the watchdog always wins, and the same spurious-orphan turn burns the
full turn timeout before crashing with zero criteria evaluated.

The cycle cap is the SOLE bound when a task sets no timeout at all. It is deliberately not
The cycle cap applies alongside that deadline, whichever comes first, and is the SOLE
bound when a task sets no timeout at all. Under a long `turn_timeout` (1800 s gives a
1440 s deadline) the deadline alone let a never-ending command (a dev server, an
unanswered prompt) idle the turn for 24 minutes. It is deliberately not
"break after N consecutive empty polls": `receive_steps()` returns identically empty
whether a backgrounded job is still running or will never resolve, and there is no signal
that tells the two apart except waiting. A count small enough to matter would abort real
Expand All @@ -447,6 +450,36 @@ answer in a headless eval), CANCELED and UNKNOWN — none of which the closed se
done, and none of which the poll loop should wait out, since they will never become DONE
on their own.

It is also an allowlist on the TOOL: only a `run_command` can be backgrounded. Any other
tool left ACTIVE never resolved in practice: an `edit_file` on the read-only skill mount,
a `view_file` paged at a large `content_offset`, a `start_subagent`. In a 370-task
SkillSpec run, 12 of 21 turns that polled to the deadline were waiting on one of those,
after the model had already finished. `finalize` force-closes them as unresolved.

## Antigravity tool allowlist

`allowed_tools` / `disallowed_tools` map onto the harness's builtin toolset
(`CapabilitiesConfig.enabled_tools` / `disabled_tools`) through
`_CLAUDE_TO_ANTIGRAVITY_TOOL_MAP`, so an experiment's allowlist confines Antigravity the
way it confines Claude Code. Without it Antigravity ran with every builtin, including
`start_subagent` and `search_web`, under a list that gave Claude Code neither. `Skill` and
`TodoWrite` have no builtin (skills load through `skills_paths`) and are skipped. `finish`
stays on under any allowlist: it returns structured output, not a capability.
`enable_subagents` is a separate switch from the toolset and follows whether
`start_subagent` survives. The SDK takes an allowlist OR a denylist, so with both set the
denied tools are removed from the allowlist.

## Antigravity non-interactive commands

`run_command` runs in a real terminal: a command that blocks becomes a background task the
model can check on and type into. A command that stops to ask (`npx` installing a missing
package, git credentials, an apt/pip confirmation, a pager) and that the model then leaves
behind keeps the turn waiting on a prompt no one answers. `_NONINTERACTIVE_ENV`
(`CI`, `npm_config_yes`, `GIT_TERMINAL_PROMPT=0`, `DEBIAN_FRONTEND`, `PIP_NO_INPUT`,
`PAGER`/`GIT_PAGER=cat`) rides the per-agent `env` seam, each variable only where the
environment does not already set it, so those commands answer themselves or fail fast. It
cannot close the terminal: a bare `read` or a server still runs until the poll cap.

## The receive_steps re-entrancy window

`receive_steps()` is two nested async generators: the public one delegates to the
Expand Down
39 changes: 24 additions & 15 deletions docs/agents/ANTIGRAVITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ working directory — both required for an unattended eval.
pip install 'coder-eval[antigravity]'
```

This pulls in `google-antigravity` (pinned to `0.1.18`), whose wheel bundles the
This pulls in `google-antigravity` (pinned to `0.1.20`), whose wheel bundles the
platform `localharness` binary. As with the other agents the SDK is imported lazily
— a base install without the extra still runs end-to-end; Antigravity tasks fail at
dispatch with a clear hint to install the extra.
Expand Down Expand Up @@ -136,18 +136,28 @@ can't be resolved or if zero skills are discovered.

## Permissions & tools — important differences

**Antigravity ignores `permission_mode`, `allowed_tools`, and `disallowed_tools`.**
The local harness runs in a single unconditional mode: every tool call (including
`run_command`) is approved via an allow-all policy, and file tools are restricted to
the configured `workspaces` (the sandbox working directory plus any skill roots).
**Antigravity ignores `permission_mode`.** The local harness runs in a single
unconditional mode: every tool call (including `run_command`) is approved via an
allow-all policy, and file tools are restricted to the configured `workspaces` (the
sandbox working directory plus any skill roots).

**`allowed_tools` / `disallowed_tools` are enforced** by mapping the Claude tool names
onto the harness's builtin tools (`Bash` → `run_command`, `Read` → `view_file`, `Write` →
`create_file`, `Edit` → `edit_file`, `Glob` → `find_file`, `Grep` → `search_directory`,
`Task` → `start_subagent`, `WebSearch` → `search_web`, `WebFetch` → `read_url_content`).
`Skill` has no builtin and is skipped; `finish` always stays on.

`run_command` also gets non-interactive environment variables (`CI=1`,
`npm_config_yes=true`, `GIT_TERMINAL_PROMPT=0`, `DEBIAN_FRONTEND=noninteractive`,
`PIP_NO_INPUT=1`, `PAGER=cat`), each only where the environment does not already set it.

The trust boundary for an Antigravity run is therefore the **sandbox**, not the
agent config. Run untrusted tasks under the [Docker driver](../DOCKER_ISOLATION.md);
the `tempdir` driver is not a security boundary. This mirrors the reality that the
`bypassPermissions`-equivalent behavior is always on for this backend.

Those inherited fields still exist on the config for schema uniformity but have no
runtime effect here — don't rely on them to gate Antigravity.
`permission_mode` still exists on the config for schema uniformity but has no runtime
effect here — don't rely on it to gate Antigravity.

## Telemetry

Expand Down Expand Up @@ -187,19 +197,18 @@ as every other agent.
4. **`permission_mode` does not confine the harness.** Every mode runs
`policy.allow_all()`; coder_eval's write boundary is the sandbox driver, and a
headless eval has no human to approve anything.
5. **`allowed_tools` / `disallowed_tools` are not read.** The harness runs with its
full builtin tool set, so an Antigravity run has tools (web search, subagents,
URL fetch) that the same task file denies on Claude Code and Codex.
6. **`max_turns` is counted by the harness.** One `communicate()` is a single SDK turn
5. **`max_turns` is counted by the harness.** One `communicate()` is a single SDK turn
here, so the harness counts model API calls itself (a MODEL step at a new
`step_index` opens one) and enforces the cap on the step loop. See
[Run-Limit Parity](HARNESS_PARITY.md).
7. **Shell commands over ~10s are moved to the background.** The localharness has a
6. **Shell commands over ~10s are moved to the background.** The localharness has a
10-second maximum synchronous wait; past it the command becomes a background task
and the model gets a task id, not a result. The turn polls for that result instead
of finalizing on an idle step stream, so slow work does complete — but the wait is
bounded by 80% of `turn_timeout`, and a job that outlives it is force-closed as
`result_status: unknown` and graded as an ordinary low score rather than a timeout.
of finalizing on an idle step stream, so slow work does complete — but only an
orphaned `run_command` is waited on, the wait is bounded by 10 minutes or 80% of
`turn_timeout` (whichever is shorter), and a job that outlives it (typically a server
the model left running) is force-closed as `result_status: unknown` and graded
normally rather than as a timeout.
Measured in [Run-Limit Parity](HARNESS_PARITY.md).

## Running in Docker
Expand Down
7 changes: 4 additions & 3 deletions docs/agents/HARNESS_PARITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -718,9 +718,10 @@ both. See [OpenCode](OPENCODE.md) and [Pi § plugins](PI.md#known-limitations).
lowercase (`bash`/`read`/`write`/`edit`/`grep`/`find`/`ls`), but the shared config
default (`experiments/default.yaml`) sets Claude-namespaced names
(`Bash`/`Read`/`Write`/…). Forwarding those to `--tools` would allowlist tools that
do not exist in Pi and strip the agent of ALL tools — so, like OpenCode (drops them),
Codex (forwards `disallowed_tools` without SDK enforcement), and Antigravity (does not
read them), Pi ignores them and runs with its full native toolset. A task that needs a
do not exist in Pi and strip the agent of ALL tools — so, like OpenCode (drops them)
and Codex (forwards `disallowed_tools` without SDK enforcement), Pi ignores them and
runs with its full native toolset. (Antigravity maps them onto its builtin tools; see
[Antigravity](ANTIGRAVITY.md).) A task that needs a
restricted Pi toolset would have to name Pi's lowercase tools — a documented follow-up.
- **`permission_mode` is NOT enforced** — Pi headless print mode auto-runs tools and
exposes only project-file trust (`--approve` / `--no-approve`), no tool-approval
Expand Down
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -131,7 +131,7 @@ codex = [
# Without this extra the framework still installs and runs; Antigravity-dependent
# code paths fail at start() with a clear hint pointing back here.
antigravity = [
"google-antigravity==0.1.18",
"google-antigravity==0.1.20",
]
# Optional extra that enables OpenCode agent support.
#
Expand Down
Loading
Loading