Skip to content

fix(providers): cut per-call cost of CLI-provider scans - #808

Open
rng1995 wants to merge 13 commits into
mainfrom
fix/claude-cli-session-cost
Open

rng1995 wants to merge 13 commits into
mainfrom
fix/claude-cli-session-cost

Conversation

@rng1995

@rng1995 rng1995 commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

Why

A user on r/ClaudeAI (thread) asked Claude Code to scan their 15 skills with SkillSpector. The agent picked claude_cli on its own and ran five scans in parallel. Because every LLM call starts a separate claude -p session, that came to 434 sessions in about 6 minutes, 22.4M tokens (about 34k of startup context per session, on Opus), and 98% of their 5-hour plan window. It also left 1,144 folders in ~/.claude/projects.

#298 (for #295) already keeps MCP servers, hooks, and settings out of the child. Three costs were still left: the built-in tool schemas, Claude Code's default agent system prompt, and a persisted transcript for every call.

Changes

  1. Lower per-call cost (claude argv). Add --tools "", so the built-in tool definitions are no longer sent. --allowed-tools "" only blocks tool use (CLI reference). Also replace the default system prompt with a short fixed one via --system-prompt; the analyzer instructions still arrive on stdin.
  2. No session clutter. Add --no-session-persistence. Each call runs in a fresh temp cwd, so before this change every call left its own project folder.
  3. Smaller default fan-out and a warning. When SKILLSPECTOR_MAX_LLM_CONCURRENCY is unset and the active provider is a CLI provider, the default is now 2 instead of 10. Once per process, when a CLI provider is available, SkillSpector warns that every LLM call is a separate session on the user's plan and points to --no-llm.
  4. skill-inspector guidance. Keep --no-llm unless the user asks for LLM analysis. Never choose or change SKILLSPECTOR_PROVIDER without asking. Don't run CLI-provider scans in parallel.

Not in this PR

  • Printing the expected number of LLM calls before a scan, or asking the user to confirm. The call count isn't known up front, and a prompt would break non-interactive runs. The warning plus the lower default cover the immediate risk.
  • A README section on CLI-provider cost.

Testing

  • uv run python -m pytest tests/unit/test_agent_cli.py tests/unit/test_providers.py tests/nodes/test_llm_analyzer_base.py: 539 passed, 10 skipped. New tests cover the argv flags, the CLI concurrency default (and that an explicit value still wins), and the warn-once behaviour.
  • ruff check and ruff format --check are clean.
  • The new flags were checked against the current Claude Code CLI reference, not run against a live claude binary. Before merge, a maintainer with claude installed should run one SKILLSPECTOR_PROVIDER=claude_cli scan to confirm that auth still works and the output parses. Older claude builds that lack --tools or --no-session-persistence will fail closed with a non-zero exit.

🤖 Generated with Claude Code

rng1995 and others added 13 commits October 8, 2026 12:05
Each LLM call with claude_cli starts a separate `claude -p` session. A
user scanning 15 skills (five scans in parallel) started 434 sessions in
about 6 minutes and used 98% of their 5-hour plan window, at roughly 34k
tokens of startup context per session.

- claude argv: add `--tools ""` so built-in tool schemas are not sent
  (`--allowed-tools ""` only denies use), replace the default agent
  system prompt with a short fixed one, and pass
  `--no-session-persistence` so calls stop leaving one
  ~/.claude/projects folder each.
- Default LLM fan-out to 2 for CLI providers when
  SKILLSPECTOR_MAX_LLM_CONCURRENCY is unset (HTTP providers keep 10).
- Warn once per process when a CLI provider is available that every LLM
  call is a separate session on the user's plan, and point to --no-llm.
- skill-inspector: keep --no-llm unless the user asks, never pick a
  provider on the agent's own, and don't run CLI-provider scans in
  parallel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant