An advanced, autonomous, Claude-Code-style AI agent execution loop built completely from scratch in TypeScript and running under Bun.
It natively orchestrates tool-calling, multi-layered context memory, tree-structured session persistence, multi-layer security guardrails, Model Context Protocol (MCP) extensibility, non-destructive workspace snapshots, and a comprehensive agent evaluation benchmark suite without using external agent frameworks (like LangChain or LangGraph).
π¬ Watch Coding Harness in Action:
Coding-Harness.4.mp4
Coding-Harness.4.mp4
- πΉ Demo
- β¨ Features
- ποΈ High-Level Architecture
- π§ Context Management Engine
- πΎ Session Management & Tree Persistence
- π‘οΈ Multi-Layer Guardrails & Safety Engine
- 1. Decision Hierarchy (
DENY>ASK>ALLOW) - 2. Path Confinement & Secret File Protection (
PathPolicy) - 3. Command Inspection & Auto-Approval (
CommandPolicy) - 4. Secret Scanner & Redaction (
SecretScanner) - 5. Resource & Payload Protection (
ResourcePolicy) - 6. Network Isolation & Exfiltration Defense (
NetworkPolicy) - 7. Interactive Unified Diff Authorization
- 1. Decision Hierarchy (
- πΈ Workspace Snapshots & Checkpoints
- π Model Context Protocol (MCP) Integration
- π Evaluation & Benchmarking Suite
- π οΈ Tool Ecosystem
- π€ Multi-Provider LLM Support
- π Execution Modes
- β‘ Quick Start
- π Repository Structure
- π License
- π» Dynamic Interactive CLI REPL: Live streaming of assistant responses, formatted
π Thinkingblocks, tool call parameters, execution results, and runtime token dashboard. - βοΈ Headless Mode: Non-interactive automation entry point that accepts
--taskand--cwdand outputs structured JSON results for script/CI integration. - π§ Advanced Context Engineering: Automatic file read staleness invalidation (tombstoning), tool-call microcompaction, LLM summarization compaction, and
cache_controlbreakpoint injection. - π³ Tree-Structured Session Storage: Append-only JSONL event-log format supporting linear history, parent-pointer branching, leaf rewinding, and session resume capabilities.
- π‘οΈ Enterprise-Grade Guardrails: 5-tier security engine enforcing path traversal prevention, destructive command denial, regex secret scanning with redaction, payload size limits, and socket/network exfiltration defense with live unified diff authorization.
- πΈ Non-Destructive Snapshots & Checkpoints: Git-plumbing-backed workspace snapshots (using alternate index files without touching working branch or
HEAD), file-archive fallbacks, clean rollbacks with automatic safety snapshots, and full runtime checkpoints. - π Model Context Protocol (MCP): Native MCP client support over stdio transport, dynamic
.mcp.jsonserver registry, schema normalization (including Gemini sanitization), and built-in servers (e.g. DuckDuckGo web search). - π Benchmarking & Evaluation Suite: Standardized coding benchmark framework with 9 specialized evaluators (build, typecheck, lint, unit tests, diffs, file rules, security, requirements) and detailed terminal/JSON reporting.
- π Dual Tool Execution Modes: Switch dynamically between Parallel execution (running independent read/write calls concurrently via
Promise.all) and Sequential execution. - π Range-Targeted Editing with Drift Recovery: Line-targeted find-and-replace (
startLine/endLine) with sliding-window offset recovery (Β±10 lines tolerance) and mismatch diagnostics. - π Fast Search Capabilities: Ripgrep-backed (
rg) recursive text pattern matching with Bun-glob fallback and wildcard glob scanning. - π€ Provider Agnostic: Seamlessly switch between local Ollama models (
qwen3:14b,qwen3:8b,llama3.1:8b) and cloud Gemini models (gemini-1.5-flash,gemini-1.5-pro,gemini-2.5-flash). - π°οΈ Read-Only Sub-Agent Dispatch: Isolated sub-agent worker context for performing background research without mutating workspace files.
The framework is decoupled into modular layers, separating execution control, context lifecycle, persistence, safety guardrails, MCP extensibility, snapshots, evaluation, and model connectivity:
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β CLI Layer β
β Interactive REPL β Headless Runner β MCP CLI β Snapshot/Checkpoint CLI β Evaluation CLI β
βββββββββββββββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββββββββ
β
ββββββββββΌβββββββββ
β Agent Core β
β (src/agent.ts) β
ββββββββββ¬βββββββββ
β
ββββββββββββββββββββββ¬βββββββββββββββΌβββββββββββββββ¬ββββββββββββββββββββ
β β β β β
βββββββββΌβββββββββ βββββββββΌβββββββ βββββββΌββββββββ ββββββΌββββββββββββββ βββββΌβββββββββββββ
β ContextManager β β SessionStore β βPolicyEngine β β SnapshotManager β β McpManager β
β - History β β - JSONL Tree β β - Path β β - Git Plumbing β β - Stdio Client β
β - Tombstoning β β - Branching β β - Command β β - Rollback Engineβ β - Tool Adapter β
β - Compaction β β - Path Resolvβ β - Secret β β - Checkpoints β β - .mcp.json Regβ
βββββββββ¬βββββββββ ββββββββββββββββ β - Resource β ββββββββββββββββββββ ββββββββββ¬ββββββββ
β β - Network β β
β βββββββββββββββ β
β β β
β ββββββββββββββββΌβββββββββββββββββ β
βββββββββββββββββββββΊβ ToolRegistry ββββββββββββββββββββββββ
β β - 13 Built-in Core Tools β
β β - Dynamic MCP Tools β
β ββββββββββββββββ¬βββββββββββββββββ
β β
β ββββββββββββββββ΄ββββββββββββββββ
β β β
β ββββββββββΌβββββββββ ββββββββββΌβββββββββ
β β Sub-Agent Engineβ β Tool Execution β
β β (Read-only) β β Sequential / β
β βββββββββββββββββββ β Parallel β
β βββββββββββββββββββ
βββββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β ChatModelClient Interface β
β OllamaClient (Local) β GeminiClient (API) β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
ββββββββββΌβββββββββ
β Eval Runner β
β (9 Evaluators) β
βββββββββββββββββββ
The ContextManager (src/context/contextManager.ts) handles memory representation, context limits, and message cleanup across turns.
When a tool modifies a file (write_file or edit_file), any earlier read_file output for that same file path residing in message history becomes outdated and potentially misleading to the model.
ContextManager.invalidateStaleReads(mutatedPath)scans history for pastread_filetool results matching the mutated file.- It overwrites the old content with a tombstone message:
[File content of <path> has been modified by a subsequent edit/write tool call. This read result is now stale and has been invalidated to save context space.] - Benefits: Prevents context window bloat, eliminates stale code references, and reduces token cost.
Continuous passive optimization (microcompact() in src/context/compaction.ts) runs before every turn. It scans history and deduplicates redundant tool operations (such as repeated identical file reads or superseded checks), replacing duplicated payload blocks with compact references.
When total session tokens exceed the configured threshold (default: 8,000 tokens):
compactIfNeeded()invokessummarizeHistory().- The LLM summarizes the oldest 50% of the message transcript into a structured system summary block.
- The original old turns are removed from the active context window, and the summary is inserted as a system block.
- Active tool call/result pairing structures are preserved to maintain model validation constraints.
ContextManager.getPayload() automatically attaches cache_control: { type: "ephemeral" } metadata to:
- The base System Prompt (containing runtime environmental rules and tool specifications).
- The
AGENT.mdProject Memory block.
This enables models supporting prompt caching (e.g. Anthropic / Gemini) to skip redundant prompt processing overhead across chat turns.
At startup, Agent checks for an AGENT.md file in the workspace root. If present, its contents are injected into system context as persistent project memory (coding conventions, architecture guidelines, forbidden files).
Session persistence (src/session/sessionStore.ts) uses an append-only JSONL format to guarantee crash-resilient storage and support non-linear conversation branching.
Sessions are saved in ~/.harness/agent/session/--<encoded-cwd>--/<sessionId>.jsonl.
Each line in the file represents a single SessionEntry with strict metadata:
export interface BaseEntry {
id: string; // Unique entry UUID/timestamp identifier
parentId: string | null; // Pointer to preceding entry ID
timestamp: string; // ISO timestamp
}SessionMessageEntry: User, assistant, or tool interaction messages.ThinkingLevelChangeEntry: Model thinking parameter changes.ModelChangeEntry: Model or provider switches.CompactionEntry: System summary generated during context compaction.BranchSummaryEntry: Context state preserved when creating alternative branches.CustomEntry/CustomMessageEntry: Metadata extensions and hook payload data.LabelEntry/SessionInfoEntry: User annotations and session summaries.
Because every entry explicitly references a parentId, history forms a Directed Acyclic Graph (DAG) / Tree:
ββββ Entry 3 (Branch A) βββ Entry 4A
Entry 1 ββ Entry 2
ββββ Entry 3 (Branch B) βββ Entry 4B <-- leafId
branch(branchFromId): Rewinds the session's activeleafIdback to an earlier entry. Subsequent entries are appended as children ofbranchFromId, creating a side branch without altering original history.branchWithSummary(branchFromId, summary): Creates a branch and records aBranchSummaryEntrycapturing context from the abandoned branch.resetLeaf(): Rewinds the leaf pointer to root, allowing full conversation restarts.
When preparing messages to send to the LLM:
buildSessionPath(entries, leafId)starts atleafIdand walks backwards viaparentIdpointers to the root. It reverses the list to reconstruct the active linear path.buildContextEntries(path)inspects the resolved path for anyCompactionEntry. If found, it drops raw messages prior tofirstKeptEntryIdand prefixes context with the compaction summary.sessionEntryToContextMessages()converts the entries into standard modelMessageobjects.
To prevent creating empty session files when users open and immediately close the CLI, Session uses deferred flushing:
- New sessions accumulate entries in memory (
flushed = false). - Disk creation and bulk writing (
flush()) occur only when the first assistant message is generated. - After flushing, subsequent entries are appended synchronously (
fs.appendFileSync).
Session files include a version header and automatically upgrade legacy formats upon loading:
- v1 β v2: Upgrades flat message arrays to tree nodes (
id,parentId) and converts index pointers to entry IDs. - v2 β v3: Normalizes legacy roles (
hookMessageβcustom).
SessionStore.getLatestSessionId()readslatest_id.txtto identify the most recent session for the current workspace.- The REPL CLI automatically detects existing sessions and prompts the user to resume or start fresh upon launch.
- Dynamic
/modelschanges,/modetoggling,/session list|switch|new|clear, andclearcommands immediately sync session state.
The Guardrail subsystem (src/guardrails/) provides strict, multi-tiered safety gating for all agent actions. Rather than a basic binary prompt, it enforces granular security policies before any tool execution occurs.
When a tool is invoked, PolicyEngine evaluates the action against all active policies concurrently:
DENY(Critical / Terminal): If any policy issues aDENY, execution is immediately aborted with an explanation.DENYtakes absolute precedence overASKandALLOW, bypassingautoConfirm.ASK(Interactive Prompt): Requires explicit user confirmation[y/N]. In headless mode withautoConfirm: true, safeASKdecisions are auto-approved while logging risk levels.ALLOW(Automatic): Low-risk, non-destructive operations proceed without user intervention.
- Directory Traversal Defense: Blocks path traversal attacks (
../,/etc/passwd, root escaping) attempting to operate outside the workspace root. - Sensitive File Shielding: Blocks reading or modifying critical secrets, credentials, and environment configuration:
.env,.env.*,*.pem,*.key,id_rsa,credentials.json,.aws/credentials.
- Auto-Approve Safe Read Commands: Safe commands (
git status,git log,git diff,ls,dir,bun test,npm test) are auto-approved. - Confirm Mutating Commands: Modifications (
npm install,git commit,git push, build commands) require confirmation. - Strictly Block Dangerous Commands: Destructive system calls (
rm -rf,mkfs, fork bombs, system privilege escalation) are strictly denied.
- Uses regex pattern matching to detect hardcoded secrets in tool parameters, generated code, and file writes.
- Detects API keys (OpenAI, Anthropic, Gemini, GitHub tokens, AWS access keys, private RSA keys, Bearer tokens).
- Automatically redacts exposed tokens before logging or sending to models and triggers authorization warnings.
- Enforces maximum payload thresholds on file writes and edits (e.g. 500 KB limit).
- Prevents buffer overflow and denial-of-service from runaway tool calls or excessively large command line inputs.
- Detects and strictly denies reverse shell patterns (e.g.
nc -e,bash -i >& /dev/tcp,telnet). - Blocks credential exfiltration payloads attempting to curl or post sensitive environment variables to remote endpoints.
- Whitelists safe local network endpoints (
localhost,127.0.0.1) while gating unlisted outbound hostnames.
When a write or edit tool requires user approval in interactive REPL mode, the policy engine computes a line-by-line unified diff and prints a colorized preview in the terminal before prompting:
β οΈ Guardrail Authorization Request [HIGH Risk]
Reason: Mutating workspace file
Tool: edit_file
Arguments: { "path": "src/auth.ts", ... }
--- Diff Preview ---
- const secret = "hardcoded_token";
+ const secret = process.env.AUTH_SECRET;
--------------------
Authorize this action? [y/N]:
The Snapshots subsystem (src/snapshots/ and SNAPSHOTS.md) provides atomic, non-destructive state capture and rollback capabilities for autonomous coding agents.
Rather than creating intrusive git branches or lossy stashes:
- Uses an isolated alternate index file (
GIT_INDEX_FILE). - Writes tree and commit objects directly to Git's object database.
- Updates snapshot references under
refs/harness/snapshots/<id>. - Zero Working Tree Contamination: Never alters the user's active branch, working index, or
HEAD. - Automatically excludes secrets (
.env), build directories (dist,node_modules,coverage), and private keys.
For directories that are not Git repositories, SnapshotStore falls back to an automated tar/file archive strategy under .harness/snapshots/ with identical metadata and rollback semantics.
When rolling back to a previous snapshot:
- Automatically takes a Safety Snapshot before modifying any files, ensuring every rollback can be undone.
- Restores modified and deleted files to their exact previous state.
- Automatically detects and deletes untracked files created by the agent after the snapshot.
Checkpoints encapsulate both the workspace files and the complete agent execution state:
- Workspace file snapshot reference
- Full conversation message history
- Active task and plan
- Evaluation state and metrics
- Standalone CLI:
bun run snapshot create "Pre-refactor checkpoint" bun run snapshot list bun run snapshot diff <snapshot-id> bun run snapshot restore <snapshot-id>
- REPL Slash Commands:
/snapshot create [desc],/snapshot list,/snapshot restore <id>,/snapshot diff <id>/checkpoint create [desc],/checkpoint list,/checkpoint restore <id>
- Agent Built-in Tools:
create_checkpoint: Agent automatically creates checkpoints prior to risky refactoring.rollback_checkpoint: Restores workspace and context if verification tests fail.list_checkpoints: Discovers available restore points.
The harness natively implements an MCP client (src/mcp/) following the standard JSON-RPC 2.0 specification over stdio transport.
Manage external tool providers through a project-level .mcp.json configuration file:
{
"mcpServers": {
"duckduckgo": {
"command": "bun",
"args": ["run", "src/mcp/servers/duckduckgo.ts"]
}
}
}- Spawns subprocesses and establishes standard bidirectional stdio channels.
- Full protocol lifecycle: Handshake initialization (
initialize), capability negotiation, tool discovery (tools/list), and execution (tools/call). - Built-in request timeouts, automatic error classification, and connection recovery.
- Discovered MCP tools are adapted seamlessly into native harness
Tooldefinitions. - Automatic schema normalization: Cleans unsupported JSON Schema constructs and strips incompatible metadata (e.g.
x-mcp-headerfor Google Gemini models). - MCP tool execution passes through the Guardrails
PolicyEnginebefore reaching the external server.
Includes an out-of-the-box MCP server (src/mcp/servers/duckduckgo.ts) providing live web search results with link scraping and markdown extraction without requiring external API keys.
# List configured servers
bun run mcp list
# Connect to servers and display discovered tools
bun run mcp tools
# Register a new MCP server
bun run mcp add <server-id> <command> [args...]
# Test connection and tool discovery
bun run mcp test <server-id>
# Remove an MCP server
bun run mcp remove <server-id>The Evaluation framework (src/evals/ and src/cli/eval.ts) allows rigorous benchmarking and grading of autonomous agent runs across realistic engineering tasks.
Tasks are organized in benchmark datasets (e.g. evals/datasets/coding-harness-v1/):
- Categories:
bug-fix,feature,refactor,debugging,testing,security,repo-navigation,multi-step. - Difficulty: Graduated difficulty tiers (
L1toL5). - Hidden Tests & Constraints: Optional hidden validation suites and forbidden path/command constraints.
{
"id": "BUG-001",
"category": "bug-fix",
"difficulty": "L1",
"description": "Fix off-by-one error in search pagination",
"prompt": "Fix pagination offset handling so the last page does not duplicate records.",
"verification": {
"testCommand": "bun test src/search.test.ts"
}
}Runs a battery of specialized evaluators against the final workspace state:
build: Verifies project build compilation.typecheck: Runs TypeScript/language compiler checks.lint: Ensures code adhering to linting rules.tests: Executes unit and integration test suites.diff: Computes file additions, modifications, and deletion ratios.files: Enforces required and forbidden path boundaries.requirements: Evaluates specific functional requirements.security: Verifies that no credentials leaked and no guardrail violations occurred.
Tracks granular operational metrics for each benchmark run:
- Total iterations & tool calls
- Token consumption (prompt + completion)
- Wall-clock execution duration
- Files modified
- Rollback & recovery attempts
- Guardrail blocks triggered
- Detailed console reports and
--jsonmachine-readable outputs
# List benchmark tasks in dataset
bun run eval list
# Run all tasks in the benchmark suite
bun run eval run
# Run specific task or category
bun run eval run --task BUG-001
bun run eval run --category bug-fix
# Run and output machine-readable JSON results
bun run eval run --json
# View report for a previous run
bun run eval report <run-id>The harness provides 13 core built-in tools in src/tools/, plus dynamically registered MCP tools:
| Tool Name | Type | Description |
|---|---|---|
read_file |
Read-only | Reads text files with line numbering, offset paging, and line range slicing. |
write_file |
Mutating | Creates new files or overwrites existing files completely. Gated by Guardrails. |
edit_file |
Mutating | Range-targeted find-and-replace using startLine/endLine, with sliding-window drift recovery (Β±10 lines tolerance) and exact line mismatch diagnostics. |
run_command |
Mutating | Executes shell commands on the host machine. Gated by PolicyEngine permission checks. |
check_syntax |
Read-only | Validates JavaScript/TypeScript files using Bun's internal bundler compiler to report syntax errors prior to execution. |
glob |
Read-only | Performs fast wildcard pattern file and directory scanning across the workspace. |
grep |
Read-only | Executes workspace text searches using system ripgrep (rg) with a native Bun glob fallback. |
todo_read |
Read-only | Reads the persistent checklist file (.todo.md). |
todo_write |
Mutating | Updates and manages the project task checklist (.todo.md). |
dispatch_subagent |
Read-only | Spawns an isolated, read-only sub-agent to perform deep research tasks without file mutation access. |
create_checkpoint |
Mutating | Creates a snapshot and saves the full runtime context (messages, task, plan) before risky edits. |
rollback_checkpoint |
Mutating | Restores both workspace files and conversation history back to a previous checkpoint. |
list_checkpoints |
Read-only | Lists all saved checkpoints with reasons and descriptions. |
| MCP Tools | Dynamic | Discovered from configured MCP servers (e.g. duckduckgo_search). |
The harness abstracts LLM integrations behind a unified ChatModelClient interface (src/providers/types.ts):
export interface ChatModelClient {
chatStream(
messages: Message[],
tools: ToolDefinition[],
onChunk: (chunk: { content: string; thinking: string; toolCalls: ToolCall[] }) => void
): Promise<Message>;
}- Ollama (
src/providers/ollama.ts):- Native streaming, thinking tag extraction (
<think>...</think>), raw tool payload parsing, and token usage reporting. - Built-in support for
qwen3:14b,qwen3:8b, andllama3.1:8b.
- Native streaming, thinking tag extraction (
- Gemini (
src/providers/gemini.ts):- Google Generative AI REST API streaming integration.
- Handles thought signatures (
thought_signature), structured tool declarations, and schema sanitization. - Built-in support for
gemini-1.5-flash,gemini-1.5-pro, andgemini-2.5-flash.
Model switching can be done interactively during REPL sessions using the /models command.
Launch using bun start or bun run src/cli/repl.ts:
- Features colored streaming responses and thinking visualization.
- Displays live tool invocation summaries, diff previews, and execution results.
- Includes interactive command shortcuts:
/models: Interactive model selection menu./mode: Toggle betweenparallelandsequentialtool execution./session: Session management (/session list,/session switch <name>,/session new,/session clear)./snapshot: Workspace snapshots (/snapshot create,/snapshot list,/snapshot restore <id>,/snapshot diff <id>)./checkpoint: Full runtime checkpoints (/checkpoint create,/checkpoint list,/checkpoint restore <id>).clear: Reset history and delete current session.exit/quit: Terminate the REPL.
- Displays token usage metrics and context window percentage after every turn.
Launch using bun run src/cli/headless.ts:
bun run src/cli/headless.ts --task "Fix bug in search parser" --cwd "/path/to/repo" --max-iterations 30- Redirects logs to
stderrand prints clean JSON results tostdout:
{
"status": "success",
"output": "Task completed successfully...",
"filesChanged": ["src/parser.ts"]
}The Agent loop supports two tool execution modes:
- Parallel Mode (Default): When the model outputs multiple tool calls in a single turn, permissions are checked sequentially, and all approved tool executions run concurrently via
Promise.all. - Sequential Mode: Executes tool calls one by one in serial order.
Pressing Ctrl+C (SIGINT) while the agent is running tool cycles triggers mid-run steering:
- The agent loop pauses after the current tool execution completes.
- Prompts the user for a steering instruction:
steer instruction (or press Enter to resume)>. - Injects
[User Steering Instruction]: <input>into conversation history without destroying session context. - Pressing
Ctrl+Ca second time forces an immediate program exit.
- Install Bun (v1.0+):
powershell -c "irm bun.sh/install.ps1 | iex" - Install and launch Ollama (optional if using Gemini API key):
ollama pull qwen3:14b ollama serve
- (Optional) Configure Gemini API key in
src/.envor project root.env:GEMINI_API_KEY=your_gemini_api_key_here
git clone https://github.com/raghuttama-dev/Coding-harness.git
cd Coding-harness
bun installbun startbun run src/cli/headless.ts --task "Refactor search utility to use async/await" --cwd "."# Capture workspace snapshot
bun run snapshot create "Pre-refactoring state"
# List snapshots
bun run snapshot list
# Diff workspace against snapshot
bun run snapshot diff <snapshot-id>
# Restore workspace
bun run snapshot restore <snapshot-id># List configured servers
bun run mcp list
# Discover tools from active servers
bun run mcp tools
# Test server connection
bun run mcp test duckduckgo# List benchmark tasks
bun run eval list
# Run evaluation suite
bun run eval run
# Run specific task with JSON output
bun run eval run --task BUG-001 --jsonExecute the Vitest-compatible Bun test suite covering tools, context compaction, staleness tracking, session storage, guardrails, MCP, snapshots, and evals:
bun testCoding-harness/
βββ src/
β βββ agent.ts # Core execution loop, steering, and turn orchestrator
β βββ client.ts # Provider export bridge
β βββ cli/
β β βββ repl.ts # Interactive terminal REPL interface
β β βββ headless.ts # Non-interactive JSON automation CLI entry point
β β βββ snapshot.ts # Snapshot & Checkpoint management CLI
β β βββ mcp.ts # Model Context Protocol CLI
β β βββ eval.ts # Benchmark evaluation runner CLI
β βββ context/
β β βββ contextManager.ts # History state, staleness tombstoning, token tracking
β β βββ compaction.ts # Microcompaction & LLM summarization compaction
β βββ guardrails/
β β βββ index.ts # Guardrail exports
β β βββ types.ts # Guardrail context, policy, and decision contracts
β β βββ policyEngine.ts # Coordinator enforcing DENY/ASK/ALLOW and diff preview
β β βββ pathPolicy.ts # Directory traversal and sensitive file shielding
β β βββ commandPolicy.ts # Command classification (safe / mutate / destructive)
β β βββ secretScanner.ts # Regex credential detection and token redaction
β β βββ resourcePolicy.ts # Payload size and argument length limits
β β βββ networkPolicy.ts # Reverse shell defense and socket isolation
β βββ snapshots/
β β βββ index.ts # Snapshot exports
β β βββ types.ts # Snapshot & Checkpoint type definitions
β β βββ snapshotManager.ts # Coordinator for snapshots, checkouts, and diffs
β β βββ snapshotStore.ts # Git plumbing & file archive storage engine
β β βββ rollback.ts # Clean rollback & safety snapshot generator
β β βββ diff.ts # Unified diff & LCS comparator
β βββ mcp/
β β βββ index.ts # MCP client exports
β β βββ types.ts # JSON-RPC & MCP protocol types
β β βββ client.ts # Stdio transport JSON-RPC client
β β βββ manager.ts # Multi-server connection & lifecycle manager
β β βββ serverRegistry.ts # .mcp.json persistent server configuration
β β βββ toolAdapter.ts # Adapts MCP tools to native harness Tool contracts
β β βββ servers/
β β βββ duckduckgo.ts # Built-in DuckDuckGo search MCP server
β βββ evals/
β β βββ types.ts # Benchmark task schema, metrics, and evaluator types
β β βββ taskLoader.ts # Dataset task discovery and parser
β β βββ evalRunner.ts # Benchmark execution orchestrator
β β βββ evaluatorRegistry.ts # Registry of evaluation criteria
β β βββ execHelper.ts # Isolated task execution runner
β β βββ evaluators/ # Evaluator implementations
β β β βββ build.ts # Project build verification
β β β βββ typecheck.ts # Type-check compiler validation
β β β βββ lint.ts # Code style linter checks
β β β βββ tests.ts # Unit and integration test verification
β β β βββ diff.ts # Patch and file change inspection
β β β βββ files.ts # Required/forbidden file constraint validation
β β β βββ requirements.ts # Requirement specification matching
β β β βββ security.ts # Guardrail & secret leak validation
β β βββ reports/
β β βββ reportGenerator.ts # Terminal & JSON evaluation report formatter
β βββ permissions/
β β βββ permissionGate.ts # Backward-compatible permission gate bridge
β βββ providers/
β β βββ types.ts # ChatModelClient, Message, and ToolCall interface types
β β βββ ollama.ts # Ollama API client implementation
β β βββ gemini.ts # Gemini REST API client implementation
β βββ session/
β β βββ sessionStore.ts # Append-only JSONL tree persistence, branching, and migrations
β βββ tools/
β β βββ index.ts # Unified ToolRegistry (13 built-ins + MCP)
β β βββ read.ts # Range-sliced file reader with line numbers
β β βββ write.ts # File creator and overwriter
β β βββ edit.ts # Targeted find-replace editor with sliding drift recovery
β β βββ bash.ts # Command execution tool
β β βββ checkSyntax.ts # JS/TS syntax validator (Bun build compiler)
β β βββ glob.ts # Wildcard pattern file scanner
β β βββ grep.ts # Ripgrep-backed workspace search tool
β β βββ todo.ts # Task list checklist management (.todo.md)
β β βββ subagent.ts # Read-only background sub-agent dispatcher
β β βββ checkpoint.ts # Checkpoint tools (create, rollback, list)
β β βββ activeClient.ts # Active LLM client reference container
β β βββ types.ts # Tool interface contracts
β βββ tests/
β βββ context.test.ts # Staleness and compaction test suite
β βββ executionMode.test.ts # Parallel vs sequential mode test suite
β βββ gemini.test.ts # Gemini provider test suite
β βββ search.test.ts # Glob and Grep test suite
β βββ tools.test.ts # File edit, read, syntax, and todo tool test suite
β βββ v4.test.ts # Session tree, diff, policy engine, and subagent test suite
β βββ guardrails.test.ts # Guardrails 5-policy comprehensive test suite
β βββ mcp.test.ts # MCP client, registry, adapter, and server test suite
β βββ snapshots.test.ts # Git plumbing snapshots and rollback test suite
β βββ evals.test.ts # Evaluation framework test suite
βββ evals/
β βββ datasets/
β βββ coding-harness-v1/ # Benchmark evaluation dataset
βββ AGENT.md # Workspace project memory rules file
βββ SNAPSHOTS.md # Workspace snapshots & checkpoints specification
βββ agent-harness-architecture.md # Architecture specification document
βββ pi-agent-session-storage.md # Session storage specification document
βββ package.json # Dependencies and run scripts
βββ tsconfig.json # TypeScript compiler configuration
MIT License. Built for autonomous AI agent research and development.