███╗ ███╗████████╗███████╗ ██████╗███╗ ██╗██╗ ██████╗
████╗ ████║╚══██╔══╝██╔════╝██╔════╝████╗ ██║██║██╔════╝
██╔████╔██║ ██║ █████╗ ██║ ██╔██╗ ██║██║██║
██║╚██╔╝██║ ██║ ██╔══╝ ██║ ██║╚██╗██║██║██║
██║ ╚═╝ ██║ ██║ ███████╗╚██████╗██║ ╚████║██║╚██████╗
╚═╝ ╚═╝ ╚═╝ ╚══════╝ ╚═════╝╚═╝ ╚═══╝╚═╝ ╚═════╝
Tools for people running AI on hardware they own — serve it, measure it, shrink it, build with it, drive it.
No cloud. No API keys. No rented GPUs. Every number below came off a machine in my house.
Two threads run through everything here: agents that ship finished, validated software — not snippets — and making big models cheaper to run on hardware you can actually buy. Around both sits the unglamorous layer nobody writes: the thermal logging, the fan curves, the launch topology, the session dashboards. Results get published either way, including the ones that didn't work.
The lab: 2× DGX Spark · two 4× RTX 3090 boxes · a third box with 3× RTX 3090 + an RTX PRO 5000 Blackwell (72 GB) · an RTX 5090. Eleven 3090s is how you end up publishing benchmarks about interconnect topology instead of guessing about it — and why the thermal and fan-curve tooling below exists at all.
Coming out of the private pile. The stuff you only build after running a home lab long enough to get burned by it. nvidiacp and muxter are out now — they've moved down into the sections below.
| Repo | What it does |
|---|---|
| tempmon | Thermal watchdog for inference boxes, with the metric nvidia-smi can't give you: per-card GDDR6X VRAM junction temperature — the thing that actually cooks a 3090 (throttles ~105°C) and returns N/A on GeForce. Crash-proof by design: every cycle is flush()ed and fsync()ed, so when the box thermally hard-locks you still have proof of how hot it got. |
| ramp | Real Agents Maximizing Productivity — a local-first coding agent harness, built against a LAN of vLLM boxes and hardened by the runs that went wrong. See the cockpit below. |
Serving, benchmarking, and knowing what your hardware is actually doing.
| Repo | What it does | ★ |
|---|---|---|
| nvidiacp | Full NVIDIA control panel for headless Linux — power limits, app clocks, ECC, MIG, compute mode, persistence. Fan control without X (no coolbits), a closed-loop temperature-curve fan daemon, and a vLLM launch optimizer that sizes tensor-parallel plans to your actual VRAM across 5 workload profiles. Settings survive reboots via systemd. Pure stdlib Python + vendored NVML — zero pip installs. | |
| vllm-topology-bench | Replicas > tensor parallelism. Everyone reaches for --tensor-parallel-size 4 to "use all the GPUs." On no-NVLink 3090s that's the slowest feasible option: two TP=2 replicas beat one TP=4 at every concurrency — +136% throughput at 128 requests (1,437 vs 610 tok/s) and TTFT 1.4 s vs 3.9 s. And "one copy per card" doesn't even fit — 4× TP=1 OOMs. Reproducible harness included. |
|
| model-chat-cli | Terminal command center for local AI servers. Auto-discovers Ollama, LM Studio, and vLLM on your LAN — no config. Chat with real TTFT and decode metrics, run blind multi-model battles with auto-judged tournaments, load-test to saturation, and score models on a 45-task agentic benchmark across 6 difficulty tiers. |
Smaller models, fewer tokens, same behavior — with the measurements to prove it.
| Repo | What it does | ★ |
|---|---|---|
| tokopt | Post-hoc BPE tokenizer adaptation for already-deployed LLMs — script-aware pruning, continued BPE extension, embedding-only calibration. −4.9% bits/char, −4.2% tokens, −18% INT4 model size (4.4 → 3.6 GB), HumanEval within noise, greedy generation byte-identical. Reproduces in ~3 hours on a single RTX 3090. Ships a ~9,000-word PAPER.md and a 13-section REPRODUCE.md. | |
| research-test-Qwen3-Coder-Next-REAP-AWQ | REAP expert pruning + AWQ on a 149 GB BF16 MoE coder model. 512 → 410 experts per layer via 4-dataset saliency calibration with super-expert preservation, then W4A16 @ group_size 32 — to get a frontier-size coder onto consumer GPUs. | |
| ctx | CTX (Context Transfer Format) — an interchange format for LLM web consumption. A 1.2 MB Wikipedia page becomes 150 KB of structure-preserving CTX: −87% bytes, ~90% fewer tokens, citations and hierarchy intact. CLI, Python library, and a FastAPI service with Redis caching + transparent proxy. |
Autonomy is easy. Finished is the hard part.
| Repo | What it does | ★ |
|---|---|---|
| cadillac | Autonomous coding agent: a sentence in, a validated app out. 12-stage pipeline (SPEC → … → CRITIC → RUNTIME → PACKAGE) with operational gates, a completeness CRITIC, runtime flow verification, and surgical-mode stuck-loop recovery. Works against any OpenAI-compatible endpoint. 616 tests · 146 apps built · 403 lessons accumulated. | |
| cadillac-builds | The receipts: 26 runnable applications built unattended with zero human edits, published from a scan of ~99 build workspaces. Every project's README lists exactly which validation checks passed — and which didn't. | |
| graphx | Describe a pipeline in English, get a real agentic workflow, run it in your terminal — all on your own model. Pregel-style supersteps, cyclic graphs with loops, per-step SQLite checkpointing (kill a run, resume continues), per-node retries + model fallback chains + budgets, human approval gates, 12 credential-wired connectors, and secret:// refs that never reach logs or checkpoints. 298 tests. |
|
| workerAI | ReAct-style autonomous agent on vLLM + LangGraph — plans, reasons, and executes multi-step tasks against a library of specialized tools. |
Driving a multi-node lab from one chair.
| Repo | What it does | ★ |
|---|---|---|
| clusterspace | Tiled workspace (Electron + React + TS) for terminals, embedded Chromium, and SSH panes auto-wrapped in tmux — sessions survive disconnects, restarts, and closing the app. Optional AI co-pilot gets 71 callable tools (including ~50 browser-automation tools and a pair of vision tools that judge what's on screen after an action) and can drive any pane toward a goal. Tracks eight concurrent agents live. | |
| muxter | Every tmux session on the machine, in one terminal, live. Read-only by default so you can watch a long build or an agent mid-task without risking a keystroke; press i to step in, Esc to step back out. Attaches as an ordinary tmux client — nothing to install on the sessions. Check an agent fleet from a phone SSH app. Pairs with clusterspace, which names its tmux sessions clusterspace-pane-… — type /clusterspace and the list narrows to exactly those. |
|
| code-atlas | Any repo as a navigable 3D world — city (height = LOC, glow = git churn), dependency galaxy, and molecule view of one file's symbols; the three morph into each other. Tarjan-SCC cycle detection, hotspot/ownership/coverage lenses, git time-scrub, local-LLM integration. Rendered vLLM — ~6,000 files, 1.5M LOC, 23,574 resolved imports — in seconds. | |
| ramp | Real Agents Maximizing Productivity — a local-first coding agent harness. One transport, one swappable wire per vendor: every local OpenAI-compatible server is just a URL, and Anthropic / Gemini / Bedrock are spoken in their own dialects behind the same seam. Most of what's in it is there because a specific run went wrong — the budget module exists because one afternoon produced 273 context overflows in 52 minutes; max_tokens is derived from the window because 5,949 recorded tool calls showed what responses actually need. ~4,000 hermetic tests, random order. |
Every figure from a repo you can clone, on hardware I own:
| Result | Where |
|---|---|
| +136% throughput and ~2× lower latency from less tensor parallelism (2×TP=2 vs 1×TP=4, 4× RTX 3090, no NVLink) | vllm-topology-bench |
| −4.9% bits/char and −4.2% tokens on held-out code — HumanEval within noise, greedy generation byte-identical | tokopt |
| −18% INT4 model size (4.4 → 3.6 GB) after tokenizer surgery + calibration, for ~25 min of one-time cost | tokopt |
| 20% of MoE experts removed (512 → 410/layer) from a 149 GB coder model, then W4A16 | REAP-AWQ |
| −87% bytes / ~90% fewer tokens on real web pages, hierarchy and citations preserved | ctx |
| 146 apps built autonomously; 26 published unedited with per-check validation status | cadillac · builds |
| 273 context overflows in 52 minutes — the failure that became a budget module | ramp |
- Reproduce or it didn't happen. Research repos ship a REPRODUCE.md you can actually follow, not a citation to a vibe.
- Negative results get published. tokopt's paper includes the method that didn't beat the baseline (hierarchical merge-tree init) right next to the ones that did.
4× TP=1OOMing is a headline finding, not a footnote. - Validation status is public. cadillac-builds lists every failed check, per project. Nothing is dressed up to look better than it is.
- Runs on hardware you can own. Consumer GPUs throughout. No cloud dependency, no API key required to reproduce anything above.
- Point it at your own model. Every tool here takes a URL. If it only works with someone else's API, it isn't finished.
🎲 Other things I've built — not local-AI, still fun
| Repo | What it is | ★ |
|---|---|---|
| homestead | Sims-style house builder in the browser. Draw walls, rooms detect themselves, then paint, furnish, and walk around inside. three.js + TS, 4 runtime deps, 475 tests, zero art assets — everything generated at runtime. | |
| neon | Cyberpunk trading game where the currency is hardware. Build rigs from GPUs, RAM and CPUs; run an empire. | |
| matrix-doom | Matrix-themed ASCII FPS raycaster in pygame — built autonomously by an LLM code generator, shipped unedited. | |
| airowling-novels | 16 novels · 553,959 words · 221 chapters, written end-to-end by an AI authoring pipeline. All first-pass, no human rewriting. | |
| cardboard | Full-stack social platform for sports card collectors — share, buy, sell, trade. | |
| koding | CodeQuest — gamified visual Python learning for kids 9+. | |
| cblchat | Real-time enterprise chat with LDAP / Active Directory auth. |
wAIve.online — self-hosted AI platform · Fox Valley AI Foundation — open tooling for everyone else
Oshkosh, Wisconsin.

