Skip to content
#

single-gpu

Here are 27 public repositories matching this topic...

Run Qwen3.8-Flash-Next on ONE RTX 3090 (24 GB) + 64 GB RAM: 128K context, up to 2,100 tok/s prefill, 43–51 tok/s decode. vLLM runtime with hot MoE experts on the GPU and cold experts computed on the CPU, INT8 KV cache, Docker, OpenAI-compatible API.

  • Updated Sep 30, 2026
  • Python

Engine for two MoE models only - DeepSeek-V4.1-Flash (main) and GLM-5.3-Flash - on one 96 GB GPU, experts offloaded to CPU RAM, 262K context. Pieced together from what we had (DDR4, PCIe 4); DDR5 would do better. Sleeps/wakes in seconds to share the GPU. OpenAI/Anthropic API, works behind LiteLLM.

  • Updated Oct 4, 2026
  • C++

A lightweight, end-to-end implementation of Stable Diffusion built from first principles on a single T4 GPU. Features a custom 192-channel U-Net, VAE, and a CLIP encoder, optimized for consumer hardware and trained on approx. 168k images.

  • Updated Feb 18, 2026
  • TypeScript

Add this topic to your repo

To associate your repository with the single-gpu topic, visit your repo's landing page and select "manage topics."

Learn more