verl for a single consumer GPU. PPO, GRPO and on-policy distillation on NVIDIA GPUs.
-
Updated
Sep 15, 2026 - Python
verl for a single consumer GPU. PPO, GRPO and on-policy distillation on NVIDIA GPUs.
Non-intimidating guide to create a KVM GPU Passthrough via libvirt/virt-manager on systems with only one GPU.
Train Dense Passage Retriever (DPR) with a single GPU
A 124.6M LLM trained from scratch on 13B tokens on a single RTX 4090 — tokenizer, pretraining, post-training, evaluation, and reproducible inference.
Seamless NVIDIA GPU hot handoff between Proxmox host and VM — bind/unbind nvidia ⇆ vfio-pci safely, no reboots.
A no-code browser-based tool that enables domain experts to fine-tune AI language models using their own knowledge with nothing more than a CSV file.
🔌 单卡 GPU LLM 推理网关 · 模型即插件 · 三态 GPU · 9 云端预设 · macOS Dashboard
Run Qwen3.8-Flash-Next on ONE RTX 3090 (24 GB) + 64 GB RAM: 128K context, up to 2,100 tok/s prefill, 43–51 tok/s decode. vLLM runtime with hot MoE experts on the GPU and cold experts computed on the CPU, INT8 KV cache, Docker, OpenAI-compatible API.
🚀 Achieve rapid training of NanoGPT (GPT-2 124M) on a single RTX 4090, targeting a validation loss below 3.28 with FineWeb-Edu data.
Cog Single GPU Quantized Implementation of Step-Video-T2V
Reproducibly turn a compatible Swift-Qwen3.8-27B W4A16 checkpoint into a syv-optimised single-GPU fast variant: target-calibrated draft vocabulary, GPTQ int4 lm_head/MTP, structural verification, RTX 3090 benchmarks.
GPT-2-class language models trained from scratch in PyTorch on one RTX 3090, with 10B-token data curation and full GPT-2 comparisons.
Autonomous research stack for continuously improving LLM training through automated experimentation. Single-GPU research labs. Karpathy-inspired.
Engine for two MoE models only - DeepSeek-V4.1-Flash (main) and GLM-5.3-Flash - on one 96 GB GPU, experts offloaded to CPU RAM, 262K context. Pieced together from what we had (DDR4, PCIe 4); DDR5 would do better. Sleeps/wakes in seconds to share the GPU. OpenAI/Anthropic API, works behind LiteLLM.
Evidence-first collection of from-scratch language models trained on a single NVIDIA L20
A lightweight, end-to-end implementation of Stable Diffusion built from first principles on a single T4 GPU. Features a custom 192-channel U-Net, VAE, and a CLIP encoder, optimized for consumer hardware and trained on approx. 168k images.
AI agents running research on single-GPU nanochat training automatically
Whole slide training with single GPU
Neurosymbolic AI organism built from scratch on a single GTX 1080: latent-space reasoning, frozen growth stages, knowledge as graph edges, automated hypothesis discovery. 15 months of honest pre-registered research journals (RU).
🧠 Minimal, hackable Group Relative Policy Optimization (GRPO) for LLM alignment — the algorithm behind DeepSeek-R1. Train reasoning models on a single GPU.
To associate your repository with the single-gpu topic, visit your repo's landing page and select "manage topics."