frit
work in progress

GPU reliability engineering at homelab scale. Running the full inference stack with real SLOs, load tests, and chaos, practicing the patterns that hold at 1000.

milestones
M0GPU Foundationshipped driver · DCGM · k3s · GPU-in-k8s · vLLM on the L4 — bare VM → stack → bare, one command
M1Full-Stack Observabilityshipped DCGM (GPU) + Ray Serve + vLLM engine metrics → Prometheus → Grafana · serve / serve-llm / DCGM dashboards, via Flux
M2Inference Layer + Token Pathshipped vLLM + Ray Serve + LiteLLM + Open WebUI · OpenAI-compatible token path · multi-model serving matrix
M3Load Testing + Benchmarksshipped guidellm harness, GitOps-triggered · TTFT / ITL / TPOT + throughput · 4-precision sweeps (INT4 / FP8 / FP16) · public benchmarks page
M4Multi-Node Autoscalingactive Ray Serve autoscale across L4 nodes · hold TTFT under burst · real horizontal scale-out
M5 – M8  ·  SLOs, chaos, postmortems, OSS cadence