inference benchmarks

guidellm load tests against models served on frit's GPU. Each row is one model × precision × profile combination, open its report for the full guidellm dashboard (latency distributions, token stats, SLOs).

Profile is the load shape: synchronous = one request at a time (latency floor), throughput saturates the GPU (ceiling), sweep traces the curve between. Open a report for the full dashboard: latency distributions (TTFT / ITL / TPOT), token stats, throughput, and SLOs.
ModelPrecisionGPUProfileConcurrencyReport
loading…