Extremely fast model loads from HTTP/HTTPS, Redis, and S3 endpoints. GPT-J (20GB) loads at wire-speed (~5GB/s) on a 40GbE network, and is only bottlenecked by the Linux kernel TCP stack. CoreWeave and ...
If that reaches /health, then the remaining problem is specifically the autotune dummy decode path. If it still fails during CUDA graph capture or first decode with the same ...