Nowhere to Put the Canary: Argo Rollouts + vLLM
Running vLLM on a single GPU with room for two pods, and what that does to progressive delivery when there’s no capacity to spare for a canary.
Running vLLM on a single GPU with room for two pods, and what that does to progressive delivery when there’s no capacity to spare for a canary.