Great blog from Kevin Lu at @vllm_project about quality control at the inference engine. Unfortunately, @AnushElangovan has not provided enough stable AMD QA clusters to vLLM, which leads to an order of magnitude worse software quality on AMD versus CUDA.
Multiple AMD CI vLLM fleet-wide outages happen every month.