Serving MedGemma 27B on Modal: FP8, vLLM sleep mode and 21-second cold starts
How I self-host MedGemma 27B as a scale-to-zero, OpenAI-compatible API on Modal: FP8 quantisation, GPU snapshots, two vLLM bugs and Gemma 3 tool calling.
10 min read#llm#vllm#self-hosting