Inside vLLM - 01: How Model Weights Reach GPU MemoryFrom PyTorch's CachingAllocator to vLLM's CuMemAllocator: the exact path that places model weights in GPU memory. Aug 20, 2026 AI, vLLM