<feed xmlns="http://www.w3.org/2005/Atom"> <id>https://leinfinitr.github.io/</id><title>Jialong Liu</title><subtitle>Jialong Liu's personal website. M.S. student at Shanghai Jiao Tong University working on systems for heterogeneous accelerators, AI infrastructure, scheduling, and memory management.</subtitle> <updated>2026-08-20T21:35:24+08:00</updated> <author> <name>Jialong Liu</name> <uri>https://leinfinitr.github.io/</uri> </author><link rel="self" type="application/atom+xml" href="https://leinfinitr.github.io/feed.xml"/><link rel="alternate" type="text/html" hreflang="en" href="https://leinfinitr.github.io/"/> <generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator> <rights> © 2026 Jialong Liu </rights> <icon>/assets/img/favicons/favicon.ico</icon> <logo>/assets/img/favicons/favicon-96x96.png</logo> <entry><title>Inside vLLM - 01: How Model Weights Reach GPU Memory</title><link href="https://leinfinitr.github.io/posts/inside-vllm-01/" rel="alternate" type="text/html" title="Inside vLLM - 01: How Model Weights Reach GPU Memory" /><published>2026-08-20T19:00:00+08:00</published> <updated>2026-08-20T19:00:00+08:00</updated> <id>https://leinfinitr.github.io/posts/inside-vllm-01/</id> <content type="text/html" src="https://leinfinitr.github.io/posts/inside-vllm-01/" /> <author> <name>Jialong Liu</name> </author> <category term="AI" /> <category term="vLLM" /> <summary>From PyTorch's CachingAllocator to vLLM's CuMemAllocator: the exact path that places model weights in GPU memory.</summary> </entry> </feed>
