William Hu

William Hu

I am currently at Modal working on LLM inference. Most recently, I was a graduate researcher at Stanford working with Simran Arora and Simon Guo as part of the Hazy Research and Scaling Intelligence labs.

My work focuses on understanding how hardware and software co-design can unlock new capabilities in efficient AI computation. Previously, I wrote GPU kernels at Together AI, worked on GPU compilers at Qualcomm, and built distributed systems at Amazon.

Publications