I am currently at Modal working on LLM inference. Most recently, I was a graduate researcher at Stanford working with Simran Arora and Simon Guo as part of the Hazy Research and Scaling Intelligence labs.
My work focuses on understanding how hardware and software co-design can unlock new capabilities in efficient AI computation. Previously, I wrote GPU kernels at Together AI, worked on GPU compilers at Qualcomm, and built distributed systems at Amazon.
-
HipKittens: Fast and Furious AMD Kernels — MLSys 2026 (Oral) (Blog, Technical Deep Dive)
William Hu, Drew Wadsworth, Sean Siddens, Stanley Winata, Daniel Y. Fu, Ryan Swann, Muhammad Osama, Christopher Ré, Simran Arora
-
DSL-Monkeys: Self-Generated In-Context Examples for Low-Resource GPU DSL Kernels — ICLR 2026 DATA-FM Workshop
Nathan Paek, Simon Guo, Vishnu Sarukkai, Willy Chan, William Hu, Ethan Boneh, Simran Arora, Ludwig Schmidt, Kayvon Fatahalian, Azalia Mirhoseini
-
Cylon: Asynchronous Linear Attention — COLM 2026
Alexander Waitz, William Hu, Benjamin F. Spector, Atri Rudra, Christopher Ré, Simran Arora
-
KernelBench: Can LLMs Write Efficient GPU Kernels? — ICML 2025 · ICLR 2025 DL4C (Best Paper) & SSI-FM Workshops
Anne Ouyang, Simon Guo, Simran Arora, Alex L. Zhang, William Hu, Christopher Ré, Azalia Mirhoseini