Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club
At our latest YC Paper Club, researchers and builders presented on multi-GPU kernel optimization, intelligence per watt for local inference, AI-generated GPU kernels and benchmarking, heterogeneous inference infrastructure design, and GPU-accelerated game engines for reinforcement learning. Thanks to the following presenters: Stuart Sol (Stanford / Cursor), John (Stanford), Mark (PyTorch / GPU Mode / CoreAuto), Misha (Marlo), and Brennan (Stanford) Chapters: 0:00 – Francois Chaubard: The case for chip and kernel specialization 7:16 – Stuart Sul: Parallel Kittens - Systematic and Practical Simp