Tiling and Collective Operations
The two data-movement primitives behind training performance: tiling inside a device, collectives between devices.
The two data-movement primitives behind training performance: tiling inside a device, collectives between devices.
torch.profiler integrates PyTorch operator attribution, optional source stacks, and CUDA activity in one trace. Its numbers mean what you think only if you control the capture window, the labels and the host syncs.
How successive FlashAttention implementations respond to the bottleneck exposed by each new GPU generation.
FlashInfer treats LLM attention as a serving systems problem spanning data layout, kernel generation, and runtime scheduling.