PyTorch Distributed Training Internals

The mechanism behind step 5 of the training loop: bucketed gradient all-reduce triggered from autograd hooks, what NCCL does with it, and how to diagnose the result.

August 12, 2026 · 13 min · Sandeep Kumar