Abstract
Nsight Compute (ncu) looks inside one kernel at a time. It reads hundreds of hardware counters and reports what fraction of the GPU’s peak memory and compute throughput the kernel reached, how many warps were resident, and why warps stalled. From that you can tell whether a kernel is memory-bound, compute-bound, or latency-bound. The GPU cannot read all of these counters in one run, so ncu replays the kernel many times. To make the replays comparable, it locks the GPU clocks, flushes the caches before each pass, and serializes kernel launches by default. These are the right conditions for measuring counters, but they are not the conditions the kernel runs under in production. We profile the GEMM kernels of the small PyTorch model from the torch.profiler post, show how to read the main sections, and measure how far ncu’s default duration moves from the duration seen in nsys.
1. Introduction
The problem. nsys or torch.profiler has pointed at one kernel that takes most of the GPU time. Now we need to know why it is slow: is it waiting on memory, saturating the math units, or simply too small to fill the GPU?
What readers usually assume. ncu’s “Duration” is the kernel’s real runtime, and its numbers can be compared directly with a timeline.
What we show. ncu is the right tool for the why, but its durations are taken under controlled conditions that differ from production. Use ncu to find the limiter, and use nsys to measure how long the kernel actually takes.
2. How ncu collects data
- Kernel replay. The GPU can only read a limited set of counters per pass. ncu saves the kernel’s memory state, runs the kernel, restores the memory, and runs it again with the next counter group, until all requested counters are collected. A large section set can need dozens of passes per kernel [1].
- Serialization. ncu runs profiled kernels one at a time, so kernels that overlapped on different streams no longer overlap.
- Controlled conditions. By default, ncu locks the GPU clocks (
--clock-control) and flushes the caches before each pass (--cache-control all), so that results are repeatable across runs [1] [VERIFY: clock-control default per version]. - Sampling (optional). Program-counter (PC) sampling records where warps are stalled, at the level of SASS instructions or, with
-lineinfo, source lines.
3. A minimal example
We use the same training script as in the torch.profiler post, without any profiler code, and save it as train.py. The linear layers run as GEMM kernels from cuBLAS. ncu filters on kernel names, so we first list them, for example with nsys stats -r cuda_gpu_kern_sum from the nsys post, and then profile a few matching launches after warmup:
ncu --set basic \
-k regex:gemm \
--launch-skip 10 --launch-count 3 \
-o simple_ncu \
python train.py
ncu -i simple_ncu.ncu-rep --page details
What each flag does:
--set basic. Collects a small default group of sections: speed-of-light throughput, launch configuration and occupancy.--list-setsshows the others.fullcollects everything, at many more replays per kernel.-k regex:gemm. Profiles only kernels whose names matchgemm[VERIFY: actual kernel names for this model, which depend on the cuBLAS version and GPU]. Without a filter, ncu profiles every kernel the program launches.--launch-skip 10 --launch-count 3. Skips the first 10 matching launches (warmup) and profiles the next 3.-o simple_ncu. Writes a report file.ncu -i ... --page detailsprints it as text; the GUI shows the same sections with charts.
3.1 The output
4. Reading the sections
| Section | What it reports | How to read it |
|---|---|---|
| GPU Speed Of Light | Memory and compute throughput as a % of the GPU’s peak, and the kernel duration | High memory %, low compute %: memory-bound. The reverse: compute-bound. Both low: latency-bound or too little work |
| Roofline | The kernel’s arithmetic intensity (FLOPs per byte) against the GPU’s peak bandwidth and compute | Left of the ridge point: limited by memory bandwidth. Right of it: limited by compute |
| Launch Statistics | Grid and block size, registers and shared memory per thread block | A grid with fewer blocks than the GPU has SMs leaves SMs idle |
| Occupancy | Theoretical and achieved resident warps per SM | Low achieved occupancy leaves too few warps to hide memory latency |
| Warp State Statistics | Average cycles warps spend in each stall reason | The largest stall reason points to the limiter, e.g. waiting on memory vs. waiting on a math pipe |
| Source Counters | Stalls and instruction counts per SASS instruction or source line | Only useful when you can change the code; needs -lineinfo for source lines |
Observation 1. [MEASURE: Speed Of Light Memory % and Compute % for the GEMM kernels of SimpleModel, and the number of thread blocks vs. the number of SMs.]
Insight. With a batch of 1024 and layers of 5–200 features, each GEMM has very little work. We expect both throughput numbers to be low and the grid to cover only a few SMs, which means the kernel is latency-bound: it finishes before the GPU is anywhere near either roof.
Implication. For small kernels, the fix is more work per launch (larger batches, fused kernels), not a faster kernel. ncu shows this directly; a timeline only shows that the kernel is short.
5. Why ncu’s duration is not the production duration
Observation 2. [MEASURE: duration of the same GEMM launch in (a) nsys, (b) ncu with defaults, (c) ncu with
--clock-control none --cache-control none.]Insight. Three defaults move ncu away from production. Locked clocks run the GPU at a fixed frequency instead of its boost clock. Flushed caches force every pass to start cold, while in a real loop the weights of a small model may stay in L2. Serialization removes any overlap with other streams.
Implication. Use ncu’s throughput percentages and stall reasons to find the limiter. Take the kernel’s real duration from nsys. If you need ncu durations closer to production, turn off clock and cache control, and then expect more run-to-run variation.
6. Options reference
| Option | What it does | Why it matters |
|---|---|---|
--set basic / full (--list-sets shows all sets) | Selects groups of sections | full can need dozens of replays per kernel |
--section <name> / --metrics <list> | Collects only named sections or metrics | Fewer replays |
-k, --kernel-name regex:<pat> | Profiles only matching kernels | Essential on real workloads |
--kernel-name-base function / demangled / mangled | Which form of the name -k matches against | Templated names differ between forms |
-s, --launch-skip / -c, --launch-count | Skips the first N matches, then profiles M | Skips warmup |
--nvtx --nvtx-include <range> | Profiles only kernels inside an NVTX range | Ties ncu to the nsys view |
--replay-mode kernel / application / range / app-range | Replays one kernel, the whole program, or a range | application avoids state save/restore but needs a deterministic program |
--cache-control all / none | Flushes caches before each replay (default all) | Cold-cache numbers by default |
--clock-control base / none | Locks clocks to base (default base) [VERIFY] | Numbers are not at boost clock |
--graph-profiling node / graph | Profiles kernels inside CUDA graphs individually, or the whole graph as one unit | Needed for graph-heavy programs |
--target-processes all | Follows child processes | Needed for multi-process programs |
--import-source yes (build with -lineinfo) | Adds per-line source view | Only useful with source code |
Run ncu --help for the exact flags in your version [2].
7. Pitfalls
- Permissions. By default, reading GPU performance counters needs administrator rights. Without them, ncu fails with
ERR_NVGPUCTRPERM. The fix is a driver setting that an administrator must enable [3]. - No filter. Without
-kand a launch count, ncu profiles every kernel, each replayed many times. A program that runs in seconds can take hours. - Child processes. ncu only profiles the process it launched unless you pass
--target-processes all. - CUDA graphs. Kernels inside a graph are only profiled individually with
--graph-profiling node[VERIFY: default per version]. - Counter conflicts. Only one client can hold the GPU counters. A running DCGM profiling job or nsys GPU-metrics sampling can block ncu [VERIFY: exact behaviour].
- Timeouts and collectives. Serializing and replaying one GPU’s kernels stalls any collective that waits on it. In multi-GPU programs this can trigger timeouts.
8. What it cannot see
- Time between kernels. ncu looks at one kernel in isolation. Launch gaps, synchronization and overlap between streams are invisible. Use nsys.
- Which code launched the kernel. ncu knows the kernel name and launch configuration, not the Python operator. Use torch.profiler or NVTX ranges.
9. Conclusion
ncu answers “why is this one kernel slow?” by measuring it against the GPU’s peak memory and compute throughput. It does this under controlled conditions (locked clocks, cold caches, serialized launches) that make the counters repeatable but the durations different from production. Use ncu for the limiter and nsys for the time.
References
- NVIDIA. Kernel Profiling Guide. Nsight Compute documentation. https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html
- NVIDIA. Nsight Compute CLI. Nsight Compute documentation. https://docs.nvidia.com/nsight-compute/NsightComputeCli/index.html
- NVIDIA. ERR_NVGPUCTRPERM: Permission issue with Performance Counters. https://developer.nvidia.com/ERR_NVGPUCTRPERM