Abstract

Nsight Compute (ncu) looks inside one kernel at a time. It reads hundreds of hardware counters and reports what fraction of the GPU’s peak memory and compute throughput the kernel reached, how many warps were resident, and why warps stalled. From that you can tell whether a kernel is memory-bound, compute-bound, or latency-bound. The GPU cannot read all of these counters in one run, so ncu replays the kernel many times. To make the replays comparable, it locks the GPU clocks, flushes the caches before each pass, and serializes kernel launches by default. These are the right conditions for measuring counters, but they are not the conditions the kernel runs under in production. We profile the GEMM kernels of the small PyTorch model from the torch.profiler post, show how to read the main sections, and measure how far ncu’s default duration moves from the duration seen in nsys.

1. Introduction

The problem. nsys or torch.profiler has pointed at one kernel that takes most of the GPU time. Now we need to know why it is slow: is it waiting on memory, saturating the math units, or simply too small to fill the GPU?

What readers usually assume. ncu’s “Duration” is the kernel’s real runtime, and its numbers can be compared directly with a timeline.

What we show. ncu is the right tool for the why, but its durations are taken under controlled conditions that differ from production. Use ncu to find the limiter, and use nsys to measure how long the kernel actually takes.

2. How ncu collects data

  • Kernel replay. The GPU can only read a limited set of counters per pass. ncu saves the kernel’s memory state, runs the kernel, restores the memory, and runs it again with the next counter group, until all requested counters are collected. A large section set can need dozens of passes per kernel [1].
  • Serialization. ncu runs profiled kernels one at a time, so kernels that overlapped on different streams no longer overlap.
  • Controlled conditions. By default, ncu locks the GPU clocks (--clock-control) and flushes the caches before each pass (--cache-control all), so that results are repeatable across runs [1] [VERIFY: clock-control default per version].
  • Sampling (optional). Program-counter (PC) sampling records where warps are stalled, at the level of SASS instructions or, with -lineinfo, source lines.

3. A minimal example

We use the same training script as in the torch.profiler post, without any profiler code, and save it as train.py. The linear layers run as GEMM kernels from cuBLAS. ncu filters on kernel names, so we first list them, for example with nsys stats -r cuda_gpu_kern_sum from the nsys post, and then profile a few matching launches after warmup:

ncu --set basic \
    -k regex:gemm \
    --launch-skip 10 --launch-count 3 \
    -o simple_ncu \
    python train.py

ncu -i simple_ncu.ncu-rep --page details

What each flag does:

  • --set basic. Collects a small default group of sections: speed-of-light throughput, launch configuration and occupancy. --list-sets shows the others. full collects everything, at many more replays per kernel.
  • -k regex:gemm. Profiles only kernels whose names match gemm [VERIFY: actual kernel names for this model, which depend on the cuBLAS version and GPU]. Without a filter, ncu profiles every kernel the program launches.
  • --launch-skip 10 --launch-count 3. Skips the first 10 matching launches (warmup) and profiles the next 3.
  • -o simple_ncu. Writes a report file. ncu -i ... --page details prints it as text; the GUI shows the same sections with charts.

3.1 The output

4. Reading the sections

SectionWhat it reportsHow to read it
GPU Speed Of LightMemory and compute throughput as a % of the GPU’s peak, and the kernel durationHigh memory %, low compute %: memory-bound. The reverse: compute-bound. Both low: latency-bound or too little work
RooflineThe kernel’s arithmetic intensity (FLOPs per byte) against the GPU’s peak bandwidth and computeLeft of the ridge point: limited by memory bandwidth. Right of it: limited by compute
Launch StatisticsGrid and block size, registers and shared memory per thread blockA grid with fewer blocks than the GPU has SMs leaves SMs idle
OccupancyTheoretical and achieved resident warps per SMLow achieved occupancy leaves too few warps to hide memory latency
Warp State StatisticsAverage cycles warps spend in each stall reasonThe largest stall reason points to the limiter, e.g. waiting on memory vs. waiting on a math pipe
Source CountersStalls and instruction counts per SASS instruction or source lineOnly useful when you can change the code; needs -lineinfo for source lines

Observation 1. [MEASURE: Speed Of Light Memory % and Compute % for the GEMM kernels of SimpleModel, and the number of thread blocks vs. the number of SMs.]

Insight. With a batch of 1024 and layers of 5–200 features, each GEMM has very little work. We expect both throughput numbers to be low and the grid to cover only a few SMs, which means the kernel is latency-bound: it finishes before the GPU is anywhere near either roof.

Implication. For small kernels, the fix is more work per launch (larger batches, fused kernels), not a faster kernel. ncu shows this directly; a timeline only shows that the kernel is short.

5. Why ncu’s duration is not the production duration

Observation 2. [MEASURE: duration of the same GEMM launch in (a) nsys, (b) ncu with defaults, (c) ncu with --clock-control none --cache-control none.]

Insight. Three defaults move ncu away from production. Locked clocks run the GPU at a fixed frequency instead of its boost clock. Flushed caches force every pass to start cold, while in a real loop the weights of a small model may stay in L2. Serialization removes any overlap with other streams.

Implication. Use ncu’s throughput percentages and stall reasons to find the limiter. Take the kernel’s real duration from nsys. If you need ncu durations closer to production, turn off clock and cache control, and then expect more run-to-run variation.

6. Options reference

OptionWhat it doesWhy it matters
--set basic / full (--list-sets shows all sets)Selects groups of sectionsfull can need dozens of replays per kernel
--section <name> / --metrics <list>Collects only named sections or metricsFewer replays
-k, --kernel-name regex:<pat>Profiles only matching kernelsEssential on real workloads
--kernel-name-base function / demangled / mangledWhich form of the name -k matches againstTemplated names differ between forms
-s, --launch-skip / -c, --launch-countSkips the first N matches, then profiles MSkips warmup
--nvtx --nvtx-include <range>Profiles only kernels inside an NVTX rangeTies ncu to the nsys view
--replay-mode kernel / application / range / app-rangeReplays one kernel, the whole program, or a rangeapplication avoids state save/restore but needs a deterministic program
--cache-control all / noneFlushes caches before each replay (default all)Cold-cache numbers by default
--clock-control base / noneLocks clocks to base (default base) [VERIFY]Numbers are not at boost clock
--graph-profiling node / graphProfiles kernels inside CUDA graphs individually, or the whole graph as one unitNeeded for graph-heavy programs
--target-processes allFollows child processesNeeded for multi-process programs
--import-source yes (build with -lineinfo)Adds per-line source viewOnly useful with source code

Run ncu --help for the exact flags in your version [2].

7. Pitfalls

  • Permissions. By default, reading GPU performance counters needs administrator rights. Without them, ncu fails with ERR_NVGPUCTRPERM. The fix is a driver setting that an administrator must enable [3].
  • No filter. Without -k and a launch count, ncu profiles every kernel, each replayed many times. A program that runs in seconds can take hours.
  • Child processes. ncu only profiles the process it launched unless you pass --target-processes all.
  • CUDA graphs. Kernels inside a graph are only profiled individually with --graph-profiling node [VERIFY: default per version].
  • Counter conflicts. Only one client can hold the GPU counters. A running DCGM profiling job or nsys GPU-metrics sampling can block ncu [VERIFY: exact behaviour].
  • Timeouts and collectives. Serializing and replaying one GPU’s kernels stalls any collective that waits on it. In multi-GPU programs this can trigger timeouts.

8. What it cannot see

  • Time between kernels. ncu looks at one kernel in isolation. Launch gaps, synchronization and overlap between streams are invisible. Use nsys.
  • Which code launched the kernel. ncu knows the kernel name and launch configuration, not the Python operator. Use torch.profiler or NVTX ranges.

9. Conclusion

ncu answers “why is this one kernel slow?” by measuring it against the GPU’s peak memory and compute throughput. It does this under controlled conditions (locked clocks, cold caches, serialized launches) that make the counters repeatable but the durations different from production. Use ncu for the limiter and nsys for the time.

References

  1. NVIDIA. Kernel Profiling Guide. Nsight Compute documentation. https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html
  2. NVIDIA. Nsight Compute CLI. Nsight Compute documentation. https://docs.nvidia.com/nsight-compute/NsightComputeCli/index.html
  3. NVIDIA. ERR_NVGPUCTRPERM: Permission issue with Performance Counters. https://developer.nvidia.com/ERR_NVGPUCTRPERM