Abstract

Nsight Systems (nsys) records CPU threads, CUDA API calls, GPU kernels, memory copies and user labels (NVTX ranges) on one shared clock. It hooks in through CUPTI and the operating system, so it needs neither the source code nor a framework, only the ability to launch the program. It answers one question very well: is the GPU busy, and if not, what was the CPU doing instead? It does not tell you why a kernel is slow. We profile the same small PyTorch training loop used in the torch.profiler post, explain the options that change what nsys records, show how to read gaps on the timeline, and list the pitfalls: unbounded captures, child processes, CUDA graphs, and competition for GPU performance counters.

1. Introduction

The problem. A GPU program is slower than its kernels should allow. Either the kernels themselves are slow, or the GPU spends time waiting between kernels. These two cases need different fixes, and a per-operator table cannot tell them apart.

What readers usually assume. nsys is for CUDA kernel developers, and PyTorch users only need torch.profiler.

What we show. nsys is the cheapest way to see the gaps between kernels and the CPU activity that explains each one. That makes it the right first tool, before torch.profiler or ncu, even for framework users.

2. How nsys collects data

  • CUDA tracing (CUPTI). Start and end time of every CUDA API call on the CPU and every kernel and memory copy on the GPU, with correlation IDs that link them.
  • OS tracing and sampling. OS runtime calls (pthread, poll, read, …), thread scheduling, and periodic CPU call-stack samples.
  • NVTX. User-defined ranges, such as “forward” or “step 12”, that appear on the timeline. Libraries such as NCCL and PyTorch can emit them too.
  • GPU metrics (optional). Sampled per-GPU counters such as SM activity, memory bandwidth and interconnect traffic.

All of this lands on one timeline per process, and processes started under the same session share the same clock.

3. A minimal example

We reuse the model from the torch.profiler post and replace record_function with NVTX ranges. torch.cuda.profiler.start() / stop() call cudaProfilerStart / cudaProfilerStop, which lets nsys record only the steps we care about [2].

import torch
import torch.nn as nn

device = "cuda"

class SimpleModel(nn.Module):
    def __init__(self):
        super().__init__()
        self.layer_stack = nn.Sequential(
            nn.Linear(5, 100), nn.ReLU(),
            nn.Linear(100, 200), nn.ReLU(),
            nn.Linear(200, 5),
        )

    def forward(self, x):
        return self.layer_stack(x).type(torch.float)

X = torch.randn(1024, 5, device=device)
y = torch.randint(0, 5, (1024,), device=device)
model = SimpleModel().to(device)
loss_fn = nn.CrossEntropyLoss()
optimizer = torch.optim.SGD(model.parameters(), lr=0.01)

for i in range(20):
    # Steps 0-4 are warmup and are not recorded.
    if i == 5:
        torch.cuda.profiler.start()
    with torch.cuda.nvtx.range(f"step{i}"):
        with torch.cuda.nvtx.range("forward"):
            logits = model(X)
        with torch.cuda.nvtx.range("loss_calc"):
            loss = loss_fn(logits, y)
        optimizer.zero_grad()
        with torch.cuda.nvtx.range("backprop"):
            loss.backward()
        with torch.cuda.nvtx.range("optimize_step"):
            optimizer.step()
        print(f"loss: {loss:.2f}")
torch.cuda.profiler.stop()

Run it under nsys and print summary tables:

nsys profile \
  --trace=cuda,nvtx,osrt \
  --capture-range=cudaProfilerApi --capture-range-end=stop \
  -o simple_nsys \
  python train_nvtx.py

nsys stats -r cuda_gpu_kern_sum,cuda_api_sum,nvtx_sum simple_nsys.nsys-rep

What each flag does:

  • --trace=cuda,nvtx,osrt. Record CUDA API calls and kernels, NVTX ranges, and OS runtime calls.
  • --capture-range=cudaProfilerApi. Record only between cudaProfilerStart and cudaProfilerStop. Without it, nsys records from process start, including imports, model loading and warmup.
  • --capture-range-end=stop. Stop the collection at the first cudaProfilerStop. With repeat, every start/stop pair becomes a separate capture.
  • nsys stats. Prints summary tables from the report: time per kernel (cuda_gpu_kern_sum), time per CUDA API call (cuda_api_sum) and time per NVTX range (nvtx_sum). You can work without the GUI.

3.1 The output

4. Reading the timeline

The core skill with nsys is reading the gaps in the GPU kernel row. For each gap, look straight up at the CPU rows at the same moment. A few common patterns:

What the CPU is doing during the GPU gapWhat it meansTypical fix
Many short cudaLaunchKernel calls, back to backLaunch-bound: kernels finish faster than the CPU can issue themFewer, larger kernels: CUDA graphs, fusion, torch.compile
cudaStreamSynchronize / cudaMemcpy device-to-hostThe host is waiting on a GPU value (.item(), print)Remove or defer the host read
Python or OS activity with no CUDA calls (sampled stacks, osrt calls)CPU work between steps: data loading, scheduling, loggingOverlap it with GPU work, or move it off the critical path
Thread not running (context-switch row)The thread was descheduled by the OSPin threads, reduce CPU contention

Observation. [MEASURE: in the example, the gap structure inside one step and the share of step time where the GPU is idle.]

Insight. A per-operator table adds time up per operator. Idle time belongs to no operator, so it never shows up in such a table, but it is the first thing visible on a timeline.

Implication. Before optimizing any kernel, check how much of the step the GPU is actually busy. If it is mostly idle, kernel work will not help.

5. Options reference

OptionWhat it addsCost
-t, --trace=cuda,nvtx,osrt,cublas,cudnn,mpi,ucxWhich API families to traceGrows with the number of traced calls
-s, --sample=cpu / noneCPU call-stack samplingLow; none removes it
--cpuctxsw=process-tree / noneOS thread scheduling, which shows when a CPU thread was descheduledLow
--cuda-graph-trace=graph / nodeOne range per graph launch, or every kernel inside the graphnode costs more [VERIFY: default per version]
--cuda-memory-usage=trueGPU memory use over timeMedium
--gpu-metrics-devices=allSampled SM, memory and interconnect activity per GPULow; uses the GPU counters [VERIFY: flag name per version]
--capture-range=cudaProfilerApi / nvtxRecords only between cudaProfilerStart/Stop or inside a named NVTX rangeLowers total cost
-y, --delay / -d, --durationTime-based capture windowLowers total cost
--trace-fork-before-exec=trueFollows child processes started with forkNone by itself
--python-sampling=true / --pytorch=autograd-nvtxPython stack samples, or automatic NVTX ranges for PyTorch operatorsMedium [VERIFY: available versions]

The costs are qualitative; they depend on how many events the program generates. Run nsys profile --help for the exact flags in your version [1].

Post-processing.

  • nsys stats -r <report> prints any built-in report. nsys stats --help-reports lists them.
  • nsys export -t sqlite writes the trace to a database that you can query with SQL.

6. Pitfalls

  • Unbounded capture. Without a capture range, delay or duration, nsys records everything from process start. The report becomes huge, and most of it is startup.
  • Child processes. Servers often start GPU work in worker processes. If nsys does not follow them, the GPU rows stay empty. Check that the workers appear in the report.
  • CUDA graphs. A replayed graph is one cudaGraphLaunch on the CPU side. With --cuda-graph-trace=graph, the GPU row shows one bar per graph. Use node to see the kernels inside it.
  • Counter conflicts. GPU metrics sampling uses the GPU’s performance counters, which only one client can hold at a time. A running DCGM profiling job or a parallel ncu session can make it fail [VERIFY: exact behaviour].
  • NVTX ranges are CPU ranges. An NVTX range marks when the CPU issued the work, not when the GPU ran it. Under asynchronous launch, the GPU work for “forward” can finish well after the range closes.

7. What it cannot see

  • Why a kernel is slow. nsys reports a duration, but not whether the kernel was bound by memory, compute or latency. For that, use ncu.
  • Python-level attribution. Without NVTX ranges or Python sampling, you see CUDA calls but not which Python code made them. torch.profiler gives that link directly.

8. Conclusion

nsys answers “is the GPU busy, and if not, why?” It does so cheaply and without touching the program, which makes it the right first tool for almost any GPU performance question. Bound the capture, follow the processes that do GPU work, and read every gap on the GPU row against the CPU rows above it. Once you know a specific kernel is the problem, move to ncu.

References

  1. NVIDIA. Nsight Systems User Guide. https://docs.nvidia.com/nsight-systems/UserGuide/index.html
  2. PyTorch. torch.cuda.nvtx and torch.cuda.profiler. PyTorch documentation. https://docs.pytorch.org/docs/stable/cuda.html