Abstract
Nsight Systems (nsys) records CPU threads, CUDA API calls, GPU kernels, memory copies and user labels (NVTX ranges) on one shared clock. It hooks in through CUPTI and the operating system, so it needs neither the source code nor a framework, only the ability to launch the program. It answers one question very well: is the GPU busy, and if not, what was the CPU doing instead? It does not tell you why a kernel is slow. We profile the same small PyTorch training loop used in the torch.profiler post, explain the options that change what nsys records, show how to read gaps on the timeline, and list the pitfalls: unbounded captures, child processes, CUDA graphs, and competition for GPU performance counters.
1. Introduction
The problem. A GPU program is slower than its kernels should allow. Either the kernels themselves are slow, or the GPU spends time waiting between kernels. These two cases need different fixes, and a per-operator table cannot tell them apart.
What readers usually assume. nsys is for CUDA kernel developers, and PyTorch users only need torch.profiler.
What we show. nsys is the cheapest way to see the gaps between kernels and the CPU activity that explains each one. That makes it the right first tool, before torch.profiler or ncu, even for framework users.
2. How nsys collects data
- CUDA tracing (CUPTI). Start and end time of every CUDA API call on the CPU and every kernel and memory copy on the GPU, with correlation IDs that link them.
- OS tracing and sampling. OS runtime calls (
pthread,poll,read, …), thread scheduling, and periodic CPU call-stack samples. - NVTX. User-defined ranges, such as “forward” or “step 12”, that appear on the timeline. Libraries such as NCCL and PyTorch can emit them too.
- GPU metrics (optional). Sampled per-GPU counters such as SM activity, memory bandwidth and interconnect traffic.
All of this lands on one timeline per process, and processes started under the same session share the same clock.
3. A minimal example
We reuse the model from the torch.profiler post and replace record_function with NVTX ranges. torch.cuda.profiler.start() / stop() call cudaProfilerStart / cudaProfilerStop, which lets nsys record only the steps we care about [2].
import torch
import torch.nn as nn
device = "cuda"
class SimpleModel(nn.Module):
def __init__(self):
super().__init__()
self.layer_stack = nn.Sequential(
nn.Linear(5, 100), nn.ReLU(),
nn.Linear(100, 200), nn.ReLU(),
nn.Linear(200, 5),
)
def forward(self, x):
return self.layer_stack(x).type(torch.float)
X = torch.randn(1024, 5, device=device)
y = torch.randint(0, 5, (1024,), device=device)
model = SimpleModel().to(device)
loss_fn = nn.CrossEntropyLoss()
optimizer = torch.optim.SGD(model.parameters(), lr=0.01)
for i in range(20):
# Steps 0-4 are warmup and are not recorded.
if i == 5:
torch.cuda.profiler.start()
with torch.cuda.nvtx.range(f"step{i}"):
with torch.cuda.nvtx.range("forward"):
logits = model(X)
with torch.cuda.nvtx.range("loss_calc"):
loss = loss_fn(logits, y)
optimizer.zero_grad()
with torch.cuda.nvtx.range("backprop"):
loss.backward()
with torch.cuda.nvtx.range("optimize_step"):
optimizer.step()
print(f"loss: {loss:.2f}")
torch.cuda.profiler.stop()
Run it under nsys and print summary tables:
nsys profile \
--trace=cuda,nvtx,osrt \
--capture-range=cudaProfilerApi --capture-range-end=stop \
-o simple_nsys \
python train_nvtx.py
nsys stats -r cuda_gpu_kern_sum,cuda_api_sum,nvtx_sum simple_nsys.nsys-rep
What each flag does:
--trace=cuda,nvtx,osrt. Record CUDA API calls and kernels, NVTX ranges, and OS runtime calls.--capture-range=cudaProfilerApi. Record only betweencudaProfilerStartandcudaProfilerStop. Without it, nsys records from process start, including imports, model loading and warmup.--capture-range-end=stop. Stop the collection at the firstcudaProfilerStop. Withrepeat, every start/stop pair becomes a separate capture.nsys stats. Prints summary tables from the report: time per kernel (cuda_gpu_kern_sum), time per CUDA API call (cuda_api_sum) and time per NVTX range (nvtx_sum). You can work without the GUI.
3.1 The output
4. Reading the timeline
The core skill with nsys is reading the gaps in the GPU kernel row. For each gap, look straight up at the CPU rows at the same moment. A few common patterns:
| What the CPU is doing during the GPU gap | What it means | Typical fix |
|---|---|---|
Many short cudaLaunchKernel calls, back to back | Launch-bound: kernels finish faster than the CPU can issue them | Fewer, larger kernels: CUDA graphs, fusion, torch.compile |
cudaStreamSynchronize / cudaMemcpy device-to-host | The host is waiting on a GPU value (.item(), print) | Remove or defer the host read |
Python or OS activity with no CUDA calls (sampled stacks, osrt calls) | CPU work between steps: data loading, scheduling, logging | Overlap it with GPU work, or move it off the critical path |
| Thread not running (context-switch row) | The thread was descheduled by the OS | Pin threads, reduce CPU contention |
Observation. [MEASURE: in the example, the gap structure inside one step and the share of step time where the GPU is idle.]
Insight. A per-operator table adds time up per operator. Idle time belongs to no operator, so it never shows up in such a table, but it is the first thing visible on a timeline.
Implication. Before optimizing any kernel, check how much of the step the GPU is actually busy. If it is mostly idle, kernel work will not help.
5. Options reference
| Option | What it adds | Cost |
|---|---|---|
-t, --trace=cuda,nvtx,osrt,cublas,cudnn,mpi,ucx | Which API families to trace | Grows with the number of traced calls |
-s, --sample=cpu / none | CPU call-stack sampling | Low; none removes it |
--cpuctxsw=process-tree / none | OS thread scheduling, which shows when a CPU thread was descheduled | Low |
--cuda-graph-trace=graph / node | One range per graph launch, or every kernel inside the graph | node costs more [VERIFY: default per version] |
--cuda-memory-usage=true | GPU memory use over time | Medium |
--gpu-metrics-devices=all | Sampled SM, memory and interconnect activity per GPU | Low; uses the GPU counters [VERIFY: flag name per version] |
--capture-range=cudaProfilerApi / nvtx | Records only between cudaProfilerStart/Stop or inside a named NVTX range | Lowers total cost |
-y, --delay / -d, --duration | Time-based capture window | Lowers total cost |
--trace-fork-before-exec=true | Follows child processes started with fork | None by itself |
--python-sampling=true / --pytorch=autograd-nvtx | Python stack samples, or automatic NVTX ranges for PyTorch operators | Medium [VERIFY: available versions] |
The costs are qualitative; they depend on how many events the program generates. Run nsys profile --help for the exact flags in your version [1].
Post-processing.
nsys stats -r <report>prints any built-in report.nsys stats --help-reportslists them.nsys export -t sqlitewrites the trace to a database that you can query with SQL.
6. Pitfalls
- Unbounded capture. Without a capture range, delay or duration, nsys records everything from process start. The report becomes huge, and most of it is startup.
- Child processes. Servers often start GPU work in worker processes. If nsys does not follow them, the GPU rows stay empty. Check that the workers appear in the report.
- CUDA graphs. A replayed graph is one
cudaGraphLaunchon the CPU side. With--cuda-graph-trace=graph, the GPU row shows one bar per graph. Usenodeto see the kernels inside it. - Counter conflicts. GPU metrics sampling uses the GPU’s performance counters, which only one client can hold at a time. A running DCGM profiling job or a parallel ncu session can make it fail [VERIFY: exact behaviour].
- NVTX ranges are CPU ranges. An NVTX range marks when the CPU issued the work, not when the GPU ran it. Under asynchronous launch, the GPU work for “forward” can finish well after the range closes.
7. What it cannot see
- Why a kernel is slow. nsys reports a duration, but not whether the kernel was bound by memory, compute or latency. For that, use ncu.
- Python-level attribution. Without NVTX ranges or Python sampling, you see CUDA calls but not which Python code made them. torch.profiler gives that link directly.
8. Conclusion
nsys answers “is the GPU busy, and if not, why?” It does so cheaply and without touching the program, which makes it the right first tool for almost any GPU performance question. Bound the capture, follow the processes that do GPU work, and read every gap on the GPU row against the CPU rows above it. Once you know a specific kernel is the problem, move to ncu.
References
- NVIDIA. Nsight Systems User Guide. https://docs.nvidia.com/nsight-systems/UserGuide/index.html
- PyTorch. torch.cuda.nvtx and torch.cuda.profiler. PyTorch documentation. https://docs.pytorch.org/docs/stable/cuda.html