Performance¶
Tuning a single-card kernel: making the execution visible first, then working through the places time actually goes.
Prerequisites: Tasks and Ordering and Scopes and Placement.
The shape of the problem¶
A PyPTO kernel's wall-clock time is spent in three different machines, and they fail in different ways:
host orchestration AICore
│ │ │
├─ copies ├─ task dispatch ├─ the kernel itself
│ (06-host) │ (01, 02, 03) │ (04-incore)
│ │ │
└────────────── memory: on-chip buffers + runtime rings (05-memory) ──────┘
Most first-time tuning goes straight to the third column — the arithmetic inside the kernel — and finds that the first two were the problem. The chapter is ordered so you meet them in the order they usually bite.
Contents¶
| Page | Covers |
|---|---|
| Reading the swimlane | Capturing the L2 swimlane and opening it — the one view that shows where time went |
| Task granularity | Dispatch is not free; growing and merging InCore functions without starving the cores |
| Runtime overhead | Mixed kernels, SPMD, allow_early_resolve, in-kernel syncall |
| Managing dependencies | Why the runtime serializes work that could overlap, and how to say otherwise |
| Tuning the InCore function | Double buffering, algorithmic splits, L0 instruction traces, hardware granularity, external kernels |
| Memory | The four scope-depth rings, scope placement, and ring sizing |
| Host | Keeping resident data resident |
How to use it¶
Every technique below is written with the same four fields, because a speedup with an unstated cost is not a result:
| Field | What it answers |
|---|---|
| When it applies | The symptom that makes this the right move |
| How | The code change |
| Cost | What it spends — memory, generality, or a correctness obligation you now own |
| How to confirm | The artifact that shows it worked, and what should change in it |
There are no speedup numbers in this chapter. They depend on your shapes, platform, and toolchain versions, and a stale number is worse than none because you cannot tell it has gone stale. The confirmation step is the transferable part: run it on your kernel and you have your own number.
Before you tune anything¶
Two cheap checks come before any of this, and both are already done for you:
report/perf_hints.login the build output. The compiler writes what it noticed during compilation — undersized transfers, matmuls it could not tile, a pipeline depth that did not fit. One summary line goes to stderr on every compile.- The host/device split.
run()does not return a timing object — itsexecution_timeis total wall clock, compile and golden included. Usepypto.runtime.benchmark, whoseBenchmarkStatscarriesdevice_wall_usandhost_wall_usseparately. If the time is on the host, nothing in pages 00–05 will move it — go to Host.
See also¶
- Tuning the schedule — the same ground as a hands-on walkthrough.
- Precision — the equivalent treatment for "the result is wrong".