Performance Tuning¶
A practical guide for tuning pypto-lib kernels on Ascend NPU (A3 / 910C). The flow is two-tiered: first balance the inter-kernel schedule on the AICPU side (chip swimlane), then optimize each kernel's internal pipeline (L1/L0 swimlane + PMU).
For the underlying levels see simpler's Hierarchical Level Runtime: L2 = one chip (AICPU + AIC/AIV cores), L1 = die / L2 cache, L0 = single compute core.
Measuring — the benchmark loop (PYPTO_BENCH)¶
Tuning needs a number before and after. Set PYPTO_BENCH=1 and every
run call in the process times the kernel on device after its
correctness dispatch — no --benchmark flag, no edit to the model file:
Effective is the framework's post-graph-build execution window on
device (orch ∪ sched — the old device-log "Total"), recovered from the
runtime's [STRACE] markers. Quote mean= — this field of this line is the
per-case number, so two runs are comparable only when both quote it.
Requirements: a real device — a *sim platform prints
effective_us unavailable: no device-domain spans — and a runtime built
with SIMPLER_PROFILING. A runtime_dir= replay benchmarks the replayed
build, so a hand-edited .cpp can be timed without recompiling; only a spec
with a stepped scalar skips it, with a [RUN] benchmark skipped note.
Multi-card (L3) output¶
A distributed program adds a per-rank breakdown and a context line:
[RUN] effective_us (100 rounds) min=520.1 median=538.4 mean=539.9 max=602.0
[RUN] rank 10: eff_us min=500.0 median=510.0 mean=511.0 max=520.0
[RUN] rank 11: eff_us min=520.1 median=538.4 mean=539.9 max=602.0
[RUN] benchmark kernel=moe_ep2 l3_resident=1 rounds=100 ranks=2 host_union_mean_us=900 host_mean_us=950
- The headline is the per-round max across ranks — the round ends when
the slowest card finishes. The
eff_uslines expose the cross-card imbalance that max hides; a persistent gap between ranks is a load-balance problem, not a kernel problem. - A
rank N: eff_usline sums that card's dispatches within a round (a card runs them serially). When a card dispatches more than once per round, each rank line gains a nestedslotline per dispatch so you can see which dispatch owns the time:
[RUN] rank 10: eff_us min=500.0 median=510.0 mean=510.0 max=520.0
[RUN] slot 0 (prefill_orch): eff_us min=200.0 median=205.0 mean=205.0 max=210.0
[RUN] slot 1 (decode_orch): eff_us min=300.0 median=305.0 mean=305.0 max=310.0
[RUN] rank 11: eff_us min=520.1 median=561.0 mean=561.0 max=602.0
[RUN] slot 0 (decode_orch): eff_us min=520.1 median=561.0 mean=561.0 max=602.0
slot is the dispatch's position within its rank's round (slot 1 is the same
dispatch in every round), and the name in parentheses is the orchestration
function it runs. Once the slot lines appear, every rank's dispatches are
listed — including single-dispatch ranks like 11 above, whose slot line
necessarily restates its rank line — so the breakdown stays a complete tree.
The slot lines are omitted entirely when every card dispatches exactly once per
round, and also when a card's dispatch order varies between rounds — a slot
then names no single callable, so pypto reports no per-dispatch view rather
than mislabelling it.
- host_union_mean_us is the cross-rank host-timeline window
(max(end) - min(start)), so it captures start skew and overlap, but
includes host dispatch overhead.
- fallback_flattened=1 means a rank's dispatch count was not divisible by
warmup + rounds (a non-deterministic dispatch shape), so per-round
segmentation was abandoned and the numbers are a pooled per-dispatch
sample — treat them as indicative only.
Knobs¶
| Env | Default | Effect |
|---|---|---|
PYPTO_BENCH |
off | Enables the timed loop. Any value except "" / 0 / false / False is on. |
PYPTO_BENCH_ROUNDS |
100 |
Timed rounds. 100 rounds is ~0.1 s of device time for a decode step but minutes for a long prefill or a multi-card run — drop it while iterating. |
PYPTO_BENCH_WARMUP |
5 |
Leading launches discarded before measurement. The resident L3 path always keeps ≥ 1 (its first warmup launch doubles as the validation dispatch). |
PYPTO_BENCH_RAW |
off | Prints every measured dispatch's Effective sample, one line per rank, in dispatch order. Use it when a summary looks suspicious — start-up drift, a bimodal rank, one card lagging. |
A malformed or out-of-range value warns and falls back to the default rather than failing the run. The 100 / 5 default is the baseline every reported number should come from; if you change the loop sizes, compare only against other runs with the same sizes.
# Quick iteration on a long prefill, with the raw per-dispatch samples.
PYPTO_BENCH=1 PYPTO_BENCH_ROUNDS=10 PYPTO_BENCH_WARMUP=2 PYPTO_BENCH_RAW=1 \
python models/deepseek_v4_flash_mtp/prefill_fwd.py -p a2a3 -d 0
When only the timing changes between iterations — not the numerics — save the
golden once and replay it via golden_data=, cutting the torch recompute out
of every later run. See
Save and Replay Golden Data.
Part 1 — L2 tuning (inter-kernel schedule)¶
Capture¶
Run the case with --enable-chip-swimlane. The runtime writes raw per-task
chip swimlane records under the build directory and, on a real-device platform,
converts them to a merged swimlane:
build_output/<ProgramName>_<ts>/dfx_outputs/
├── chip_swimlane_records.json
├── deps.json # real-device graph pass
└── merged_swimlane_<ts>.json # real device only; open this
The flag takes a capture level, and a bare flag means level 1 — per-task
AICore timing, which is what reading the L2 schedule needs. Raise it only when
the question requires it; each level records more and perturbs the timing it
measures. Gap attribution and early-dispatch proofs need
--enable-chip-swimlane 4 — see
Capture levels.
Two viewers work:
- Open
merged_swimlane_<ts>.jsonin https://ui.perfetto.dev/. - Or open
chip_swimlane_records.jsondirectly with the pypto-toolkit VSCode extension.
Simulator platforms emit chip_swimlane_records.json but intentionally skip
the merged conversion because their records do not yet include the task
metadata the converter requires. Use a real-device capture when you need the
merged Perfetto view and dependency arrows.
The trace shows one lane per AICPU / AIC / AIV with task name, duration and dependency edges — gaps and stalls are visible directly.
Timing one stage of a full network — task-timing slots¶
A swimlane answers "where did every task go". A narrower question comes up
constantly on a full network: what does this one stage cost inside the whole
program? simpler answers that with selective task-timing slots — 16 fixed
slots into which the Scheduler folds a tagged task's AICPU dispatch→finish
window, at the same boundaries the swimlane's finish_time uses.
| Task-timing slot | Chip swimlane | |
|---|---|---|
| Covers | up to 16 tagged tasks | every task |
| Switch | tagging in the generated orchestration .cpp — no env var, no compile gate, works in SIMPLER_DFX=0 builds |
--enable-chip-swimlane |
| Cost on untagged tasks | one cache-hot sentinel compare | per-task records and collector threads |
| Output | [STRACE] spans named chip.run.runner_run.device_wall.task_slot_<N> (clk=dev, ts / dur in ns) |
merged Perfetto JSON |
Reach for a slot when a whole-network swimlane is too large to read or perturbs the schedule being measured, and you already know which stage you care about. Reach for the swimlane when the question is why that stage is slow.
They are also how a change the benchmark loop distorts gets attributed. An L2
warm is the standard example: PYPTO_BENCH replays the same weights every round,
so its L2 is already warm and the end-to-end delta is both flattered and diluted
— see L2 Prefetch.
PyPTO exposes no DSL surface for the tag, so the workflow patches the
generated orchestration C++ and replays it — the same
runtime_dir loop
used for any generated-code edit:
- Compile once (
--compile-only, or reuse an existing build directory). - Open the orchestration source —
<work_dir>/orchestration/<prog>.cppfor an L2 program,<work_dir>/next_levels/<prog>/orchestration/<prog>.cppfor an L3 one. Every submit block carries a comment naming its scope and kernel, which is how a stage is located:
// Spmd w1_mm_spmd: w1_mm
CoreTaskArgs params_t0;
params_t0.add_input(ext_recv_x);
params_t0.add_output(ext_gate_i32);
params_t0.set_task_timing_slot(0); // <- the tag
params_t0.launch_spec.set_block_num(8);
rt_submit_aic_task(0, params_t0);
- Tag one iteration, not all. A
pl.rangelayer loop becomes a real C++forwhose induction variable keeps the DSL's own name, so an unguarded tag merges all 20 layers into one useless window. Guard it:
- Replay the patched build. Editing the
.cppis the only signal the harness needs — do not delete the sibling.o/.so:
PYPTO_BENCH=1 python models/deepseek_v4_flash_mtp/decode_fwd.py -p a2a3 -d 0,1 --ep 2 \
--runtime-dir build_output/<ProgramName>_<ts>
- Read the
[RUN] task slotsblock the benchmark prints (below). WithoutPYPTO_BENCHthe one correctness dispatch prints its spans to the runtime log instead, at the defaulttiminglevel — passPYPTO_RUNTIME_LOG=timingif the entry raised the threshold, then grep fortask_slot_. The benchmark captures stderr itself, so underPYPTO_BENCHthose log lines never appear.
Reading the numbers¶
- Each slot reduces to
min(dispatch)/max(finish). Reusing one slot across several tasks yields a single merged window from the earliest tagged dispatch to the latest tagged finish — which is exactly how a multi-kernel stage is measured. Distinct slots keep each task's own window, so tooling can recoverfinish(B) − dispatch(A). A MIX task's AIC/AIV0/AIV1 subtasks and an SPMD task's blocks all fold into the one tagged slot. - Read finish-to-finish, not span length.
dispatchis the speculative publication, so underallow_early_resolvethe windows overlap heavily and a slot's ownduris not that stage's cost. Take(slot_{k+1}.ts + slot_{k+1}.dur) − (slot_k.ts + slot_k.dur), which means tagging the preceding stage too — otherwise the first stage has no anchor. - Slots reset every run. A plain replay therefore yields one sample per
slot. Add
PYPTO_BENCH=1to the replay for a distribution: every measured round reports its slots, and the harness prints one line per rank and slot —
[RUN] task slots: ranks=2
[RUN] rank 1872915 task_slot 4: n=100 fin_us=2177.4 dur_us=8.6 dfin_us=713.9
[RUN] rank 1872915 task_slot 5: n=100 fin_us=2599.3 dur_us=7.9 dfin_us=428.1
fin_us is the slot's finish from the run's device-clock origin, dur_us
its own window, and dfin_us the per-round finish minus the previous slot's
finish — the finish-to-finish stage cost when slots are numbered in stage
order. All three are medians; PYPTO_BENCH_RAW=1 adds each round's fin_us.
- The patch lives in build_output/ only. Recompiling regenerates the
orchestration .cpp and silently drops every tag — which is also how the
instrumentation is removed.
What to look for¶
Look for these shapes on the swimlane that indicate a problem:
| Symptom | Likely cause | Fix |
|---|---|---|
| Cores idle while AICPU lane is solid | Kernels too small; AICPU scheduling is the bottleneck | Make kernels larger (item 2) |
| Long tail on a single AIC/AIV | One kernel is too big and serializes | Split it (item 3) |
| Cube / vector unit utilization low even though kernel is busy | Tile size under-fills the user-visible on-chip buffers | Re-tile against Mat / Acc for cube or Vec for vector work (item 4) |
| Cube lane busy while vector lane idle (or vice versa) | Vec/cube epilogue is split into separate kernels | Merge into a mixed kernel (item 2c) |
| A stage re-reads the same weights every layer, MTE2-bound, with a large unrelated stage in between | The weights are evicted from L2 before the next use | Warm them with pl.prefetch (item 5) |
| Sequential AICPU dispatch trail per region | Region issues one kernel per iteration | Use pl.spmd to dispatch a block fan-out once (item 6) |
A gap on this trace is not automatically a scheduling problem: the interval before a task splits into producer-FIN detection, ready-but-undispatched scheduler delay, and post-dispatch pickup, and each has a different fix. See Dependencies and Scheduling for how edges are formed, what the four per-task timestamps mean, and how to attribute a gap without guessing.
Tuning rules¶
1. Use pl.range vs. pl.parallel correctly¶
pl.parallel declares iterations are independent — the compiler may
distribute them across cores. pl.range is strict sequential and forces
a dependency chain. Use pl.parallel whenever there is no carried state,
and reserve pl.range for accumulators or stateful loops.
A pl.range over an independent dimension forces the swimlane into a
single lane; switching to pl.parallel is usually the largest single
win at this stage.
2. Kernels too small — make each kernel do more¶
When the swimlane shows cores idling while the AICPU lane is fully saturated, the AICPU dispatcher is the bottleneck. Target ~50 µs per kernel on A3 / 910C (smaller kernels add dispatch overhead that the AICPU can't hide). Three ways to grow each kernel:
a. Fold outer iterations into the core. Move part of an outer
pl.range / pl.parallel's iterations into the pl.at region as an
inner pl.range, so each dispatched kernel processes a tile of iterations
instead of one:
# Before: one kernel per outer iteration — many tiny dispatches
for b in pl.parallel(0, BATCH):
with pl.at(level=pl.Level.CORE_GROUP, name_hint="step"):
...
# After: fold BATCH_TILE iterations into each kernel via an inner pl.range
for b0 in pl.parallel(0, BATCH, BATCH_TILE):
with pl.at(level=pl.Level.CORE_GROUP, name_hint="step"):
for b in pl.range(b0, b0 + BATCH_TILE):
...
b. Merge consecutive pl.at blocks. Adjacent pl.at regions in the
same scope each become a separate kernel with an AICPU hand-off between
them. Fuse back-to-back regions into one pl.at so a single kernel covers
the whole sequence:
# Before: two adjacent regions → two kernels + a hand-off
with pl.at(level=pl.Level.CORE_GROUP, name_hint="rmsnorm"):
...
with pl.at(level=pl.Level.CORE_GROUP, name_hint="q_proj"):
...
# After: one region → one kernel
with pl.at(level=pl.Level.CORE_GROUP, name_hint="rmsnorm_q_proj"):
... # rmsnorm, then q_proj
c. Merge cube + vector into a mixed kernel. When a matmul (cube) and
its epilogue (cast / add / norm — vector) sit in separate pl.at regions,
every projection generates two kernels and an AICPU hand-off between them.
Place both inside the same pl.at and the compiler co-schedules cube
and vector on the right unit internally, removing the hand-off:
with pl.at(level=pl.Level.CORE_GROUP, name_hint="q_proj"):
for kb in pl.pipeline(0, HIDDEN // K_STEP, stage=2):
...
q_acc = pl.matmul_acc(q_acc, tile_a, tile_b) # cube
q_bf16 = pl.cast(q_acc, target_type=pl.BF16) # vector
q_proj[b0:b0 + BATCH_TILE, q0:q0 + Q_OUT_CHUNK] = q_bf16
3. Kernels too big — split and parallelize¶
When one kernel dominates the swimlane and the rest of the chip waits on
it, the kernel is too coarse. Pull a pl.range out of the pl.at and
convert it to a pl.parallel chunk loop so each chunk becomes its own
InCore kernel scheduled across cores:
# Before: one giant InCore region over all q_out blocks
with pl.at(level=pl.Level.CORE_GROUP, name_hint="q_proj"):
for q0 in pl.range(0, hidden, Q_OUT_CHUNK):
...
# After: each q-chunk is its own kernel, parallel across cores
for q0 in pl.parallel(0, hidden, Q_OUT_CHUNK):
with pl.at(level=pl.Level.CORE_GROUP, name_hint="q_proj"):
...
4. Tiling — fill the core-internal buffers¶
Each AIC / AIV core has fixed on-chip buffers. At the DSL level, cube tiles
directly consume Mat (L1 operand storage) and Acc (L0C accumulator
storage), while vector tiles consume Vec (UB). The tile sizes declared in
your pl.slice / pl.matmul (typically BATCH_TILE, K_STEP,
Q_OUT_CHUNK, …) control these spaces.
- Too small → buffers are under-utilized, cube/vector throughput drops proportionally, MTE2 issues many small loads.
- Too large → tile spills, the compiler falls back to smaller transfer units, or compile-time shape checks fail.
Check actual occupancy. Every compile writes a per-kernel buffer report to
listing, for each compute function, how full each on-chip space runs
against its hardware limit (on the illustrated 910C configuration: vector
Vec has a 184 KB compiler-safe limit within the 192 KB physical UB; cube
Mat is 512 KB, Left / Right are 64 KB each, and Acc is 128 KB):
--- gather_kv ---
Space | Used | Limit | Usage | MemRefs
-------+-------------+-------------+----------+---------
Vec | 129.0 KB | 184.0 KB | 70.1% | 2
--- kv_proj_matmul ---
Space | Used | Limit | Usage | MemRefs
-------+-------------+-------------+----------+---------
Mat | 80.0 KB | 512.0 KB | 15.6% | 4
Left | 32.0 KB | 64.0 KB | 50.0% | 1
Right | 16.0 KB | 64.0 KB | 25.0% | 1
Acc | 4.0 KB | 128.0 KB | 3.1% | 1
Scan the Usage column for Mat, Acc, and Vec. These are the
user-visible constraints affected by the M/N/K or vector fragment. Left
and Right report the L0A/L0B staging chosen by the compiler for the L1
fragment; they can routinely read close to 100% and are not independent DSL
tile budgets. Do not shrink a tile merely to reduce a Left or Right
percentage.
Grow the space that limits the task: Mat for operand fragments, Acc for
the output fragment, or Vec for vector working data. An oversized plain
matmul output may be compiler-subtiled through L0C, so Acc is often a
performance boundary rather than an immediate compile failure; extra tiles
still add FIXPIPE drains. The exact build report and compile result are
authoritative.
Practical procedure:
- Start from the natural problem dimensions (
BATCH,HIDDEN, …). - Pick
K_STEPand the output-chunk size soMatandAccstay within the intended bounds without forcing inefficient compiler sub-tiling. - Sweep one tile dim up/down by 2× and re-measure with PMU — keep the size that pushes the cube (or vector) unit closer to 100 %.
The K loop is then driven by pl.pipeline(stage=2 or 4) so the next
tile's MTE2 overlaps the current tile's compute (see Part 2 item 2).
For the complete M/N/K constraint model and empirical sweep method, see
Cube Tile Tuning.
5. Warm L2 for a weight set the next stage evicts¶
When a stage re-reads a fixed weight set every layer and something between two
layers evicts it, an SDMA cache warm (pl.prefetch) can hide the reload behind
compute that is already running. It writes no tensor, so it is free to try and
free to delete — but a partial or oversized warm costs more than it saves. See
L2 Prefetch.
6. Stream a weight that has no reuse: NZ layout and CachePolicy.BYPASS¶
A weight read once per dispatch and never re-read is streamed, not cached, and two declarations make that read cheaper. Both are properties of the weight, so both belong on the parameter and the fixture that fills it — neither is a compiler inference.
| Declaration | What it changes | What it costs |
|---|---|---|
pl.Tensor[[E, N, K], pl.INT8, pl.NZ] |
the cube loads the weight NZ→NZ instead of reformatting ND on the way into L1, so the contiguous run is one fractal column block rather than one K-tile row | the fixture must write the bytes in fractal order (utils.pack_nz) and the golden must read them back (utils.unpack_nz) |
pl.set_cache_policy(w, pl.CachePolicy.BYPASS) at the top of the reading scope |
the load is issued against the device's uncached alias, so it does not evict what does have reuse | nothing to the kernel; the author asserts nobody writes those bytes while it runs |
The ND penalty is worth seeing concretely. A logical [E, N, K] INT8 weight
sliced by a K tile gives the DMA K_TILE contiguous bytes before the next row
is K bytes away — 512 B runs at K_TILE = 512. The same weight in NZ order
gives 16 rows × one C0 line, an 8 KB run, with the column-block stride between
them. Same bytes, same footprint, 16× the contiguous run.
Where NZ fits, and where it does not. pto-isa declares NZ at a fixed rank-5 shape with one batch slot, so the layout is only addressable when the stacked axis is a leading one:
[E, N, K](routed experts) and[LAYERS * O_GROUPS, O_LORA, O_GROUP_IN](wo_a,wo_b) work as they stand — a per-layerpl.slicenarrows a leading axis and the trailing matrix stays whole.[LAYERS * D, Q_LORA](wq_a,wq_b,wkv) does not. One layer's rows sit inside every fractal column block, so the window is not contiguous and the compiler refuses it by name. Those are declared[LAYERS, D, C]instead — a leading stacked axis — which is what makes NZ available to them.- Only one leading axis may be narrowed at a time. The leading axes fold into the single batch slot row-major, so a window on an axis that a spanning axis precedes selects a set no contiguous run describes; the compiler refuses it.
- An NZ tensor is read-only, but it can be flattened whole for an SDMA warm
(rule 5): a rank-1 view of every element is the same byte range either way.
wo_aandwo_bare both NZ and both warmed. - The trailing extents must be static, rows a multiple of 16, and columns a whole C0 line (32 B); a slice offset must be a provably non-negative multiple of those, which rules out an offset built by subtraction — and a remainder or quotient is only non-negative once its dividend is, since both truncate toward zero on device.
Measure both separately. They are independent, and which one pays depends on
whether the ND load is already at its bandwidth ceiling: at a coarse blocking
where ND already saturates, NZ can be worth nothing while the bypass still pays,
and at a finer blocking the order reverses. One PYPTO_BENCH run per variant on
the same frozen inputs settles it.
7. pl.spmd for parallel sub-kernel dispatch¶
pl.spmd(N) dispatches N blocks of an InCore body in parallel from
one AICPU schedule entry, instead of N successive pl.parallel +
pl.at dispatches. When a region has many parallel chunks and AICPU
overhead is visible per iteration on the swimlane, replace the explicit
for ... in pl.parallel: with pl.at: ... pattern with pl.spmd:
# qwen3_14b/decode_layer_a8w8.py: one AICPU dispatch fans out every q_proj block
for q_grid in pl.spmd(Q_ON * N_SUB, name_hint="q_proj_fused_dequant"):
...
Collapsing N dispatches into one schedule entry cuts AICPU scheduling
overhead sharply. The win is largest for MPMD-shaped regions with
heavy fan-in / fan-out — where each block depends on (or feeds) many
others, the AICPU would otherwise track a dependency edge per block, and
pl.spmd replaces that whole fan with a single dispatch and its barrier.
Use pl.spmd once the per-iteration body is self-contained and the
AICPU lane shows a dispatch trail; keep the explicit form when you need
to nest named sub-regions inside the chunk.
Part 2 — L1 / L0 tuning (intra-kernel)¶
Once L2 is balanced, individual kernels become the bottleneck. Two artifacts drive intra-kernel tuning:
Capture¶
PMU counters per kernel:
python models/deepseek_v4_flash_mtp/decode_sparse_attn.py -p a2a3 -d 0 --enable-pmu 2
# → build_output/<...>/dfx_outputs/pmu.csv
Not every kernel exposes --enable-pmu; a kernel that does not can still be
captured by passing config={"enable_pmu": 2} to its run call (the
RunConfig carries it to the runtime as a DFX option).
For a per-kernel intra-core swimlane, use
In-Core Simulator Profiling.
It explains how the repository workflow builds a standalone single-core
testcase from the generated .cpp and sibling .pto, runs it under
msprof op simulator, validates that data-dependent work actually executed,
and cleans the Insight trace for Perfetto.
For phase timing inside a multi-core extern on real hardware, use
cce-incore-profiling.md. It covers
per-core on-device timestamps, collective-barrier interpretation, and exact
partitions that reconcile internal phases with the L2 task total.
Tuning rules¶
1. Fix tile-shape MTE hints from perf_hints.log¶
Every compile writes a perf-hint log next to the memory report:
The compiler flags every tile.load / tile.store whose innermost
(trailing) dimension is smaller than the 512 B L2 cache line — the case
that forces MTE into many short, cache-line-straddling transfers. Each
hint carries the exact source location:
[perf_hint PH001] TileInnermostDimGranularity: tile.load has innermost
dim = 256B; recommended >= 512B for backend a2a3 (L2 cache line = 512B).
Consider increasing tile shape on the innermost axis.
at models/deepseek_v4_flash_mtp/qkv_proj_rope.py:68:4
Walk the log and widen the trailing tile dimension at each flagged site
so the innermost slice is a multiple of 512 B (item 3 gives the per-dtype
element counts). Bringing every flagged tile.load / tile.store up to
≥ 512 B is usually the single biggest MTE-efficiency win at this level.
2. pl.pipeline for ping-pong on the K loop¶
Inside a pl.at region, the reduction loop of a matmul (the K loop)
should be pl.pipeline(..., stage=2 or 4). The compiler replicates the
loop body stage times for ping-pong buffering, so MTE2 (load) overlaps
with cube/vec compute on alternating tiles.
# stage=4 for the largest input-projection K dim
for kb in pl.pipeline(HIDDEN // K_STEP, stage=4):
...
# stage=2 is the common default
for kb in pl.pipeline(0, hidden // K_STEP, stage=2):
...
A pl.range here forces strictly serial K iterations — the cube unit
will stall on every load. Always prefer pl.pipeline in the K loop.
3. Watch pl.slice / pl.assemble granularity¶
MTE transfers prefer 512-byte aligned addresses and lengths on A3 / 910C. Pick the trailing-dim tile size so the slice is a multiple of 512 B:
- BF16 (2 B/element) → trailing dim multiple of 256 elements
- FP32 (4 B/element) → trailing dim multiple of 128 elements
- INT8 (1 B/element) → trailing dim multiple of 512 elements
Misaligned slices fall back to slower paths visible as long MTE2 bars in
the kernel-insight swimlane. In the qwen3-14b kernels, all K_STEP /
Q_OUT_CHUNK constants are picked to keep the inner load 512 B aligned.
4. Read PMU utilization¶
Recommended PMU counters to collect per kernel:
pmu_total_cycles
vec_busy_cycles cube_busy_cycles scalar_busy_cycles
mte1_busy_cycles mte2_busy_cycles mte3_busy_cycles
fixpipe_cycles
What each pipe means in context:
| Counter | Cube kernel (AIC) | Vector kernel (AIV) |
|---|---|---|
mte1_busy_cycles |
L1 → L0 (operand staging into cube) | — |
mte2_busy_cycles |
GM → L1 (operand load from device memory) | GM → UB (input load) |
mte3_busy_cycles |
— | UB → GM (output store) |
fixpipe_cycles |
L0C → GM (cube result write-out) | — |
cube_busy_cycles |
cube compute | — |
vec_busy_cycles |
— | vector compute |
The bottleneck pipe should sit near 100 % of pmu_total_cycles; the
others run overlapped underneath it. Targets:
- Cube kernel:
max(mte2_busy_cycles, cube_busy_cycles) / pmu_total_cycles ≈ 100 %. Either the L1 load or the cube compute is saturated — whichever the shape is bound by. - Vector kernel:
max(mte2_busy_cycles, vec_busy_cycles) / pmu_total_cycles ≈ 100 %. Either the GM→UB load or the vector compute is saturated. For very store-heavy kernels,mte3_busy_cyclescan be the bottleneck instead.
If both compute and MTE2 are well below 100 %, open the kernel-insight
swimlane: gaps usually mean (a) a missing pl.pipeline on the K loop,
(b) suboptimal instruction scheduling, or (c) incorrectly placed
synchronization barriers.