Compile and Runtime Workflow¶
What usually happens when you run python <kernel>.py -p <platform>.
Examples and model harnesses go through golden.run, which takes a kernel of
either form — a module-level @pl.jit function or a @pl.program class — and
picks its compile path from it. A few specialized smoke, artifact-regeneration, and
external-runtime drivers call PyPTO's compile/runtime APIs directly; their
__main__ blocks are the authority for what a command actually validates.
CLI shape¶
A typical model __main__ block parses three flags and dispatches into
the harness:
parser.add_argument("-p", "--platform", choices=["a2a3", "a2a3sim", "a5", "a5sim"])
parser.add_argument("-d", "--device", type=int, default=0)
parser.add_argument("--enable-chip-swimlane", type=int, nargs="?", const=1, default=0, choices=range(5))
args = parser.parse_args()
result = run(
fn=qwen3_decode, # module-level @pl.jit function
specs=build_tensor_specs(...), # TensorSpec / ScalarSpec, in fn's param order
golden_fn=golden_qwen3_decode, # PyTorch reference
config=dict(dump_passes=True, platform=args.platform,
device_id=args.device,
enable_chip_swimlane=args.enable_chip_swimlane),
rtol=3e-3, atol=3e-3,
)
A kernel built as a @pl.program class passes that program as fn=
instead. Both forms share tensor specs, golden computation, runtime dispatch,
and validation, and both read the same config; only the compile step
differs, see Compile configuration.
| Flag | Purpose |
|---|---|
-p / --platform |
Target backend. a2a3 is Ascend 910B/C; a5 is Ascend 950 — both run on real NPU. a2a3sim / a5sim are the matching simulators. |
-d / --device |
Device ID for multi-card hosts. |
--enable-chip-swimlane |
Capture level 0-4 forwarded to the runtime; collects per-task chip swimlane records into the build_output (see Runtime DFX flags). A bare flag means level 1; dispatch analysis needs an explicit 4. |
a2a3* maps to BackendType.Ascend910B; a5* maps to
BackendType.Ascend950.
Multi-card kernels¶
Most kernels take a single -d <id>. A kernel that needs several NPUs is an
EP/TP program that parses -d as a comma-separated device list, and it runs
at its default world size unless an explicit --ep / --tp argument
says otherwise — commonly EP2 for the distributed DeepSeek entries.
Multi-card kernels use HCCL, which silent-crashes inside docker: run them on
the host, in a shell that has entered the Python environment and sourced
CANN's set_env.sh, e.g.
python models/deepseek_v4_flash_mtp/decode_moe.py -p a2a3 --ep 2 -d 0,1. Ring
sizing is not part of that environment — it is per task now, see
Ring Heap and Scope Stats.
Phases inside the Golden Harness¶
The harness prints [RUN] <stage> ... / [RUN] <stage> done (Xs) around
each phase, so the console log is the authoritative trace of what ran:
1. Compile (pypto)¶
Driven by the pypto repo. run normalizes config into one
pypto.runtime.RunConfig. For a @pl.program kernel it calls
pypto.ir.compile(program, **config.compile_kwargs()), the mapping PyPTO
itself owns. For a @pl.jit kernel it calls fn.compile(..., config=...) and
the JIT layer applies that same mapping, specializing the function before
entering the same IR compiler. Both paths run a pass pipeline followed by
a codegen pipeline and normally write
build_output/<ProgramName>_<timestamp>/.
1a. Pass pipeline¶
PassManager.get_strategy(strategy).run_passes(program, ...) runs an
ordered sequence of passes that progressively rewrites the IR. The exact
pass list changes often — consult the pypto repo for the current pipeline,
and look at passes_dump/ when dump_passes is enabled. The direct
@pl.program path inherits ir.compile's enabled default; the @pl.jit path
inherits RunConfig's disabled default.
The end state, regardless of which passes ran, is the same:
- exactly one orchestration function (
FunctionType.Orchestration), - plus one InCore function per outlined
pl.at/pl.spmdregion.
A pl.at region that mixes cube and vector ops is split into two
InCore functions during outlining: one cube-only kernel (matmul,
matmul_acc, …) and one vector-only kernel (cast, add, row_sum, …). The
orchestration function calls them in dependency order.
The InCore / Orchestration boundary the frontend left implicit becomes explicit at this stage.
1b. Codegen pipeline¶
pypto.backend.pto_backend.generate(...) walks the transformed program
and emits files in three streams:
- InCore kernels →
.pto→ C++ wrapper. Each kernel function (or group thereof) goes throughPTOCodegento produce an MLIR text file (.pto) underptoas/. Thenptoas(the external assembler/optimizer toolchain) compiles each.ptoto a C++ kernel wrapper underkernels/aic/(cube) orkernels/aiv/(vector). The ptoas invocations run in a thread pool since each is an independent subprocess.skip_ptoas=Truekeeps the raw.ptofiles and skips the C++ wrapper step (useful for inspecting pure MLIR output or for isolating whether a regression came from pypto's IR→MLIR or from ptoas). - Orchestration → C++.
generate_orchestrationemits oneorchestration/<orch_name>.cppthat drives the kernels through the simpler runtime API (task graph build, scheduling, dependencies). - Config →
kernel_config.py. Records each kernel's name, runtime ID, and core type (cube / vector) for the runtime to load.
When PTOAS_ROOT is set, PyPTO searches only
$PTOAS_ROOT/ptoas and then $PTOAS_ROOT/bin/ptoas; it deliberately does
not fall back to a potentially mismatched binary on PATH. When
PTOAS_ROOT is unset, PyPTO searches PATH.
Output directory layout¶
build_output/<ProgramName>_<ts>/
├── passes_dump/ # IR after each pass, when dump_passes is enabled
├── ptoas/ # raw .pto MLIR + ptoas intermediates
├── kernels/
│ ├── aic/ # cube kernel C++ wrappers from ptoas
│ └── aiv/ # vector kernel C++ wrappers from ptoas
├── orchestration/ # generated AICPU orchestration C++ (compiled into .so)
├── kernel_config.py
├── report/ # memory allocation + scheduling reports
├── data/ # populated by later phases (in/, out/)
└── dfx_outputs/ # runtime DFX artefacts (any --enable-* flag)
Compile configuration¶
config is one dict of pypto.runtime.RunConfig keyword arguments — the same
keys for both kernel forms, because both go through RunConfig.compile_kwargs().
The compile-side ones:
config field |
Compiler mapping |
|---|---|
platform |
Selects the target; ir.compile derives backend_type from it. |
dump_passes |
Write pass IR under passes_dump/. Defaults to False. |
strategy |
Select the optimization strategy. |
save_kernels_dir |
Maps to ir.compile(output_dir=...) — override build_output/<name>_<timestamp>/. |
compile_profiling |
Maps to ir.compile(profiling=...) — compile-stage timing reports under report/. |
dump_ptoas_passes |
Dump full-module IR after every ptoas pass. |
diagnostic_phase, disabled_diagnostics |
Configure compiler diagnostics. |
distributed_config, analyze_auto_scopes_for_deps, memory_planner |
Configure distributed lowering, AUTO-scope dependency analysis, and memory planning; forwarded only when set. |
An unknown key is RunConfig's own TypeError naming it.
A few ir.compile parameters have no RunConfig field — skip_ptoas,
verification_level, emit_source_loc — so config cannot carry them. They
are reachable by calling ir.compile directly on a @pl.program kernel.
To stop after compile without touching the device, see compile_only under
Skipping phases.
2. Generate inputs¶
Each entry of specs is a TensorSpec (named tensor, shape, dtype,
direction) or a ScalarSpec (named scalar, dtype, value); see
golden/spec.py. The list is ordered to match the parameter order of the
top opaque function — single-chip and distributed alike. The harness compares
the spec names against the compiled artifact's parameters element by element
and fails with compiled parameter ABI mismatch (parameter order ...) before
allocating anything; it never rebinds a mis-ordered list by name. For each
entry, allocate a torch tensor:
- Pure inputs and inout initial values are filled via
spec.create_tensor()(init_value=Nonecreates zeros; random data requires an explicit factory such astorch.randn). - Pure outputs are zero-initialised.
- Scalars become 0-D tensors carrying the spec value.
See Golden Harness for the complete
TensorSpec initialization contract.
When save_data=True, the input snapshot is written to data/in/<name>.pt
so the same inputs can be replayed later; this is off by default, so no
snapshot is written unless you opt in. If golden_data=<dir> is passed
instead, the harness loads <dir>/in/*.pt rather than generating fresh
data — useful for deterministic regression checks.
3. Compute golden¶
The golden runs before device execution: it depends only on the input
snapshot, not on the runtime, so the reference is ready for validation. With
save_data=True, a later runtime crash still leaves the persisted
data/out/ snapshot.
If golden_fn is provided, run builds a scratch dict with cloned
inputs and zero-init outputs, calls golden_fn(scratch) (which fills the
output entries in place), and — when save_data=True — writes the result
to data/out/<name>.pt.
If golden_data=<dir> is set, the harness loads <dir>/out/*.pt instead
of recomputing — golden_data always wins over golden_fn.
If neither is provided, validation is skipped and the run reports
PASS (validation skipped).
Golden PyTorch operations use 16 intra-op CPU threads by default. Set
PYPTO_GOLDEN_NUM_THREADS to a positive integer to override the repository
default for a run:
4. Runtime (simpler)¶
Driven by the simpler repo. The harness orders the arguments according to
specs and calls the compiled object with the RunConfig —
compiled(*args, config=...), one call shape for a single-chip
CompiledProgram and an L3 DistributedCompiledProgram alike; resident-weight
L3 programs use the prepared-worker path instead. A runtime_dir replay
rebuilds that same handle from the build directory's metadata sidecar via
from_dir. Tensors are mutated in place, so outputs land in the same Python
tensors after dispatch.
The dispatch reads the same RunConfig the compile did — PyPTO takes the
platform, device, ring sizes and aicpu_thread_num off run_options() and the
five DFX toggles off dfx_options(). The one key the harness intercepts is
log_level, which is not a RunConfig field: it configures PyPTO's runtime
logger and never reaches the config.
Runtime DFX flags¶
PyPTO surfaces simpler's five runtime DFX (Design For X) sub-features as
independent toggles on config. They share the same output
directory and can be enabled in any combination. CLI spellings are
entry-specific; the table lists the common spelling when a script exposes it.
| Kwarg | CLI flag | Artefact under dfx_outputs/ |
|---|---|---|
enable_chip_swimlane=<N> (int 0-4, 0=off) |
--enable-chip-swimlane [N] (bare = 1) |
chip_swimlane_records.json; onboard runs also attempt merged_swimlane_*.json |
enable_dump_args=<N> (int, 0=off) |
--dump-args [N] (bare = 1) |
args_dump/{args_dump.json,args.bin} |
enable_pmu=<N> (int, 0=off) |
--enable-pmu [N] (bare = 2) |
pmu.csv |
enable_dep_gen=True |
--enable-dep-gen |
deps.json |
enable_scope_stats=True |
--enable-scope-stats |
scope_stats/scope_stats.jsonl |
Args-dump level 1 captures only arguments selected with pl.dump_tag or a
dumps= list; level 2 captures every task's tensor payloads and scalar
values. Level 3 captures the same argument metadata without writing tensor
payloads or args.bin.
Chip-swimlane level 1 records AICore start/end per task, 2 adds AICPU
dispatch/finish, 3 adds scheduler phases, and 4 adds orchestrator phases —
see
Capture levels.
Every entry in this repository declares the flag identically, so a bare
--enable-chip-swimlane always means level 1 and every level through 4 is
accepted everywhere; gap attribution and early-dispatch proofs need an explicit
--enable-chip-swimlane 4.
For an onboard chip swimlane run, PyPTO first attempts a dependency-graph
capture and then a clean timing capture so the converter can add dependency
arrows without perturbing the timing pass. Open the generated
merged_swimlane_*.json at
ui.perfetto.dev to visualize per-task
execution on each AICPU / AIC / AIV lane and inspect kernel duration,
gaps, and dependency stalls. Simulator runs retain
chip_swimlane_records.json but do not generate the merged trace because the
required task metadata is unavailable.
The raw dependency and scope-stat files can also be rendered offline:
python -m simpler_setup.tools.deps_viewer \
build_output/<name>/dfx_outputs/deps.json --format html --engine sfdp
python -m simpler_setup.tools.scope_stats_plot \
build_output/<name>/dfx_outputs/scope_stats/scope_stats.jsonl
For kernel-internal swimlanes and MindStudio Insight traces, see
In-Core Simulator Profiling.
The repository workflow can reuse an existing build or drive a case
end-to-end. It writes the export root below
build_output/<ProgramName>_<ts>/kernel_insight_all_funcs_<ts>/.
See pypto's docs/en/dev/03-runtime-dfx.md and the simpler reference at
runtime/docs/dfx/{chip-swimlane-profiling,args-dump,pmu-profiling,dep_gen,scope-stats}.md
for full per-flag details. There is no runtime_profiling /
--runtime-profiling compatibility alias in the current Python API; use
enable_chip_swimlane or the entry's --enable-chip-swimlane flag.
4b. Benchmark (opt-in, before validation)¶
With PYPTO_BENCH=1 in the environment, the harness re-dispatches the
compiled program in a timed loop right after the correctness dispatch and
before validation, and prints
[RUN] effective_us (N rounds) min=… median=… mean=… max=…. It is
env-gated only — no model file needs a flag. PYPTO_BENCH_ROUNDS /
PYPTO_BENCH_WARMUP (default 100 / 5) size the loop and PYPTO_BENCH_RAW
dumps the per-dispatch samples. See
performance-tuning.md
for the output format, the multi-card breakdown, and what the number means.
5. Validate¶
golden.validation.validate_golden compares each device output against
the golden using torch.allclose(rtol, atol) by default. Override
per-output with the compare_fn={"out_name": custom_callable} argument.
golden.validation ships five ready-made gates:
| Comparator | Use case |
|---|---|
topk_pair_compare(vals_name) |
Top-k index outputs whose ordering is implementation-dependent — checks the paired value tensor matches after sort, tolerating legal tie-break swaps. |
ratio_allclose(atol, rtol, max_error_ratio=0.005) |
Quantized kernels where a small outlier fraction may exceed per-point atol + rtol·|expected|. NaN/Inf always fail. |
ratio_reldiff(diff_thd, pct_thd, max_diff_hd=inf) |
cann-recipes-infer-style relative-diff check: per-point rdiff > diff_thd bad-point ratio capped by pct_thd, with optional single-point max_diff_hd cap. |
mapped_pool_ratio_allclose(mapping_name, mapping_shape=…, block_size=…) |
A block-major paged pool (KV cache, compressor state) written through a slot mapping: allocator-mapped rows take the ratio_allclose check, every unmapped row must stay exactly equal to its golden snapshot, so a write outside the mapping fails. leading_rank_axis=True for a [ranks, blocks, block_size, …] pool. |
mapped_pool_ratio_reldiff(mapping_name, mapping_shape=…, block_size=…) |
The same mapped-row and exact-unmapped checks using ratio_reldiff, for composed quantized paths. |
A sixth helper, error_distribution(), is a measurement rather than a gate
— it always passes and prints the error shape. See
Precision Tuning.
Both ratio comparators also take valid_rows / valid_axis / zero_tail:
valid_rows=n compares only the leading n entries along valid_axis
(default axis 0), so a packed buffer's inactive token tail cannot dilute the
error ratio, and zero_tail=True additionally fails the check when that
dropped tail is not all zeros. valid_rows=0 compares nothing and passes.
The harness returns RunResult(passed=True) on success. Validation mismatches
and a small set of harness-level setup errors return
RunResult(passed=False, error=...). Compiler errors, golden-function errors,
and ordinary runtime exceptions generally propagate instead of being converted
to RunResult. A CLI should check a returned result and exit nonzero when it is
false; an uncaught compile/runtime exception is already a nonzero failure.
Skipping phases¶
run knobs that short-circuit the pipeline:
| Knob | Effect |
|---|---|
compile_only=True |
Stops after the compile phase. Useful for a smoke test that just checks the program lowers cleanly. |
runtime_dir="<path>" |
Skips compile and reuses an existing build_output/<...> directory. Useful when iterating on golden_fn or validation logic without recompiling. |
golden_data="<path>" |
Loads inputs from <path>/in/ and goldens from <path>/out/ instead of generating them. golden_data overrides golden_fn. With golden_fn=None and no <path>/out/, only the inputs are replayed and validation is skipped. Useful for deterministic regressions: a previous run leaves these files in its data/ dir, so passing that dir reproduces the exact failing inputs. |
save_data=True (default False) |
Writes the data/in/ + data/out/ snapshot so the exact inputs/goldens can be replayed later via golden_data. Off by default: runs skip the snapshot and validate against the in-memory golden only. Opt in when you need replay; full-model kernels like models/qwen3_14b/{prefill_fwd,decode_fwd}.py expose it as --save-data. |
For diagnosing compile errors, runtime hangs, and precision mismatches, see debugging.md.