Skip to content

Compile and Runtime Workflow

What usually happens when you run python <kernel>.py -p <platform>. Examples and model harnesses go through golden.run, which takes a kernel of either form — a module-level @pl.jit function or a @pl.program class — and picks its compile path from it. A few specialized smoke, artifact-regeneration, and external-runtime drivers call PyPTO's compile/runtime APIs directly; their __main__ blocks are the authority for what a command actually validates.

CLI shape

A typical model __main__ block parses three flags and dispatches into the harness:

parser.add_argument("-p", "--platform", choices=["a2a3", "a2a3sim", "a5", "a5sim"])
parser.add_argument("-d", "--device", type=int, default=0)
parser.add_argument("--enable-chip-swimlane", type=int, nargs="?", const=1, default=0, choices=range(5))
args = parser.parse_args()

result = run(
    fn=qwen3_decode,                          # module-level @pl.jit function
    specs=build_tensor_specs(...),            # TensorSpec / ScalarSpec, in fn's param order
    golden_fn=golden_qwen3_decode,            # PyTorch reference
    config=dict(dump_passes=True, platform=args.platform,
                device_id=args.device,
                enable_chip_swimlane=args.enable_chip_swimlane),
    rtol=3e-3, atol=3e-3,
)

A kernel built as a @pl.program class passes that program as fn= instead. Both forms share tensor specs, golden computation, runtime dispatch, and validation, and both read the same config; only the compile step differs, see Compile configuration.

Flag Purpose
-p / --platform Target backend. a2a3 is Ascend 910B/C; a5 is Ascend 950 — both run on real NPU. a2a3sim / a5sim are the matching simulators.
-d / --device Device ID for multi-card hosts.
--enable-chip-swimlane Capture level 0-4 forwarded to the runtime; collects per-task chip swimlane records into the build_output (see Runtime DFX flags). A bare flag means level 1; dispatch analysis needs an explicit 4.

a2a3* maps to BackendType.Ascend910B; a5* maps to BackendType.Ascend950.

Multi-card kernels

Most kernels take a single -d <id>. A kernel that needs several NPUs is an EP/TP program that parses -d as a comma-separated device list, and it runs at its default world size unless an explicit --ep / --tp argument says otherwise — commonly EP2 for the distributed DeepSeek entries.

Multi-card kernels use HCCL, which silent-crashes inside docker: run them on the host, in a shell that has entered the Python environment and sourced CANN's set_env.sh, e.g. python models/deepseek_v4_flash_mtp/decode_moe.py -p a2a3 --ep 2 -d 0,1. Ring sizing is not part of that environment — it is per task now, see Ring Heap and Scope Stats.

Phases inside the Golden Harness

The harness prints [RUN] <stage> ... / [RUN] <stage> done (Xs) around each phase, so the console log is the authoritative trace of what ran:

1. Compile (pypto)

Driven by the pypto repo. run normalizes config into one pypto.runtime.RunConfig. For a @pl.program kernel it calls pypto.ir.compile(program, **config.compile_kwargs()), the mapping PyPTO itself owns. For a @pl.jit kernel it calls fn.compile(..., config=...) and the JIT layer applies that same mapping, specializing the function before entering the same IR compiler. Both paths run a pass pipeline followed by a codegen pipeline and normally write build_output/<ProgramName>_<timestamp>/.

1a. Pass pipeline

PassManager.get_strategy(strategy).run_passes(program, ...) runs an ordered sequence of passes that progressively rewrites the IR. The exact pass list changes often — consult the pypto repo for the current pipeline, and look at passes_dump/ when dump_passes is enabled. The direct @pl.program path inherits ir.compile's enabled default; the @pl.jit path inherits RunConfig's disabled default.

The end state, regardless of which passes ran, is the same:

  • exactly one orchestration function (FunctionType.Orchestration),
  • plus one InCore function per outlined pl.at / pl.spmd region.

A pl.at region that mixes cube and vector ops is split into two InCore functions during outlining: one cube-only kernel (matmul, matmul_acc, …) and one vector-only kernel (cast, add, row_sum, …). The orchestration function calls them in dependency order.

The InCore / Orchestration boundary the frontend left implicit becomes explicit at this stage.

1b. Codegen pipeline

pypto.backend.pto_backend.generate(...) walks the transformed program and emits files in three streams:

  • InCore kernels → .pto → C++ wrapper. Each kernel function (or group thereof) goes through PTOCodegen to produce an MLIR text file (.pto) under ptoas/. Then ptoas (the external assembler/optimizer toolchain) compiles each .pto to a C++ kernel wrapper under kernels/aic/ (cube) or kernels/aiv/ (vector). The ptoas invocations run in a thread pool since each is an independent subprocess. skip_ptoas=True keeps the raw .pto files and skips the C++ wrapper step (useful for inspecting pure MLIR output or for isolating whether a regression came from pypto's IR→MLIR or from ptoas).
  • Orchestration → C++. generate_orchestration emits one orchestration/<orch_name>.cpp that drives the kernels through the simpler runtime API (task graph build, scheduling, dependencies).
  • Config → kernel_config.py. Records each kernel's name, runtime ID, and core type (cube / vector) for the runtime to load.

When PTOAS_ROOT is set, PyPTO searches only $PTOAS_ROOT/ptoas and then $PTOAS_ROOT/bin/ptoas; it deliberately does not fall back to a potentially mismatched binary on PATH. When PTOAS_ROOT is unset, PyPTO searches PATH.

Output directory layout

build_output/<ProgramName>_<ts>/
├── passes_dump/    # IR after each pass, when dump_passes is enabled
├── ptoas/          # raw .pto MLIR + ptoas intermediates
├── kernels/
│   ├── aic/        # cube kernel C++ wrappers from ptoas
│   └── aiv/        # vector kernel C++ wrappers from ptoas
├── orchestration/  # generated AICPU orchestration C++ (compiled into .so)
├── kernel_config.py
├── report/         # memory allocation + scheduling reports
├── data/           # populated by later phases (in/, out/)
└── dfx_outputs/    # runtime DFX artefacts (any --enable-* flag)

Compile configuration

config is one dict of pypto.runtime.RunConfig keyword arguments — the same keys for both kernel forms, because both go through RunConfig.compile_kwargs(). The compile-side ones:

config field Compiler mapping
platform Selects the target; ir.compile derives backend_type from it.
dump_passes Write pass IR under passes_dump/. Defaults to False.
strategy Select the optimization strategy.
save_kernels_dir Maps to ir.compile(output_dir=...) — override build_output/<name>_<timestamp>/.
compile_profiling Maps to ir.compile(profiling=...) — compile-stage timing reports under report/.
dump_ptoas_passes Dump full-module IR after every ptoas pass.
diagnostic_phase, disabled_diagnostics Configure compiler diagnostics.
distributed_config, analyze_auto_scopes_for_deps, memory_planner Configure distributed lowering, AUTO-scope dependency analysis, and memory planning; forwarded only when set.

An unknown key is RunConfig's own TypeError naming it.

A few ir.compile parameters have no RunConfig field — skip_ptoas, verification_level, emit_source_loc — so config cannot carry them. They are reachable by calling ir.compile directly on a @pl.program kernel.

To stop after compile without touching the device, see compile_only under Skipping phases.

2. Generate inputs

Each entry of specs is a TensorSpec (named tensor, shape, dtype, direction) or a ScalarSpec (named scalar, dtype, value); see golden/spec.py. The list is ordered to match the parameter order of the top opaque function — single-chip and distributed alike. The harness compares the spec names against the compiled artifact's parameters element by element and fails with compiled parameter ABI mismatch (parameter order ...) before allocating anything; it never rebinds a mis-ordered list by name. For each entry, allocate a torch tensor:

  • Pure inputs and inout initial values are filled via spec.create_tensor() (init_value=None creates zeros; random data requires an explicit factory such as torch.randn).
  • Pure outputs are zero-initialised.
  • Scalars become 0-D tensors carrying the spec value.

See Golden Harness for the complete TensorSpec initialization contract.

When save_data=True, the input snapshot is written to data/in/<name>.pt so the same inputs can be replayed later; this is off by default, so no snapshot is written unless you opt in. If golden_data=<dir> is passed instead, the harness loads <dir>/in/*.pt rather than generating fresh data — useful for deterministic regression checks.

3. Compute golden

The golden runs before device execution: it depends only on the input snapshot, not on the runtime, so the reference is ready for validation. With save_data=True, a later runtime crash still leaves the persisted data/out/ snapshot.

If golden_fn is provided, run builds a scratch dict with cloned inputs and zero-init outputs, calls golden_fn(scratch) (which fills the output entries in place), and — when save_data=True — writes the result to data/out/<name>.pt.

If golden_data=<dir> is set, the harness loads <dir>/out/*.pt instead of recomputing — golden_data always wins over golden_fn.

If neither is provided, validation is skipped and the run reports PASS (validation skipped).

Golden PyTorch operations use 16 intra-op CPU threads by default. Set PYPTO_GOLDEN_NUM_THREADS to a positive integer to override the repository default for a run:

PYPTO_GOLDEN_NUM_THREADS=8 \
  python models/deepseek_v4_flash_mtp/decode_csa.py -p a2a3 -d 0

4. Runtime (simpler)

Driven by the simpler repo. The harness orders the arguments according to specs and calls the compiled object with the RunConfig — compiled(*args, config=...), one call shape for a single-chip CompiledProgram and an L3 DistributedCompiledProgram alike; resident-weight L3 programs use the prepared-worker path instead. A runtime_dir replay rebuilds that same handle from the build directory's metadata sidecar via from_dir. Tensors are mutated in place, so outputs land in the same Python tensors after dispatch.

The dispatch reads the same RunConfig the compile did — PyPTO takes the platform, device, ring sizes and aicpu_thread_num off run_options() and the five DFX toggles off dfx_options(). The one key the harness intercepts is log_level, which is not a RunConfig field: it configures PyPTO's runtime logger and never reaches the config.

Runtime DFX flags

PyPTO surfaces simpler's five runtime DFX (Design For X) sub-features as independent toggles on config. They share the same output directory and can be enabled in any combination. CLI spellings are entry-specific; the table lists the common spelling when a script exposes it.

Kwarg CLI flag Artefact under dfx_outputs/
enable_chip_swimlane=<N> (int 0-4, 0=off) --enable-chip-swimlane [N] (bare = 1) chip_swimlane_records.json; onboard runs also attempt merged_swimlane_*.json
enable_dump_args=<N> (int, 0=off) --dump-args [N] (bare = 1) args_dump/{args_dump.json,args.bin}
enable_pmu=<N> (int, 0=off) --enable-pmu [N] (bare = 2) pmu.csv
enable_dep_gen=True --enable-dep-gen deps.json
enable_scope_stats=True --enable-scope-stats scope_stats/scope_stats.jsonl

Args-dump level 1 captures only arguments selected with pl.dump_tag or a dumps= list; level 2 captures every task's tensor payloads and scalar values. Level 3 captures the same argument metadata without writing tensor payloads or args.bin.

Chip-swimlane level 1 records AICore start/end per task, 2 adds AICPU dispatch/finish, 3 adds scheduler phases, and 4 adds orchestrator phases — see Capture levels. Every entry in this repository declares the flag identically, so a bare --enable-chip-swimlane always means level 1 and every level through 4 is accepted everywhere; gap attribution and early-dispatch proofs need an explicit --enable-chip-swimlane 4.

For an onboard chip swimlane run, PyPTO first attempts a dependency-graph capture and then a clean timing capture so the converter can add dependency arrows without perturbing the timing pass. Open the generated merged_swimlane_*.json at ui.perfetto.dev to visualize per-task execution on each AICPU / AIC / AIV lane and inspect kernel duration, gaps, and dependency stalls. Simulator runs retain chip_swimlane_records.json but do not generate the merged trace because the required task metadata is unavailable.

The raw dependency and scope-stat files can also be rendered offline:

python -m simpler_setup.tools.deps_viewer \
  build_output/<name>/dfx_outputs/deps.json --format html --engine sfdp
python -m simpler_setup.tools.scope_stats_plot \
  build_output/<name>/dfx_outputs/scope_stats/scope_stats.jsonl

For kernel-internal swimlanes and MindStudio Insight traces, see In-Core Simulator Profiling. The repository workflow can reuse an existing build or drive a case end-to-end. It writes the export root below build_output/<ProgramName>_<ts>/kernel_insight_all_funcs_<ts>/.

See pypto's docs/en/dev/03-runtime-dfx.md and the simpler reference at runtime/docs/dfx/{chip-swimlane-profiling,args-dump,pmu-profiling,dep_gen,scope-stats}.md for full per-flag details. There is no runtime_profiling / --runtime-profiling compatibility alias in the current Python API; use enable_chip_swimlane or the entry's --enable-chip-swimlane flag.

4b. Benchmark (opt-in, before validation)

With PYPTO_BENCH=1 in the environment, the harness re-dispatches the compiled program in a timed loop right after the correctness dispatch and before validation, and prints [RUN] effective_us (N rounds) min=… median=… mean=… max=…. It is env-gated only — no model file needs a flag. PYPTO_BENCH_ROUNDS / PYPTO_BENCH_WARMUP (default 100 / 5) size the loop and PYPTO_BENCH_RAW dumps the per-dispatch samples. See performance-tuning.md for the output format, the multi-card breakdown, and what the number means.

5. Validate

golden.validation.validate_golden compares each device output against the golden using torch.allclose(rtol, atol) by default. Override per-output with the compare_fn={"out_name": custom_callable} argument. golden.validation ships five ready-made gates:

Comparator Use case
topk_pair_compare(vals_name) Top-k index outputs whose ordering is implementation-dependent — checks the paired value tensor matches after sort, tolerating legal tie-break swaps.
ratio_allclose(atol, rtol, max_error_ratio=0.005) Quantized kernels where a small outlier fraction may exceed per-point atol + rtol·|expected|. NaN/Inf always fail.
ratio_reldiff(diff_thd, pct_thd, max_diff_hd=inf) cann-recipes-infer-style relative-diff check: per-point rdiff > diff_thd bad-point ratio capped by pct_thd, with optional single-point max_diff_hd cap.
mapped_pool_ratio_allclose(mapping_name, mapping_shape=…, block_size=…) A block-major paged pool (KV cache, compressor state) written through a slot mapping: allocator-mapped rows take the ratio_allclose check, every unmapped row must stay exactly equal to its golden snapshot, so a write outside the mapping fails. leading_rank_axis=True for a [ranks, blocks, block_size, …] pool.
mapped_pool_ratio_reldiff(mapping_name, mapping_shape=…, block_size=…) The same mapped-row and exact-unmapped checks using ratio_reldiff, for composed quantized paths.

A sixth helper, error_distribution(), is a measurement rather than a gate — it always passes and prints the error shape. See Precision Tuning.

Both ratio comparators also take valid_rows / valid_axis / zero_tail: valid_rows=n compares only the leading n entries along valid_axis (default axis 0), so a packed buffer's inactive token tail cannot dilute the error ratio, and zero_tail=True additionally fails the check when that dropped tail is not all zeros. valid_rows=0 compares nothing and passes.

The harness returns RunResult(passed=True) on success. Validation mismatches and a small set of harness-level setup errors return RunResult(passed=False, error=...). Compiler errors, golden-function errors, and ordinary runtime exceptions generally propagate instead of being converted to RunResult. A CLI should check a returned result and exit nonzero when it is false; an uncaught compile/runtime exception is already a nonzero failure.

Skipping phases

run knobs that short-circuit the pipeline:

Knob Effect
compile_only=True Stops after the compile phase. Useful for a smoke test that just checks the program lowers cleanly.
runtime_dir="<path>" Skips compile and reuses an existing build_output/<...> directory. Useful when iterating on golden_fn or validation logic without recompiling.
golden_data="<path>" Loads inputs from <path>/in/ and goldens from <path>/out/ instead of generating them. golden_data overrides golden_fn. With golden_fn=None and no <path>/out/, only the inputs are replayed and validation is skipped. Useful for deterministic regressions: a previous run leaves these files in its data/ dir, so passing that dir reproduces the exact failing inputs.
save_data=True (default False) Writes the data/in/ + data/out/ snapshot so the exact inputs/goldens can be replayed later via golden_data. Off by default: runs skip the snapshot and validate against the in-memory golden only. Opt in when you need replay; full-model kernels like models/qwen3_14b/{prefill_fwd,decode_fwd}.py expose it as --save-data.

For diagnosing compile errors, runtime hangs, and precision mismatches, see debugging.md.