PyPTO-Lib¶
PyPTO-Lib holds PyPTO kernels and end-to-end LLM models for Ascend NPUs, with the harness that compiles, runs, and validates them.
Use this documentation to move from a first simulator run to validated model kernels and systematic precision or performance tuning.
Choose a path¶
Run a kernel¶
Start with installation and environment setup, then run your first kernel.
Write a kernel¶
Use the PyPTO Coding chapter for the canonical kernel style and hand-written CCE extern-kernel conventions.
Examples¶
Use the example catalog for focused, self-contained kernels organized by learning level.
Models¶
Use the model pages for end-to-end and component-level model implementations.
Run and Validate¶
Read the compile and runtime workflow and the Golden Harness overview to understand how a script compiles, executes, and checks its result.
Diagnose and optimize¶
Begin with the debugging playbook, then choose the precision or performance workflow for the problem at hand. DeepSeek V4 decode optimization follows one model's decode path end to end and records which levers paid. Qwen3-14B optimization records how a single-card dense model's kernels were tuned, in the order the work happened.
Ecosystem¶
- PyPTO documentation covers the programming model, language semantics, and compiler.
- simpler documentation covers the runtime.
- PyPTO-Lib on GitHub contains the source code and issue tracker.