Skip to content

Model implementation

models/ holds end-to-end LLM kernels, one flat directory per model build. Files ending in _draft.py are work in progress and are not part of the tracked runnable set.

Directory What it implements pypto-serving
qwen3_14b Qwen3-14B BF16 prefill and decode with the serving contract, plus A8W8 and TurboQuant variants and the sampling components Supported — one-card A2/A3 accuracy
deepseek_v4_flash_mtp DeepSeek V4-Flash at MTP = 1, batch 4 per card: operators, layer and MTP compositions, prefill/decode full forwards Supported — eight-card accuracy job on relevant PRs
deepseek_v4_flash_dspark The same V4-Flash checkpoint at batch 64 per card and S = 8 DSpark speculation, with TP-sharded DSA-CP attention; under development Not supported
deepseek_v4_pro Ascend A5 DeepSeek V4-Pro with an optional Flash preset, quantized Hybrid MXFP8-MXFP4 Not supported
deepseek_v4_1_flash DeepSeek V4.1 Flash TP4/DP2/EP8 scaffold with low-bit SWA/C2A/C1A caches, hierarchical indexer, mHC, and MoE Not supported
glm5_3_flash GLM-5.3-Flash on a2a3 at TP16/EP16 with W8A8 INT8: the hybrid KDA + NoPE-MLA backbone, the kpool DSA indexer, mHC and MoE. Staging area — configuration, goldens and kernel ABIs only Not supported

Each page covers that directory's deployment configuration and how its files compose. The Qwen3-14B and V4-Flash MTP pages also carry the optimization history of their tuned path — Qwen3-14B optimization and V4-Flash decode optimization — which record which levers moved the number and what each one cost.

Entry points take script-specific platform and device arguments; inspect --help, the platform table, and the Golden Harness guide.

python models/qwen3_14b/decode_fwd.py -p a2a3 -d 0
python models/deepseek_v4_flash_mtp/decode_layer.py -p a2a3 --ep 2 -d 0,1