Model implementation¶
models/ holds end-to-end LLM kernels, one flat directory per model build.
Files ending in _draft.py are work in progress and are not part of the
tracked runnable set.
| Directory | What it implements | pypto-serving |
|---|---|---|
| qwen3_14b | Qwen3-14B BF16 prefill and decode with the serving contract, plus A8W8 and TurboQuant variants and the sampling components | Supported — one-card A2/A3 accuracy |
| deepseek_v4_flash_mtp | DeepSeek V4-Flash at MTP = 1, batch 4 per card: operators, layer and MTP compositions, prefill/decode full forwards | Supported — eight-card accuracy job on relevant PRs |
| deepseek_v4_flash_dspark | The same V4-Flash checkpoint at batch 64 per card and S = 8 DSpark speculation, with TP-sharded DSA-CP attention; under development | Not supported |
| deepseek_v4_pro | Ascend A5 DeepSeek V4-Pro with an optional Flash preset, quantized Hybrid MXFP8-MXFP4 | Not supported |
| deepseek_v4_1_flash | DeepSeek V4.1 Flash TP4/DP2/EP8 scaffold with low-bit SWA/C2A/C1A caches, hierarchical indexer, mHC, and MoE | Not supported |
| glm5_3_flash | GLM-5.3-Flash on a2a3 at TP16/EP16 with W8A8 INT8: the hybrid KDA + NoPE-MLA backbone, the kpool DSA indexer, mHC and MoE. Staging area — configuration, goldens and kernel ABIs only | Not supported |
Each page covers that directory's deployment configuration and how its files compose. The Qwen3-14B and V4-Flash MTP pages also carry the optimization history of their tuned path — Qwen3-14B optimization and V4-Flash decode optimization — which record which levers moved the number and what each one cost.
Entry points take script-specific platform and device arguments; inspect
--help, the platform table, and the
Golden Harness guide.