Runner documentation
Every page under docs/ in the repository, by file name. Pages that exist on this site link here; the rest link to GitHub.
| file | title |
|---|---|
| adaptation-engine.md | The adaptation engine: scoring, adapters, and training in the serving binary |
| afmoe-cert-goal-2026-08-05.md | Goal: certify afmoe (Trinity-Nano) on this box — 2026-08-05 evening |
| afmoe-cert-report-2026-08-05.md | afmoe (Trinity-Nano) certification report — 2026-08-05 |
| afmoe-divergence-triage-2026-08-05.md | afmoe greedy-identity failure: root cause found — 2026-08-05 evening |
| afmoe-sensitivity-floor-2026-08-05.md | afmoe (Trinity-Nano) sensitivity-floor run — 2026-08-05 |
| agent-compatibility.md | Coding-agent compatibility evidence |
| agent-profile.md | Xyntetik agent profile metadata |
| agent-torture.md | The agent torture suite |
| atem-tool-calling-plan.md | Build plan — atem constrained tool calling |
| bench-2026-08-01-3070.md | Cross-engine benchmarks — RTX 3070 box, 2026-08-01 |
| benchmarks.md | GPU benchmarks — Runner vs llama.cpp (CUDA) |
| blackwell-results-2026-08-17.md | Blackwell Q3_K rung results — 2026-08-17 |
| cert-matrix-2026-08-05.md | Cert-matrix detailed report — GPT-OSS x Gemma 4 derivative ecosystems |
| cert-matrix-goal-2026-08-05.md | Goal: artifact certification matrix — GPT-OSS × Gemma 4 derivative ecosystems |
| cert-matrix-status.md | Cert-matrix status — GPT-OSS x Gemma 4 derivative ecosystems |
| compatibility-program.md | Compatibility program |
| context-drafts.md | Context-grounded drafts: --draft-lookup |
| cuda-gptoss-divergence-2026-08-19.md | The remaining gpt-oss CPU/CUDA divergence is chaotic amplification, not a wrong op |
| cuda-gptoss-router-bias-2026-08-18.md | The CUDA fused MoE path routed gpt-oss without its router bias |
| cuda-microbatch-identity-2026-08-18.md | The CUDA decode microbatch is not bit-identical on quantized models |
| determinism-scope.md | What the determinism claim covers, exactly |
| gemma4-31b-metal-m5max-2026-09-01.md | gemma-4-31B dense Metal validation — M5 Max, 2026-09-01 |
| gpt-oss-120b-metal-m5max-2026-08-31.md | gpt-oss-120B fully resident on Metal — M5 Max, 2026-08-31 |
| gpt-oss-harmony-2026-08-14.md | gpt-oss Harmony chat — before and after, 2026-08-14 |
| granite-cert-2026-08-11.md | granite (IBM Granite 4.1 dense) certification — 2026-08-11 |
| idle-coexistence-120b-m5max-2026-09-01.md | Idle coexistence at 120B — M5 Max, 2026-09-01 |
| llama33-70b-metal-m5max-2026-08-31.md | Llama 3.3 70B dense Metal validation — M5 Max, 2026-08-31 |
| metal-bigfile-limits-m5max-2026-09-01.md | Big-file limits exercised for real — M5 Max, 2026-09-01 |
| metal-decode-dispatch-budget-2026-09-01.md | Where the 120B’s decode milliseconds go — M5 Max, 2026-09-01 |
| metal-decode-kv-traffic-2026-08-15.md | What the long-context decode loss actually is: the KV read, at 1.5 GB/s |
| metal-dispatch-census-2026-08-13.md | Metal dispatch census: where the batch dimension actually collapses |
| metal-fallback.md | Metal runtime fallback ownership |
| metal-full-router-trace-2026-08-31.md | Full Metal router-logit trace |
| metal-gemma4-moe-divergence-2026-08-31.md | gemma-4-26B-A4B generated token 0 on Metal: an unclamped GELU in the routed-expert kernel |
| metal-long-context-decode-2026-08-14.md | Metal decode at long context: the chunked path already captures it |
| metal-microbatch-decode-2026-09-01.md | Metal microbatch decode — M5 Max, 2026-09-01 |
| metal-moe-grouped-mma-2026-09-01.md | Grouped-MMA MoE prefill: the win, and the instrument that judges it — M5 Max, 2026-09-01 |
| metal4-tensor-measurement-2026-08-31.md | Metal 4 tensor-op GEMM: first M5 measurement, and what it does to the premise |
| model-scope.md | Model scope — geographic focus: Europe & US (standing policy, 2026-07-29) |
| moe-support.md | Sparse-MoE support — implementation and test report |
| muse-atem-cert-2026-08-11.md | Muse Glimmer native atem certification — 2026-08-11 |
| muse-glimmer-cert-2026-08-11.md | muse-glimmer (Meta Muse Glimmer 30B) certification — 2026-08-11 |
| negative-result-expert-cache.md | Negative result: an expert-residency cache tier for MoE models larger than RAM |
| negative-result-harmony-analysis-bound.md | Negative result: lifting the Harmony analysis bound on the auto branch |
| negative-result-metal-gemm-occupancy.md | Negative result: Metal prefill GEMM occupancy and tile shape |
| negative-result-metal-moe-expert-major.md | Negative result: expert-major MoE prefill kernels — M5 Max, 2026-09-01 |
| negative-result-metal-multirow-matvec.md | Negative result: multi-row-per-simdgroup Metal decode matvec |
| negative-result-muse-glimmer-selective-2026-08-20.md | Negative result: Runner-built selective quants of Muse Glimmer 30B |
| negative-result-nemotron-lightning-artifacts-2026-08-20.md | Negative result: three of the four planned Nemotron-3.5-Lightning artifacts |
| ornith-reference.md | Ornith 1.0 9B reference gate |
| performance.md | Performance: closing the CPU/GPU gap |
| quant-fidelity.md | Quant-vs-tool-call-fidelity harness |
| qwen3-235b-metal-m5max-2026-09-01.md | Qwen3 235B-A22B Q2_K-mix Metal validation — M5 Max, 2026-09-01 |
| qwen3-30b-a3b-metal-m5max-2026-08-31.md | Qwen3 30B-A3B Q8_0 Metal validation — M5 Max, 2026-08-31 |
| qwen35-precision-2026-08-21.md | Qwen3.5-4B precision-vs-fidelity — BF16 vs upstream Q4_K_M (2026-08-21) |
| reproducible-lora-training-receipts.md | Reproducible LoRA training with receipts |
| review-2026-08-12.md | Bug-hunt and simplification review — 2026-08-12 |
| review-2026-08-13.md | External-evaluation bug sweep — 2026-08-13 |
| sensitivity-floors-m5max-2026-09-01.md | Self-sensitivity floors: 235B and gpt-oss — M5 Max, 2026-09-01 |
| sublayer-removal.md | Sublayer removal: --remove-sublayer |
| tool-choice-boundary-lane.md | Tool-choice decision-boundary lane (unlabeled) |
| train-lora-on-quantized-gguf.md | Train a LoRA directly on the quantized GGUF you serve |
| training-floor-m5max-2026-09-01.md | The local training floor, re-measured — M5 Max, 2026-09-01 |
| tray-controller.md | Tray / menu-bar controller |
| tray-v1-report.md | Tray controller v1 — validation report (branch tray-controller) |
| truncation-benchmark.md | Truncation-recovery benchmark |
| truncation-safe-tool-calling.md | Tool calls that survive the token limit |
| windows-remote-checks.md | Windows remote-check protocol (learned the hard way, 2026-08-10) |
| write-side-stall.md | Real write-side stall experiment |