Xyntetik Runner 1.0, a Zenova AB product
Run AI on your own machine. Know exactly what you are getting.
Running an AI model yourself is mostly guesswork. Will it fit? Did shrinking it break it? Why is the laptop out of memory when nothing is running? Runner is a free, open engine that replaces those guesses with answers you can check.
The mark is an ensö, the Zen circle drawn in one stroke and often left open. One stroke: a single program with nothing hidden underneath. Left open: what is not yet claimed is written down as carefully as what is.
See it answer
It read 16 MB of a 5 GB model and already knew.
Before you download anything, Runner looks at the first few megabytes of a model file and tells you whether it will run on your machine, how much lands on the graphics card, and what each memory setting would buy you.
Checked against what really happened on load, it named the right split in 70 of 72 tests and was never too optimistic.
How to ask before you download →New to this?
What an inference engine is, in three steps.
An AI model is a file.
A few gigabytes of numbers that you download, the way you would download a film. On its own it does nothing.
An engine is the player.
It opens the file and turns your question into an answer, on your own computer. Nothing you type is sent anywhere.
Runner is a player that shows its work.
It tells you whether the file will fit, what was lost when someone shrank it, and it can prove later exactly what the model said.
Your documents never leave the building, there is no bill per question, and it keeps working without an internet connection.
What it can do
Six answers that are guesswork everywhere else.
Each one is a measurement you can repeat, and each links to how it was made.
Know whether a model fits before you download it.
Runner reads the start of the file and predicts how it will run on your machine. It called the real result in 70 of 72 tests and was one step cautious in the other two.
Ask before you download → Smaller and closerShrink a model and see what it cost.
One command and a one-line plan built a Qwen3-30B that is smaller than the official 4-bit file and closer to the original: 99.5% agreement where that file reads 94.75%.
Make your own model file → 41%Read longer documents in the same memory.
The notes a model keeps while it reads can outgrow your memory. Runner stores them in about 41% of the usual space, with full agreement on models from 4B to 30B.
More context, less memory → 35 MBLeave it running.
A 63 GB model sitting idle holds 35 MB of your memory under Runner. On the same Mac, llama-server holds 60.8 GB. Runner hands the rest back until you ask again.
A server that gives memory back → 1 of 6Tool calls that survive running out of words.
Six engines were cut off in the middle of a tool call. Only Runner handed back something a program could still run, at every budget tested.
What six engines sent back → Same bytes, twiceTeach the exact file you run.
Runner trains a small add-on directly on the compressed model you already use. Two runs wrote the identical file, and an outside tester got the same result on their own hardware.
Train what you serve →Do not take our word for it
Two of those measurements, drawn.
Who it is for
One engine, four reasons to care.
Curious
I want to try AI that stays on my computer.
Download one file, start a model, ask it something. It takes about a minute and nothing leaves your machine.
Start here →Building
I want my tools to just connect.
Runner speaks the OpenAI and Anthropic APIs, so Claude Code, Codex, OpenCode and your own apps point at it without changes.
Connect a coding agent →Leading a team
I have to answer for what the model did.
Keep data in the building, know whether a model fits your machines before you commit to it, and hand a reviewer a signed record they can replay.
Records a reviewer can check →Backing
I want to know what is behind it.
An open engine, a research lab that publishes its failures, a measurement behind every claim, and a Swedish company that builds it.
The company behind it →The Xyntetik-Lab on Hugging Face
Everything the lab makes is published, failures included.
Models, the measurements behind them and the experiments that did not work, all in the open.
Can a model half the size do its parent's job?
Kvist-14B was distilled from a 30B model for one thing, tool-calling agents on a 24 GB card, and it solves 57 of the 60 held-out tasks its parent solves.
Meet Kvist →What survives when you cut pieces out of a 30B model?
Nine preregistered surgeries on one frozen Muse-Glimmer-30B. Two of the results run with up to 6% of the decoder gone and still pass the fidelity bar.
See the surgery →A 3-bit Qwen3.8-27B that fits a 16 GB card whole.
Another lab's file with every block scale retrained against the original. It is the lab's most downloaded file so far, and its card shows it side by side with the source.
Get the file →Can a model learn with no human text at all?
GENESIS: eighteen preregistered experiments on tiny models that only ever saw programs they ran themselves. Read what survived, and everything that did not.
Read the record →Also in the box
Everything the one file does.
Serves your apps
OpenAI and Anthropic APIs on your own machine, several requests at once, with a /metrics endpoint and models fetched and checked by one flag.
Answers in the shape you asked for
JSON and JSON Schema are guaranteed, not hoped for, and the model can say how sure it was so unsure answers go to a person.
Structured output →Long agent turns that finish
A cap on how long a model thinks out loud and a guard that notices when it starts going round in circles.
Keep agents on track →Proof of what was said
Signed records that replay, evidence packs a reviewer checks offline, and a signature check on the model file itself.
Receipts →One graphics card, several models
Programs queue for the card instead of crashing each other, and an idle model unloads itself after a while.
Share a GPU →Shadow mode
Keep the coding assistant you pay for. Shadow retries your own finished tasks with a local model under your own tests and keeps score.
How shadow mode works →How it compares
Measured side by side, on the same machine.
Runner uses the same model files as llama.cpp, so trying it costs you nothing. These are the places where the measurements differ.
| Question | Runner | Others measured |
|---|---|---|
| A tool call is cut short by the token limit. What does the program get? | A call it can run, at every budget from 1 to 16 tokens | Nothing usable from vLLM, llama.cpp, Ollama, TensorRT-LLM or SGLang below the control budget |
| How much memory does an idle model keep from the rest of the machine? | 35 MB with a 63 GB model; 8 MB on an 8 GB Mac | llama-server: 60.8 GB and 3,819 MB on the same two machines |
| How close is the output to the model publisher's own implementation? | Closer than llama.cpp on 21 of 22 rows across twelve model families | llama.cpp, given its stronger configuration at each size |
| Can you train the compressed file you actually run, and get the same result twice? | Yes, byte for byte, with a provenance record | The adapter Runner writes scores the same in stock llama.cpp |
| Can someone else verify what the model said? | Signed, chained records that replay to VERIFIED, DIVERGED or UNVERIFIABLE | Not claimed by the engines measured |
Tool calls: 2026-08-19, docs/truncation-benchmark.md. Idle memory: 2026-09-01 and the M1 measurement, docs/idle-coexistence-120b-m5max-2026-09-01.md. Publisher reference: 2026-09-07, docs/golden-pass-2026-09-07.md. Speed tables and every other measurement are on the evidence page.