Xyntetik xyntetik

Xyntetik Runner 1.0, a Zenova AB product

Run AI on your own machine. Know exactly what you are getting.

Running an AI model yourself is mostly guesswork. Will it fit? Did shrinking it break it? Why is the laptop out of memory when nothing is running? Runner is a free, open engine that replaces those guesses with answers you can check.

Free forever, Apache 2.0 One file, nothing to install Mac, Windows and Linux NVIDIA and Apple graphics, or none at all v1.1.0, released 2026-10-07

The mark is an ensö, the Zen circle drawn in one stroke and often left open. One stroke: a single program with nothing hidden underneath. Left open: what is not yet claimed is written down as carefully as what is.

Runner asked whether Granite 4.2 8B fits an RTX 3070: FITS, 35 of 40 layers on the GPU
Real output, Runner on an RTX 3070 (October 2026), shortened to fit. Full capture

See it answer

It read 16 MB of a 5 GB model and already knew.

Before you download anything, Runner looks at the first few megabytes of a model file and tells you whether it will run on your machine, how much lands on the graphics card, and what each memory setting would buy you.

Checked against what really happened on load, it named the right split in 70 of 72 tests and was never too optimistic.

How to ask before you download →

New to this?

What an inference engine is, in three steps.

1

An AI model is a file.

A few gigabytes of numbers that you download, the way you would download a film. On its own it does nothing.

2

An engine is the player.

It opens the file and turns your question into an answer, on your own computer. Nothing you type is sent anywhere.

3

Runner is a player that shows its work.

It tells you whether the file will fit, what was lost when someone shrank it, and it can prove later exactly what the model said.

Why run a model yourself at all?

Your documents never leave the building, there is no bill per question, and it keeps working without an internet connection.

What it can do

Six answers that are guesswork everywhere else.

Each one is a measurement you can repeat, and each links to how it was made.

Do not take our word for it

Two of those measurements, drawn.

Wired memory with a 63 GB model loaded and idle: Runner 35 MB on a 0 to 40 MB scale, llama-server 60.8 GB off the scale
Memory an idle model keeps from you. gpt-oss-120b, 63 GB, loaded and idle on a 128 GB M5 Max; on an 8 GB M1 the same pair is 8 MB against 3,819 MB. Measured 2026-09-01. Source: docs/idle-coexistence-120b-m5max-2026-09-01.md.
One truncated tool call across six engines: only Runner returns a usable call below the control budget
One tool call, cut short, on six engines. Same machine, same prompt and tool, token budgets from 1 to 64; the last column is the control where every engine completes. Measured 2026-08-19. Source: docs/truncation-benchmark.md.

How it compares

Measured side by side, on the same machine.

Runner uses the same model files as llama.cpp, so trying it costs you nothing. These are the places where the measurements differ.

QuestionRunnerOthers measured
A tool call is cut short by the token limit. What does the program get? A call it can run, at every budget from 1 to 16 tokens Nothing usable from vLLM, llama.cpp, Ollama, TensorRT-LLM or SGLang below the control budget
How much memory does an idle model keep from the rest of the machine? 35 MB with a 63 GB model; 8 MB on an 8 GB Mac llama-server: 60.8 GB and 3,819 MB on the same two machines
How close is the output to the model publisher's own implementation? Closer than llama.cpp on 21 of 22 rows across twelve model families llama.cpp, given its stronger configuration at each size
Can you train the compressed file you actually run, and get the same result twice? Yes, byte for byte, with a provenance record The adapter Runner writes scores the same in stock llama.cpp
Can someone else verify what the model said? Signed, chained records that replay to VERIFIED, DIVERGED or UNVERIFIABLE Not claimed by the engines measured

Tool calls: 2026-08-19, docs/truncation-benchmark.md. Idle memory: 2026-09-01 and the M1 measurement, docs/idle-coexistence-120b-m5max-2026-09-01.md. Publisher reference: 2026-09-07, docs/golden-pass-2026-09-07.md. Speed tables and every other measurement are on the evidence page.

About

A Zenova AB product, built in Sweden

Xyntetik is the program under which Zenova AB builds tools for running capable models on hardware you already own. Runner is the released engine, Suite is the testing and evidence layer around it, and Shade is the research lab that feeds both. The engine is free forever under Apache 2.0.

Runner, Suite, Shade and the company behind them →

Support the work

Open infrastructure, funded in the open

Runner is developed openly rather than around a subscription. Contributions pay for measurement time on real hardware, the thing every number on this site cost.