Skip to content

Splash engine: serve the MTPLX API over Inco's Apple-silicon runtime - #514

Open
ksg98 wants to merge 4 commits into
youssofal:mainfrom
ksg98:splash-engine
Open

ksg98 wants to merge 4 commits into
youssofal:mainfrom
ksg98:splash-engine

Conversation

@ksg98

@ksg98 ksg98 commented Sep 20, 2026 •

Copy link
Copy Markdown

Adds Inco's Splash as a second inference engine behind the existing API:

brew install incoai/tap/splash
mtplx serve --engine splash --model incoai/Qwen3.8-27B-Splash

Same endpoints, same dashboard, same native app, same browser chat. Only the kernels change. --engine defaults to mlx, so nothing changes unless you ask for it.

Splash is built the opposite way round from a general engine: it supports a small set of models and rebuilds itself around each one, with fused Metal kernels compiled for that model's exact shapes and a DFlash 2 draft trained for it. It's faster on the two packages it supports, and it loads only those packages.

Measured on a 64 GB M3 Max

macOS 26.6.2, Splash 1.0, incoai/Qwen3.8-27B-Splash, reasoning off, through the bridge over HTTP.

decode TTFT draft acceptance
Short code prompt, 300 tok out 66.5 tok/s 432 ms 77%
Short prose prompt, 136 tok out 30.7 tok/s 252 ms 28%

Decode tracks draft acceptance closely, and acceptance depends on how predictable the text is. Code runs about twice as fast as prose on the same machine. Over a five-request mix: min 30.7, mean 45.4, p95 62.7 tok/s.

Prefill and prefix reuse, on a 25,812-token prompt:

TTFT prefill
Cold 129.7 s 199 tok/s
Same prompt again 275 ms 25,792 of 25,812 tokens reused

That's the part worth calling out: replaying a 25K-token context returns the first token in 275 ms instead of 129.7 s, a ~470× improvement, because Splash reuses the prefix it already computed. It makes multi-turn work on a long document feel immediate after the first turn.

Memory: 21.1 GB active and peak, on a 64 GB machine.

Splash 1.0.2. The branch is now tested on Splash 1.0.2 (brew upgrade incoai/tap/splash). Spot checks on the same machine: two short code prompts decoded at 96 and 106 tok/s, with 62% and 73% of drafts accepted. These are single samples, not a rerun of the benchmark above.

For reference, Inco publishes 74 tok/s decode and 363 tok/s prefill at 32K for this model on a 48 GB M5 Pro.

How it fits together

mtplx serve --engine {mlx|splash}
   |                      |
   v                      v
MLX in-process      splash serve on a private loopback port

Splash has no runtime model-swap API, so loading a model is process lifecycle: the bridge supervises one engine process and proxies to it. Splash 1.0's CLI also hardcodes port 8000 behind an exclusive lock, so the bridge runs its inner server on a private port and keeps the public one for MTPLX.

--engine splash imports no MLX at all, so a Mac that only ever runs Splash doesn't need it installed.

The dashboard

The bridge drives mtplx.server.dashboard_state (the MLX server's own module, which has no MLX dependency) rather than imitating its output. The Live tab therefore works unchanged: progress events move the decode gauge, completed carries each request's envelope, new_max_tps fires the record toast, and the min/max/mean/p95 row comes from the same RollingMetrics.

A finished request's speeds are the engine's own: Splash's /status counters are cumulative, so the bridge reads them either side of a request and reports the difference. Only the live gauge mid-stream is timed by the bridge, because Splash publishes counters rather than a live rate. Thinking tokens count toward that live rate. Memory is memory_actual, the process footprint, not the KV pages alone.

Chat: the same UI on both engines

Browser chat. http://127.0.0.1:8000/ serves MTPLX's own chat page on both engines. The page builder (_chat_ui_html) moved from openai.py into mtplx/server/chat_page.py, which uses only the standard library, so the bridge renders the exact same page without importing MLX. openai.py now imports it; its diff is that one import and the moved function. On Splash the page fills a few engine slots differently:

  • Speculative: reads DFlash 2 · on, with the package's draft length (7 tokens, read from its manifest). Both are fixed.
  • Presence penalty: fixed at 0, because Splash rejects penalties.
  • Top P / Top K: the sliders cover only the range Splash accepts, so no setting on the page can fail a request.

Replies carry the same stats line as MLX, for example DFlash 2 71% accepted · 105.6 tok/s · 43 tokens · 37 thinking · ttft 1.03s.

Thinking selector in the chat bar, on both engines. Splash's own page had one next to Send, so switching pages would have lost it. The chat bar now has a Thinking selector bound to the same reasoning setting as the sidebar. Its choices come from the loaded model's reasoning policy:

  • Qwen 3.8: Auto / XHigh / Medium / Low / Off.
  • A model without effort levels: Auto / On / Off.
  • A model without thinking: the selector is hidden.

A chosen level syncs like every other setting and is sent with each turn. This is the one visible change on the MLX page.

Stream stats for the app's chat. The MLX server adds mtplx_progress frames and puts usage and mtplx_stats on the finish frame, and the app's chat builds its live tok/s chip and per-reply footer from them. Splash sends neither, so Splash replies rendered with no stats. The bridge now adds them:

  • a progress frame about every 200 ms;
  • the finish frame held until the stream ends, then stamped with the engine's own per-request numbers.

The reply footer also gains a cached count from usage.prompt_tokens_details.cached_tokens, which both engines send. A Stop sent with Splash's chatcmpl-… id cancels the right request.

Live settings on Splash

/v1/mtplx/settings was read-only on the bridge, and writes were dropped. The app's own chat sends no sampling fields, so its parameter panel never reached Splash. The bridge now keeps live defaults the way the MLX server does. They are shared by the browser chat, the app's parameter panel and the app's chat, and filled into any request that leaves a field out.

  • Out-of-range values are moved to the nearest value Splash runs. The reply lists what moved (adjusted) or was not applied (ignored), so every panel shows the value actually in use.
  • top_k: 0 (MTPLX's "no filter") becomes 32, Splash's widest, not a greedy 1.
  • "Hide thinking" reaches Splash as reasoning_effort: "none", the switch its template reads. It ignores enable_thinking.
  • Reasoning policy: declared in model_controls and the settings reply, with effort levels read from the package's chat template (Qwen 3.8: xhigh, medium, low; default xhigh). The app's parameter panel shows its effort picker on Splash from the same policy.

What Splash can't do, and how that's reported

MLX engine Splash engine
Models many architectures Splash packages only
Speculation native MTP, tunable depth DFlash 2, fixed
KV cache 4-bit / 8-bit / unquantized 8-bit, fixed
Sampling temperature, top_p, top_k (0 = off), presence penalty temperature 0–2, top_p above 0, top_k 1–32, no penalties
Runtime settings scheduler, batching, adaptive depth sampling defaults and reasoning only

Splash's Metal kernels read 8-bit KV directly, so the width is a property of the compiled kernels, not a setting. There is no flag, env var or engine argument for it. The bridge publishes kv_quant_policy: {supported: false, modes: ["q8"], disabled_reason: …}. Both KV controls in the app (the Settings card and the top-bar inference-params panel) lock to q8 and explain why, rather than offering a choice the engine can't honor. A test walks every view file and fails if a new surface offers a narrower width without consulting a policy.

MTP depth reads absent. Splash's draft acceptance is real and is reported per request, labelled draft: dflash2 so it is never confused with MTPLX's MTP.

In the app

  • Engine dropdown in the top bar, before the model control, since the engine decides which models are loadable.
  • The model picker lists Splash packages when Splash is selected, in the same rows as the MLX list. Splash keeps its own model setting, so switching engines never overwrites the other side's choice.
  • Settings → Engine shows each package with Download / Update / Cancel and live installer progress, via a new mtplx splash list|install|verify. Downloads and manifest verification stay with Splash's own installer.
  • Chat footer and live chip work on Splash, as described above. The new Settings strings are in all 13 localization tables.

Tests

  • Bridge contract: 63 Python tests, all passing.
  • Swift: 27 Splash tests, all passing. LocalizationTableTests now passes too; the new Settings strings had been missing from the tables.
  • MLX server suites (tests/test_server_openai.py, tests/test_dashboard_endpoints.py): all 430 pass with the page builder moved.
  • Full Swift suite: 949 tests, one failure. testDaemonSupervisorStartsFakeProcessAndKeepsLogsBounded waits 200 ms for a log line and fails 2 of 3 runs in isolation. DaemonSupervisor is untouched by this branch.
  • tests/test_public_cli.py: fails 21 tests both with and without this branch on my machine (missing [server] extras), so no regression there.

How the tests are written:

  • The Swift decoding tests run against real payloads captured from a live Splash daemon, including chat stream frames and settings replies, not hand-written approximations. A missing non-optional field can't silently blank the dashboard.
  • The Python contract tests parse the app's own Swift DTOs to derive the required fields, so the server and the app can't drift apart.
  • The browser chat was checked in a real browser against the live engine: streaming, the thinking block, Off / Low from the chat bar, a settings round trip, and the stats line.

Not covered yet

With an API key set, the MLX server's browser sign-in (/mtplx/browser-auth) is not bridged, so on Splash the browser chat only works on a local server without a key. The API itself honors the key as before.

Requirements

Splash needs an M3 or newer Mac on macOS 26.4+ with at least 36 GB of unified memory (48 GB recommended). Packages are 17.4 GB (27B) and 20.9 GB (35B-A3B), downloaded and verified on first use.

See docs/splash-engine.md.

🤖 Generated with Claude Code

`mtplx serve --engine splash --model incoai/Qwen3.8-27B-Splash` runs the
existing API, dashboard and native app over Splash instead of the MLX
runtime. Splash is specialized per model — Metal kernels compiled for its
exact shapes and a trained DFlash 2 draft — and is faster on the two
packages it supports, at the cost of running only those packages.

Nothing about the MLX path changes: this is additive, and `--engine` still
defaults to `mlx`.

How it fits together

  mtplx serve --engine {mlx|splash}
     |                      |
     v                      v
  MLX in-process      splash serve on a private loopback port

Splash has no runtime model-swap API, so loading a model is process
lifecycle; the bridge supervises one engine process and proxies to it.
Splash 1.0's CLI also hardcodes port 8000 behind an exclusive lock, so the
bridge starts its inner server on a private port and keeps the public one.

The dashboard drives the MLX server's own `dashboard_state` module rather
than imitating its output, so the Live tab works unchanged: `progress`
events move the decode gauge, `completed` carries each request's envelope,
`new_max_tps` fires the record toast, and min/max/p95 come from the same
RollingMetrics. A finished request's speeds are the engine's own — Splash's
cumulative counters read either side of the request — not recomputed here.
Only the live gauge mid-stream is timed by the bridge, because Splash
publishes counters rather than a live rate.

What Splash cannot do is reported rather than faked. Its KV cache is fixed
at 8-bit because its Metal kernels read int8 KV directly, so the bridge
publishes a kv_quant_policy of {supported: false, modes: [q8]} with a
reason, and both KV controls in the app — the Settings card and the
inference-params panel — lock to q8 and say why. MTP depth reads absent;
Splash's DFlash 2 draft acceptance is real and is reported, labelled as
DFlash 2 rather than borrowed into the MTP fields.

`--engine splash` imports no MLX, so a Mac that only runs Splash does not
need it installed.

Also adds `mtplx splash list|install|verify` as the seam the app's Settings
uses to show, download and update packages; downloads and manifest
verification stay with Splash's own installer.

Tests: 44 Python and 23 Swift. The Swift decoding tests run against real
payloads captured from a live Splash daemon, and the contract tests parse
the app's own Swift DTOs so the two cannot drift.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@ksg98
ksg98 requested a review from youssofal as a code owner September 20, 2026 17:34
ksg98 and others added 3 commits September 22, 2026 12:07
The MLX engine decorates its chat stream with mtplx_progress frames and puts
usage and mtplx_stats on the finish frame; the app's chat builds its live
tok/s chip and per-reply footer from them. The Splash bridge passed Splash's
stream through bare, so every Splash reply rendered with no stats.

- Bridge: emit mtplx_progress every ~200 ms, hold the finish frame until the
  stream ends and stamp it with usage and mtplx_stats from the engine's own
  counters. A usage-only frame reaches the client only when it asked.
- Count reasoning_content as decoded tokens (live rate read 0 while thinking).
- Cancel by Splash's chatcmpl id reaches the bridge request.
- Chat footer shows cached prompt tokens, from
  usage.prompt_tokens_details.cached_tokens, which both engines send.
- Add the engine/Splash Settings strings to all 13 localization tables.
- Fixtures and tests from a live Splash 1.0.2 stream.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On the Splash engine, http://127.0.0.1:8000/ served Inco's own chat page, so
switching engines swapped the whole browser UI. The MLX server's page now
lives in mtplx/server/chat_page.py (stdlib only; openai.py imports it, its
output unchanged apart from no-op engine hooks) and the bridge serves the
same page with Splash's slots: DFlash 2 and the package's draft length,
fixed; presence penalty fixed at 0; top_p and top_k sliders over exactly the
range Splash accepts.

The page syncs through /v1/mtplx/settings, where the bridge dropped every
write. It now keeps live sampling defaults like the MLX server: shared by the
page, the app's parameter panel and the app's chat, filled into requests that
leave a field out (the app's chat sends none, so its panel never reached
Splash), and clamped to Splash's range with the reply saying what moved.
MTPLX's top_k 0 (no filter) becomes 32, not a greedy 1. "Hide thinking" is
sent as reasoning_effort "none", which Splash reads; it ignores
enable_thinking. Replies report thinking tokens for the stats line.

Verified in a browser against the live engine: streaming, thinking block,
hide thinking, settings round trip, and the stats line.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Splash's own page had one, and switching to the MTPLX page lost it: thinking
was only in the sidebar. The chat bar now carries a Thinking selector bound
to the same reasoning setting, built from the loaded model's reasoning
policy: Auto, the model's effort levels, Off (Auto, On, Off without levels;
hidden without thinking). A chosen level is synced like every other setting
and sent with each turn.

The Splash bridge now declares that policy the MTPLX way, in model_controls
and the settings reply, with the levels read from the package's chat
template (Qwen 3.8: xhigh, medium, low; default xhigh). It stores a chosen
effort, applies it to turns that name none, folds high/max/minimal like
Splash does, and drops MTPLX's "auto", which Splash would refuse. The app's
parameter panel gets its effort picker on Splash from the same policy.

Verified in a browser: Off gives a direct answer and moves the sidebar to
"Hide thinking"; Low thinks and moves it to "Always show thinking"; the MLX
page sends reasoning_effort "medium" when Medium is picked.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@youssofal

Copy link
Copy Markdown
Owner

Thanks for building this and for the numbers. A second engine behind the API is a large change, and it can't go into 2.12.0. I'll look at it after the release.

@Lancelotbronner

Copy link
Copy Markdown

Alternatively, could some of the changes done by Splash be integrated into the MTPLX inference engine? They are open-source.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants