Conversation
`mtplx serve --engine splash --model incoai/Qwen3.8-27B-Splash` runs the
existing API, dashboard and native app over Splash instead of the MLX
runtime. Splash is specialized per model — Metal kernels compiled for its
exact shapes and a trained DFlash 2 draft — and is faster on the two
packages it supports, at the cost of running only those packages.
Nothing about the MLX path changes: this is additive, and `--engine` still
defaults to `mlx`.
How it fits together
mtplx serve --engine {mlx|splash}
| |
v v
MLX in-process splash serve on a private loopback port
Splash has no runtime model-swap API, so loading a model is process
lifecycle; the bridge supervises one engine process and proxies to it.
Splash 1.0's CLI also hardcodes port 8000 behind an exclusive lock, so the
bridge starts its inner server on a private port and keeps the public one.
The dashboard drives the MLX server's own `dashboard_state` module rather
than imitating its output, so the Live tab works unchanged: `progress`
events move the decode gauge, `completed` carries each request's envelope,
`new_max_tps` fires the record toast, and min/max/p95 come from the same
RollingMetrics. A finished request's speeds are the engine's own — Splash's
cumulative counters read either side of the request — not recomputed here.
Only the live gauge mid-stream is timed by the bridge, because Splash
publishes counters rather than a live rate.
What Splash cannot do is reported rather than faked. Its KV cache is fixed
at 8-bit because its Metal kernels read int8 KV directly, so the bridge
publishes a kv_quant_policy of {supported: false, modes: [q8]} with a
reason, and both KV controls in the app — the Settings card and the
inference-params panel — lock to q8 and say why. MTP depth reads absent;
Splash's DFlash 2 draft acceptance is real and is reported, labelled as
DFlash 2 rather than borrowed into the MTP fields.
`--engine splash` imports no MLX, so a Mac that only runs Splash does not
need it installed.
Also adds `mtplx splash list|install|verify` as the seam the app's Settings
uses to show, download and update packages; downloads and manifest
verification stay with Splash's own installer.
Tests: 44 Python and 23 Swift. The Swift decoding tests run against real
payloads captured from a live Splash daemon, and the contract tests parse
the app's own Swift DTOs so the two cannot drift.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The MLX engine decorates its chat stream with mtplx_progress frames and puts usage and mtplx_stats on the finish frame; the app's chat builds its live tok/s chip and per-reply footer from them. The Splash bridge passed Splash's stream through bare, so every Splash reply rendered with no stats. - Bridge: emit mtplx_progress every ~200 ms, hold the finish frame until the stream ends and stamp it with usage and mtplx_stats from the engine's own counters. A usage-only frame reaches the client only when it asked. - Count reasoning_content as decoded tokens (live rate read 0 while thinking). - Cancel by Splash's chatcmpl id reaches the bridge request. - Chat footer shows cached prompt tokens, from usage.prompt_tokens_details.cached_tokens, which both engines send. - Add the engine/Splash Settings strings to all 13 localization tables. - Fixtures and tests from a live Splash 1.0.2 stream. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On the Splash engine, http://127.0.0.1:8000/ served Inco's own chat page, so switching engines swapped the whole browser UI. The MLX server's page now lives in mtplx/server/chat_page.py (stdlib only; openai.py imports it, its output unchanged apart from no-op engine hooks) and the bridge serves the same page with Splash's slots: DFlash 2 and the package's draft length, fixed; presence penalty fixed at 0; top_p and top_k sliders over exactly the range Splash accepts. The page syncs through /v1/mtplx/settings, where the bridge dropped every write. It now keeps live sampling defaults like the MLX server: shared by the page, the app's parameter panel and the app's chat, filled into requests that leave a field out (the app's chat sends none, so its panel never reached Splash), and clamped to Splash's range with the reply saying what moved. MTPLX's top_k 0 (no filter) becomes 32, not a greedy 1. "Hide thinking" is sent as reasoning_effort "none", which Splash reads; it ignores enable_thinking. Replies report thinking tokens for the stats line. Verified in a browser against the live engine: streaming, thinking block, hide thinking, settings round trip, and the stats line. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Splash's own page had one, and switching to the MTPLX page lost it: thinking was only in the sidebar. The chat bar now carries a Thinking selector bound to the same reasoning setting, built from the loaded model's reasoning policy: Auto, the model's effort levels, Off (Auto, On, Off without levels; hidden without thinking). A chosen level is synced like every other setting and sent with each turn. The Splash bridge now declares that policy the MTPLX way, in model_controls and the settings reply, with the levels read from the package's chat template (Qwen 3.8: xhigh, medium, low; default xhigh). It stores a chosen effort, applies it to turns that name none, folds high/max/minimal like Splash does, and drops MTPLX's "auto", which Splash would refuse. The app's parameter panel gets its effort picker on Splash from the same policy. Verified in a browser: Off gives a direct answer and moves the sidebar to "Hide thinking"; Low thinks and moves it to "Always show thinking"; the MLX page sends reasoning_effort "medium" when Medium is picked. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner
|
Thanks for building this and for the numbers. A second engine behind the API is a large change, and it can't go into 2.12.0. I'll look at it after the release. |
|
Alternatively, could some of the changes done by Splash be integrated into the MTPLX inference engine? They are open-source. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds Inco's Splash as a second inference engine behind the existing API:
Same endpoints, same dashboard, same native app, same browser chat. Only the kernels change.
--enginedefaults tomlx, so nothing changes unless you ask for it.Splash is built the opposite way round from a general engine: it supports a small set of models and rebuilds itself around each one, with fused Metal kernels compiled for that model's exact shapes and a DFlash 2 draft trained for it. It's faster on the two packages it supports, and it loads only those packages.
Measured on a 64 GB M3 Max
macOS 26.6.2, Splash 1.0,
incoai/Qwen3.8-27B-Splash, reasoning off, through the bridge over HTTP.Decode tracks draft acceptance closely, and acceptance depends on how predictable the text is. Code runs about twice as fast as prose on the same machine. Over a five-request mix: min 30.7, mean 45.4, p95 62.7 tok/s.
Prefill and prefix reuse, on a 25,812-token prompt:
That's the part worth calling out: replaying a 25K-token context returns the first token in 275 ms instead of 129.7 s, a ~470× improvement, because Splash reuses the prefix it already computed. It makes multi-turn work on a long document feel immediate after the first turn.
Memory: 21.1 GB active and peak, on a 64 GB machine.
Splash 1.0.2. The branch is now tested on Splash 1.0.2 (
brew upgrade incoai/tap/splash). Spot checks on the same machine: two short code prompts decoded at 96 and 106 tok/s, with 62% and 73% of drafts accepted. These are single samples, not a rerun of the benchmark above.For reference, Inco publishes 74 tok/s decode and 363 tok/s prefill at 32K for this model on a 48 GB M5 Pro.
How it fits together
Splash has no runtime model-swap API, so loading a model is process lifecycle: the bridge supervises one engine process and proxies to it. Splash 1.0's CLI also hardcodes port 8000 behind an exclusive lock, so the bridge runs its inner server on a private port and keeps the public one for MTPLX.
--engine splashimports no MLX at all, so a Mac that only ever runs Splash doesn't need it installed.The dashboard
The bridge drives
mtplx.server.dashboard_state(the MLX server's own module, which has no MLX dependency) rather than imitating its output. The Live tab therefore works unchanged:progressevents move the decode gauge,completedcarries each request's envelope,new_max_tpsfires the record toast, and the min/max/mean/p95 row comes from the sameRollingMetrics.A finished request's speeds are the engine's own: Splash's
/statuscounters are cumulative, so the bridge reads them either side of a request and reports the difference. Only the live gauge mid-stream is timed by the bridge, because Splash publishes counters rather than a live rate. Thinking tokens count toward that live rate. Memory ismemory_actual, the process footprint, not the KV pages alone.Chat: the same UI on both engines
Browser chat.
http://127.0.0.1:8000/serves MTPLX's own chat page on both engines. The page builder (_chat_ui_html) moved fromopenai.pyintomtplx/server/chat_page.py, which uses only the standard library, so the bridge renders the exact same page without importing MLX.openai.pynow imports it; its diff is that one import and the moved function. On Splash the page fills a few engine slots differently:Replies carry the same stats line as MLX, for example
DFlash 2 71% accepted · 105.6 tok/s · 43 tokens · 37 thinking · ttft 1.03s.Thinking selector in the chat bar, on both engines. Splash's own page had one next to Send, so switching pages would have lost it. The chat bar now has a Thinking selector bound to the same reasoning setting as the sidebar. Its choices come from the loaded model's reasoning policy:
A chosen level syncs like every other setting and is sent with each turn. This is the one visible change on the MLX page.
Stream stats for the app's chat. The MLX server adds
mtplx_progressframes and putsusageandmtplx_statson the finish frame, and the app's chat builds its live tok/s chip and per-reply footer from them. Splash sends neither, so Splash replies rendered with no stats. The bridge now adds them:The reply footer also gains a cached count from
usage.prompt_tokens_details.cached_tokens, which both engines send. A Stop sent with Splash'schatcmpl-…id cancels the right request.Live settings on Splash
/v1/mtplx/settingswas read-only on the bridge, and writes were dropped. The app's own chat sends no sampling fields, so its parameter panel never reached Splash. The bridge now keeps live defaults the way the MLX server does. They are shared by the browser chat, the app's parameter panel and the app's chat, and filled into any request that leaves a field out.adjusted) or was not applied (ignored), so every panel shows the value actually in use.top_k: 0(MTPLX's "no filter") becomes 32, Splash's widest, not a greedy 1.reasoning_effort: "none", the switch its template reads. It ignoresenable_thinking.model_controlsand the settings reply, with effort levels read from the package's chat template (Qwen 3.8: xhigh, medium, low; default xhigh). The app's parameter panel shows its effort picker on Splash from the same policy.What Splash can't do, and how that's reported
Splash's Metal kernels read 8-bit KV directly, so the width is a property of the compiled kernels, not a setting. There is no flag, env var or engine argument for it. The bridge publishes
kv_quant_policy: {supported: false, modes: ["q8"], disabled_reason: …}. Both KV controls in the app (the Settings card and the top-bar inference-params panel) lock to q8 and explain why, rather than offering a choice the engine can't honor. A test walks every view file and fails if a new surface offers a narrower width without consulting a policy.MTP depth reads absent. Splash's draft acceptance is real and is reported per request, labelled
draft: dflash2so it is never confused with MTPLX's MTP.In the app
mtplx splash list|install|verify. Downloads and manifest verification stay with Splash's own installer.Tests
LocalizationTableTestsnow passes too; the new Settings strings had been missing from the tables.tests/test_server_openai.py,tests/test_dashboard_endpoints.py): all 430 pass with the page builder moved.testDaemonSupervisorStartsFakeProcessAndKeepsLogsBoundedwaits 200 ms for a log line and fails 2 of 3 runs in isolation.DaemonSupervisoris untouched by this branch.tests/test_public_cli.py: fails 21 tests both with and without this branch on my machine (missing[server]extras), so no regression there.How the tests are written:
Not covered yet
With an API key set, the MLX server's browser sign-in (
/mtplx/browser-auth) is not bridged, so on Splash the browser chat only works on a local server without a key. The API itself honors the key as before.Requirements
Splash needs an M3 or newer Mac on macOS 26.4+ with at least 36 GB of unified memory (48 GB recommended). Packages are 17.4 GB (27B) and 20.9 GB (35B-A3B), downloaded and verified on first use.
See
docs/splash-engine.md.🤖 Generated with Claude Code