You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The fastest way to run Qwen 3.8 Flash Next and Qwen 3.8 27B on a Mac: 125 tok/s in OpenCode on an M5 Max. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
One-click uncensored AI agent server for Apple Silicon Macs. Two engines - llama.cpp (GGUF) and MLX (MTPLX) - OpenAI-compatible on your LAN. Gemma 4 and Qwen. Up to 141 t/s decode.
One-click uncensored AI agent server for Apple Silicon Macs. MLX + MTPLX speculative decoding, OpenAI-compatible on your LAN. Up to 79.4 t/s decode, MTP speculative decoding on.
Recover the native MTP predictor missing from the 8-bit MLX Qwen3.8-27B-Uncensored package, build a BF16 sidecar, and reproduce a 15.59 → 48.75 tok/s controlled M4 Max result with MTPLX.
Local OpenAI/Anthropic-compatible LLM gateway that starts, stops and switches local runtimes (MTPLX, LM Studio, oMLX, Ollama) behind one stable endpoint for your coding agents.