Releases: ROCm/FastFlowLM
Release list
🚀 FastFlowLM v1.0.7 — Qwen3.8-27B with Speculative Decoding
This release adds qwen3.8-mtp:27b — the largest dense model FastFlowLM has shipped to date, and the first to run speculative decoding on the NPU through its built-in MTP draft head. It also renames hy-mt2:1.8b to hy-mt2-flash:1.8b to match the flash naming convention, and picks up four community fixes to the OpenAI usage contract, CLI exit codes, flm validate, and a stray file in the release packages.
📦 New Model Support
🧠 Qwen3.8-27B
FastFlowLM now supports qwen3.8-mtp:27b, a 27B reasoning model with tool calling — the largest dense model FastFlowLM has shipped, and the first to run speculative decoding on the NPU through its built-in MTP draft head.
- Tag:
qwen3.8-mtp:27b
Run in CLI mode:
flm run qwen3.8-mtp:27bRun in server mode:
flm serve qwen3.8-mtp:27b💾 This is by far the largest dense model FastFlowLM has shipped. Check that you have the disk space for the pull and enough system memory to hold it resident before running it.
⚡ What MTP Speculative Decoding Means
Qwen3.8 ships with a multi-token prediction (MTP) head — a small draft model trained alongside the main network. FastFlowLM now puts it to work:
- The MTP head drafts several candidate tokens in one shot.
- The full model verifies them in a single pass.
- Accepted drafts are kept; the first rejected one is replaced by the base model's own choice, and drafting restarts from there.
The practical consequences:
- Faster decoding, same output. Verification is an exact compare against what the base model would have emitted, so every token you receive is a token the base model would have produced on its own. Speculation changes throughput, not quality.
- Gains are prompt-dependent. Predictable, low-entropy stretches — code, structured output, long reasoning chains — see the highest draft acceptance. Highly creative or surprising text accepts fewer drafts and converges toward ordinary decode speed.
- Nothing to configure. Speculation is driven by the engine; there is no flag to set and no API change.
Every other model is unaffected — speculation is opt-in at the engine level, so the rest of the lineup decodes exactly as it did in v1.0.6.
For model details, context limits, and measured speedups, see the model card and benchmark results.
🔄 hy-mt2:1.8b → hy-mt2-flash:1.8b
The multilingual translation model introduced in v1.0.5 became single-turn in v1.0.6. Its name now says so: it is hy-mt2-flash:1.8b.
This is a rename only — same weights, same 1k context, same single-turn behavior, same translation quality. It simply brings the tag in line with gemma4e-flash and qwen3vl-flash, so that "flash" consistently signals the same contract: single-turn, fixed context, optimized kernels.
Action required: update any script, config, or client that pins the old tag.
- flm run hy-mt2:1.8b
+ flm run hy-mt2-flash:1.8b- "model": "hy-mt2:1.8b"
+ "model": "hy-mt2-flash:1.8b"The recommended prompt shape from the v1.0.5 guide is unchanged:
{"role": "system", "content": "将以下文本翻译为英语,注意只需要输出翻译后的结果,不要额外解释。输出必须全部使用英语,不要输出源语言或原文"},
{"role": "user", "content": "{TEXT}"}For more details, see the model card and benchmark results.
🐛 Fixes from the Community
Four fixes in this release came from outside contributors. Thank you — these are exactly the kind of sharp, well-scoped reports that make the project better. 🙏
📊 OpenAI usage Now Reports the Full Prompt Length
PR #729 — thanks to @Javinator9889 (Javier Alonso)
The prompt cache strips the matching prefix before prefill, so usage.prompt_tokens was reporting only the newly evaluated suffix rather than the whole input. The first turn of a conversation looked correct; every turn after that reused the history as a cached prefix and reported just the new message.
That broke agentic front-ends badly. Tools like OpenCode size their context from usage.prompt_tokens, so the context appeared to reset on every request — built-in compaction never triggered, and the client eventually hit a context overflow, re-fed the conversation, and overflowed again in a loop.
What changed:
usage.prompt_tokensnow counts the entire input, cached prefix included — matching the OpenAI specification.usage.prompt_tokens_details.cached_tokensis now exposed on the OpenAI endpoints, so clients can see how much of that input was served from cache.prefill_speed_tpsnow divides by the tokens actually evaluated, so a cache hit no longer inflates the reported prefill speed.- Ollama-compatible
prompt_eval_countis unchanged — it continues to report evaluated tokens only, matching upstream Ollama.
If you drive FLM from an agent framework that manages its own context window, this is the fix to upgrade for.
↩️ flm help, flm version, and flm port Now Exit 0
PR #723 — thanks to @jtuyls (Jorn Tuyls)
These three commands succeeded and then exited with status 1, because they stopped argument parsing the same way a usage error does. Any script, CI job, or shell with set -e that ran flm version treated a perfectly good invocation as a failure.
They are now handled as successful commands — and handled before the model list is loaded, so flm help and flm version no longer pay that startup cost. Genuine usage errors still exit 1.
✅ flm validate No Longer Hard-Checks the Kernel Version
PR #738 — thanks to @superm1 (Mario Limonciello)
flm validate required kernel 6.17 or newer. Several distributions backport the NPU driver to considerably older kernels, so working setups were being reported as invalid.
The check is now removed. Validation already checks the firmware version, which implicitly requires a recent enough driver — making the kernel version test both redundant and wrong for backported kernels.
🧹 Stray Backup File Removed from the Release Packages
src/lib/xrt/libq4_npu_eXpress.so.bak-20260826 — a backup copy of a shared library — was committed by accident and had been shipping inside the release packages, despite being used by neither the build nor the executable. It has been deleted, so the Linux artifacts are a little smaller.
🙏 Acknowledgements
| Contributor | Contribution |
|---|---|
| @Javinator9889 | #729 — full prompt length in OpenAI usage |
| @jtuyls | #723 — exit 0 for help, version, and port |
| @superm1 | #738 — drop the kernel check in flm validate |
| @yorickvP | #747 — remove the stray .so.bak file |
🌟 Summary
| Highlight | |
|---|---|
| 📦 | New model: Qwen3.8-27B (qwen3.8-mtp:27b) — reasoning and tool calling, the largest dense model FastFlowLM has shipped |
| ⚡ | First model with MTP speculative decoding: a draft head proposes, the full model verifies — faster decode, identical output |
| 🔄 | hy-mt2:1.8b renamed to hy-mt2-flash:1.8b — rename only; update pinned tags |
| 📊 | usage.prompt_tokens reports the full input, cached_tokens is now exposed, and prefill_speed_tps is no longer inflated by cache hits |
| ↩️ | flm help, flm version, and flm port exit 0 instead of 1 |
| ✅ | flm validate no longer rejects backported kernels older than 6.17 |
| 🧹 | Stray .so.bak backup file removed from the release packages |
Thanks for your support — see you in the next one! 🚀
🚀 FastFlowLM v1.0.6 — Flash Models, tool_choice, and Tool-Call Fixes
This release introduces three flash models — kernel-optimized runtimes over the existing weights, so no re-download is needed — makes hy-mt2:1.8b single-turn, adds tool_choice support in server mode, fixes a tool-call finish_reason bug, and changes where the MSI installer writes FLM_MODEL_PATH.
⚡ New Flash Models
FastFlowLM now supports three flash models. These are not new checkpoints — they run the same weights as their standard counterparts through optimized kernels, so there is no weights update and nothing new to download:
| Tag | Prefill @ 128 ctx | Model card |
|---|---|---|
gemma4e-flash:e2b |
~390 tokens/s | Gemma 4 E2B-IT · Flash |
gemma4e-flash:e4b |
~256 tokens/s | Gemma 4 E4B-IT · Flash |
qwen3vl-flash:4b |
~350 tokens/s | Qwen3-VL 4B-Instruct · Flash |
Run in CLI mode:
flm run gemma4e-flash:e2bRun in server mode:
flm serve qwen3vl-flash:4b📖 What "Flash" Means
Flash models trade multi-turn flexibility for speed: the same weights run through optimized kernels, under a set of constraints that make those kernels possible. Please note the following behavior before deploying them:
- Same weights, no re-download. Flash models reuse the weights you already have — no
flm pullrequired. - Single-turn only. Any request containing an assistant message is rejected. Send a system prompt (optional) and a single user message.
- System KV cache supported. The system prompt is prefilled once and reused across requests, so repeated calls sharing a system prompt skip that prefill cost.
- 1k maximum context length. Requests over 1k tokens are rejected with a warning — they are not silently truncated. Size your prompts accordingly.
🖼️ Image & Audio Handling
Flash models apply a fixed media budget. No per-request tuning is required — oversized input is reduced automatically rather than rejected:
| Model | Images | Audio |
|---|---|---|
qwen3vl-flash:4b |
Resized so the longer side is 256 pixels | Not supported |
gemma4e-flash:e2b / :e4b |
Resized to 70 tokens | Truncated to the first 30 seconds |
⚠️ Note the difference from the context limit: oversized media is silently reduced (images downscaled, audio cut at 30 s), while an over-1k prompt is rejected outright. Audio longer than 30 seconds will not raise an error — the model simply never sees the remainder, so split long clips yourself if you need full coverage.
🐍 Example: Python + OpenAI SDK (Streaming)
Flash models speak the standard OpenAI chat-completions API, so the official openai Python client works as-is — just point it at your local FLM server.
pip install openai
flm serve gemma4e-flash:e2bVision, streaming, with a pinned system prompt:
Below, the same system prompt is pinned across requests that each caption a different image. The system-prompt KV is prefilled once on the first request; every request after that reuses it, so only the new image and text need to be prefilled — expect noticeably lower latency from the second request onward.
import time
import base64
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:52625/v1",
api_key="dummykey", # FLM runs locally; the key is not checked
)
# Keep everything fixed in the system prompt — it is prefilled once and
# reused across requests. Vary only the user message and image.
SYSTEM_PROMPT = "You are a concise assistant. Describe the image in one short sentence."
image_paths = [
r"C:\Users\info\OneDrive\Desktop\FLM\image_test\image0.jpg",
r"C:\Users\info\OneDrive\Desktop\FLM\image_test\image1.png",
]
for path in image_paths:
with open(path, "rb") as image_file:
image = base64.b64encode(image_file.read()).decode("utf-8")
print(f"\n> {path}")
start = time.time()
stream = client.chat.completions.create(
model="qwen3vl-flash:4b",
messages=[
{"role": "system", "content": SYSTEM_PROMPT}, # same every time — its KV is cached after the first call
{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image."},
{"type": "image_url", "image_url": {"url": f"data:image/jpg;base64,{image}"}},
],
},
],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
print(f"\n[time to first response: {time.time() - start:.2f}s]")
# 1st image: pays the full system-prompt prefill cost.
# 2nd image onward: system-prompt KV is reused from cache — expect lower latency.
# cleanup
del stream, client
import gc
gc.collect()Vision, streaming:
Serve the vision flash model with flm serve qwen3vl-flash:4b, then pass an image as a base64 data URL. Resizing is automatic, so there is no need to downscale beforehand:
import base64
from openai import OpenAI
image_path = r"C:\path\to\image.png" # <-- edit this
client = OpenAI(base_url="http://127.0.0.1:52625/v1", api_key="dummykey")
with open(image_path, "rb") as image_file:
image = base64.b64encode(image_file.read()).decode("utf-8")
stream = client.chat.completions.create(
model="qwen3vl-flash:4b",
messages=[
{"role": "system", "content": "Describe images in one sentence."},
{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{
"type": "image_url",
"image_url": {"url": f"data:image/png;base64,{image}"},
},
],
},
],
stream=True,
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
print()Note: the image counts against the same 1k context budget as your text. Keep prompts short when sending an image.
Audio, streaming:
Audio goes to the same endpoint with the same message structure — only the content part changes. Use a gemma4e-flash model, since qwen3vl-flash:4b does not accept audio:
import base64
from openai import OpenAI
audio_path = r"C:\path\to\audio.wav" # <-- edit this
client = OpenAI(base_url="http://127.0.0.1:52625/v1", api_key="dummykey")
with open(audio_path, "rb") as audio_file:
audio = base64.b64encode(audio_file.read()).decode("utf-8")
stream = client.chat.completions.create(
model="gemma4e-flash:e2b",
messages=[
{"role": "system", "content": "Transcribe and summarize audio briefly."},
{
"role": "user",
"content": [
{"type": "text", "text": "What is said in this clip?"},
{
"type": "input_audio",
"input_audio": {"data": audio},
},
],
},
],
stream=True,
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
print()
⚠️ Only the first 30 seconds are processed. Longer clips are truncated without an error, so split them client-side if you need full coverage.
Audio + image in one request:
gemma4e-flash accepts both in the same content array, so a single call can reason over audio and an image together:
import base64
from openai import OpenAI
audio_path = r"C:\path\to\audio.wav" # <-- edit this
image_path = r"C:\path\to\image.png" # <-- edit this
client = OpenAI(base_url="http://127.0.0.1:52625/v1", api_key="dummykey")
with open(audio_path, "rb") as audio_file:
audio = base64.b64encode(audio_file.read()).decode("utf-8")
with open(image_path, "rb") as image_file:
image = base64.b64encode(image_file.read()).decode("utf-8")
stream = client.chat.completions.create(
model="gemma4e-flash:e2b",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Finish two tasks: 1. Summarize the audio. 2. Describe the image."},
{
"type": "input_audio",
"input_audio": {"data": audio},
},
{
"type": "image_url",
"image_url": {"url": f"data:image/png;base64,{image}"},
},
],
}
],
stream=True,
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
print()🔄 hy-mt2:1.8b Is Now Single-Turn
hy-mt2:1.8b, the multilingual translation model added in v1.0.5, is now a single-turn model as well. Like the flash models, it rejects any request containing an assistant message.
This matches how a dedicated translation model is actually used: every translation is independent, and prior turns add prefill cost without improving the result.
What this means for you: if you were replaying conversation history back to hy-mt2:1.8b, remove it — send only the instruction and the source text. The recommended shape from the v1.0.5 prompt guide is unchanged and remains the best option:
{"role": "system", "content": "将以下文本翻译为英语,注意只需要输出翻译后的结果,不要额外解释。输出必须全部使用英语,不要输出源语言或原文"},
{"role": "user", "content": "{TEXT}"}
``...🚀 FastFlowLM v1.0.5 — Hy-MT2-1.8B
This release adds support for a new model, Hy-MT2-1.8B — a small, quick, multilingual translation model — and fixes a flm bench bug causing degraded performance when benchmarking from long context down to short context.
📦 New Model Support
🌐 Hy-MT2-1.8B
FastFlowLM now supports hy-mt2:1.8b, a small and quick multilingual translation model.
- Tag:
hy-mt2:1.8b
Run in CLI mode:
flm run hy-mt2:1.8bRun in server mode:
flm serve hy-mt2:1.8b📖 Prompt Guide
Prompt Format
Hy-MT2 is a dedicated translation model, not a general-purpose chat model — it has no default system prompt. There are two ways to prompt it:
Option 1: instruction + text in a single user message
将以下文本翻译为{TARGET_LANG},注意只需要输出翻译后的结果,不要额外解释:
{TEXT}
or, in English:
Translate the following segment into {TARGET_LANG}, without additional explanation.
{TEXT}
Option 2: instruction as the system prompt, text as the user message
Pin the translation instruction as the system prompt so it isn't re-prefilled every turn, then send only the source text as the user message:
{"role": "system", "content": "将以下文本翻译为英语,注意只需要输出翻译后的结果,不要额外解释。输出必须全部使用英语,不要输出源语言或原文"},
{"role": "user", "content": "{TEXT}"}This is the recommended format in server mode for multi-turn or repeated translation calls (e.g. batch translating subtitle lines), since the instruction's prefill cost is paid once instead of once per request.
For more details, see the model card and benchmark results.
🐛 Bug Fix: flm bench Performance Degradation Across Context Lengths
Fixed a bug in flm bench where running benchmarks from long context down to short context resulted in progressively poorer performance numbers.
🌟 Summary
| Highlight | |
|---|---|
| 📦 | New model: Hy-MT2-1.8B (hy-mt2:1.8b) — small, quick, multilingual translation model |
| 🐛 | Fixed flm bench performance degradation when benchmarking from long context to short context |
Thanks for your support — see you in the next one! 🚀
🚀 FastFlowLM v1.0.4 — Gemma4-12B-IT
This release adds a new model and fixes an image-max-tokens bug affecting the Gemma4 family.
📦 New Model Support
🌎 Gemma4-12B-IT
FastFlowLM now supports gemma4-it:12b for language, vision, audio workloads, including concurrent multimodal input for omni-model use cases.
- Tag:
gemma4-it:12b
Run in CLI mode:
flm run gemma4-it:12bRun in server mode:
flm serve gemma4-it:12bFor more details, see the model card and benchmark results.
🐛 Bug Fix: Gemma4 Family image-max-tokens
Fixed a bug where the image-max-tokens custom parameter was not workable in previous versions for the Gemma4 family. Image token limits are now correctly applied.
See Add FLM Custom Parameters for how to set it via Open WebUI.
🌟 Summary
| Highlight | |
|---|---|
| 📦 | New model: Gemma4-12B-IT (gemma4-it:12b) — language, vision, audio, and concurrent multimodal support |
| 🐛 | Fixed image-max-tokens not working for the Gemma4 family in previous versions |
Thanks for your support — see you in the next one! 🚀
🚀 FastFlowLM v1.0.3 — Higher-Accuracy Qwen3.5 & Qwen3.6-MoE Weights
This release upgrades the quantization for the entire Qwen3.5 family and Qwen3.6-MoE from Q4_1 to Q4_K, improving accuracy — plus a heads-up on a required weight update.
📥 Weights Update Required
Models in this release are re-quantized by FLM itself. If you're upgrading to v1.0.3, you'll need to re-download weights for the Qwen3.5 family and Qwen3.6-MoE — existing local copies from prior versions (Q4_1) are not compatible.
flm pull qwen3.5:0.8b
flm pull qwen3.5:2b
flm pull qwen3.5:4b
flm pull qwen3.5:9b
flm pull qwen3.6-moe:35b-a3b🌟 Summary
| Highlight | |
|---|---|
| 📥 | Weights update required for Qwen3.5 family and Qwen3.6-MoE — re-pull models before running v1.0.3 |
Thanks for your support — see you in the next one! 🚀
🚀 FastFlowLM v1.0.2 — Faster Qwen3.5 & Qwen3.6-MoE
This release brings a solid decoding and prefill speed boost across the entire Qwen3.5 family and Qwen3.6-MoE, plus a heads-up on a required weight update.
📥 Weights Update Required
Models in this release are quantized by FLM itself. If you're upgrading to v1.0.2, you'll need to re-download weights for the Qwen3.5 family and Qwen3.6-MoE — existing local copies from prior versions are not compatible.
flm pull qwen3.5:0.8b
flm pull qwen3.5:2b
flm pull qwen3.5:4b
flm pull qwen3.5:9b
flm pull qwen3.6-moe:35b-a3b⚡ Performance Boost: Qwen3.5 Family & Qwen3.6-MoE
Both prefill and decoding throughput have been improved across all context lengths (1k–32k) for Qwen3.5 (0.8B, 2B, 4B, 9B) and Qwen3.6-MoE (35B-A3B).
| Model | Decoding Gain (avg / peak) | Prefill Gain (avg / peak) |
|---|---|---|
| Qwen3.5 0.8B | +9.5% / +14.5% | +34.8% / +44.3% |
| Qwen3.5 2B | +10.2% / +11.8% | +27.2% / +33.3% |
| Qwen3.5 4B | +11.5% / +13.1% | +31.5% / +37.4% |
| Qwen3.5 9B | +12.0% / +13.0% | +25.5% / +29.2% |
| Qwen3.6-MoE 35B-A3B | +11.4% / +13.0% | +15.9% / +19.7% |
Gains are largest on smaller models and at mid-to-long context lengths, with prefill benefiting more than decoding across the board.
🌟 Summary
| Highlight | |
|---|---|
| 📥 | Weights update required for Qwen3.5 family and Qwen3.6-MoE — re-pull models before running v1.0.2 |
| ⚡ | Qwen3.5 family (0.8B/2B/4B/9B): up to +14.5% decoding and +44.3% prefill throughput |
| ⚡ | Qwen3.6-MoE 35B-A3B: up to +13.0% decoding and +19.7% prefill throughput |
Thanks for your support — see you in the next one! 🚀
🚀 FastFlowLM v1.0.1 — Windows Installer Switch & SmolVLA Benchmark
📦 Windows Installer: flm-setup.exe → flm-setup.msi
The Windows installer has been switched from flm-setup.exe to flm-setup.msi. This change is reflected across:
- Release artifacts
- Download links
- Documentation
- Website
🚀 SmolVLA Benchmark
We've provided the benchmark numbers for SmolVLA. Full results: https://fastflowlm.com/docs/benchmarks/smolvla_results/
Test System: AMD Ryzen™ AI 9 370 (Strix Point) with 32 GB DRAM; performance is comparable to other Strix Point and Strix Halo Point systems.
Inference Latency (ms per inference, with different camera input counts)
| Model | HW | 1 image | 2 images | 3 images |
|---|---|---|---|---|
| SmolVLA | NPU (FLM) | 298 | 363 | 430 |
Thanks for your support — see you in the next one! 🚀
🚀 FastFlowLM v1.0.0 — First Release Under ROCm
This is a big one — v1.0.0 marks our first general release under the ROCm organization. Here's what's new 🎉
🏠 General Release v1.0.0 in the ROCm Org
FastFlowLM is now officially maintained under ROCm/FastFlowLM. Everything from the old repo — issues, pull requests, and history — has been transferred over, so nothing is lost in the move. All future releases, issues, and contributions happen there — please make sure your bookmarks, forks, and remotes point to the new home.
🤖 New Model: SmolVLA
FLM now supports SmolVLA, a vision-language-action (VLA) model, adding robotics support to the lineup alongside our existing LLM and VLM models.
Model card and usage details:
- Hugging Face: https://huggingface.co/FastFlowLM/smolvla
- ModelScope: https://modelscope.cn/models/amd/smolvla
⚙️ Fine-Grained Control for flm bench
To keep benchmarking fast by default, flm bench now runs 2 iterations at each context length from 1k to 32k.
Need more data points? Override the iteration count with a flag:
flm bench --bench-iterations 4🐛 Bug Fix: Qwen3-VL Two-Image Handling
Fixed an issue in qwen3vl-it where passing two images in a single request could cause the model to fail to recognize either image. Multi-image prompts now resolve correctly.
🌟 Summary
| Highlight | |
|---|---|
| 🏠 | General release v1.0.0 — first release under the ROCm/FastFlowLM org |
| 🤖 | New model: SmolVLA support |
| ⚙️ | flm bench now defaults to 2 iterations per context length, configurable via --bench-iterations |
| 🐛 | Fixed qwen3vl-it failing to recognize images when given two at once |
Thanks for your support — see you in the next one! 🚀
🚀 FastFlowLM v0.9.46 — We're Moving!
🏠 FastFlowLM is Now an Official AMD Project
🎉 FastFlowLM has joined the ROCm organization and is now officially maintained by AMD. This is the last release under FastFlowLM/FastFlowLM — starting from v1.0.0, everything moves to ROCm/FastFlowLM. Please update your bookmarks, forks, and remotes accordingly. See you there!
🌐 ModelScope Support
Models are pulled from HuggingFace by default. You can now opt into ModelScope as an alternative source with a single flag.
Pull a model from ModelScope:
flm pull llama3.2:1b --modelscope 1Auto-pull from ModelScope in CLI mode:
flm run llama3.2:1b --modelscope 1Auto-pull from ModelScope in server mode:
flm serve --modelscope 1Check the compatibility of your local model with ModelScope:
flm check llama3.2:1b --modelscope 1🖼️ More Image Resize Levels for Qwen-VL Models
Fine-grained control over input image resolution is now available via the -r flag:
flm serve -r <img-pre-resize-level>| Level | Resolution |
|---|---|
| 0 | Original size |
| 1 | Height = 480 |
| 2 | Height = 720 |
| 3 | Height = 1080 |
| 4 | Height = 1440 |
| 5 | Height = 2160 |
| 6 | Height = 2880 |
| 7 | Height = 3240 |
| 8 | Height = 4320 |
⚡ Speed Improvements for Qwen3.6-MoE
Both prefill and decoding throughput have been improved for Qwen3.6-MoE across all context lengths.
Decoding throughput (tokens/s):
| Context | Old | New | Gain |
|---|---|---|---|
| 1k | 12.41 | 13.65 | +9.99% |
| 2k | 12.26 | 13.41 | +9.38% |
| 4k | 11.96 | 13.09 | +9.45% |
| 8k | 11.38 | 12.51 | +9.93% |
| 16k | 10.40 | 11.24 | +8.08% |
| 32k | 8.88 | 9.51 | +7.09% |
Prefill throughput (tokens/s):
| Context | Old | New | Gain |
|---|---|---|---|
| 1k | 75.18 | 78.98 | +5.05% |
| 2k | 109.85 | 118.04 | +7.46% |
| 4k | 150.93 | 156.43 | +3.64% |
| 8k | 181.56 | 197.93 | +9.02% |
| 16k | 214.46 | 218.84 | +2.04% |
| 32k | 219.72 | 221.96 | +1.02% |
🌟 Summary
| Highlight | |
|---|---|
| 🏠 | FastFlowLM is now an official AMD project — repo moved to ROCm/FastFlowLM |
| 🌐 | ModelScope support: pull, serve, and check model compatibility |
| 🖼️ | 9-level image resize control for Qwen-VL models |
| ⚡ | Up to ~+10% decoding and ~+9% prefill speedup for Qwen3.6-MoE |
🚀 FastFlowLM v0.9.45 — Qwen3.6-35B-A3B & Smoother KV Cache Through Multi-Backend Support
Here's what's new 🎉
🤖 New Model: Qwen3.6-35B-A3B
Say hello to Qwen3.6-35B-A3B — the second MoE model in FLM, joining GPT-OSS. It packs 35B total parameters with only 3B activated per forward pass, so you get strong reasoning quality at a fraction of the compute cost.
Tag: qwen3.6-moe:35b-a3b
Run in CLI mode:
flm run qwen3.6-moe:35b-a3b
Run in server mode:
flm serve qwen3.6-moe:35b-a3b
Check out the model card and benchmark results for more details.
⚡ Smoother KV Cache Through Multi-Backend Support
FLM now supports per-round KV cache checks, making context management more precise and robust when mixing inference backends.
Here's a real example — imagine a Lemonade user running gemma4-it:e2b on both NPU (via FLM) and GPU (via llama.cpp):
- They send an initial prompt to the NPU and get a response.
- They continue on the GPU and get a second response.
- They switch back to the NPU with the full conversation history.
Previously, the NPU would see gaps from the GPU round, fail the KV cache check, and re-prefill everything from scratch. Now, with per-round KV cache checks, it knows the first round is already cached — so it only prefills the GPU round and the new prompt. Mixing backends is no longer a headache! 🙌
🌟 Summary
- New model:
Qwen3.6-35B-A3B— our second MoE model, with 35B total parameters and just 3B activated per token 🧠 - Smarter KV cache — per-round checks let you mix FLM and other backends without losing cache or re-prefilling the whole context 🔄
Thanks for your support — more good stuff is on the way. See you in the next one! 🚀