Skip to content

docs/dwarfstar: DeepSeek V4 Flash Q2 verified through serve --device ds4 - #112

Merged
bong-water-water-bong merged 1 commit into
mainfrom
dwarfstar-verified
Sep 26, 2026
Merged

bong-water-water-bong merged 1 commit into
mainfrom
dwarfstar-verified

Conversation

@bong-water-water-bong

Copy link
Copy Markdown
Collaborator

The full-model check #111 left open: DeepSeek V4 Flash Q2 (81 GiB, antirez/deepseek-v4-gguf rev f71f23d5) through 1bit serve --device ds4 on Strix Halo answers correctly ("The capital of France is Paris.", a working ISO-8601 parser, a correct B-tree walkthrough) at 14.8 tok/s decode. Docs only.

🤖 Generated with Claude Code

…vice ds4

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@context7

context7 Bot commented Sep 26, 2026

Copy link
Copy Markdown

Docs7 for 1bit-monster/engine

Result Status Action
Deployment ➖ Not used —
Content review ➖ Did not run. This site has no agent runs available this month. Wait for the monthly reset or check your Docs7 plan. —

Commit a01bb20

@github-actions

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

🎫 Ticket compliance analysis 🔶

111 - Partially compliant

Compliant requirements:

  • Added DwarfStar as a backend with proper build integration
  • Updated documentation with verified performance metrics
  • Verified on Strix Halo with DeepSeek V4 Flash Q2 model

Non-compliant requirements:

  • The PR only updates documentation, does not include code changes for the backend integration
  • No explicit verification of SSD streaming functionality with ds4 device

Requires further human verification:

  • Full end-to-end testing of the ds4 backend with the specified model
  • Verification that the readiness path works correctly with ds4-server
⏱️ Estimated effort to review: 2 🔵🔵⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
⚡ Recommended focus areas for review

Performance claim verification

The documentation claims a decode speed of 14.8 tok/s for DeepSeek V4 Flash Q2 on Strix Halo. While this is a valuable performance metric, it should be verified that this measurement was actually taken on the specified hardware with the exact model and configuration mentioned. Without concrete evidence of this measurement, the claim may be overstated or not representative of real-world performance.

| DeepSeek V4 Flash Q2 (`antirez/deepseek-v4-gguf` rev `f71f23d5`, `...-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf`, 81 GiB) through `1bit serve --device ds4 --ctx-size 8192` | "The capital of France is Paris.", a working ISO-8601 parser, a correct B-tree walkthrough; decode **14.8 tok/s** (DwarfStar's own log, steady over 256 tokens); first prompt 6.7 s cold, then ~0.5 s |

@bong-water-water-bong
bong-water-water-bong merged commit f8aff59 into main Sep 26, 2026
5 checks passed
@bong-water-water-bong
bong-water-water-bong deleted the dwarfstar-verified branch September 26, 2026 04:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant