DeepSeek-V4-Flash-0731 — ROCmFPX mixed-precision (gu106 repair)

Mixed-precision GGUF of DeepSeek-V4-Flash-0731 for AMD Strix Halo (gfx1151) at the exact 102.32 GB deployment budget (128 GB LPDDR5X − 11.3 GB DSpark draft − headroom).

ds4-0731-gu106-v1.gguf (102,321,006,592 B) + ds4-0731-gu106-v1.gguf.gumix.bin — BOTH files are required. The .gumix.bin sidecar carries the gate/up codebooks; the loader registers 86 qtype-106 tensors from it (down codebooks are embedded in-GGUF).

Allocation

family format bpw
routed gate/up experts (57.7 GB) Q2_1_ROCMFP2_MIX (106) — learned per-expert codebooks, importance-weighted, gain-corrected 2.50
routed down experts Q3_1_ROCMFP3_MIX (105) — adaptive, ds4chat-calibrated 3.50
attention / dense / shared Q4_0_ROCMFP4_FAST (101) 4.25
scaffolding etc. 101 / F32 passthrough

Every family is allocated from measured layer-output damage under a whole-file byte ceiling (geo-quant auto). An earlier revision of this repo shipped gate/up as uniform absmax (107) — a family-level GGUF bisection identified that as the cause of a 9/17→17/17 quality gap, and this artifact is the repair: same bytes, +8 questions. Q8_0 attention was also built and measured (+2.4 GB): lossless on perplexity, worthless at task level (78 vs 79/92) — this smaller artifact wins.

Evaluation (dflash serving, fused decode off, serial HC; all items held out)

suite score
COMPSEC 17 (think budget 15488) 17/17
COMPSEC 17 (forced-close, think 2000) 16/17 answered-17
full ds4-eval 92 (COMPSEC+AIME2025+GPQA-D+SuperGPQA, think 15488) 79/92 (88/92 answered)

Calibration (c4:train importance for gate/up, ds4chat-v2 for down) excludes all 92 eval items. For comparison, the published 86.7 GB IQ2_XXS reference scores 82/92 — with 75 of the 92 items present in its own imatrix calibration set; on the genuinely held-out COMPSEC slice both score 17/17.

Serving

Requires a dflash build with qtype-105/106 decode (2026-07-28+). Set DFLASH_DS4_HC_SERIAL=1 (large-footprint stability). Keep the sidecar next to the GGUF. --ds4-fused-decode only on gfx1151 single-device; never for evals on CUDA builds.

Downloads last month
782
GGUF
Model size
284B params
Architecture
deepseek4
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Geometric-AI/DeepSeek-V4-Flash-0731-ROCmFP3-MIX

Quantized
(104)
this model