What the two trackers actually compared
Two independent benchmark sites — BenchLM.ai and llm-stats.com — put DeepSeek’s V4 Flash 0731 update head-to-head with Alibaba’s Qwen3.6-27B this week. Both are open-weight (Flash under MIT, Qwen under Apache 2.0), both are pitched at the same buyer: someone with a high-RAM workstation who wants the smartest local model they can squeeze into memory.
The trackers read the question differently. BenchLM focused on a small set of shared benchmarks — only the tests both models have actually been scored on — and refused to declare a winner where the underlying test sets differed. LLM-stats ran a broader sweep and concluded that DeepSeek V4 Flash beats Qwen3.6-27B on most benchmarks and costs roughly a twelfth per token. Both pages were updated on 5 August 2026.
12× — DeepSeek V4 Flash’s per-token cost advantage over Qwen3.6-27B on a 3:1 input/output mix, per llm-stats (5 August 2026).
Where Flash wins, where Qwen holds on
On the five tests both trackers share, Qwen3.6-27B leads on all five. The gaps are not close: 43.5 points on HMMT Feb 2026, 16.6 on GPQA, 15.9 on HLE, 10.2 on Terminal-Bench 2.0, and 4.4 on SWE-bench Pro. Read one way, Qwen3.6-27B is the clearly better model.
Read another way — the way the broader llm-stats sweep frames it — DeepSeek V4 Flash leads on most other coding and agent benchmarks, including Terminal-Bench 2.1 (82.7%), CyberGym (76.7%), Toolathlon (70.3%), and SWE Multilingual (69.7%). Qwen’s wins cluster on a different benchmark family: CountBench (97.8%), AIME 2026 (94.1%), HMMT 2025 (93.8%).
The fair reading: both models are strong, and the leader depends on which tests you trust. BenchLM’s shared-only approach is the more rigorous one; llm-stats’s broader sweep is the more flattering to DeepSeek.
What they cost and what they fit
DeepSeek V4 Flash is roughly 12× cheaper per token than Qwen3.6-27B on a 3:1 input/output mix — the cost gap that decides production agent loops. BenchLM’s three fixed-cost scenarios tell the same story in dollar terms: a single chat turn costs Flash a fraction of a US cent, a 50K-token repository review stays under one cent, and a 200K-token cache-heavy agent loop is comparable. Qwen3.6-27B is hosted via Novita but at roughly 7× Flash’s per-token rate, so Flash remains the production-API pick on cost.
The catch is size. DeepSeek V4 Flash is roughly eleven times larger than Qwen3.6-27B in raw model terms. Both fit a high-RAM desktop after quantisation — the compression trick that lets a model run on less memory at a small quality cost — but Flash needs a heavier quant while Qwen3.6-27B runs comfortably at a lighter one. DeepSeek’s larger context window — how much text the model can read in one go — is around four times Qwen’s, which matters for long documents. Qwen supports image inputs, Flash does not.
What to run on a high-RAM desktop this afternoon
The benchmark scores matter less than the question the community keeps asking: which is the smartest model you can actually run at home on a 128GB workstation? The answer is still DeepSeek V4 Flash, even after quantisation.
Three practical paths for a UK small team:
- Run Qwen3.6-27B first if you want the smoothest ride. It quantises lightly, runs without drama, leads on the shared reasoning tests both trackers publish, and accepts image inputs out of the box. Best for a team that wants multimodal at home and a model that boots in minutes rather than hours.
- Run DeepSeek V4 Flash if agent and coding work matters more. Quantise aggressively, accept the slower tokens, and you get the model that tops Terminal-Bench 2.1, CyberGym, and SWE Multilingual — the agent and coding tests that decide real automation. We covered this same model on a 5090 earlier in the year (A 5090 desktop runs DeepSeek V4 Flash 0731), and the Codex integration keeps it fast.
- Skip the local run and use the API if cost is the lever. Flash at the cheapest hosted rate is the credible choice for production agent loops; Qwen’s hosted tier only makes sense if you specifically need its multimodal input, and running Qwen3.6-27B locally is the more sensible path if you don’t.
The benchmark gap is real but narrow on the tests both trackers share, and Flash dominates on the broader sweep. For a UK team with one workstation and a weekend to spare, Flash is still the smart pick — quantised, careful, and on your own metal. For a team that needs images-in today, Qwen3.6-27B is the smoother answer.
Sources & quotes
Every quotation in this article is verbatim from a named source — click any 1 to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify →


