Hands-on · Local & Open

Test your Qwen 3.8 27B quants

Single-shot tests such as GPQA Diamond and IFBench will rate every Qwen 3.8 27B weight file within a hair of every other. Here is the in-house agentic bench your procurement conversations actually need.

R
RAR Editor
Published October 2026 · 5 min read

Drafted by an AI agent · reviewed and approved by a human editor before publication. How this works.

The Quick Version
  • Public benchmarks flatten every Qwen 3.8 27B quant into a similar band
  • The split lives on long-running coding work where leaderboards can't see it
  • Matt at Kelcode's tool-eval-bench is a working recipe — fork the rubric, change only the quant
  • Completion and human-repair turns beat headline scores for local agent decisions

Two serving recipes from the Qwen 3.8 family land on a public leaderboard within a band so tight the numbers stop meaning anything. The split lives on the long-running coding work your team actually runs.

Where the leaderboards flatten the answer

Last month, Matt at Kelcode put three serving variants of Qwen3.8-Flash-Next and a Qwen3.8-27B control through the same agentic benchmark — a tool-use test that runs a coding agent end-to-end. The benchmark returned clean, separate scores for each serving recipe; standard multiple-choice tests had flattened them. The MyClaw side-by-side of the lighter Qwen 3.8 build against the dense one tells the same story: the widest margins sit on agent benchmarks — a long-horizon developer-test suite scores a 22-point gap, a code-fix tool scores a 16-point gap. The narrowest margins sit on single-shot reasoning.

We’ve hit the theme before in r/LocalLLaMA beats the benchmarks and benchmarks miss what Gemma 4 actually does.

An afternoon bench, step by step

The Kelcode approach is the cleanest worked example. Three serving variants of the lighter build went through identical prompts, tool lists and retry policies, with the dense 27B as a control. Translating that into a compression-level bench for your own team is one afternoon:

  • Pick a real task you actually run. Matt’s example was a 24-hour autonomous coding prompt that produces a small browser game — a finish-able, score-able objective any small firm can swap for a real internal ticket: a CSV-to-typed-report, a refactor, a smoke test against an internal app.
  • Fix everything except the compression level. Same prompt, temperature, reasoning effort, tool list, retry policy and seed. Only the weight file changes. Test two options only — most small teams realistically choose between two tiers on the kit they own.
  • Score completion and corrections, not a single headline. Capture completion (did it finish?) and corrections (turns taken, human-repair minutes after the run). A version that finishes in 80% of retries often beats one that finishes in 100% but eats your evening.

The full prompt used in the Kelcode test is at the bottom of their write-up — fork it into a private repo and you have the bones of a working bench.

What to do this afternoon

  • Don’t pick from a public leaderboard. It tells you almost nothing about finish rate on your agents. Pick the version you ran through your own one-task bench, twice.
  • Keep the rubric tight. Completion, turns-to-finish and human-repair minutes is enough — pick the two or three that map to your team’s actual pain.
  • Start with the smallest option that fits. Aggressive compression will surprise you on small tasks where the model has room to retry. Save the largest variant for the headless run that needs headroom.
  • Log everything in a CSV. Compression level, completion, turns, repair minutes, GPU temperature, average tokens-per-second. That CSV is your procurement-grade evidence when someone asks why this model on this kit?

The bench you’d run to pick a compression level is the same bench you’d run to pick a model. If you’ve got one running on your kit, fork the prompt, swap weights, and the numbers fall out this afternoon.

−50%tokens Qwen3.8-Flash-Next used to finish the same 24-hour coding task as the dense 27B on the same rubric — the same rubric, run on your quant files, surfaces whatever gap is actually there.

Sources & quotes

Every quotation in this article is verbatim from a named source — click any 1 to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify →

  1. Qwen3.8 Flash: Is Faster Better Than Smarter?
  2. Qwen 3.8 Flash-Next vs Qwen 3.8 27B: Which Model Should You Use?
Filed under Hands-on · Local & Open

Continue Reading