News · Models

A writers' leaderboard, not a coder's

Coding benchmarks dominate the AI conversation. A new writing leaderboard, judged by 5,000 blind tests from working writers, has a different shortlist — and a more useful one for most readers.

R
RAR Editor
Published October 2026 · 5 min read

Drafted by an AI agent · reviewed and approved by a human editor before publication. How this works.

The Quick Version
  • Hemingway-bench, published 29 September 2026, ran 5,000+ blind pairwise comparisons judged by professional writers.
  • Gemini 3 Flash, Gemini 3 Pro and Claude Opus 4.5 took the top three slots; GPT-5.2, Qwen3, Kimi K2, Grok, Llama 4 Maverick and Nova filled the rest.
  • The expert panel found EQ-Bench's autograder agreed with them as little as 43% of the time, rewarding metaphor-stuffed prose over coherence.
  • Qwen3 was the most original open-weight entry but the most error-prone; Llama 4 Maverick was the most verbose; Kimi K2 was best on business.
  • Match the model to the job — the runner-up often beats the leader for everyday emails, business docs and voice-led pieces.

A new AI writing benchmark, judged by working writers rather than by an autograder, has put Google’s Gemini 3 Flash and Gemini 3 Pro and Anthropic’s Claude Opus 4.5 at the top of its table. Hemingway-bench, published on 29 September by Surge AI, was built from more than 5,000 blind pairwise comparisons run by professional screenwriters, poets, speechwriters and copyeditors. The goal, the team said, was to score the qualities a reader actually cares about: voice, restraint, originality and the ability to follow a brief.

What the judges found

The full shortlist covered ten named models. Closed flagships took the first three places, with Gemini 3 Flash described as a master wordsmith that loved creative constraints, Gemini 3 Pro building rich worlds with specific detail (afternoon light in a 1950s Philippine kitchen, the way a grandmother made adobo), and Opus 4.5, which the judges described as having the most human-like voice of the field.

The runners-up, in the judges’ own personas, were where the picture got useful:

  • Qwen3 — the most original model on the list, but the most prone to factual errors and forced humour. A high-risk, high-reward pick for creative work.
  • GPT-5.2 Chat — solid for practical everyday tasks, from coordination emails to friendly advice, but not flashy on creative briefs.
  • Kimi K2 — competent on business writing, strained on creative. One judge called it an executive assistant who had studied creative writing long ago.
  • Grok — informal and trope-heavy; worked for casual emails, not much else.
  • Llama 4 Maverick — the most verbose model on the list and the one most given to placeholder phrasing on creative prompts.
  • Nova — structure in place, but a deeper spark missing.

Why the autograder got it wrong

The Surge team built Hemingway-bench in part because they had watched EQ-Bench’s autograder get reward-hacked. In their analysis, the automated grader’s verdicts aligned with expert writers’ verdicts on the same pairs as little as 43% of the time, the Hemingway-bench team said. The worked example they published was a short-story prompt that asked for a Rita Mae Brown voice. The model EQ-Bench ranked second overall, the judges wrote, leaned so hard on metaphor that the prose became incoherent; the model that EQ-Bench scored lower, Gemini 3 Flash, won unanimously among the human judges.

43%of the time, EQ-Bench’s autograder agreed with expert writers on which model wrote better.

The wider argument is that automated rubrics reward the surface features of good writing — many literary devices, varied sentence structure, dense vocabulary — while readers respond to the cumulative effect. Hemingway-bench scores each response on a handful of sub-dimensions (full list in the tech box) and then takes a holistic view on top.

How to use this for your own writing

The clearest read of the Hemingway-bench results is that the right model depends on the job, and the runner-up often beats the leader for the thing you actually need to write today.

  • Creative writing — short stories, scripts, poetry. Gemini 3 Flash or Opus 4.5, with Qwen3 as a free local option if you can stomach the factual-error risk and have the hardware to run it (see the tech box) (Qwen 3.8 27B is ready to download for heavier briefs).
  • Business and professional writing — emails, marketing copy, internal docs. GPT-5.2 Chat, with Kimi K2 a credible open alternative.
  • Heartfelt, voice-led pieces — speeches, personal notes, obituaries. Opus 4.5, with Gemini 3 Pro a strong second.
  • Long-form and translation-heavy work — models with longer-context windows earn their keep on documents the smaller ones cannot hold (specific token budgets in the tech box). For a worked example of the open-weights end, see our Llama 4 Scout piece. MindStudio’s text-generation guide is a useful general primer if you want to map these strengths across more use cases.

Try the leaderboard prompts yourself this afternoon — a bedtime story, a thank-you email to a school, a Rita Mae Brown pastiche — and grade the output with your own rubric. The model you keep coming back to is the one for you; the one that scores well on a checklist is a different question, and one the Hemingway-bench team is asking for a reason. For a practical stack reference, our free AI tiers guide and our LLM tier roundup still hold: match the model to the task, and don’t pay flagship prices for a thank-you note.

Sources & quotes

Every quotation in this article is verbatim from a named source — click any 1 to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify →

  1. Hemingway-bench: The AI Writing Leaderboard Judged by Expert Writers
  2. Choosing the Right AI Model for Text Generation | MindStudio
Filed under News · Models

Continue Reading