The honest answer to "what's the best LLM for creative writing?" in October 2026 is that it depends on which judge you trust — because the three kinds of leaderboard crown three different winners. On the automated rubric board most people cite for fiction, EQ-Bench Creative Writing, Anthropic's creative tier Claude Fable 5.1 trades the top with GPT-6 Astra. On the human-vote boards, Google's Gemini line wins. And the model with some of the best benchmark scores — GPT-6 Astra — is the one experienced writers most often tell you to avoid for prose. So before the picks, the one idea that makes this whole question tractable:

The boards disagree because they ask different judges. A rubric-tuned model can top EQ-Bench and still read like a competent robot; a human-vote winner can feel alive and miss half the brief. Pick the board whose judge is closest to your reader.

Here's the decision, by what you're actually writing:

That's the pick. The rest is why the boards split and the three settings that beat a model upgrade — which is where most people are leaving quality on the table.

Why one model can be #2 and #20 at the same time#

Three kinds of judge, three kinds of answer:

A model can top the rubric and lose the human vote for feeling mechanical. So the ranking isn't noise — it's a signal about which kind of good a model is. Want structural competence and reliable instruction-following? Trust the rubric board. Want prose a person enjoys? Trust the human vote. If you're writing fiction for human readers, I'd weight the human-vote and human-expert boards over the rubric — which is exactly why Fable and Gemini, not the top rubric score, lead the picks above. (For everyday non-fiction — drafts, docs, marketing copy — the calculus is different and the model matters less; we cover that in the best LLM for writing.)

The three settings that beat a model upgrade#

Most people pick a model and leave it on defaults tuned for chat. For creative work, that's backwards — the settings carry more of the quality than the last half-tier of model does.

  1. Temperature 0.8–1.1. The ~0.7 chat default is too conservative; nudging it up buys more surprising word choice. Too high and it loses the thread — which is why the next lever matters.
  2. Temperature + Min-P, not top-p. The 2026 local-writing consensus is pairing a higher temperature with Min-P sampling (~0.05–0.1), which keeps the output coherent at temperatures where top-p would wander. A common baseline: temp 0.95, Min-P 0.05.
  3. A DRY repetition penalty. Over a long piece, models fall into repeated phrasings and tics. DRY (in llama.cpp, Ollama, and most runners) suppresses repeated n-grams far better than a flat frequency penalty — the difference between prose that develops and prose that loops.

Two more rules on top. Turn reasoning off for drafting. Thinking modes help you outline and untangle plot logic, but on the sentence itself they tend to produce stiffer, over-explained prose — good writing is fluency, not deliberation. Reason about the plot, then draft without it. And write a voice system prompt: POV, tense, rhythm, a short list of banned clichés, and one sample paragraph in the target style. In blind tests, a strong voice profile on a mid-tier model beats a top-tier model running naked. The model you can steer beats the model that merely scores — so pick for steerability, set it up right, and spend your real effort on the brief.