Best AI for Creative Writing in 2026: Ranked by EQBench and Writing Benchmarks
Which AI writes the best fiction, poetry, and screenplays in 2026? We rank the top models by EQBench creative writing scores and the Lechmazur benchmark.
Two independent benchmarks now give us the clearest picture yet of which AI models actually write well, not just technically correct text, but fiction with emotional weight, consistent characters, and earned narrative turns. The results in February 2026 largely agree: Claude Opus 4.6 leads, but the gap to second place is narrower than it has ever been, and a handful of Chinese models are now competitive at a fraction of the cost.
Here is what the data shows, what the benchmarks actually measure, and which model you should use for each type of creative work.
How we ranked them
Rankings are drawn from two primary sources:
Lechmazur Writing Benchmark: measures a model's ability to weave 10 mandatory story elements into a cohesive short story, scored by an LLM judge panel across an 18-question rubric (github.com/lechmazur/writing).
EQBench Creative Writing: measures emotional intelligence in creative writing across 32 prompts, 3 iterations each, evaluated on 14 dimensions. It also measures a "Slop Score" (GPT-isms and clichés; lower is better) and a "Repetition Score" (eqbench.com/creative_writing.html).
Neither benchmark tests grammar or fluency, every frontier model clears that bar. They test the harder things: whether characters feel real, whether emotion is earned, whether the prose avoids the flattened corporate voice that plagues AI-generated text.
What EQBench measures
EQBench Creative Writing uses 32 prompts with 3 iterations each (96 total samples per model). Each output is scored across 14 dimensions (eqbench.com):
- Nuanced characters: multi-dimensional, believable people, not archetypes
- Emotionally engaging: evokes genuine response in the reader
- Compelling plot: structure that earns its resolution
- Coherence: internal logic across the piece
- Well-earned lightness or darkness: tonal shifts that feel motivated, not manipulative
- Characters consistent with profile: behavior matches established characterization
- Avoids clichés: the Slop Score; lower is better
- Avoids repetition: no recycled phrases or sentence patterns
- Vivid imagery: concrete sensory detail over abstract description
- Authentic voice: distinct stylistic identity
- Emotional depth: subtext and implication, not stated feeling
- Thematic resonance: ideas that linger after the piece ends
- Pacing: rhythm appropriate to the subject matter
- Dialogue quality: exchanges that reveal character and advance plot
The Slop Score is particularly useful for comparing models. It quantifies how often a model reaches for the same tired constructions that flood AI writing: "a testament to," "dance of," "tapestry of," "in the realm of," and similar hollow filler. Claude models score consistently low on this metric.
What Lechmazur measures
The Lechmazur benchmark gives each model 10 mandatory story elements drawn from different categories: a character, an object, a concept, an attribute, an action, a method, a setting, a timeframe, a motivation, and a tone. The model must incorporate all 10 into a single cohesive short story.
An 18-question rubric evaluates the result. Judge panels are themselves LLMs, typically an ensemble of top models to reduce individual model bias.
This is deliberately hard. Weaving 10 specific constraints into a story that reads naturally, not like a checklist, requires genuine narrative intelligence: the ability to motivate why a character in a Victorian conservatory at dusk, with a pocket watch and a sense of unease, would be performing a specific action for a specific reason (github.com/lechmazur/writing).
Top AI for creative writing in 2026
1. Claude Opus 4.6: the benchmark leader
Lechmazur score: 8.561 (thinking mode, 16K) / 8.533 (standard)
Claude Opus 4.6 sits at the top of both major creative writing benchmarks. On Lechmazur, it scores 8.561 in thinking mode with 16K extended reasoning tokens, and 8.533 in standard mode, the highest scores on the board. On EQBench, it leads the creative writing leaderboard with the lowest Slop Score among frontier models.
What makes Claude Opus 4.6 work for creative writing is its handling of emotional subtext. It tends to show rather than tell, trusts the reader, and avoids the AI default of resolving tension too neatly. The characters behave consistently. The prose has a distinct voice that holds across a piece.
Pricing: $5 per million input tokens, $25 per million output tokens. Not cheap, but no current model beats it on quality (claude5.com).
Best for: literary fiction, emotionally complex narratives, character-driven stories, long-form work where consistency across thousands of words matters.
2. GPT-5.2 (medium reasoning): close second
Lechmazur score: 8.511
GPT-5.2 in medium reasoning mode scores 8.511 on Lechmazur, 0.022 behind Claude Opus 4.6 standard, and 0.050 behind Opus in thinking mode. That is a narrow gap on a 10-point scale. In practice, the difference between GPT-5.2 and Claude Opus 4.6 on most creative tasks will come down to stylistic preference rather than objective quality.
GPT-5.2 tends toward cleaner structural execution. Its stories are well-paced and plot-coherent. Where it trails Claude Opus is emotional depth and the Slop Score, it reaches for familiar constructions more often.
Best for: genre fiction, plot-driven stories, users already in the OpenAI ecosystem.
3. GPT-5 Pro
Lechmazur score: 8.474
GPT-5 Pro is the non-reasoning version of the flagship GPT-5 line, sitting at 8.474 on Lechmazur. Solid all-around performance. A reliable choice when you do not want to wait for reasoning tokens and just need good output fast.
4. GPT-5.1 (medium reasoning)
Lechmazur score: 8.438
GPT-5.1 represents the prior generation of OpenAI's reasoning models. Still competitive. The step up to GPT-5.2 is meaningful if you are doing extended creative projects, less so for one-off pieces.
5. Kimi K2-0905: best Chinese model for creative writing
Lechmazur score: 8.331
Kimi K2-0905, from Moonshot AI, places 7th on Lechmazur with a score of 8.331. This is notable: it is the first Chinese model to break into the top 10 of a rigorous creative writing benchmark. On EQBench, Kimi K2 Instruct shows per-dimension scores in the 18-19 range on several criteria, competitive with the top Western models (eqbench.com).
Kimi K2-0905 is an open-weight model (1 trillion total parameters, 32B active) available to download and self-host. At API pricing, it is substantially cheaper than Claude or GPT-5. For cost-sensitive creative workloads that still need above-average quality, it is the value pick.
Kimi K2.5 (the more recent version) scores 8.068 at rank 16 on Lechmazur. The older K2-0905 checkpoint actually outperforms the newer one on this particular benchmark, suggesting the training updates optimized for other dimensions.
Best for: high-volume creative work (blog fiction, social content), writers who want an open-weight model they can fine-tune, cost-conscious users who need better than average output.
6. Mistral Medium 3.1
Lechmazur score: 8.201 (rank 10)
Mistral Medium 3.1 is European-developed and ranks 10th on Lechmazur with 8.201. It punches above its price point. For users who need GDPR-compliant infrastructure with a European provider, Mistral Medium 3.1 is the top creative writing option.
7. DeepSeek V3.2
Lechmazur score: 7.601 (rank 21)
DeepSeek V3.2 is a solid workhorse model that performs well on coding and reasoning but trails the field on creative writing. At rank 21 on Lechmazur with 7.601, it produces competent prose but lacks the emotional precision of the top models. The Slop Score is noticeably higher, DeepSeek defaults to familiar constructions more readily.
It is, however, spectacularly cheap: approximately $0.27 per million input tokens. For high-volume content where "good enough" creative quality is acceptable, product descriptions, basic marketing copy, first drafts for human revision, DeepSeek V3.2 is cost-effective.
Best for: first drafts, high-volume content generation, workflows where a human will substantially revise the output.
8. MiniMax M2.5
MiniMax M2.5 does not rank among the top creative writing models on either benchmark. It is a coding-optimized model that dominated OpenRouter usage charts in February 2026 for technical work, not narrative writing. Do not use it as your primary creative writing tool. Its per-dimension EQBench scores trail the models above significantly.
9. GLM-5
Lechmazur score: 7.452 (rank 25)
GLM-5 from Zhipu AI ranks 25th on Lechmazur with 7.452. It is a strong open-weight model for coding, reasoning, and Chinese-language tasks, but it is not the top choice for English creative writing. The prose tends toward functional rather than evocative.
MiniMax M2.1 (an older generation) scores 7.777 at rank 19 on Lechmazur, slightly above GLM-5 for creative writing quality.
Best AI by creative writing type
Best for fiction
Claude Opus 4.6 is the clear choice for literary fiction and character-driven narrative. Its EQBench scores on Nuanced Characters, Emotional Depth, and Avoids Clichés are the highest in class. For genre fiction where pacing and plot matter more than emotional subtlety, GPT-5.2 is a close alternative.
For writers on a budget who still want genuine creative quality, Kimi K2-0905 (open-weight, self-hostable) is the best value option.
Best for screenwriting
Screenwriting has specific format requirements and dialogue-heavy structure. Claude Opus 4.6 leads on Dialogue Quality in EQBench. GPT-5.2 is competitive and has better default adherence to Final Draft-style formatting conventions.
For television spec scripts and short film work, Claude Sonnet 4.6 is the practical pick, near-Opus quality at substantially lower cost, which matters when you are iterating through multiple drafts.
Best for poetry
Poetry evaluation is harder to benchmark, no standard rubric covers formal verse, free verse, and prose poetry equally. Anecdotally, and from community testing on AfricanAI, Claude Opus 4.6 produces the most surprising line breaks and genuinely non-generic imagery. It avoids the rhyme-forced rhythm that plagues most AI poetry. GPT-5 Pro is the runner-up for longer prose poems.
For traditional formal poetry (sonnets, villanelles, ghazals), Claude Opus 4.6 handles constraint-weaving better than any other model, exactly the skill the Lechmazur benchmark tests.
Free vs paid
Paid models worth the cost
- Claude Opus 4.6: the benchmark leader; worth paying for literary-quality work
- GPT-5.2: close second; reasonable if you prefer OpenAI
Strong free and low-cost options
- Kimi K2.5: free tier available; 8.068 on Lechmazur; open-weight
- DeepSeek V3.2: ~$0.27/M tokens; rank 21; strong for drafts
- MiniMax M2.5: ~$0.30/M tokens; not for creative writing quality, but lowest-cost tier
- Mistral Medium 3.1: competitive pricing; rank 10 Lechmazur; GDPR-friendly
On AfricanAI, you can access Claude Opus 4.6, Kimi K2.5, DeepSeek V3.2, and MiniMax M2.5 through a single interface and compare outputs directly.
Verdict
If creative writing quality is the only variable, Claude Opus 4.6 wins both benchmarks, Lechmazur and EQBench, by a meaningful margin. It earns that position through emotional intelligence, low cliché density, and consistent characterization, not just surface-level fluency.
GPT-5.2 is the closest competitor and the right choice if you are already invested in the OpenAI ecosystem. The quality gap is real but not enormous.
The most interesting data point in 2026 is Kimi K2-0905 at rank 7 on Lechmazur. It is the first open-weight model to genuinely compete with the top commercial models on creative writing. As fine-tuning tooling matures, expect community fine-tunes optimized for specific genres to close that gap further.
For writers who want the best output and can afford it: use Claude Opus 4.6. For everyone else: Kimi K2.5 or DeepSeek V3.2 for drafts, then refine with Opus for the final version.