Please wait...
I tested the big three AI models with real-world writing tasks — emails, code, fiction, and summaries. Here's which one writes best, and which one I'd actually pay for.

By Vera Maxwell
I've spent the last three months putting the big three AI writing models through their paces. Not with benchmarks. Not with academic prompts. With the same messy, real-world tasks I actually need done — emails, blog posts, code documentation, creative fiction, technical explainers, and the dreaded "summarize this 40-page report."
Here's my honest take on which one writes best, which one I'd actually pay for, and which one might surprise you.
Claude 3.5 Sonnet wins on writing quality. GPT-4 wins on versatility. Gemini Ultra wins on research.
If I could only keep one subscription? Claude. But I wouldn't be happy about losing the other two.
Let's be clear about what we're comparing. All three companies ship multiple models, but I tested the flagship tier of each:
Same price, same promise: "the best AI we've got." Let's see who delivers.
I asked each model to write a polite but firm email declining a partnership offer that doesn't align with our current priorities.
Claude nailed the tone immediately. The language was warm but unambiguous — no hedge words, no passive-aggressive qualifiers. Two paragraphs, done.
GPT-4 was solid but slightly corporate. That particular brand of professional prose that reads like it came from a template library.
Gemini was the weakest here. The email was too long, too eager to soften the blow, and included a hollow paragraph about "potential future opportunities."
Winner: Claude, by a clear margin.
I asked each model to write 800 words explaining how database indexing works, aimed at junior developers.
GPT-4 excelled here. The structure was logical, the examples were practical, and it naturally included common pitfalls. Its training data depth in technical content is unmatched.
Claude was close behind, but with a notably different flavor. Where GPT-4 wrote like a senior engineer explaining to a junior, Claude wrote more like a technical writer crafting documentation.
Gemini had the best factual accuracy — it correctly cited B-tree structures and included PostgreSQL-specific details. But the prose was dry.
Winner: GPT-4, narrowly. Claude for readability, Gemini for accuracy.
I asked each model to write the opening page of a noir detective story set in a rain-soaked Tokyo.
Claude absolutely ran away with this one. The prose was atmospheric without being purple. The dialogue had rhythm. The sensory details were specific — not generic "neon lights" stuff, but precise observations that felt authored.
GPT-4 produced competent genre fiction. All the right beats were there, but it felt assembled from parts rather than written whole.
Gemini struggled the most. The creative output was noticeably flatter — fewer risks, more clichés, less personality.
Winner: Claude, decisively. This isn't close.
I fed each model the same 12,000-word policy document and asked for a 500-word executive summary.
Gemini won here, and it wasn't subtle. The summary was accurate, well-structured, and clearly separated key findings from recommendations.
Winner: Gemini, especially for anything where factual completeness matters more than prose style.
I asked each model to write a React component for a sortable data table with pagination. All three produced functional code.
GPT-4 generated the most production-ready code. TypeScript types were correct, edge cases were handled, and it included accessibility attributes without being asked.
Claude wrote cleaner, more readable code. Fewer lines, better variable names, more thoughtful abstractions.
Winner: GPT-4 for production code. Claude for code you'll enjoy reading.
Choose Claude if you care about writing quality above all else. It's the only model where I regularly forget I'm reading AI output.
Choose GPT-4 if you need a versatile all-rounder with the best ecosystem. Plugins, DALL-E integration, browsing, Code Interpreter — the platform is unmatched.
Choose Gemini if your work is research-heavy, factual, and analytical. Gemini's tight integration with Google Search makes it ideal for analysis and research synthesis.
The real story isn't which AI "wins." It's that all three are good enough to be genuinely useful, and different enough that your choice actually matters. That's the sign of a maturing market — and it means the best is still ahead.
Vera Maxwell is Slopthing's resident tech reviewer — testing products in real-world conditions and giving straight answers about what's worth your money.

Sign in to join the conversation.